champollion-mcp-server 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +468 -0
- package/README.md +135 -82
- package/bin/server.js +4 -3
- package/instructions.md +188 -232
- package/package.json +9 -5
- package/src/index.js +1364 -297
- package/src/tools/args.js +90 -0
- package/src/tools/contests.js +474 -0
- package/src/tools/corpora.js +423 -0
- package/src/tools/forge.js +1076 -78
- package/src/tools/harness-fst.js +404 -0
- package/src/tools/harness.js +1918 -188
- package/src/tools/job-supervisor.js +96 -0
- package/src/tools/jobs.js +408 -0
- package/src/tools/language-card.js +697 -0
- package/src/tools/languages.js +1361 -42
- package/src/tools/metrics-plan.js +287 -0
- package/src/tools/output-trim.js +177 -0
- package/src/tools/overview.js +423 -0
- package/src/tools/plan-notes.js +203 -0
- package/src/tools/plural.js +14 -0
- package/src/tools/publish-preview.js +191 -0
- package/src/tools/publish-report.js +382 -0
- package/src/tools/queue.js +413 -82
- package/src/tools/register-corpus-hint.js +77 -0
- package/src/tools/releasable.js +80 -0
- package/src/tools/reliability.js +114 -21
- package/src/tools/results.js +187 -20
- package/src/tools/run-plan.js +608 -0
- package/src/tools/state.js +40 -0
- package/src/tools/training.js +76 -18
- package/src/tools/translate.js +1612 -77
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,468 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.2.0 (2026-10-03)
|
|
4
|
+
|
|
5
|
+
Built around one flow: someone tells their agent "let's build a Cree model
|
|
6
|
+
for our school" (or "an Atya phrasebook for the clinic") and the agent can
|
|
7
|
+
take them from discovery to a measured, deployable model. Synthetic-user
|
|
8
|
+
testing against the 0.1.x npm install found the gaps this release closes.
|
|
9
|
+
|
|
10
|
+
- **Round 13 (synthetic-user findings, 2026-10-04).**
|
|
11
|
+
- New `preview_publish`: the READ-ONLY half of `publish_report` — the
|
|
12
|
+
harness's own `mt-eval publish <report> --dry-run`, WHAT GETS PUBLISHED,
|
|
13
|
+
the exact `publish_ack` and the exact `publish_report` call that would
|
|
14
|
+
publish it. It has no `confirm` or `publish_ack` (refused by name) and
|
|
15
|
+
carries the MCP annotation `readOnlyHint: true`, so an agent host that
|
|
16
|
+
gates writes can allow it alone (a host had blocked publish_report's
|
|
17
|
+
preview as a production deploy). `publish_report` is annotated
|
|
18
|
+
`destructiveHint` / `openWorldHint`; without `confirm` it still returns
|
|
19
|
+
the preview. 34 tools.
|
|
20
|
+
- `run_benchmark`: `metricx` (+ `metricx_model`) and `fuse` pass the
|
|
21
|
+
harness's opt-in `--metricx` / `--metricx-model` / `--fuse`; `comet`
|
|
22
|
+
requires COMET (the harness computes it whenever unbabel-comet is
|
|
23
|
+
installed — there is no run flag). The plan says, from the harness's
|
|
24
|
+
Python, whether each is installed, what to install and what it
|
|
25
|
+
downloads; a confirmed run that asks for a metric the harness cannot
|
|
26
|
+
compute is REFUSED, never run without it. Item and corpus runs only.
|
|
27
|
+
- `run_benchmark`: a corpus inside a folder `mt-eval contest prepare`
|
|
28
|
+
marks releasable (`.champollion-releasable.json` in it or above it, or a
|
|
29
|
+
`public/` beside `local/manifest.json`) runs into `<contest>/runs/` —
|
|
30
|
+
results and cache — never into the released folder; the plan's
|
|
31
|
+
`Results:` line says where and why. No output folder inside a releasable
|
|
32
|
+
folder is ever passed to the harness.
|
|
33
|
+
- The coached plan gives ONE verdict on the coaching file, the harness's
|
|
34
|
+
sentence: ✓ it names the language (as "…"), ⚠ it names neither, or not
|
|
35
|
+
checked — the 'Not sent: …' line is gone.
|
|
36
|
+
- A method plugin's cost says why, in the harness's words
|
|
37
|
+
(`method_loader.plugin_cost_basis`, mirrored and compared), in the plan
|
|
38
|
+
and the start message.
|
|
39
|
+
- `language_overview`: speaker counts keep their scope notes (ELCat's
|
|
40
|
+
"10-99" for Plains Cree is British Columbia only) and say when two
|
|
41
|
+
records of one source have nothing on the card telling them apart. Its
|
|
42
|
+
FST line and the run plan read the analyzer and the pyhfst runtime from
|
|
43
|
+
the harness's `config.fst_state` (forge's reading), and name the Python
|
|
44
|
+
they asked.
|
|
45
|
+
- The forge tools relay mt-eval's `score_caveats` as forge gives them
|
|
46
|
+
(never reworded or recomputed): `forge_export`'s summary (and each
|
|
47
|
+
twin-free sibling's), `forge_status`'s exports
|
|
48
|
+
(`summary.export_caveats`), `forge_compare` (`summary.score_caveats`)
|
|
49
|
+
and `forge_lint`'s `R9-harness-score-caveat`; a MAJOR one leads the next
|
|
50
|
+
step — a score is never "the number to quote" without it. No caveat
|
|
51
|
+
list → nothing said. `forge_export` also returns `hypotheses` (the file
|
|
52
|
+
`forge_compare` takes) and its compare command; `forge_status` in
|
|
53
|
+
`no-dev-set` says to call `get_training_guardrails` once, then
|
|
54
|
+
`forge_split`.
|
|
55
|
+
- **Round 12 (synthetic-user findings, 2026-10-04).**
|
|
56
|
+
- `run_benchmark`: a method plugin whose method.json declares dependency
|
|
57
|
+
class S or O with no `gateway` / `external-api` dependency is "$0 API
|
|
58
|
+
cost (runs on this machine)" in the plan, by the harness's rule (it was
|
|
59
|
+
"unknown (the plugin prices its own calls)"); it still needs the user's
|
|
60
|
+
attestation for a local-only corpus. A coached run's plan on a
|
|
61
|
+
local-only corpus says why the coaching file's first line is withheld.
|
|
62
|
+
- `get_run_status`: the Results block lists every score caveat the report
|
|
63
|
+
records (the harness's new near-constant-output caveat among them).
|
|
64
|
+
- `publish_report`: for a local-only corpus, WHAT GETS PUBLISHED relays
|
|
65
|
+
which of its facts go public with the score, which stay on this
|
|
66
|
+
machine, and how others read a score on a set they cannot open.
|
|
67
|
+
- `forge_export` returns a summary like the other forge tools;
|
|
68
|
+
`forge_status` names the real next step (export, or the re-audit when
|
|
69
|
+
the twin-free corpus predates the dev split) and a `serving` state;
|
|
70
|
+
`forge_register_eval` refuses a missing file with the path it looked
|
|
71
|
+
for and the fix.
|
|
72
|
+
- The `register-corpus` command the tools suggest for a test set carries
|
|
73
|
+
`--role test`; `search_languages` docs say a location is shown only with
|
|
74
|
+
its source, otherwise a link to its Glottolog record.
|
|
75
|
+
- **Round 11 (synthetic-user findings, 2026-10-04).**
|
|
76
|
+
- `run_benchmark` names a language given as a code ("sme" → "Northern
|
|
77
|
+
Sami", through the card) instead of prompting with the code; the plan
|
|
78
|
+
shows the prompt the model gets (a coaching file by its first line and
|
|
79
|
+
hash, and that it REPLACES the built-in prompt); for a card with several
|
|
80
|
+
scripts it reports the references' script share (counts only) and the
|
|
81
|
+
script it prompts for; it names the translation cache folder; and it
|
|
82
|
+
prints the exact `compare` command for runs on a registered corpus id.
|
|
83
|
+
- `search_languages` matches the start of a name word ("North Sami" finds
|
|
84
|
+
Northern Sami), ranked below exact matches; a candidate whose location
|
|
85
|
+
cannot be cited links its Glottolog record instead.
|
|
86
|
+
- `language_overview` takes its runnable open models from the CLI's
|
|
87
|
+
`recommend` (one rule, `-ct2` and ONNX repos excluded).
|
|
88
|
+
- `translate`: each row is marked `cache: pair|fallback` and every count
|
|
89
|
+
is derived from the rows; a discarded cached answer names what it
|
|
90
|
+
collided with; an unpriced or unknown model reports "cost unknown —
|
|
91
|
+
<why>".
|
|
92
|
+
- **Round 10 (synthetic-user findings, 2026-10-04).** Next steps, register
|
|
93
|
+
hints and the run plan put the forge steps (register, leak-audit,
|
|
94
|
+
predictions) before any baseline; `local-model` requires a model;
|
|
95
|
+
`allow_model_pair_mismatch`; `publish_report` repeats the harness's method
|
|
96
|
+
lines; `get_results` says publishing is explicit; locations are shown only
|
|
97
|
+
with a source.
|
|
98
|
+
- **Round 9 (synthetic-user findings, 2026-10-04).**
|
|
99
|
+
- New `forge_compare`: compares two systems on a registered eval set and
|
|
100
|
+
prints the near-twin caveat when either was trained on near-copies of
|
|
101
|
+
the test rows ("never the winner alone"). instructions.md now lists
|
|
102
|
+
which forge commands have a tool and which run in a terminal.
|
|
103
|
+
- `forge_prereg` takes `config_hash`; a test checks that every flag
|
|
104
|
+
forge's advice names is an argument of the tool that runs that command.
|
|
105
|
+
- `forge_leak_audit` takes `companion_config` and relays the twin-free
|
|
106
|
+
model's config and train command; `forge_status`,
|
|
107
|
+
`forge_register_eval` and `forge_leak_audit` name the preregistration
|
|
108
|
+
as the next step, before any benchmark, once a test set is registered.
|
|
109
|
+
- `get_training_guardrails` points to the public diagnosing-training
|
|
110
|
+
page, never a monorepo path.
|
|
111
|
+
- `run_benchmark` passes the resolved `--target-lang-code` (a
|
|
112
|
+
private-use `qaa` included), takes `script` for two-script targets
|
|
113
|
+
(`--target-script`), and shell-quotes every printed argument.
|
|
114
|
+
- `language_overview` on a private-use code (qaa–qtz) explains what the
|
|
115
|
+
code is and what still applies, instead of a dead end.
|
|
116
|
+
- `translate` reports `estimated_api_cost_label` ("$0 API cost (runs on
|
|
117
|
+
this machine)" for a local model) instead of "0 USD".
|
|
118
|
+
- `get_metric_reliability` no longer prints internal review wording.
|
|
119
|
+
- **Round 8 (synthetic-user findings, 2026-10-04).**
|
|
120
|
+
- `get_run_status`: the output trim keeps warning and notice lines (⚠,
|
|
121
|
+
✗, ❌, `[WARN`, `Warning:`, `Note:`) and their indented continuation
|
|
122
|
+
lines from the trimmed middle, in place, up to 2,000 characters. Every
|
|
123
|
+
other line is still counted in a marker. A run's "COMET not computed"
|
|
124
|
+
notice used to sit exactly where the trim cut.
|
|
125
|
+
- `publish_report` passes `--anonymous` to the harness's dry run as well,
|
|
126
|
+
and relays the preview's trust tier (self-benchmarked, `unverified`)
|
|
127
|
+
and score lane.
|
|
128
|
+
- `run_benchmark` plans: a missing FST alone reads "the run PROCEEDS
|
|
129
|
+
without it" (harness 2026-10-04+), with the setup command and the
|
|
130
|
+
`mt-eval test <run log>` re-score. The overview's FST line uses the
|
|
131
|
+
harness's own sentence and never says the FST "downloads on the first
|
|
132
|
+
evaluation".
|
|
133
|
+
- `search_languages` matches the Roman-letter form of a card's
|
|
134
|
+
non-Latin endonym, produced by the CLI's script converter and labelled
|
|
135
|
+
as derived (`nêhiyawêwin` finds crk).
|
|
136
|
+
- `run_benchmark` no longer asks for `target_language` when one was
|
|
137
|
+
passed for a code with no card (private-use `qaa`).
|
|
138
|
+
- `translate` says when a cached answer was discarded by the
|
|
139
|
+
shared-output check, and does not count it as served.
|
|
140
|
+
- **Round 7 (synthetic-user findings, 2026-10-04).**
|
|
141
|
+
- `translate` with `project_dir` resolves language NAMES to the project's
|
|
142
|
+
own locale codes ("English"/"French" → en/fr, through the CLI's
|
|
143
|
+
`resolveLanguageInput`) before matching the pair and before any cache
|
|
144
|
+
read or write; an ambiguous name refuses with the choices. Names used to
|
|
145
|
+
miss the pair, run the engine default model and write cache entries
|
|
146
|
+
under a locale literally named "French".
|
|
147
|
+
- `publish_report`: publish a finished run's `*_report.json`, scores-only
|
|
148
|
+
if the user wants (`scores_only`, `redact_coaching`, `anonymous`), behind
|
|
149
|
+
the same gate as `run_benchmark` — every call shows the harness's own
|
|
150
|
+
`mt-eval publish --dry-run` preview and the exact `publish_ack`; only
|
|
151
|
+
`confirm` with those words publishes (`--prod` for production).
|
|
152
|
+
- `run_benchmark` plans open with the target language's `EVAL PACK:`
|
|
153
|
+
status (missing — with the setup command — / ready / none needed) and
|
|
154
|
+
name the corpus licence and `do_not_train` term the run accepts by
|
|
155
|
+
passing `--yes` ("unknown" when nobody states them), asked of the
|
|
156
|
+
harness's registry and its own eval-pack gate. New `skip_fst` /
|
|
157
|
+
`skip_eval_standard` (`--skip-fst` / `--skip-eval-standard`) score
|
|
158
|
+
without them, marked not computed; a run the harness stops for a missing
|
|
159
|
+
pack says so and names both.
|
|
160
|
+
- `forge_split` takes `near_dupe` (`--near-dupe`) and `max_group`
|
|
161
|
+
(`--max-group`, with `near_dupe`) and relays forge's near-twin advice
|
|
162
|
+
instead of promising `--near-dupe 0.6`.
|
|
163
|
+
- `get_run_status` keeps a long log's first and last lines and trims the
|
|
164
|
+
middle with `… [N lines trimmed] …` (it began mid-sentence); the
|
|
165
|
+
harness's `EVAL PACK` lines lead the answer.
|
|
166
|
+
- The register-corpus command the guidance prints now runs as printed
|
|
167
|
+
(`--yes`, `--domain`; it writes the local-only sidecar itself).
|
|
168
|
+
- Significance wording: a small Δ can be significant yet not meaningful
|
|
169
|
+
(check the CI on Δ and metric reliability; p-values are per metric,
|
|
170
|
+
uncorrected) — no more "probably noise" beside a significant result.
|
|
171
|
+
- Counts agree with their nouns: "1 is marked do_not_train", "1 report".
|
|
172
|
+
- **Round 6 (synthetic hospital, school and researcher personas, 2026-10-04).**
|
|
173
|
+
- `run_benchmark { publish: true }` shows WHAT GETS PUBLISHED before
|
|
174
|
+
anything runs — every row with its sentence text, or scores only; the
|
|
175
|
+
prompt published (a coaching file in full), redacted, or none; the
|
|
176
|
+
target — read from the harness's own publish gates. A real publish needs
|
|
177
|
+
`publish_ack` in the exact words the plan prints; without them (or when
|
|
178
|
+
the harness cannot be asked) it is REFUSED and nothing runs.
|
|
179
|
+
- One attestation rule everywhere: the harness's own `method:
|
|
180
|
+
"local-model"` runs in-process and needs no attestation (an
|
|
181
|
+
`attest_local_transport` for it is refused as meaningless); MT engines
|
|
182
|
+
and `method_dir` plugins still need the user's. The server instructions,
|
|
183
|
+
the refusal message and the tool descriptions used to give two answers.
|
|
184
|
+
- `model` with `method_dir` is passed to the plugin (`-m`, its own naming),
|
|
185
|
+
as the methods spec and the CLI say; `provider` with a method stays
|
|
186
|
+
refused. A local model directory is accepted as `model` for
|
|
187
|
+
`local-model` too.
|
|
188
|
+
- `source_language` for corpus mode (`--source-lang`); when the steward's
|
|
189
|
+
sidecar or the corpus card it names states the pair, its source code is
|
|
190
|
+
passed (`--source-code`) and the harness names the language — the plan
|
|
191
|
+
says which. The run card's source language used to be blank.
|
|
192
|
+
- One cost rule: a loopback or in-process run is "$0 API cost (runs on
|
|
193
|
+
this machine)" in the plan as in the start message and the run card (the
|
|
194
|
+
plan said "unknown, never $0").
|
|
195
|
+
- `forge_prereg_verdict`: record the USER's verdict on a free-text
|
|
196
|
+
prediction (ledgered with who, when and a note; shown as a human verdict).
|
|
197
|
+
- `forge_status` after `forge_init` reports `initialized` and points at
|
|
198
|
+
`forge_split` (it said "discover, then init" again); `summary.runs` lists
|
|
199
|
+
every run with its checkpoint, dev score and exports, and
|
|
200
|
+
`summary.warnings` carries a saturated dev set. `forge_report` on an
|
|
201
|
+
exported run includes the test result.
|
|
202
|
+
- `translate` with `project_dir` runs the project pair's `fallback` exactly
|
|
203
|
+
as `champollion sync` does (the CLI's own `translateWithFallback`) and
|
|
204
|
+
marks each text the fallback produced, with its method.
|
|
205
|
+
- `search_languages` matches a multi-word query part by part
|
|
206
|
+
("Plains Cree nêhiyawêwin" → crk), ranked by how much of the query each
|
|
207
|
+
language's recorded names cover.
|
|
208
|
+
- `language_overview` separates "an FST exists (the card)" from "the
|
|
209
|
+
installed harness has a pin for it and can use it here".
|
|
210
|
+
|
|
211
|
+
- **`translate` with `project_dir` shares `champollion sync`'s cache —
|
|
212
|
+
exactly.** The tool keyed entries itself: the resolved code (`fra`) for the
|
|
213
|
+
project's `fr`, a hash of the register text for the card preset sync keys
|
|
214
|
+
by name (`formal-vous`), and an endpoint suffix sync never adds. A persona
|
|
215
|
+
got 0 of 2 hits on strings sync had just cached. With `project_dir` the
|
|
216
|
+
pair now comes from the CLI's own config and pair code — what
|
|
217
|
+
`champollion sync --method <method> [--model <model>]` runs there — so
|
|
218
|
+
model, register, script, cache key and locale code are sync's, in both
|
|
219
|
+
directions. The answer names the project pair and the key. Without
|
|
220
|
+
`project_dir` the server's own cache keeps its keys. A register given as a
|
|
221
|
+
card preset name is now sent as that preset's instructions (the name itself
|
|
222
|
+
used to reach the model).
|
|
223
|
+
- **`translate` never picks an orthography.** A target with two real writing
|
|
224
|
+
systems (Plains Cree: SRO or Syllabics) is refused until one is chosen, as
|
|
225
|
+
`champollion sync` refuses: pass the new `script` argument (`"Latn"`,
|
|
226
|
+
`"Cans"`), or let the `project_dir` pair's `script` decide. A chosen display
|
|
227
|
+
script is produced from the cached working-script translation, exactly as
|
|
228
|
+
sync writes it.
|
|
229
|
+
- **`translate` reports `isError: true` when no text was translated** (an
|
|
230
|
+
unreachable engine, say: it answered "0 of 2" as a success). A partial
|
|
231
|
+
result stays a normal answer that lists each failure.
|
|
232
|
+
- **`search_languages` tells same-named languages apart.** Every result says
|
|
233
|
+
where the language is spoken (countries, Glottolog's point, macroarea) and
|
|
234
|
+
its other names, each with its source ("Atya" → six Ayta languages, all in
|
|
235
|
+
the Philippines: only their points differ). Ties are called out. In an npm
|
|
236
|
+
install, name-only results are filled in from their published cards (first
|
|
237
|
+
10, within 6 s, cached).
|
|
238
|
+
- **One argument name for one language.** Every tool that takes one language
|
|
239
|
+
(`search_languages`, `get_language`, `language_overview`,
|
|
240
|
+
`get_metric_reliability`, `forge_discover`, `forge_init`) also accepts
|
|
241
|
+
`language`; `get_metric_reliability {"language": "crk"}` no longer fails
|
|
242
|
+
with -32602. The original names still work; a missing or conflicting
|
|
243
|
+
language is a tool error naming both arguments. README, instructions.md and
|
|
244
|
+
the champollion.dev MCP page list every tool's arguments, and tests hold
|
|
245
|
+
them equal to the schemas.
|
|
246
|
+
- **`forge_status` knows the `training` state.** While a run holds the
|
|
247
|
+
workspace run lock, nmt-forge reports `training`; the tool's description
|
|
248
|
+
lists it and its hint says wait — never start another run — instead of the
|
|
249
|
+
generic `nmt-forge run config.json` example.
|
|
250
|
+
- **`language_overview` shows every source on a disputed fact.** The summary
|
|
251
|
+
line kept three unattributed endangerment values (Plains Cree has five,
|
|
252
|
+
from three sources) and joined disputed families without their sources; it
|
|
253
|
+
now lists each value with its source and says the sources differ.
|
|
254
|
+
|
|
255
|
+
- **New `language_overview`** — the stage-1 answer. One honest page per
|
|
256
|
+
language: what the index knows, which benchmarks exist, published results,
|
|
257
|
+
which methods can run here and with what evidence (the CLI's own
|
|
258
|
+
`champollion recommend` logic), tooling (FSTs, dictionaries, keyboards),
|
|
259
|
+
licence/consent constraints read from the corpora, contests, and numbered
|
|
260
|
+
next steps that each name the exact tool or command (protect your data →
|
|
261
|
+
baseline → build → prove → deploy). Every section that cannot be reached
|
|
262
|
+
says "unavailable: why" instead of sinking the answer.
|
|
263
|
+
- **New `get_language`** — the full cited card for one language, resolved
|
|
264
|
+
exactly as the `champollion` CLI resolves it (bundled card → per-user cache
|
|
265
|
+
→ champollion.dev's published card tables, via the package's own async
|
|
266
|
+
prefetch). Every value carries its source; disagreements list every claim;
|
|
267
|
+
absent fields are named, with what absence means for the tier the card came
|
|
268
|
+
from. From an npm install, most low-resource languages used to answer
|
|
269
|
+
"family: unknown, speakers: unknown".
|
|
270
|
+
- **Fuzzy `search_languages`.** No exact or whole-word hit → the closest
|
|
271
|
+
names by restricted Damerau-Levenshtein distance (a swapped letter pair
|
|
272
|
+
costs ½; budget 1 edit for 4–5 letters, 2 beyond; names only, never codes).
|
|
273
|
+
"Atya" now surfaces the six Ayta languages first. Matching is accent- and
|
|
274
|
+
punctuation-insensitive, and a whole-word hit ("Cree" in "Plains Cree")
|
|
275
|
+
outranks a prefix ("Creek").
|
|
276
|
+
- **`run_benchmark` no longer publishes by default.** 0.1.x auto-published
|
|
277
|
+
every budget/top queue run to the PRODUCTION leaderboard unless the agent
|
|
278
|
+
passed `publish:false`. Now nothing is published unless `publish: true`;
|
|
279
|
+
every plan, start and completion message names the target (production, or
|
|
280
|
+
the non-production project in `MT_EVAL_SUPABASE_URL`), and a production
|
|
281
|
+
publish passes the harness's separate `--prod` opt-in.
|
|
282
|
+
- **`run_benchmark` corpus mode + local models.** Run ANY corpus — a registry
|
|
283
|
+
id or a test file the user holds — on any model: `provider: "local"` (an
|
|
284
|
+
OpenAI-compatible server on this machine; Ollama by default, `base_url` /
|
|
285
|
+
`LOCAL_API_BASE` otherwise), `method: "local-model"` (NLLB / OPUS-MT /
|
|
286
|
+
MADLAD weights), a hosted provider, or an MT API. Optional `max_cost`,
|
|
287
|
+
`attest_no_training`, `accept_nc_terms`, `anonymous`, field names.
|
|
288
|
+
- **Refusals stay refusals.** A data steward's `<file>.champollion.json`
|
|
289
|
+
`{"transmission":"local-only"}` mark refuses every remote model up front
|
|
290
|
+
(no job, no spawn). The harness's transmission-policy, NC-terms, prod-
|
|
291
|
+
publish and cost-cap stops come back from `get_run_status` as REFUSED with
|
|
292
|
+
what IS allowed — never as a generic FAILED, and never with advice to try
|
|
293
|
+
another provider.
|
|
294
|
+
- **Cost "unknown" is never $0.** `get_results` shows each row's cost and
|
|
295
|
+
"unknown" for unpriced/local runs (sorted last by cost); `get_run_card`
|
|
296
|
+
says so when `total_cost_usd` is null.
|
|
297
|
+
- **New read-only contest tools `list_contests` / `get_contest`.** Phases
|
|
298
|
+
(with the active window), the organizer's declared terms (the frozen
|
|
299
|
+
PROMISE keys, mirrored from `contest_policy.py` and parity-tested) with a
|
|
300
|
+
computed SHA-256 digest, results visibility, and the public ranking when one
|
|
301
|
+
is visible (the frozen final ranking with CIs and tie groups after close;
|
|
302
|
+
interim published scores when results are immediate). Bounded anon reads;
|
|
303
|
+
an entrant's sign-in email (`submitted_by`, `created_by`, `closed_by`) is
|
|
304
|
+
never read or shown; what an anonymous reader cannot see is said plainly.
|
|
305
|
+
Entering a contest remains a human-authorized CLI flow.
|
|
306
|
+
- **forge tools work from a pip install.** forge resolves as
|
|
307
|
+
`NMT_FORGE_BIN` → `CHAMPOLLION_FORGE_DIR` (a clone; set-but-wrong is an
|
|
308
|
+
error) → the monorepo sibling → `nmt-forge` on `PATH` → `python -m
|
|
309
|
+
nmt_forge.cli` from the active Python. The misleading "forge is not on
|
|
310
|
+
PyPI" message is gone; a missing forge now says `pip install nmt-forge`
|
|
311
|
+
(and `pip install 'nmt-forge[hf]'` to train), and common cold-start
|
|
312
|
+
tracebacks (missing harness, missing torch, no card directory) map to fixes.
|
|
313
|
+
- **forge tools speak nmt-forge 0.2.0's `--json` contract.** Every forge
|
|
314
|
+
tool passes `--json` and returns `{result, summary?, next?}` (forge 0.2.0
|
|
315
|
+
prints human text by default for `init`, `split`, `leak-audit` and
|
|
316
|
+
`registry add`, which the old text fallback relayed without structure). A
|
|
317
|
+
refusal — `{"error": {type, guard, message, why, fix, …}}`, exit 2 — comes
|
|
318
|
+
back as a tool error with what / why / fix plus the envelope; a pre-0.2.0
|
|
319
|
+
forge, a crash or a missing dependency comes back as the fix, never a
|
|
320
|
+
traceback. Each call is bounded (2 min; 10 min for evaluate/export): past
|
|
321
|
+
the bound forge gets SIGINT, so its own cleanup runs, then SIGKILL, and the
|
|
322
|
+
tool names the terminal command.
|
|
323
|
+
- **New `forge_export`**: score the test battery once (prereg-gated, CIs)
|
|
324
|
+
and package the model, an mt-eval RunLog + TestReport, a champollion
|
|
325
|
+
`method.json` and DEPLOY.md. **New `forge_prereg_template`**: the one
|
|
326
|
+
valid predictions format, written to edit.
|
|
327
|
+
- `forge_preflight` takes the `evaluate` / `export` / `serve` targets and a
|
|
328
|
+
`config`; a failing gate (forge exits 2) is an answer — `summary.passed`,
|
|
329
|
+
`summary.failing` — not a tool error.
|
|
330
|
+
- `forge_init` takes `model` (`cpu-tiny` default, `cpu-finetune` + `base`,
|
|
331
|
+
`nllb-600m`), `no_card` + `name` and `cards_dir`, and drops `workspace`
|
|
332
|
+
(forge's `init` never read it: the workspace is `<dir>/.forge`); `forge_split` takes
|
|
333
|
+
`test: 0`; `forge_evaluate` takes `harness_out`; corpora and eval files
|
|
334
|
+
may be `.tsv`.
|
|
335
|
+
- `forge_status` knows the `exported` state and maps its next command onto
|
|
336
|
+
tools (`summary.tools`); training and `nmt-forge serve` are terminal
|
|
337
|
+
steps, never tools.
|
|
338
|
+
- Every forge tool but `forge_init` (which creates the project at `dir`)
|
|
339
|
+
takes `project_dir`: forge runs from there, as in its
|
|
340
|
+
own `cd <project> && nmt-forge …` advice, so config.json's relative paths
|
|
341
|
+
and the `.forge` workspace resolve as `forge_init` wrote them.
|
|
342
|
+
- `forge_discover` no longer says it needs a card directory — forge 0.2.0
|
|
343
|
+
resolves cards from a pip install.
|
|
344
|
+
- **`get_metric_reliability` works from an npm install** — it now reads the
|
|
345
|
+
index the `champollion` package ships (`shared/catalogue/`), not only the
|
|
346
|
+
monorepo copy.
|
|
347
|
+
- **`translate` with the local engine works.** Keylessness is derived from
|
|
348
|
+
the method registry's loopback `default_base_url`, so method `local` no
|
|
349
|
+
longer demands an endpoint variable as if it were a key, and engines other
|
|
350
|
+
than the OpenRouter lane no longer receive the OpenRouter default model slug.
|
|
351
|
+
- **`translate` can target the model you just deployed.** A synthetic user
|
|
352
|
+
served their model with `nmt-forge serve`, passed `endpoint` to `translate`
|
|
353
|
+
(which had no such argument) — zod stripped it silently and the call ran on
|
|
354
|
+
a different local model. Now `base_url` points method `local` (or `openai`)
|
|
355
|
+
at an OpenAI-compatible server (serve's `/v1` URL), and `endpoint` drives
|
|
356
|
+
the new method `api` (the champollion API contract, serve's `/translate`;
|
|
357
|
+
key `CHAMPOLLION_API_KEY`, none needed on loopback). The engine is checked
|
|
358
|
+
to have taken the target before anything is sent. The schema is strict: an
|
|
359
|
+
unknown argument is refused by name (so is `run_benchmark`'s), and a
|
|
360
|
+
`model` sent to a machine-translation API, a `base_url` for an engine that
|
|
361
|
+
cannot use one, or an `endpoint` without method `api` is refused too. Every
|
|
362
|
+
answer has an `Engine:` line naming the engine that ran and, where it has
|
|
363
|
+
them, its model (or "engine default"), endpoint and key source — and names
|
|
364
|
+
the Translation Memory file it used. The local lanes' TM entries are keyed on the endpoint, so your model
|
|
365
|
+
is never served Ollama's cached output.
|
|
366
|
+
- **`translate` TM location is explicit; `project_dir` uses a project's.**
|
|
367
|
+
The tool's own TM is `~/.champollion-mcp/.champollion/tm.json`, separate
|
|
368
|
+
from every project's (`CHAMPOLLION_MCP_HOME` moves it). `project_dir` reads
|
|
369
|
+
and writes `<project>/.champollion/tm.json` instead (the file `champollion
|
|
370
|
+
sync` uses) and runs the engine there, with that project's coaching and
|
|
371
|
+
`.env`, as the CLI does.
|
|
372
|
+
- **`translate` failures say why, and engine output never reaches stdout.**
|
|
373
|
+
The engines report failures on the console and return nothing; the tool
|
|
374
|
+
used to answer "translation failed". Their output is now captured per call
|
|
375
|
+
(the reason becomes the failure, the last lines are shown) and echoed to
|
|
376
|
+
stderr. Progress lines that DeepL / Google / Microsoft / LibreTranslate and
|
|
377
|
+
the api engine print to stdout no longer corrupt the JSON-RPC stream.
|
|
378
|
+
- **`run_benchmark` jobs survive a server restart.** Hosts restart MCP
|
|
379
|
+
servers; a job then "never existed" although the harness had written its
|
|
380
|
+
results. Every job is now recorded in `~/.champollion-mcp/jobs.json` (the
|
|
381
|
+
newest 50; running jobs are kept) with its command, working directory,
|
|
382
|
+
where results land, status, start time and pid — never its environment.
|
|
383
|
+
mt-eval runs under a small detached supervisor that writes the run's
|
|
384
|
+
output and exit status to `jobs/<id>/`, so the run outlives the server and
|
|
385
|
+
its outcome is recorded with nobody watching. `get_run_status` on any
|
|
386
|
+
server answers from that evidence: RUNNING (supervisor alive, checked
|
|
387
|
+
against its command line so a recycled pid does not count), COMPLETED /
|
|
388
|
+
FAILED from the exit record, or INTERRUPTED with the log tail — with the
|
|
389
|
+
harness's report path and headline numbers read from the results folder.
|
|
390
|
+
Item and corpus runs get their own `--output-dir` inside the job folder;
|
|
391
|
+
queue runs report under the harness's `eval/logs/harness/queue/`. Job ids
|
|
392
|
+
are now random (`run-3f9c2a7b1e04`), unique across restarts. Corpus-mode
|
|
393
|
+
plans name the endpoint a local run will call.
|
|
394
|
+
- **New prompt `start_language_project`**; `explore_language` now points at
|
|
395
|
+
`get_language` / `language_overview`. `get_project_info` reports the
|
|
396
|
+
language count from the loaded index instead of a hand-typed "7,900+".
|
|
397
|
+
- **Docs match the server.** README tool tables, `instructions.md` (now led
|
|
398
|
+
by the north-star flow and which tool serves each stage) and the tool
|
|
399
|
+
descriptions are checked against the live tool list by
|
|
400
|
+
`test/server-surface.test.js` (in-memory MCP client).
|
|
401
|
+
- **Contract suite** (`npm run test:contract`, opt-in): packs this server and
|
|
402
|
+
the local CLI, installs both into a temp prefix, spawns the INSTALLED server
|
|
403
|
+
over stdio with the MCP SDK client, and calls every tool with realistic
|
|
404
|
+
arguments under a scrubbed environment (no API keys, temp HOME) — each must
|
|
405
|
+
return a non-error result or an error naming an actionable prerequisite; no
|
|
406
|
+
stack traces, no "[object Object]", nothing over 60 s.
|
|
407
|
+
- Requires `champollion` ^0.4.0.
|
|
408
|
+
|
|
409
|
+
## 0.1.2 (2026-09-06)
|
|
410
|
+
|
|
411
|
+
Closes the gap found by the 2026-09-06 benchmark-hosting beta: an agent could
|
|
412
|
+
list the *queue* but had no way to ask "what benchmarks exist for eng→yor?".
|
|
413
|
+
|
|
414
|
+
- **New tool `list_corpora`** (read-only, 24th tool). Lists the registered
|
|
415
|
+
evaluation corpora for a language pair and/or benchmark family from the
|
|
416
|
+
corpus registry — metadata cards only (size, license, contamination grade,
|
|
417
|
+
domain, provider) plus what the harness can actually do with each entry
|
|
418
|
+
(`fetch` on demand from its pinned upstream / `gated` behind an access
|
|
419
|
+
token with the exact accept-terms instructions / `quarantined`). Corpus
|
|
420
|
+
content is never hosted or returned. Quarantined entries are hidden by
|
|
421
|
+
default but always counted, so a catalogued-but-held pair (eng→crk) says
|
|
422
|
+
so instead of looking unsupported. Source ladder: the in-repo
|
|
423
|
+
`arena/datasets/registry.json` when running inside a checkout →
|
|
424
|
+
`champollion.dev/registry.json` (HTML-holding-page guarded) → the prod
|
|
425
|
+
`datasets` mirror over PostgREST (last, and labelled as lagging).
|
|
426
|
+
`CHAMPOLLION_CORPORA_SOURCE=registry|remote|db` forces a rung. Requires
|
|
427
|
+
at least one filter — the registry is thousands of rows and an unfiltered
|
|
428
|
+
dump is never a useful answer. Its normalizer is the JS twin of the
|
|
429
|
+
harness's `corpora_browse.normalize_entry`.
|
|
430
|
+
- **Live open-item count was under-reported 3×.** `queue_pairs` returns
|
|
431
|
+
one row per open pair (3,620 on prod) and PostgREST serves at most 1,000
|
|
432
|
+
per response; the live count summed only that first page (69,695 against
|
|
433
|
+
a true 211,082). The fetch now pages with `limit`/`offset` until a short
|
|
434
|
+
page. (A `Range` header is not honoured on this RPC by the deployed
|
|
435
|
+
PostgREST.) Ledger entry D15 in `arena/DATABASE_SCHEMA.md`.
|
|
436
|
+
|
|
437
|
+
## 0.1.1 (2026-08-27)
|
|
438
|
+
|
|
439
|
+
Fixes the queue-tool timeouts found by live testing of the 0.1.0 npm release:
|
|
440
|
+
with the queue at 211k+ open items, `list_queue`, `estimate_cost`,
|
|
441
|
+
`get_project_info`, and `run_benchmark` all exceeded MCP clients' default 60s
|
|
442
|
+
request timeout, because the DB fetch path drained the entire `queue_top`
|
|
443
|
+
ranking (~423 sequential pages ≈ 3 minutes) before answering anything.
|
|
444
|
+
|
|
445
|
+
- **Bounded, purpose-fit queue fetching.** Metadata comes from
|
|
446
|
+
queue-preview.json plus a live open-item count from the unpaged
|
|
447
|
+
`queue_pairs` RPC; ranked items are paged from `queue_top` only as deep as
|
|
448
|
+
the caller's selection needs (fetch-until-satisfied, bounded by
|
|
449
|
+
`CHAMPOLLION_QUEUE_MAX_PAGES`, default 20 pages / 10,000 rows); single
|
|
450
|
+
items are read by primary key over PostgREST with a verified-coverage
|
|
451
|
+
probe. When a bound truncates a search, the tool says how deep it looked —
|
|
452
|
+
no silent caps.
|
|
453
|
+
- **`get_queue_item` / `run_benchmark(item_id)`** now do a direct by-id (or
|
|
454
|
+
mode+priority) lookup instead of scanning a full drain, and refuse items
|
|
455
|
+
already covered by a VERIFIED run instead of re-spending on them.
|
|
456
|
+
- **Queue-mode `dry_run` is backgrounded** like a real run (the installed
|
|
457
|
+
harness loads the full queue before printing its plan): it returns a job id
|
|
458
|
+
immediately; the plan arrives via `get_run_status`.
|
|
459
|
+
- **Failure ladder hardened.** A slow-but-alive DB degrades to a
|
|
460
|
+
truncated-but-honest prefix; a dead DB falls back to the static queue.json
|
|
461
|
+
blob; when both are down the error names both causes.
|
|
462
|
+
- Harness (mt-eval, monorepo): `--top N` runs now page the DB queue only as
|
|
463
|
+
deep as selection needs, and `CHAMPOLLION_QUEUE_SOURCE=blob` is accepted as
|
|
464
|
+
a sentinel (previously read as a literal file path).
|
|
465
|
+
|
|
466
|
+
## 0.1.0 (2026-08-27)
|
|
467
|
+
|
|
468
|
+
Initial npm release.
|