@chessceo/mcp 0.49.11 → 0.50.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  # Engine usage guide
2
2
 
3
- When you call `cloud_analyse`, chess.ceo runs Stockfish and Lc0 in parallel on the user's rented combo instance and returns both engines' final read. This doc explains what each engine is good for and how to interpret the numbers you get back — the difference between "this line is a draw" and "this line is easy to draw" is central to real prep and both engines are needed.
3
+ chess.ceo offers three cloud engines: Stockfish, Lc0 and the human engine. Each runs on its own rental, and `cloud_analyse` runs any mix of them on the same positions in parallel. This doc explains what each engine is good for, how to read the numbers, and how to run the work end to end. The difference between "this line is a draw" and "this line is easy to draw" is central to real prep, and it takes more than one engine to see it.
4
4
 
5
5
  ## Grounding: don't invent, run the engine
6
6
 
@@ -10,13 +10,13 @@ When you call `cloud_analyse`, chess.ceo runs Stockfish and Lc0 in parallel on t
10
10
 
11
11
  **When you're inside a prep file, call `cloud_analyse` with `file_id`+`node_id` — never with a hand-typed FEN.** The server derives the FEN from the tree node and, critically, **auto-stores the resulting eval on that node's `ceoEval`**. This is what makes engine attribution trustworthy end-to-end:
12
12
 
13
- 1. `cloud_analyse({ id, node_id })` — runs analysis on the node's exact position, stores `{sf, lc0, nag}` on the node.
13
+ 1. `cloud_analyse({ file_id, node_ids, engines })`: runs analysis on the nodes' exact positions and merges the result into each node's `ceoEval` (`sf` as cp/mate, `lc0` and `human` as win/draw/loss %). Engines you didn't run keep their stored read.
14
14
  2. Later, before you write "engines say X on this position" in a comment, call `quote_engine_eval({ id, node_id })` — it returns the stored eval or `null`.
15
15
  3. If `quote_engine_eval` returns `null`, you have no measurement to cite. **Do NOT infer an eval for a node from siblings, from children, or from a "position that looks similar."** Either analyse the node (`cloud_analyse`) or omit the number from your prose entirely.
16
16
 
17
17
  Concrete failure this rule blocks: the LLM says *"9...Bb7: both engines 0.00"* after only calling `cloud_analyse` on the child positions (post-1.d4, post-castling). Both continuations really returned 0.00, but the Bb7 node was never analysed, and the claim reads to the user as a measurement. With this protocol, `quote_engine_eval(node_id=Bb7)` would return null and the LLM would either analyse it or reword to *"both continuations run to 0.00, so this position looks balanced"* (soft inference, honestly labelled).
18
18
 
19
- **Transposition propagation.** `cloud_analyse({file_id, node_id})` also stamps the resulting `ceoEval` on every OTHER node in the same file that reaches the same position by a different move order (3-field FEN match: pieces + side-to-move + castling). The response includes `also_stored_on: [id, id]` when this happens, and a follow-up `quote_engine_eval` on any of those twin nodes returns the same measurement — no second analysis needed. `auto_evaluate` also dedupes candidates by the same key and returns `skipped_transpositions` so you can see how much engine time it saved. See `list_transpositions` / `list_nodes({filter: "transpositions"})` for auditing where duplication exists before you start.
19
+ **Transposition propagation.** `cloud_analyse({file_id, node_id})` also stamps the resulting `ceoEval` on every OTHER node in the same file that reaches the same position by a different move order (3-field FEN match: pieces + side-to-move + castling). Each position in the response lists `stored_on: [id, ...]`, every node it was saved to, and a follow-up `quote_engine_eval` on any of those twin nodes returns the same measurement — no second analysis needed. `auto_evaluate` also analyses each position once and returns `skipped_transpositions`, so you can see how much engine time it saved. See `list_transpositions` / `list_nodes({filter: "transpositions"})` for auditing where duplication exists before you start.
20
20
 
21
21
  **Concrete failure modes to avoid:**
22
22
 
@@ -73,8 +73,9 @@ Use it to:
73
73
 
74
74
  ### Rule of thumb
75
75
 
76
- - Objective truth ("does this hold?", "is this a mate?") → **trust Stockfish**
77
- - Practical prep ("which side is easier?", "which candidate is best?") → **trust Lc0**
76
+ - Objective truth ("does this hold?", "is this a mate?") → **trust Stockfish**. Run it on every position you evaluate.
77
+ - Practical prep ("which side is easier?", "which candidate is best?") → **Lc0**
78
+ - What a human will play, and how a human would judge the position → **the human engine**
78
79
  - Both agree → high confidence, ship the recommendation
79
80
  - They disagree → look at both scores together and reason about *why*:
80
81
  - Stockfish sharply higher: probably a tactic Lc0 didn't calculate
@@ -167,14 +168,15 @@ Absolute values are useful on their own too — a large "King safety" term flags
167
168
 
168
169
  Complements `cloud_analyse` (best move + PV): different questions, two angles on the same position.
169
170
 
170
- ## Deep Stockfish thinks: `deep_analyse`
171
+ ## Deep thinks: `deep_analyse`
171
172
 
172
- `cloud_analyse` runs both engines with a movetime cap of 10 s — deliberately fast because most opening-tree questions are answered in 2–3 s. For **one specific critical position** where you want depth 35+ instead of the usual depth 22, use `deep_analyse`:
173
+ `cloud_analyse` is built for short thinks (2 s default) because most opening-tree questions are answered in 2-3 s. For **one specific critical position** where you want depth 35+ instead of the usual depth 22, use `deep_analyse`:
173
174
 
174
- - SF-only (Lc0 saturates in a handful of seconds — no benefit past ~5 s of movetime).
175
+ - One engine (`engine`, default `stockfish`). Lc0 and the human engine gain little past a few seconds, so a deep think is almost always Stockfish.
175
176
  - Movetime up to 5 min.
176
- - Async: returns a `job_id` immediately, poll `deep_analyse_status(job_id)`, cancel with `deep_analyse_cancel(job_id)` if you decide partial depth is enough.
177
- - **Holds only the Stockfish slot on the combo.** Lc0 remains callable for other positions via `cloud_analyse({ engines: ["lc0"], … })` while the deep think runs. Use that during the wait — walk other branches, sanity-check candidates.
177
+ - Async: returns a `job_id` immediately, poll `deep_analyse_status(job_id)`, cancel with `deep_analyse_cancel(job_id)` if partial depth is enough.
178
+ - **Holds only that engine.** The other engines stay callable through `cloud_analyse` while the deep think runs. Use the wait to walk other branches.
179
+ - With `file_id`+`node_id`, the result is merged into the node's stored eval; the other engines' stored reads stay.
178
180
 
179
181
  Typical movetimes:
180
182
  - `30_000 – 60_000` (30–60 s) — careful check on a candidate move
@@ -184,53 +186,49 @@ Only use it when you actually need the depth. Regular `cloud_analyse` handles th
184
186
 
185
187
  ## Per-engine `multipv` defaults
186
188
 
187
- `cloud_analyse` defaults to different candidate-line counts per engine because the two engines behave differently:
189
+ `multipv: {stockfish, lc0, human}` sets the candidate-line count per engine. Defaults:
188
190
 
189
- - **Stockfish: `stockfish_multipv=2`** — SF gets weaker as multipv grows. Each extra PV steals search bandwidth from the top choice, so the "objective best move" leg is at its strongest with a tight list. Raise only when you specifically need SF's read on a wide range of candidates (e.g. sanity-checking an unusual sideline).
190
- - **Lc0: `lc0_multipv=8`** — Lc0 doesn't degrade the same way with multipv. Use it wide by default: 8 candidates give the LLM a real slate of practical ideas to inspect ("what does Lc0 think of the fun moves here?"). Lower it only when you don't need that breadth.
191
+ - **Stockfish: 2.** SF gets weaker as multipv grows: each extra PV steals search bandwidth from the top choice. Raise only when you need SF's read on a wide range of candidates.
192
+ - **Lc0 and human: 8.** Neural-net engines don't lose strength with multipv. 8 candidates give a real slate of ideas (Lc0) or of the moves people actually play (human).
191
193
 
192
194
  Mental model:
193
195
  - **Stockfish = the checker.** "Is this line objectively good? What's actually best?" Narrow, deep, one answer.
194
- - **Lc0 = the explorer.** "What's practically interesting? What alternatives are worth a look?" Wide, breadth-first, several answers.
196
+ - **Lc0 = the explorer.** "What's practically interesting? What long-term ideas are there?"
197
+ - **Human = the opponent.** "What will a person play here, and how will they judge it?"
195
198
 
196
- Override the defaults with `stockfish_multipv` / `lc0_multipv` when a specific position warrants a different shape.
199
+ ## Running the work: rent, analyse, stop
197
200
 
198
- ## Splitting engines on `cloud_analyse`
201
+ Every analysis needs a running rental for each engine you ask for. Each rental runs exactly one engine.
199
202
 
200
- `cloud_analyse` accepts an `engines` list to run only one leg:
203
+ 1. **Plan:** `ensure_engines({engines, positions})` shows which engines are already running, the cheapest SKU and price for each missing one, and a time and cost estimate. It starts nothing.
204
+ 2. **Ask:** show the user the price of anything you'd start and get a yes. Rentals bill per minute until stopped.
205
+ 3. **Start:** `start_cloud_engine({machine_type})` for each missing engine. The engine comes from the SKU. Static engines are ready at once; others take ~1-5 min (`list_cloud_engines` shows when).
206
+ 4. **Analyse:**
207
+ - A few positions: `cloud_analyse` with `fens`, `lines`, or `file_id`+`node_ids` (up to 10 per call), e.g. `{fens: [a, b, c], engines: ["stockfish", "human"]}`.
208
+ - A whole file or subtree: `auto_evaluate({id, engines})`. It runs only the engines each node is missing, analyses transpositions once, and saves as it goes. Adding an engine later runs just that engine.
209
+ - One critical position deep: `deep_analyse`.
210
+ 5. **Read back:** `quote_engine_eval` returns the stored evals per node.
211
+ 6. **Stop:** `stop_cloud_engine` when the work is done, unless the user wants the rental kept.
201
212
 
202
- - `engines: ["lc0"]` — Lc0 only, useful while a `deep_analyse` is holding the Stockfish slot.
203
- - `engines: ["stockfish"]` — SF only, when only the objective read matters and you want to skip the Lc0 latency.
204
- - Default (omitted) — both engines. This is what you want for real prep decisions.
213
+ If two running rentals provide the same engine, the call fails with both contract ids; pass `rentals: {engine: contract_id}` to choose.
205
214
 
206
- The skipped engine's field is omitted from the response (not present as an empty object).
215
+ ### Recipe: finding ideas
207
216
 
208
- ## Renting and choosing engines
217
+ 1. Run Lc0 or the human engine with a wide multipv on the position. The human engine shows the moves people play; Lc0 with contempt (e.g. `-20` to make Black play for a win) shows ambitious tries.
218
+ 2. Collect the interesting candidates and run Stockfish on the positions after each one (`lines` makes this one call).
219
+ 3. A candidate Stockfish refutes only with a line no human would find is a practical idea. A candidate that loses to a simple refutation is not.
209
220
 
210
- Analysis needs a running rental. Start one with `start_cloud_engine`, using a `machine_type` SKU from `list_cloud_machine_options`. Each SKU provides one of three shapes, shown in its `engine` field:
221
+ ### Eval symbols are your call
211
222
 
212
- - **`stockfish`**: Stockfish only. Fits when the question is objective (is this tactic real, does this defense hold).
213
- - **`lc0`**: Lc0 only. Fits the practical, human-feel read, or running alongside a `deep_analyse` that holds the Stockfish side.
214
- - **`combo`**: both engines in one container. The default, and the right choice for prep decisions, where both reads are needed on the same position.
215
-
216
- Rule of thumb: start a `combo` rental unless the task only needs one engine. A single-engine rental gives one read per position, so a prep walk that needs both reads on a single-engine rental takes two passes.
217
-
218
- `list_cloud_machine_options` returns the price for each SKU. Show the user the price and get their confirmation before starting a rental; every rental bills per second until `stop_cloud_engine`.
219
-
220
- Routing to a rental:
221
-
222
- - `cloud_analyse`, `auto_evaluate`, and `deep_analyse` each take an optional `contract_id` (from `list_cloud_engines`).
223
- - Without `contract_id`, the call uses the only running rental that can serve the request. If several can, the error lists their contract ids, and you pick one.
224
- - Any shape works. With `engines` set, the rental must provide those engines: asking for `engines: ["lc0"]` on a stockfish-only rental fails.
225
- - Do not guess contract ids. Call `list_cloud_engines` first if more than one rental is running.
223
+ The symbol you put on a position (=, +=, ±, +-) is written for a human reader, and no tool sets it for you. Use every engine you ran: a Stockfish 0.00 can still be `+=` when Lc0 or the human engine clearly prefer one side, because the symbol describes the practical picture, not only the objective one. Don't put a symbol on a position nobody analysed.
226
224
 
227
225
  ## Worked example
228
226
 
229
227
  User is preparing Black against a 2600 opponent who plays 1.e4 c5 2.Nf3 d6 3.d4 cxd4 4.Nxd4 Nf6 5.Nc3 a6 6.Be3 e5. You want to know if 7.Nb3 or 7.Nf3 is more testing.
230
228
 
231
229
  Two calls:
232
- 1. `cloud_analyse(fen=<position after 6...e5>, multipv=2)` — get both engines' top choices with movetime=2000.
233
- 2. If Stockfish scores them equal but Lc0 prefers one by 0.10-0.20, that's your practical answer. The user will find that line harder to face.
230
+ 1. `cloud_analyse({lines: ["<moves to 6...e5> Nb3", "<moves to 6...e5> Nf3"], engines: ["stockfish", "lc0", "human"]})`: all three engines on both positions in one call.
231
+ 2. If Stockfish scores them equal but Lc0 or the human engine clearly prefer one, that's your practical answer. The user will find that line harder to face.
234
232
 
235
233
  If the user is specifically preparing to *play* the black side in a must-win, add `contempt=-30` (or up to `-60` for a harder steer) on a follow-up call to see which lines Lc0 finds most fighting for Black. Compare against Stockfish's objective read to make sure the fighting choice isn't just losing.
236
234
 
@@ -242,7 +240,7 @@ If the user is specifically preparing to *play* the black side in a must-win, ad
242
240
 
243
241
  **To see further into a line, walk the tree.** Take the position at the tail of your truncated PV, run a fresh `cloud_analyse` on it. That call gives you the multipv candidate set at THAT position — the opponent's actual options — which is what you need to decide whether to branch. Don't raise `pv_max_plies` unless you're verifying a forcing sequence (a mate, an obligated recapture chain).
244
242
 
245
- Practical: raise `lc0_multipv` (default 8) to see the candidate spread on the current position, NOT to see further down one PV. If Lc0 shows moves 1-3 within 0.15 of each other, that's a branching point — three responses need coverage, not one PV.
243
+ Practical: raise `multipv.lc0` or `multipv.human` (default 8) to see the candidate spread on the current position, NOT to see further down one PV. If Lc0 shows moves 1-3 within 0.15 of each other, that's a branching point — three responses need coverage, not one PV.
246
244
 
247
245
  ## What NOT to do
248
246
 
@@ -40,12 +40,12 @@ All mutations **auto-save** with optimistic locking. Response includes the new `
40
40
 
41
41
  ## Typical build order
42
42
 
43
- 0a. **Cloud engine running?** Call `list_cloud_engines` first. Every substantive step below needs Stockfish + Lc0 to be reachable — `cloud_analyse` at critical positions, `auto_evaluate` for the whole tree, engine-derived NAGs at endpoints, describe_position's Stockfish eval-terms breakdown. If the caller has zero running combos: STOP, tell the user prep needs an engine, list options via `list_cloud_machine_options`, get the SKU + explicit confirmation (real money per second), then `start_cloud_engine`. If they already have one, note the contract_id and continue. Never silently write prep without engines — the file ends up with placeholder NAGs the user has no way to distinguish from real ones.
43
+ 0a. **Cloud engine running?** Call `ensure_engines` first. Every substantive step below needs engines: `cloud_analyse` at critical positions, `auto_evaluate` for the whole tree, engine-backed NAGs at endpoints, describe_position's Stockfish eval-terms breakdown. Stockfish at minimum; add Lc0 or the human engine for the practical read. If what you need isn't running: STOP, tell the user prep needs engines, show the price ensure_engines returns, get explicit confirmation (rentals bill per minute), then `start_cloud_engine`. Never silently write prep without engines — the file ends up with placeholder NAGs the user has no way to distinguish from real ones.
44
44
 
45
45
  0b. **`read_docs({ docs: ["pgn-authoring", "examples/najdorf-6-f4-repertoire"] })`** — do this once per session, before writing any prose. Not optional. Log analysis shows most sessions skip the example files and produce documented anti-patterns (long PVs in prose, restating what the app renders, verbose citations). Reading the reference PGN once inoculates the LLM against those.
46
46
  1. `read_prep_file` — see what's there. Every node has an `id` you'll pass to the mutation and engine/DB tools. Use `view: "compact"` (default) plus `node_id` + `max_depth` to scope; the full tree of a 500+-node file can blow the token limit.
47
47
  2. `apply_mutations([...])` — one call with your whole intended build (a mix of `add_move` / `add_line` for structure, plus any `set_comment`/`set_annotations` you already know at author time, plus any `set_nags` where you already have a clear judgment — novelty `$146`, `!?` speculative sac, obvious `?` blunder in a sideline you're rejecting).
48
- 3. `auto_evaluate(id)` — spawns a background job that PERSISTS engine numbers on every node. Does not touch visible NAGs. Cheap way to get every position's Stockfish + Lc0 read baked into the file for later reference. Grab the returned `job_id` and either (a) poll `auto_evaluate_status(job_id)` every ~10-30s until done, or (b) fire and do useful work meanwhile (write more of the tree, walk the opponent's repertoire) and check back later — engine walk-time serialises on the per-combo semaphore, so it takes roughly `target_count × movetime_ms` in wall time.
48
+ 3. `auto_evaluate({id, engines})` — spawns a background job that PERSISTS engine numbers on every node, for the engines you name (e.g. `["stockfish", "human"]`), filling only what each node is missing. Does not touch visible NAGs. Cheap way to get every position's engine reads baked into the file for later reference. Grab the returned `job_id` and either (a) poll `auto_evaluate_status(job_id)` every ~10-30s until done, or (b) fire and do useful work meanwhile (write more of the tree, walk the opponent's repertoire) and check back later — it takes roughly `target_count × movetime_ms` in wall time (engines run in parallel per position).
49
49
  4. Once the job reports `done:true`, re-read the file and add NAGs where they carry real signal (see NAG discipline below). Use individual mutation tools for surgical follow-ups (fix one comment, add one arrow, promote a specific variation, prune a branch).
50
50
 
51
51
  The build-cost math: a 200-move file via individual `add_move` calls is 200 saves ≈ 100+ seconds of tool overhead. The same file via one `apply_mutations` call is one save ≈ 500ms. Use batch by default.
@@ -210,7 +210,7 @@ This is the single biggest quality problem in current LLM output on this system:
210
210
 
211
211
  **Correct pattern for building a variation.** At every ply:
212
212
 
213
- 1. What are the plausible replies? `get_position_stats` (frequencies), `cloud_analyse` with `lc0_multipv: 8` (candidate spread), `predict_human_move` (what people actually play).
213
+ 1. What are the plausible replies? `get_position_stats` (frequencies), `cloud_analyse` with Lc0 or the human engine (8 candidate lines by default), `predict_human_move` (what people actually play).
214
214
  2. If ≥2 are plausible → branch. `add_move` each, then recurse on each branch or `add_line` for each of the several straightforward continuations.
215
215
  3. If exactly 1 → continue linearly; note *why* it's the only move in a comment.
216
216
 
@@ -85,7 +85,7 @@ If yes to all three, ship it. If no, either add what's missing or CUT the branch
85
85
 
86
86
  The same tools as any prep file, but different order and different density.
87
87
 
88
- 0. **Cloud engine running?** A summary looks light but requires SHARPER analysis than a big file — every endpoint NAG has to be right, and there are few enough of them that a wrong one stands out. Call `list_cloud_engines` first. Zero combos running → STOP: tell the user a summary needs engines, list options via `list_cloud_machine_options`, get their SKU + explicit confirmation (real money per second), then `start_cloud_engine`. Do NOT build a summary from cached / guessed evals — the point of the summary is trust, and a placeholder `$14` at an endpoint the reader will internalize is worse than no summary.
88
+ 0. **Cloud engine running?** A summary looks light but requires SHARPER analysis than a big file — every endpoint NAG has to be right, and there are few enough of them that a wrong one stands out. Call `ensure_engines` first. Nothing running → STOP: tell the user a summary needs engines, show the price it returns, get explicit confirmation (rentals bill per minute), then `start_cloud_engine`. Do NOT build a summary from cached / guessed evals — the point of the summary is trust, and a placeholder `$14` at an endpoint the reader will internalize is worse than no summary.
89
89
  0.5. **Which colour is the reader playing?** Ask if you don't know. This decides everything downstream — the mainline is their moves, the branches are their opponent's replies, the endpoint NAGs are judged from their POV, the voice is written to them. See the "Summaries are color-oriented" section above.
90
90
  1. `create_prep_file(collection_id, name)` — name it with the colour AND "Summary" in the Event tag (`"Modern Defence — Tiger ...a6/...b5 — Black (Summary)"`, `"Petroff 6.Bd3 Bd6 — White (Summary)"`) so the reader picks it out of a list immediately.
91
91
  2. `set_comment(root, "framing")` — the file's thesis, first thing.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@chessceo/mcp",
3
- "version": "0.49.11",
3
+ "version": "0.50.0",
4
4
  "description": "Model Context Protocol server for chess.ceo — 11.7M+ games, ~1.5M FIDE player profiles, opening preparation, live broadcasts.",
5
5
  "type": "module",
6
6
  "bin": {