@chessceo/mcp 0.49.10 → 0.50.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/tools.js CHANGED
@@ -118,7 +118,7 @@ export const TOOLS = [
118
118
  "- `gm-classical` — GM classical games (both players ≥2500, real thinking-time). BEST for opening prep — every move is signal, avgElo ~2600 across all listed moves.\n" +
119
119
  "- `main` — the whole 11.7M-game DB. Widest coverage but noisiest (includes 1000-Elo blunder-fests in the move stats). Use as fallback when gm-classical's totalCount is too small to be informative.\n\n" +
120
120
  "Game movetext is trimmed to the moves AFTER the queried position (using each game's plyNumber). Saves ~70% of the bytes vs full movetext.\n\n" +
121
- "AUTO-EVAL: if a cloud engine instance is running, the response includes `.eval` with a compact read from the engine(s) it provides and the corresponding NAG. Do NOT fire cloud_analyse separately for the same FEN. When called with `file_id`+`node_id`, the eval is also auto-stored on that node's `ceoEval` — later readable via quote_engine_eval.",
121
+ "AUTO-EVAL: if a Stockfish rental is running, the response includes `.eval` with a short Stockfish read (cp/mate White's point of view, best move, PV). Not stored on any node; for a stored eval or other engines use cloud_analyse.",
122
122
  inputSchema: {
123
123
  type: "object",
124
124
  properties: {
@@ -256,27 +256,34 @@ export const TOOLS = [
256
256
  required: ["fide_id"],
257
257
  },
258
258
  },
259
+ {
260
+ name: "ensure_engines",
261
+ description: "Plan cloud-engine work before running it. For each engine you need (`stockfish`, `lc0`, `human`), says whether a rental providing it is already running (with its contract_id), or the cheapest SKU to start and its price. Pass `positions` (how many positions you will analyse) for a time and cost estimate. Starts nothing.\n\n" +
262
+ "CALL FIRST whenever a task needs cloud engines. Then show the user the price of anything to start, get a yes, and start it with start_cloud_engine. The response also carries `engine_guide` (what each engine is for) and `workflow` (the full plan → start → analyse → stop sequence).",
263
+ inputSchema: {
264
+ type: "object",
265
+ properties: {
266
+ engines: { type: "array", items: { type: "string", enum: ["stockfish", "lc0", "human"] }, description: "Engines the task needs. Default [\"stockfish\"]." },
267
+ positions: { type: "integer", minimum: 1, description: "Optional: how many positions you plan to analyse, for the estimate." },
268
+ movetime_ms: { type: "integer", minimum: 100, maximum: 300000, description: "Optional: think time per position for the estimate (default 2000)." },
269
+ },
270
+ },
271
+ },
259
272
  {
260
273
  name: "list_cloud_machine_options",
261
- description: "Returns the catalog of cloud-engine machine types the user can start — every engine shape (stockfish-only, lc0-only, and combo) they're entitled to, same visibility rules as the chess.ceo app (SKU, which engine(s) it runs, human display name, cost per hour, availability), plus an `engine_guide` explaining what each engine type is actually good/bad for so you can pick the right one for the task, not just the cheapest. ALWAYS call this before start_cloud_engine — SKU strings like 'rtx-5090-64' do not match the display names ('Stockfish 32 CPUs + Lc0 1× RTX 5090') and are NOT guessable, and a given SKU only supports the engine(s) shown in its `engine` field (a stockfish-only SKU cannot be started as lc0 or combo). Present the user the display names + prices; pass the SKU to start_cloud_engine.",
274
+ description: "Returns the catalog of cloud-engine SKUs the user can start, same visibility as the chess.ceo app: SKU (`machineType`), the one engine it runs (`engine`: stockfish, lc0 or human), display name, cost per hour, availability. Also returns `engine_guide`, `eval_scale`, `eval_symbols` and `workflow`. ensure_engines already picks the cheapest SKU per engine; use this when the user wants to choose hardware. SKU strings are NOT guessable from display names; pass them verbatim to start_cloud_engine.",
262
275
  inputSchema: { type: "object", properties: {} },
263
276
  },
264
277
  {
265
278
  name: "start_cloud_engine",
266
- description: "Rent a GPU/CPU instance on the user's chess.ceo account — stockfish-only, lc0-only, or combo (both in one container), whichever the chosen SKU provides. Real money — billed per second while running.\n\n" +
267
- "CRITICAL: `machine_type` must be an exact SKU from `list_cloud_machine_options` (e.g. 'rtx-5090-64', NOT 'rtx-5090'). Guessing SKUs will fail. Call list_cloud_machine_options first, show the user the display names + prices, get their confirmation, then pass the SKU here. `engine_type` must match what that SKU actually provides (its `engine` field from the catalog) — omit it to default to whatever the SKU is (a combo SKU defaults to \"combo\").\n\n" +
268
- "Use list_cloud_engines first to check if the user already has one running; don't start a second instance unless the user asked for it. Requires an MCP token with agent access.",
279
+ description: "Rent one cloud engine on the user's chess.ceo account. Each SKU runs exactly one engine (stockfish, lc0 or human), taken from the SKU. Real money: billed per minute from start until stop_cloud_engine.\n\n" +
280
+ "`machine_type` must be an exact SKU from ensure_engines or list_cloud_machine_options. Show the user the price and get their confirmation first. For several engines, start one rental per engine; cloud_analyse uses them together. Don't start an engine that is already running (ensure_engines / list_cloud_engines show what is).",
269
281
  inputSchema: {
270
282
  type: "object",
271
283
  properties: {
272
284
  machine_type: {
273
285
  type: "string",
274
- description: "SKU from list_cloud_machine_options (e.g. 'rtx-5090-64', 'rtx-5090', 'epyc-256'). MUST be the exact SKU, not the display name and not a guess.",
275
- },
276
- engine_type: {
277
- type: "string",
278
- enum: ["stockfish", "lc0", "combo"],
279
- description: "Which engine(s) to run. Must match the SKU's own `engine` field from list_cloud_machine_options — e.g. a stockfish-only SKU can only be started as \"stockfish\". Optional; defaults to \"combo\".",
286
+ description: "Exact SKU from ensure_engines or list_cloud_machine_options. Not a display name, not a guess.",
280
287
  },
281
288
  },
282
289
  required: ["machine_type"],
@@ -284,12 +291,12 @@ export const TOOLS = [
284
291
  },
285
292
  {
286
293
  name: "list_cloud_engines",
287
- description: "List the user's currently running cloud engines. Use before starting a new one, or to find the contract_id for stop_cloud_engine. `cloud_analyse` uses the only running rental that can serve the request unless you pass `contract_id`, so listing is only necessary when there are several.",
294
+ description: "List the user's running cloud engines, each with its `contractId` and `engines` (what it provides). A row with empty `engines` is still starting. Use the contractId for stop_cloud_engine, or in `rentals` when two running rentals provide the same engine.",
288
295
  inputSchema: { type: "object", properties: {} },
289
296
  },
290
297
  {
291
298
  name: "stop_cloud_engine",
292
- description: "Destroy a running cloud engine. Billing stops immediately. Use the contract_id from `list_cloud_engines` — don't guess.",
299
+ description: "Destroy a running cloud engine. Billing stops immediately. Use the contract_id from `list_cloud_engines`, don't guess. Stop rentals when the analysis is done unless the user wants them kept.",
293
300
  inputSchema: {
294
301
  type: "object",
295
302
  properties: {
@@ -303,66 +310,38 @@ export const TOOLS = [
303
310
  },
304
311
  {
305
312
  name: "cloud_analyse",
306
- description: "Runs a synchronous ~2s analysis on one of the user's running engine instances and returns the requested engines' final read for the FEN — depth, top-N candidate moves with scores (**scoreCp is White-POV centipawns**: +20 = White is +0.20 pawns better regardless of whose turn it is; matches the sign convention used everywhere else in this MCP, including the stored ceoEval). Mate is White-POV plies-to-mate (+5 = White mates in 5). Also returns each engine's principal variation.\n\n" +
307
- "GROUNDING: every claim you make about a position must trace back to actual engine output from a call in THIS session. Don't invent evaluations, don't name 'best moves' you haven't seen the engine list, don't fabricate variations that 'look plausible.' Compute is cheap — call this 5-10 times while walking a tree rather than pattern-matching from your training data. When you don't have data for the position, either run the tool or say so; don't fill the gap with chess prose the user can't distinguish from measured output.\n\n" +
308
- "Runs on any of the caller's running rentals. Pass `contract_id` to choose one; without it the only rental that can serve the requested engines is used, and the error lists the candidates if there are several. Pass `contract_ids` (several rentals) to analyse the same FEN on all of them in parallel: the result merges the engines, e.g. Stockfish from one rental and Lc0 from another.\n\n" +
309
- "How to read the response:\n" +
310
- "• Stockfish is objective truth — trust it for 'does this line hold?' 'is there a tactic?' 'is this endgame drawn?' A Stockfish 0.00 means 'objectively equal', NOT 'trivial draw' — one side can still be much harder to play in practice.\n" +
311
- "• Lc0 is practical eval — trust it for 'which side is easier?' 'which candidate is best when Stockfish shows several as equal?' Lc0 sees long-term positional factors Stockfish's fixed search can miss.\n" +
312
- "• When they agree → high confidence. When they disagree → look at both scores and reason WHY (Stockfish sharply higher = tactic Lc0 missed; Lc0 higher = long-term positional edge past Stockfish's horizon). Never dismiss either — the disagreement is the signal.\n\n" +
313
- "Contempt (`contempt`) skews Lc0 (only Lc0 — Stockfish always stays objective) toward White (positive) or Black (negative). Signed 0-100 strength — same scale as the web UI's ContemptStrength slider (the server multiplies by 8 to produce Lc0's internal cp bias). Typical values: ±15 for a light nudge, ±30-60 for real fighting play, ±80-100 for maximum steer. Use it to find non-objective 'practical' ideas or when the user needs to lean toward fighting/solid lines with a specific colour. Do NOT quote a contempt-biased eval as objective — cross-check with Stockfish.\n\n" +
314
- "Also useful: pass `moves` on top of `fen` to explore a variation without computing FENs yourself (e.g. fen='<tabiya>', moves='b4 a5 c3'). And the flip-side-to-move threat check documented in the guide is a great free trick.\n\n" +
315
- "**PVs are capped at 6 plies by default (3 full moves), and lines that got truncated are marked with `pv_truncated: true`.** This is deliberate: the tail of a PV is where the engine's confidence collapses, AND pasting a long PV into `add_line` as if it were prepared repertoire is the #1 documented anti-pattern of this MCP — a 15-move PV is one line of engine output through positions where both sides had real choices, not a repertoire. To see further, don't raise `pv_max_plies`; instead, walk the tree one branch at a time with a fresh `cloud_analyse` at each position where the opponent has real alternatives — that's what makes it prep instead of pasted output. Only raise the cap when you're verifying a forcing sequence (a mate, a forced tactical resolution), not to build lines.\n\n" +
316
- "For the full guide including worked examples, call the `read_engine_usage_guide` tool.\n\n" +
317
- "Not for casual questions — this costs real money per second. Use `get_position_stats` for anything that doesn't require deep prep.\n\n" +
318
- "**When called with `file_id`+`node_id` (preferred inside a prep file), the resulting eval is auto-stored on that node's `ceoEval` — you can then quote it with quote_engine_eval on any later call.** This is what makes engine attribution trustworthy: prose that says 'engines say X on node Y' can only be true if a call was actually made against node_id=Y.",
319
- inputSchema: {
320
- type: "object",
321
- properties: {
322
- file_id: { type: "string", description: "Prep file id. **Prefer file_id+node_id over fen** when inside a prep file — the FEN comes from the tree AND the result is stored on the node." },
323
- node_id: { type: "string", description: "Node id inside `file_id`. Root is 'r'. When set, overrides `fen`/`moves`." },
324
- fen: { type: "string", description: "Starting position as FEN. Only used if `node_id` is not set." },
325
- moves: {
326
- type: "string",
327
- description: "Optional SAN moves to apply on top of `fen` (or startpos). Only used if `node_id` is not set.",
328
- },
329
- movetime_ms: {
330
- type: "integer",
331
- minimum: 100,
332
- maximum: 10000,
333
- description: "Think time in milliseconds (default 2000).",
334
- },
335
- stockfish_multipv: {
336
- type: "integer",
337
- minimum: 1,
338
- maximum: 10,
339
- description: "Stockfish candidate lines (default 2). Kept tight because each extra PV steals search bandwidth from the top choice — SF is the 'what's objectively best' leg, use a low multipv to keep it strong. Raise only when you specifically need SF's take on a wide range of candidates.",
340
- },
341
- lc0_multipv: {
342
- type: "integer",
343
- minimum: 1,
344
- maximum: 10,
345
- description: "Lc0 candidate lines (default 8). Kept wide because multipv doesn't degrade Lc0's strength the way it does Stockfish's — Lc0 is the 'find inspiration / explore practical tries' leg, use a high multipv to get a full slate of ideas.",
346
- },
347
- contempt: {
348
- type: "integer",
349
- minimum: -100,
350
- maximum: 100,
351
- description: "Lc0 contempt bias. Signed 0-100 strength (same scale as the web UI's ContemptStrength slider — server multiplies by 8 to get the internal cp bias). 0 = objective (default). Positive favours White, negative favours Black. Typical: ±15 light nudge, ±30-60 real fighting play, ±80-100 maximum steer. Not applied to Stockfish. See engine_usage_primer for when to use.",
352
- },
353
- contract_id: { type: "string", description: "Which of your running rentals to use, from list_cloud_engines. Any engine shape works (combo, stockfish-only, lc0-only). Omit it when only one rental can serve the request; if several can, the error lists their contract ids." },
354
- contract_ids: { type: "array", items: { type: "string" }, description: "Several running rentals to analyse the same FEN in parallel, from list_cloud_engines. Results merge by engine. Use it to get Stockfish and Lc0 on one position from separate rentals in one call." },
355
- engines: {
356
- type: "array",
357
- items: { type: "string", enum: ["stockfish", "lc0"] },
358
- description: "Which engines to run. Default = both. Use `[\"lc0\"]` to skip Stockfish (e.g. while a deep_analyse job is holding the SF slot on the same rental). Use `[\"stockfish\"]` when only the objective read matters. The skipped engine's field is omitted from the response.",
313
+ description: "Analyse up to 10 positions on the engines you choose (`stockfish`, `lc0`, `human`), each engine on whichever running rental provides it. Engines run in parallel per position. Returns, per position, each engine's depth, candidate lines and best move.\n\n" +
314
+ "Positions: `fens` (list), `lines` (list of SAN move sequences from `fen` or the start), `file_id` + `node_ids` (list), or a single `fen` / `moves` / `file_id`+`node_id`. Examples: analyse three FENs with Stockfish → `{fens: [a, b, c], engines: [\"stockfish\"]}`; with the human engine and Stockfish → `engines: [\"stockfish\", \"human\"]`. For a whole file or subtree use auto_evaluate instead.\n\n" +
315
+ "Scores: `scoreCp` and `mate` are White's point of view (+20 = White +0.20; mate +5 = White mates in 5). lc0 and human lines carry `wdl`: win/draw/loss percent, White's point of view. Read the response's `eval_scale` before calling a position equal.\n\n" +
316
+ "GROUNDING: every claim about a position must trace back to engine output from this session. Don't invent evaluations, best moves or variations. When you have no data for a position, run it or say so.\n\n" +
317
+ "Storing: with `file_id`, each result is merged into the `ceoEval` of every node that reaches that position (`stored_on` in the response), keeping other engines' stored reads. quote_engine_eval cites them later.\n\n" +
318
+ "Contempt (`contempt`, lc0 only): signed -100..100, positive favours White. Use it to find ideas (e.g. -20 makes Black play for a win). Never quote a contempt eval as objective.\n\n" +
319
+ "PVs are capped at 6 plies (`pv_truncated: true` when cut). To see further, analyse the position at the end of the line; raise `pv_max_plies` only to verify a forcing line. Don't paste PVs into add_line as prep.\n\n" +
320
+ "Needs running rentals for the engines you ask for (see ensure_engines). Costs money while rentals run; use get_position_stats for casual questions.",
321
+ inputSchema: {
322
+ type: "object",
323
+ properties: {
324
+ engines: { type: "array", items: { type: "string", enum: ["stockfish", "lc0", "human"] }, description: "Engines to run. Default [\"stockfish\"]. Each needs a running rental that provides it." },
325
+ fens: { type: "array", items: { type: "string" }, description: "Positions as FENs (up to 10 in total with the other inputs)." },
326
+ lines: { type: "array", items: { type: "string" }, description: "SAN move sequences, each applied from `fen` (or the start position), e.g. [\"e4 c5 Nf3\", \"e4 e5 Nf3\"]." },
327
+ file_id: { type: "string", description: "Prep file id. With node_ids/node_id the FEN comes from the tree; with any input, results are stored on matching nodes." },
328
+ node_ids: { type: "array", items: { type: "string" }, description: "Node ids inside `file_id`." },
329
+ node_id: { type: "string", description: "One node id inside `file_id`. Root is 'r'." },
330
+ fen: { type: "string", description: "One position, or the start position for `moves` / `lines`." },
331
+ moves: { type: "string", description: "SAN moves applied on top of `fen` (or the start position) for a single position." },
332
+ movetime_ms: { type: "integer", minimum: 100, maximum: 15000, description: "Think time per position in ms (default 2000)." },
333
+ multipv: {
334
+ type: "object",
335
+ properties: { stockfish: { type: "integer", minimum: 1, maximum: 10 }, lc0: { type: "integer", minimum: 1, maximum: 10 }, human: { type: "integer", minimum: 1, maximum: 10 } },
336
+ description: "Candidate lines per engine. Defaults: stockfish 2 (extra lines cost it strength), lc0 and human 8 (wide slate of ideas, no strength cost).",
359
337
  },
360
- pv_max_plies: {
361
- type: "integer",
362
- minimum: 1,
363
- maximum: 40,
364
- description: "Cap each returned PV to this many plies (default 6 = 3 full moves). PVs beyond ~6 plies are speculative and are the anti-pattern behind pasted-engine-line 'prep' — don't raise unless you're specifically checking a forcing tactic or verifying a mate. When a line was truncated, the response marks it with `pv_truncated: true`.",
338
+ contempt: { type: "integer", minimum: -100, maximum: 100, description: "lc0 only. Positive favours White, negative Black. 0 = objective (default)." },
339
+ rentals: {
340
+ type: "object",
341
+ properties: { stockfish: { type: "string" }, lc0: { type: "string" }, human: { type: "string" } },
342
+ description: "Only when two running rentals provide the same engine: engine → contract_id from list_cloud_engines.",
365
343
  },
344
+ pv_max_plies: { type: "integer", minimum: 1, maximum: 40, description: "Cap each PV (default 6 plies). Raise only to verify a forcing line." },
366
345
  },
367
346
  },
368
347
  },
@@ -723,25 +702,27 @@ export const TOOLS = [
723
702
  },
724
703
  {
725
704
  name: "auto_evaluate",
726
- description: "Walk the tree from `node_id` (default `'r'` = whole file) and populate the persistent `ceoEval` on every descendant via cloud_analyse. Requires a running cloud engine instance (any shape; pass contract_id to choose).\n\n" +
727
- "**Async job — returns immediately.** Response: `{ job_id, target_count, status: 'running', estimated_seconds }`. Then poll `auto_evaluate_status(job_id)` until `done: true`. Cancel a run with `auto_evaluate_cancel(job_id)` — partial progress is preserved. Do useful other work between polls (write more of the tree, walk the opponent's repertoire) — the engine runs in the background.\n\n" +
728
- "Progress is checkpointed to the prep file every 8 successfully-evaluated nodes, so a cancel / crash / MCP restart mid-run leaves the tree partially populated rather than losing everything. On MCP restart the job record disappears; re-run auto_evaluate and `only_missing=true` naturally skips what was already saved.\n\n" +
729
- "**Does NOT set visible NAGs.** NAG placement is your call, not the engine's — an opening tree full of 0.00 positions doesn't need a `$10` (=) glyph on every move. Use quote_engine_eval on individual nodes before writing prose that references engine numbers.\n\n" +
730
- "Costs real money — one cloud_analyse per node. A 200-node walk at default movetime is ~5 min of engine time (calls serialise on the per-engine semaphore in the backend).",
705
+ description: "Analyse a whole prep file, or the subtree under `node_id`, on the engines you choose and store the evals on every node. Example: \"analyse the whole tree with Stockfish and the human engine and save the evals\" → `{id, engines: [\"stockfish\", \"human\"]}`.\n\n" +
706
+ "Each node gets only the requested engines it is missing (`only_missing`, default true), so adding an engine later runs just that engine. Transpositions are analysed once and stamped on every node that reaches the position. Stored evals merge per engine; other engines' stored reads are kept. Results are saved after every batch of up to 10 positions.\n\n" +
707
+ "**Async job, returns immediately:** `{ job_id, target_count, already_complete, estimated_seconds }`. Poll auto_evaluate_status(job_id); cancel with auto_evaluate_cancel. Do other work between polls.\n\n" +
708
+ "Needs running rentals for every engine you ask for (ensure_engines). Wall time ≈ target_count × movetime (engines run in parallel per position). Does not set visible NAGs: the eval symbol is your call.",
731
709
  inputSchema: {
732
710
  type: "object",
733
- properties: { contract_id: { type: "string", description: "Which of your running rentals to use, from list_cloud_engines. Any engine shape works (combo, stockfish-only, lc0-only). Omit it when only one rental can serve the request; if several can, the error lists their contract ids." },
734
- id: { type: "string" },
711
+ properties: {
712
+ id: { type: "string", description: "Prep file id." },
713
+ engines: { type: "array", items: { type: "string", enum: ["stockfish", "lc0", "human"] }, description: "Engines to run. Default [\"stockfish\"]." },
735
714
  node_id: { type: "string", description: "Subtree root (default 'r' = whole file)." },
736
- only_missing: { type: "boolean", description: "Skip nodes that already carry a stored ceoEval (default true)." },
737
- movetime_ms: { type: "integer", minimum: 500, maximum: 5000, description: "Per-node cloud_analyse think time (default 1500)." },
715
+ only_missing: { type: "boolean", description: "Run only the engines a node has no stored eval for (default true). false re-analyses everything." },
716
+ movetime_ms: { type: "integer", minimum: 500, maximum: 15000, description: "Think time per position in ms (default 2000)." },
717
+ contempt: { type: "integer", minimum: -100, maximum: 100, description: "lc0 only. Stored evals then carry the bias; use for idea finding, not for the stored record." },
718
+ rentals: { type: "object", properties: { stockfish: { type: "string" }, lc0: { type: "string" }, human: { type: "string" } }, description: "Only when two running rentals provide the same engine: engine → contract_id." },
738
719
  },
739
720
  required: ["id"],
740
721
  },
741
722
  },
742
723
  {
743
724
  name: "auto_evaluate_status",
744
- description: "Poll the status of an auto_evaluate job. Response: `{ status: 'running' | 'done' | 'cancelled' | 'error' | 'not_found', target_count, evaluated, errored, remaining, done, error?, version? }`. When `status: 'not_found'` the job either expired (kept ~15 min after completion), never existed, or the MCP restarted since it was created — re-run auto_evaluate.\n\n" +
725
+ description: "Poll the status of an auto_evaluate job. Response: `{ status: 'running' | 'done' | 'cancelled' | 'error' | 'not_found', engines, target_count, processed, remaining, evaluated: {engine: n}, failed: {engine: n}, failed_node_ids: {engine: [ids]}, last_error?, aborted_reason?, done }`. Retry failures with cloud_analyse({file_id, node_ids}) or a new auto_evaluate. When `status: 'not_found'` the job either expired (kept ~15 min after completion), never existed, or the MCP restarted since it was created — re-run auto_evaluate.\n\n" +
745
726
  "Typical poll cadence: every 3-5 s for small walks, every 10-30 s for large ones. Don't hammer — status is a pure in-memory read but polling doesn't speed the engine up.",
746
727
  inputSchema: {
747
728
  type: "object",
@@ -753,7 +734,7 @@ export const TOOLS = [
753
734
  },
754
735
  {
755
736
  name: "auto_evaluate_cancel",
756
- description: "Ask a running auto_evaluate job to stop as soon as its current node finishes. Whatever progress was completed before cancellation is durably saved (checkpoint on cancel). Idempotent — cancelling an already-finished job is a no-op with a clear note in the response.",
737
+ description: "Ask a running auto_evaluate job to stop after the batch in flight. Every batch already analysed is saved in the file. Idempotent — cancelling an already-finished job is a no-op with a clear note in the response.",
757
738
  inputSchema: {
758
739
  type: "object",
759
740
  properties: {
@@ -764,34 +745,27 @@ export const TOOLS = [
764
745
  },
765
746
  {
766
747
  name: "deep_analyse",
767
- description: "Start a long Stockfish think on a single position (up to 5 min movetime). Returns a `job_id` immediately; poll `deep_analyse_status(job_id)` for the result, cancel with `deep_analyse_cancel(job_id)`. Runs SF only — Lc0 doesn't benefit from long thinks past a handful of seconds — and **holds only the SF engine slot on the rental, so `cloud_analyse(..., engines: [\"lc0\"])` stays available for other work in parallel**.\n\n" +
768
- "Use this when a specific critical position deserves depth — a novelty candidate, a hairy tactical shot, a difficult endgame — and you want Stockfish at depth 35+ rather than the ~depth 22 you get from a 2s cloud_analyse. Movetime is in ms; typical: 30_000-60_000 for 'careful check', 120_000-300_000 for 'find the truth'.\n\n" +
769
- "Result shape when done matches cloud_analyse's Stockfish leg (depth, top-N candidates with scoreCp/mate, best move, PV). Auto-stores the eval on `file_id`+`node_id` when both are supplied, same as cloud_analyse.",
748
+ description: "Start one long think (up to 5 min) on one position, on one engine (`engine`, default stockfish). Returns a `job_id` immediately; poll deep_analyse_status, cancel with deep_analyse_cancel. It holds only that engine, so cloud_analyse on the other engines keeps working meanwhile.\n\n" +
749
+ "Use it when a critical position deserves depth: a novelty candidate, a sharp tactic, a hard endgame. Typical movetime: 30_000-60_000 for a careful check, 120_000-300_000 to find the truth.\n\n" +
750
+ "With `file_id`+`node_id`, the result is merged into that node's ceoEval (other engines' stored reads are kept) and stamped on its transpositions.",
770
751
  inputSchema: {
771
752
  type: "object",
772
- properties: { contract_id: { type: "string", description: "Which of your running rentals to use, from list_cloud_engines. Any engine shape works (combo, stockfish-only, lc0-only). Omit it when only one rental can serve the request; if several can, the error lists their contract ids." },
773
- file_id: { type: "string", description: "Prep file id. Combine with `node_id` to derive FEN from the tree AND persist the result on the node's ceoEval." },
753
+ properties: {
754
+ engine: { type: "string", enum: ["stockfish", "lc0", "human"], description: "Default stockfish." },
755
+ file_id: { type: "string", description: "Prep file id. With `node_id`, the FEN comes from the tree and the result is stored on the node." },
774
756
  node_id: { type: "string", description: "Node id inside `file_id`. Root is 'r'. When set, overrides `fen`/`moves`." },
775
757
  fen: { type: "string", description: "Position as FEN. Only used if `node_id` is not set." },
776
758
  moves: { type: "string", description: "Optional SAN moves on top of `fen`. Only used if `node_id` is not set." },
777
- movetime_ms: {
778
- type: "integer",
779
- minimum: 5_000,
780
- maximum: 300_000,
781
- description: "Think time in ms. Default 60_000 (1 min). Max 300_000 (5 min).",
782
- },
783
- multipv: {
784
- type: "integer",
785
- minimum: 1,
786
- maximum: 10,
787
- description: "Number of candidate lines (default 2). Stockfish gets weaker as multipv grows — each extra PV steals search bandwidth from the top choice — so keep this low unless you specifically want to see several candidates ranked deep.",
788
- },
759
+ movetime_ms: { type: "integer", minimum: 5_000, maximum: 300_000, description: "Think time in ms. Default 60_000 (1 min). Max 300_000 (5 min)." },
760
+ multipv: { type: "integer", minimum: 1, maximum: 10, description: "Candidate lines. Default 2 for stockfish (extra lines cost it strength), 8 for lc0/human." },
761
+ contempt: { type: "integer", minimum: -100, maximum: 100, description: "lc0 only." },
762
+ rental: { type: "string", description: "Only when two running rentals provide this engine: the contract_id to use." },
789
763
  },
790
764
  },
791
765
  },
792
766
  {
793
767
  name: "deep_analyse_status",
794
- description: "Poll a deep_analyse job. Response: `{ status: 'running' | 'done' | 'cancelled' | 'error' | 'not_found', elapsed_ms, movetime_ms, result?, error? }`. `result` shape when done: `{ engine, depth, timeMs, bestMove, lines: [{rank, depth, scoreCp?, mate?, pv, nodes?}] }` — the SF leg of a cloud_analyse response.\n\n" +
768
+ description: "Poll a deep_analyse job. Response: `{ status: 'running' | 'done' | 'cancelled' | 'error' | 'not_found', elapsed_ms, movetime_ms, result?, error? }`. `result` when done: that engine's block from a cloud_analyse response (`{ depth, bestMove, lines: [{rank, depth, scoreCp?, mate?, wdl?, pv}] }`, PVs in SAN, wdl in percent); `stored_on` lists the nodes it was saved to.\n\n" +
795
769
  "Poll cadence: every ~15-30s for long thinks; there's no penalty for polling more often but the engine progresses at its own pace.",
796
770
  inputSchema: {
797
771
  type: "object",
@@ -864,7 +838,7 @@ export const TOOLS = [
864
838
  {
865
839
  name: "quote_engine_eval",
866
840
  description: "Return the stored engine eval for a node, or null if that node was never analysed. **Call this before writing prose or NAGs that quote engine numbers** — if it returns null, you have no measurement to cite. Do NOT infer an eval for the node from siblings or children; either analyse it (cloud_analyse with node_id) or omit the number from your prose.\n\n" +
867
- "Response: `{ ceoEval: { sf: {cp, depth}, lc0: {cp, depth}, nag } | null }`. `cp` is White-POV centipawns as an integer (+20 = +0.20). `nag` is the threshold-derived glyph as a SUGGESTION — promote to a visible NAG via set_nags only when a glyph on that move carries editorial signal.",
841
+ "Response: `{ ceoEval: { sf?: {cp | mate, depth}, lc0?: {w, d, l, depth}, human?: {w, d, l, depth} } | null }`. `cp` is White-POV centipawns (+20 = +0.20); `mate` White-POV moves to mate; `w/d/l` win/draw/loss percent from White's point of view. An engine missing from the object was never run on this position. Files analysed before 0.50 may hold an lc0 `cp` instead of w/d/l.",
868
842
  inputSchema: {
869
843
  type: "object",
870
844
  properties: {
@@ -1,6 +1,6 @@
1
1
  # Engine usage guide
2
2
 
3
- When you call `cloud_analyse`, chess.ceo runs Stockfish and Lc0 in parallel on the user's rented combo instance and returns both engines' final read. This doc explains what each engine is good for and how to interpret the numbers you get back — the difference between "this line is a draw" and "this line is easy to draw" is central to real prep and both engines are needed.
3
+ chess.ceo offers three cloud engines: Stockfish, Lc0 and the human engine. Each runs on its own rental, and `cloud_analyse` runs any mix of them on the same positions in parallel. This doc explains what each engine is good for, how to read the numbers, and how to run the work end to end. The difference between "this line is a draw" and "this line is easy to draw" is central to real prep, and it takes more than one engine to see it.
4
4
 
5
5
  ## Grounding: don't invent, run the engine
6
6
 
@@ -10,13 +10,13 @@ When you call `cloud_analyse`, chess.ceo runs Stockfish and Lc0 in parallel on t
10
10
 
11
11
  **When you're inside a prep file, call `cloud_analyse` with `file_id`+`node_id` — never with a hand-typed FEN.** The server derives the FEN from the tree node and, critically, **auto-stores the resulting eval on that node's `ceoEval`**. This is what makes engine attribution trustworthy end-to-end:
12
12
 
13
- 1. `cloud_analyse({ id, node_id })` — runs analysis on the node's exact position, stores `{sf, lc0, nag}` on the node.
13
+ 1. `cloud_analyse({ file_id, node_ids, engines })`: runs analysis on the nodes' exact positions and merges the result into each node's `ceoEval` (`sf` as cp/mate, `lc0` and `human` as win/draw/loss %). Engines you didn't run keep their stored read.
14
14
  2. Later, before you write "engines say X on this position" in a comment, call `quote_engine_eval({ id, node_id })` — it returns the stored eval or `null`.
15
15
  3. If `quote_engine_eval` returns `null`, you have no measurement to cite. **Do NOT infer an eval for a node from siblings, from children, or from a "position that looks similar."** Either analyse the node (`cloud_analyse`) or omit the number from your prose entirely.
16
16
 
17
17
  Concrete failure this rule blocks: the LLM says *"9...Bb7: both engines 0.00"* after only calling `cloud_analyse` on the child positions (post-1.d4, post-castling). Both continuations really returned 0.00, but the Bb7 node was never analysed, and the claim reads to the user as a measurement. With this protocol, `quote_engine_eval(node_id=Bb7)` would return null and the LLM would either analyse it or reword to *"both continuations run to 0.00, so this position looks balanced"* (soft inference, honestly labelled).
18
18
 
19
- **Transposition propagation.** `cloud_analyse({file_id, node_id})` also stamps the resulting `ceoEval` on every OTHER node in the same file that reaches the same position by a different move order (3-field FEN match: pieces + side-to-move + castling). The response includes `also_stored_on: [id, id]` when this happens, and a follow-up `quote_engine_eval` on any of those twin nodes returns the same measurement — no second analysis needed. `auto_evaluate` also dedupes candidates by the same key and returns `skipped_transpositions` so you can see how much engine time it saved. See `list_transpositions` / `list_nodes({filter: "transpositions"})` for auditing where duplication exists before you start.
19
+ **Transposition propagation.** `cloud_analyse({file_id, node_id})` also stamps the resulting `ceoEval` on every OTHER node in the same file that reaches the same position by a different move order (3-field FEN match: pieces + side-to-move + castling). Each position in the response lists `stored_on: [id, ...]`, every node it was saved to, and a follow-up `quote_engine_eval` on any of those twin nodes returns the same measurement — no second analysis needed. `auto_evaluate` also analyses each position once and returns `skipped_transpositions`, so you can see how much engine time it saved. See `list_transpositions` / `list_nodes({filter: "transpositions"})` for auditing where duplication exists before you start.
20
20
 
21
21
  **Concrete failure modes to avoid:**
22
22
 
@@ -73,8 +73,9 @@ Use it to:
73
73
 
74
74
  ### Rule of thumb
75
75
 
76
- - Objective truth ("does this hold?", "is this a mate?") → **trust Stockfish**
77
- - Practical prep ("which side is easier?", "which candidate is best?") → **trust Lc0**
76
+ - Objective truth ("does this hold?", "is this a mate?") → **trust Stockfish**. Run it on every position you evaluate.
77
+ - Practical prep ("which side is easier?", "which candidate is best?") → **Lc0**
78
+ - What a human will play, and how a human would judge the position → **the human engine**
78
79
  - Both agree → high confidence, ship the recommendation
79
80
  - They disagree → look at both scores together and reason about *why*:
80
81
  - Stockfish sharply higher: probably a tactic Lc0 didn't calculate
@@ -167,14 +168,15 @@ Absolute values are useful on their own too — a large "King safety" term flags
167
168
 
168
169
  Complements `cloud_analyse` (best move + PV): different questions, two angles on the same position.
169
170
 
170
- ## Deep Stockfish thinks: `deep_analyse`
171
+ ## Deep thinks: `deep_analyse`
171
172
 
172
- `cloud_analyse` runs both engines with a movetime cap of 10 s — deliberately fast because most opening-tree questions are answered in 2–3 s. For **one specific critical position** where you want depth 35+ instead of the usual depth 22, use `deep_analyse`:
173
+ `cloud_analyse` is built for short thinks (2 s default) because most opening-tree questions are answered in 2-3 s. For **one specific critical position** where you want depth 35+ instead of the usual depth 22, use `deep_analyse`:
173
174
 
174
- - SF-only (Lc0 saturates in a handful of seconds — no benefit past ~5 s of movetime).
175
+ - One engine (`engine`, default `stockfish`). Lc0 and the human engine gain little past a few seconds, so a deep think is almost always Stockfish.
175
176
  - Movetime up to 5 min.
176
- - Async: returns a `job_id` immediately, poll `deep_analyse_status(job_id)`, cancel with `deep_analyse_cancel(job_id)` if you decide partial depth is enough.
177
- - **Holds only the Stockfish slot on the combo.** Lc0 remains callable for other positions via `cloud_analyse({ engines: ["lc0"], … })` while the deep think runs. Use that during the wait — walk other branches, sanity-check candidates.
177
+ - Async: returns a `job_id` immediately, poll `deep_analyse_status(job_id)`, cancel with `deep_analyse_cancel(job_id)` if partial depth is enough.
178
+ - **Holds only that engine.** The other engines stay callable through `cloud_analyse` while the deep think runs. Use the wait to walk other branches.
179
+ - With `file_id`+`node_id`, the result is merged into the node's stored eval; the other engines' stored reads stay.
178
180
 
179
181
  Typical movetimes:
180
182
  - `30_000 – 60_000` (30–60 s) — careful check on a candidate move
@@ -184,53 +186,49 @@ Only use it when you actually need the depth. Regular `cloud_analyse` handles th
184
186
 
185
187
  ## Per-engine `multipv` defaults
186
188
 
187
- `cloud_analyse` defaults to different candidate-line counts per engine because the two engines behave differently:
189
+ `multipv: {stockfish, lc0, human}` sets the candidate-line count per engine. Defaults:
188
190
 
189
- - **Stockfish: `stockfish_multipv=2`** — SF gets weaker as multipv grows. Each extra PV steals search bandwidth from the top choice, so the "objective best move" leg is at its strongest with a tight list. Raise only when you specifically need SF's read on a wide range of candidates (e.g. sanity-checking an unusual sideline).
190
- - **Lc0: `lc0_multipv=8`** — Lc0 doesn't degrade the same way with multipv. Use it wide by default: 8 candidates give the LLM a real slate of practical ideas to inspect ("what does Lc0 think of the fun moves here?"). Lower it only when you don't need that breadth.
191
+ - **Stockfish: 2.** SF gets weaker as multipv grows: each extra PV steals search bandwidth from the top choice. Raise only when you need SF's read on a wide range of candidates.
192
+ - **Lc0 and human: 8.** Neural-net engines don't lose strength with multipv. 8 candidates give a real slate of ideas (Lc0) or of the moves people actually play (human).
191
193
 
192
194
  Mental model:
193
195
  - **Stockfish = the checker.** "Is this line objectively good? What's actually best?" Narrow, deep, one answer.
194
- - **Lc0 = the explorer.** "What's practically interesting? What alternatives are worth a look?" Wide, breadth-first, several answers.
196
+ - **Lc0 = the explorer.** "What's practically interesting? What long-term ideas are there?"
197
+ - **Human = the opponent.** "What will a person play here, and how will they judge it?"
195
198
 
196
- Override the defaults with `stockfish_multipv` / `lc0_multipv` when a specific position warrants a different shape.
199
+ ## Running the work: rent, analyse, stop
197
200
 
198
- ## Splitting engines on `cloud_analyse`
201
+ Every analysis needs a running rental for each engine you ask for. Each rental runs exactly one engine.
199
202
 
200
- `cloud_analyse` accepts an `engines` list to run only one leg:
203
+ 1. **Plan:** `ensure_engines({engines, positions})` shows which engines are already running, the cheapest SKU and price for each missing one, and a time and cost estimate. It starts nothing.
204
+ 2. **Ask:** show the user the price of anything you'd start and get a yes. Rentals bill per minute until stopped.
205
+ 3. **Start:** `start_cloud_engine({machine_type})` for each missing engine. The engine comes from the SKU. Static engines are ready at once; others take ~1-5 min (`list_cloud_engines` shows when).
206
+ 4. **Analyse:**
207
+ - A few positions: `cloud_analyse` with `fens`, `lines`, or `file_id`+`node_ids` (up to 10 per call), e.g. `{fens: [a, b, c], engines: ["stockfish", "human"]}`.
208
+ - A whole file or subtree: `auto_evaluate({id, engines})`. It runs only the engines each node is missing, analyses transpositions once, and saves as it goes. Adding an engine later runs just that engine.
209
+ - One critical position deep: `deep_analyse`.
210
+ 5. **Read back:** `quote_engine_eval` returns the stored evals per node.
211
+ 6. **Stop:** `stop_cloud_engine` when the work is done, unless the user wants the rental kept.
201
212
 
202
- - `engines: ["lc0"]` — Lc0 only, useful while a `deep_analyse` is holding the Stockfish slot.
203
- - `engines: ["stockfish"]` — SF only, when only the objective read matters and you want to skip the Lc0 latency.
204
- - Default (omitted) — both engines. This is what you want for real prep decisions.
213
+ If two running rentals provide the same engine, the call fails with both contract ids; pass `rentals: {engine: contract_id}` to choose.
205
214
 
206
- The skipped engine's field is omitted from the response (not present as an empty object).
215
+ ### Recipe: finding ideas
207
216
 
208
- ## Renting and choosing engines
217
+ 1. Run Lc0 or the human engine with a wide multipv on the position. The human engine shows the moves people play; Lc0 with contempt (e.g. `-20` to make Black play for a win) shows ambitious tries.
218
+ 2. Collect the interesting candidates and run Stockfish on the positions after each one (`lines` makes this one call).
219
+ 3. A candidate Stockfish refutes only with a line no human would find is a practical idea. A candidate that loses to a simple refutation is not.
209
220
 
210
- Analysis needs a running rental. Start one with `start_cloud_engine`, using a `machine_type` SKU from `list_cloud_machine_options`. Each SKU provides one of three shapes, shown in its `engine` field:
221
+ ### Eval symbols are your call
211
222
 
212
- - **`stockfish`**: Stockfish only. Fits when the question is objective (is this tactic real, does this defense hold).
213
- - **`lc0`**: Lc0 only. Fits the practical, human-feel read, or running alongside a `deep_analyse` that holds the Stockfish side.
214
- - **`combo`**: both engines in one container. The default, and the right choice for prep decisions, where both reads are needed on the same position.
215
-
216
- Rule of thumb: start a `combo` rental unless the task only needs one engine. A single-engine rental gives one read per position, so a prep walk that needs both reads on a single-engine rental takes two passes.
217
-
218
- `list_cloud_machine_options` returns the price for each SKU. Show the user the price and get their confirmation before starting a rental; every rental bills per second until `stop_cloud_engine`.
219
-
220
- Routing to a rental:
221
-
222
- - `cloud_analyse`, `auto_evaluate`, and `deep_analyse` each take an optional `contract_id` (from `list_cloud_engines`).
223
- - Without `contract_id`, the call uses the only running rental that can serve the request. If several can, the error lists their contract ids, and you pick one.
224
- - Any shape works. With `engines` set, the rental must provide those engines: asking for `engines: ["lc0"]` on a stockfish-only rental fails.
225
- - Do not guess contract ids. Call `list_cloud_engines` first if more than one rental is running.
223
+ The symbol you put on a position (=, +=, ±, +-) is written for a human reader, and no tool sets it for you. Use every engine you ran: a Stockfish 0.00 can still be `+=` when Lc0 or the human engine clearly prefer one side, because the symbol describes the practical picture, not only the objective one. Don't put a symbol on a position nobody analysed.
226
224
 
227
225
  ## Worked example
228
226
 
229
227
  User is preparing Black against a 2600 opponent who plays 1.e4 c5 2.Nf3 d6 3.d4 cxd4 4.Nxd4 Nf6 5.Nc3 a6 6.Be3 e5. You want to know if 7.Nb3 or 7.Nf3 is more testing.
230
228
 
231
229
  Two calls:
232
- 1. `cloud_analyse(fen=<position after 6...e5>, multipv=2)` — get both engines' top choices with movetime=2000.
233
- 2. If Stockfish scores them equal but Lc0 prefers one by 0.10-0.20, that's your practical answer. The user will find that line harder to face.
230
+ 1. `cloud_analyse({lines: ["<moves to 6...e5> Nb3", "<moves to 6...e5> Nf3"], engines: ["stockfish", "lc0", "human"]})`: all three engines on both positions in one call.
231
+ 2. If Stockfish scores them equal but Lc0 or the human engine clearly prefer one, that's your practical answer. The user will find that line harder to face.
234
232
 
235
233
  If the user is specifically preparing to *play* the black side in a must-win, add `contempt=-30` (or up to `-60` for a harder steer) on a follow-up call to see which lines Lc0 finds most fighting for Black. Compare against Stockfish's objective read to make sure the fighting choice isn't just losing.
236
234
 
@@ -242,7 +240,7 @@ If the user is specifically preparing to *play* the black side in a must-win, ad
242
240
 
243
241
  **To see further into a line, walk the tree.** Take the position at the tail of your truncated PV, run a fresh `cloud_analyse` on it. That call gives you the multipv candidate set at THAT position — the opponent's actual options — which is what you need to decide whether to branch. Don't raise `pv_max_plies` unless you're verifying a forcing sequence (a mate, an obligated recapture chain).
244
242
 
245
- Practical: raise `lc0_multipv` (default 8) to see the candidate spread on the current position, NOT to see further down one PV. If Lc0 shows moves 1-3 within 0.15 of each other, that's a branching point — three responses need coverage, not one PV.
243
+ Practical: raise `multipv.lc0` or `multipv.human` (default 8) to see the candidate spread on the current position, NOT to see further down one PV. If Lc0 shows moves 1-3 within 0.15 of each other, that's a branching point — three responses need coverage, not one PV.
246
244
 
247
245
  ## What NOT to do
248
246
 
@@ -40,12 +40,12 @@ All mutations **auto-save** with optimistic locking. Response includes the new `
40
40
 
41
41
  ## Typical build order
42
42
 
43
- 0a. **Cloud engine running?** Call `list_cloud_engines` first. Every substantive step below needs Stockfish + Lc0 to be reachable — `cloud_analyse` at critical positions, `auto_evaluate` for the whole tree, engine-derived NAGs at endpoints, describe_position's Stockfish eval-terms breakdown. If the caller has zero running combos: STOP, tell the user prep needs an engine, list options via `list_cloud_machine_options`, get the SKU + explicit confirmation (real money per second), then `start_cloud_engine`. If they already have one, note the contract_id and continue. Never silently write prep without engines — the file ends up with placeholder NAGs the user has no way to distinguish from real ones.
43
+ 0a. **Cloud engine running?** Call `ensure_engines` first. Every substantive step below needs engines: `cloud_analyse` at critical positions, `auto_evaluate` for the whole tree, engine-backed NAGs at endpoints, describe_position's Stockfish eval-terms breakdown. Stockfish at minimum; add Lc0 or the human engine for the practical read. If what you need isn't running: STOP, tell the user prep needs engines, show the price ensure_engines returns, get explicit confirmation (rentals bill per minute), then `start_cloud_engine`. Never silently write prep without engines — the file ends up with placeholder NAGs the user has no way to distinguish from real ones.
44
44
 
45
45
  0b. **`read_docs({ docs: ["pgn-authoring", "examples/najdorf-6-f4-repertoire"] })`** — do this once per session, before writing any prose. Not optional. Log analysis shows most sessions skip the example files and produce documented anti-patterns (long PVs in prose, restating what the app renders, verbose citations). Reading the reference PGN once inoculates the LLM against those.
46
46
  1. `read_prep_file` — see what's there. Every node has an `id` you'll pass to the mutation and engine/DB tools. Use `view: "compact"` (default) plus `node_id` + `max_depth` to scope; the full tree of a 500+-node file can blow the token limit.
47
47
  2. `apply_mutations([...])` — one call with your whole intended build (a mix of `add_move` / `add_line` for structure, plus any `set_comment`/`set_annotations` you already know at author time, plus any `set_nags` where you already have a clear judgment — novelty `$146`, `!?` speculative sac, obvious `?` blunder in a sideline you're rejecting).
48
- 3. `auto_evaluate(id)` — spawns a background job that PERSISTS engine numbers on every node. Does not touch visible NAGs. Cheap way to get every position's Stockfish + Lc0 read baked into the file for later reference. Grab the returned `job_id` and either (a) poll `auto_evaluate_status(job_id)` every ~10-30s until done, or (b) fire and do useful work meanwhile (write more of the tree, walk the opponent's repertoire) and check back later — engine walk-time serialises on the per-combo semaphore, so it takes roughly `target_count × movetime_ms` in wall time.
48
+ 3. `auto_evaluate({id, engines})` — spawns a background job that PERSISTS engine numbers on every node, for the engines you name (e.g. `["stockfish", "human"]`), filling only what each node is missing. Does not touch visible NAGs. Cheap way to get every position's engine reads baked into the file for later reference. Grab the returned `job_id` and either (a) poll `auto_evaluate_status(job_id)` every ~10-30s until done, or (b) fire and do useful work meanwhile (write more of the tree, walk the opponent's repertoire) and check back later — it takes roughly `target_count × movetime_ms` in wall time (engines run in parallel per position).
49
49
  4. Once the job reports `done:true`, re-read the file and add NAGs where they carry real signal (see NAG discipline below). Use individual mutation tools for surgical follow-ups (fix one comment, add one arrow, promote a specific variation, prune a branch).
50
50
 
51
51
  The build-cost math: a 200-move file via individual `add_move` calls is 200 saves ≈ 100+ seconds of tool overhead. The same file via one `apply_mutations` call is one save ≈ 500ms. Use batch by default.
@@ -210,7 +210,7 @@ This is the single biggest quality problem in current LLM output on this system:
210
210
 
211
211
  **Correct pattern for building a variation.** At every ply:
212
212
 
213
- 1. What are the plausible replies? `get_position_stats` (frequencies), `cloud_analyse` with `lc0_multipv: 8` (candidate spread), `predict_human_move` (what people actually play).
213
+ 1. What are the plausible replies? `get_position_stats` (frequencies), `cloud_analyse` with Lc0 or the human engine (8 candidate lines by default), `predict_human_move` (what people actually play).
214
214
  2. If ≥2 are plausible → branch. `add_move` each, then recurse on each branch or `add_line` for each of the several straightforward continuations.
215
215
  3. If exactly 1 → continue linearly; note *why* it's the only move in a comment.
216
216
 
@@ -85,7 +85,7 @@ If yes to all three, ship it. If no, either add what's missing or CUT the branch
85
85
 
86
86
  The same tools as any prep file, but different order and different density.
87
87
 
88
- 0. **Cloud engine running?** A summary looks light but requires SHARPER analysis than a big file — every endpoint NAG has to be right, and there are few enough of them that a wrong one stands out. Call `list_cloud_engines` first. Zero combos running → STOP: tell the user a summary needs engines, list options via `list_cloud_machine_options`, get their SKU + explicit confirmation (real money per second), then `start_cloud_engine`. Do NOT build a summary from cached / guessed evals — the point of the summary is trust, and a placeholder `$14` at an endpoint the reader will internalize is worse than no summary.
88
+ 0. **Cloud engine running?** A summary looks light but requires SHARPER analysis than a big file — every endpoint NAG has to be right, and there are few enough of them that a wrong one stands out. Call `ensure_engines` first. Nothing running → STOP: tell the user a summary needs engines, show the price it returns, get explicit confirmation (rentals bill per minute), then `start_cloud_engine`. Do NOT build a summary from cached / guessed evals — the point of the summary is trust, and a placeholder `$14` at an endpoint the reader will internalize is worse than no summary.
89
89
  0.5. **Which colour is the reader playing?** Ask if you don't know. This decides everything downstream — the mainline is their moves, the branches are their opponent's replies, the endpoint NAGs are judged from their POV, the voice is written to them. See the "Summaries are color-oriented" section above.
90
90
  1. `create_prep_file(collection_id, name)` — name it with the colour AND "Summary" in the Event tag (`"Modern Defence — Tiger ...a6/...b5 — Black (Summary)"`, `"Petroff 6.Bd3 Bd6 — White (Summary)"`) so the reader picks it out of a list immediately.
91
91
  2. `set_comment(root, "framing")` — the file's thesis, first thing.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@chessceo/mcp",
3
- "version": "0.49.10",
3
+ "version": "0.50.0",
4
4
  "description": "Model Context Protocol server for chess.ceo — 11.7M+ games, ~1.5M FIDE player profiles, opening preparation, live broadcasts.",
5
5
  "type": "module",
6
6
  "bin": {