@chessceo/mcp 0.49.10 → 0.50.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/analysis/auto.js +173 -182
- package/dist/analysis/cloud.js +182 -0
- package/dist/analysis/deep.js +39 -66
- package/dist/analysis/file_handle.js +43 -36
- package/dist/analysis/response.js +180 -142
- package/dist/index.js +28 -65
- package/dist/pgn/exporter.js +13 -11
- package/dist/pgn/mutations.js +18 -5
- package/dist/pgn/parser.js +14 -15
- package/dist/pgn/types.js +2 -0
- package/dist/tools.js +75 -101
- package/docs/engine-usage.md +38 -40
- package/docs/pgn-authoring.md +3 -3
- package/docs/summary-authoring.md +1 -1
- package/package.json +1 -1
package/dist/tools.js
CHANGED
|
@@ -118,7 +118,7 @@ export const TOOLS = [
|
|
|
118
118
|
"- `gm-classical` — GM classical games (both players ≥2500, real thinking-time). BEST for opening prep — every move is signal, avgElo ~2600 across all listed moves.\n" +
|
|
119
119
|
"- `main` — the whole 11.7M-game DB. Widest coverage but noisiest (includes 1000-Elo blunder-fests in the move stats). Use as fallback when gm-classical's totalCount is too small to be informative.\n\n" +
|
|
120
120
|
"Game movetext is trimmed to the moves AFTER the queried position (using each game's plyNumber). Saves ~70% of the bytes vs full movetext.\n\n" +
|
|
121
|
-
"AUTO-EVAL: if a
|
|
121
|
+
"AUTO-EVAL: if a Stockfish rental is running, the response includes `.eval` with a short Stockfish read (cp/mate White's point of view, best move, PV). Not stored on any node; for a stored eval or other engines use cloud_analyse.",
|
|
122
122
|
inputSchema: {
|
|
123
123
|
type: "object",
|
|
124
124
|
properties: {
|
|
@@ -256,27 +256,34 @@ export const TOOLS = [
|
|
|
256
256
|
required: ["fide_id"],
|
|
257
257
|
},
|
|
258
258
|
},
|
|
259
|
+
{
|
|
260
|
+
name: "ensure_engines",
|
|
261
|
+
description: "Plan cloud-engine work before running it. For each engine you need (`stockfish`, `lc0`, `human`), says whether a rental providing it is already running (with its contract_id), or the cheapest SKU to start and its price. Pass `positions` (how many positions you will analyse) for a time and cost estimate. Starts nothing.\n\n" +
|
|
262
|
+
"CALL FIRST whenever a task needs cloud engines. Then show the user the price of anything to start, get a yes, and start it with start_cloud_engine. The response also carries `engine_guide` (what each engine is for) and `workflow` (the full plan → start → analyse → stop sequence).",
|
|
263
|
+
inputSchema: {
|
|
264
|
+
type: "object",
|
|
265
|
+
properties: {
|
|
266
|
+
engines: { type: "array", items: { type: "string", enum: ["stockfish", "lc0", "human"] }, description: "Engines the task needs. Default [\"stockfish\"]." },
|
|
267
|
+
positions: { type: "integer", minimum: 1, description: "Optional: how many positions you plan to analyse, for the estimate." },
|
|
268
|
+
movetime_ms: { type: "integer", minimum: 100, maximum: 300000, description: "Optional: think time per position for the estimate (default 2000)." },
|
|
269
|
+
},
|
|
270
|
+
},
|
|
271
|
+
},
|
|
259
272
|
{
|
|
260
273
|
name: "list_cloud_machine_options",
|
|
261
|
-
description: "Returns the catalog of cloud-engine
|
|
274
|
+
description: "Returns the catalog of cloud-engine SKUs the user can start, same visibility as the chess.ceo app: SKU (`machineType`), the one engine it runs (`engine`: stockfish, lc0 or human), display name, cost per hour, availability. Also returns `engine_guide`, `eval_scale`, `eval_symbols` and `workflow`. ensure_engines already picks the cheapest SKU per engine; use this when the user wants to choose hardware. SKU strings are NOT guessable from display names; pass them verbatim to start_cloud_engine.",
|
|
262
275
|
inputSchema: { type: "object", properties: {} },
|
|
263
276
|
},
|
|
264
277
|
{
|
|
265
278
|
name: "start_cloud_engine",
|
|
266
|
-
description: "Rent
|
|
267
|
-
"
|
|
268
|
-
"Use list_cloud_engines first to check if the user already has one running; don't start a second instance unless the user asked for it. Requires an MCP token with agent access.",
|
|
279
|
+
description: "Rent one cloud engine on the user's chess.ceo account. Each SKU runs exactly one engine (stockfish, lc0 or human), taken from the SKU. Real money: billed per minute from start until stop_cloud_engine.\n\n" +
|
|
280
|
+
"`machine_type` must be an exact SKU from ensure_engines or list_cloud_machine_options. Show the user the price and get their confirmation first. For several engines, start one rental per engine; cloud_analyse uses them together. Don't start an engine that is already running (ensure_engines / list_cloud_engines show what is).",
|
|
269
281
|
inputSchema: {
|
|
270
282
|
type: "object",
|
|
271
283
|
properties: {
|
|
272
284
|
machine_type: {
|
|
273
285
|
type: "string",
|
|
274
|
-
description: "SKU from
|
|
275
|
-
},
|
|
276
|
-
engine_type: {
|
|
277
|
-
type: "string",
|
|
278
|
-
enum: ["stockfish", "lc0", "combo"],
|
|
279
|
-
description: "Which engine(s) to run. Must match the SKU's own `engine` field from list_cloud_machine_options — e.g. a stockfish-only SKU can only be started as \"stockfish\". Optional; defaults to \"combo\".",
|
|
286
|
+
description: "Exact SKU from ensure_engines or list_cloud_machine_options. Not a display name, not a guess.",
|
|
280
287
|
},
|
|
281
288
|
},
|
|
282
289
|
required: ["machine_type"],
|
|
@@ -284,12 +291,12 @@ export const TOOLS = [
|
|
|
284
291
|
},
|
|
285
292
|
{
|
|
286
293
|
name: "list_cloud_engines",
|
|
287
|
-
description: "List the user's
|
|
294
|
+
description: "List the user's running cloud engines, each with its `contractId` and `engines` (what it provides). A row with empty `engines` is still starting. Use the contractId for stop_cloud_engine, or in `rentals` when two running rentals provide the same engine.",
|
|
288
295
|
inputSchema: { type: "object", properties: {} },
|
|
289
296
|
},
|
|
290
297
|
{
|
|
291
298
|
name: "stop_cloud_engine",
|
|
292
|
-
description: "Destroy a running cloud engine. Billing stops immediately. Use the contract_id from `list_cloud_engines
|
|
299
|
+
description: "Destroy a running cloud engine. Billing stops immediately. Use the contract_id from `list_cloud_engines`, don't guess. Stop rentals when the analysis is done unless the user wants them kept.",
|
|
293
300
|
inputSchema: {
|
|
294
301
|
type: "object",
|
|
295
302
|
properties: {
|
|
@@ -303,66 +310,38 @@ export const TOOLS = [
|
|
|
303
310
|
},
|
|
304
311
|
{
|
|
305
312
|
name: "cloud_analyse",
|
|
306
|
-
description: "
|
|
307
|
-
"
|
|
308
|
-
"
|
|
309
|
-
"
|
|
310
|
-
"
|
|
311
|
-
"
|
|
312
|
-
"
|
|
313
|
-
"
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
|
|
329
|
-
|
|
330
|
-
type: "integer",
|
|
331
|
-
minimum: 100,
|
|
332
|
-
maximum: 10000,
|
|
333
|
-
description: "Think time in milliseconds (default 2000).",
|
|
334
|
-
},
|
|
335
|
-
stockfish_multipv: {
|
|
336
|
-
type: "integer",
|
|
337
|
-
minimum: 1,
|
|
338
|
-
maximum: 10,
|
|
339
|
-
description: "Stockfish candidate lines (default 2). Kept tight because each extra PV steals search bandwidth from the top choice — SF is the 'what's objectively best' leg, use a low multipv to keep it strong. Raise only when you specifically need SF's take on a wide range of candidates.",
|
|
340
|
-
},
|
|
341
|
-
lc0_multipv: {
|
|
342
|
-
type: "integer",
|
|
343
|
-
minimum: 1,
|
|
344
|
-
maximum: 10,
|
|
345
|
-
description: "Lc0 candidate lines (default 8). Kept wide because multipv doesn't degrade Lc0's strength the way it does Stockfish's — Lc0 is the 'find inspiration / explore practical tries' leg, use a high multipv to get a full slate of ideas.",
|
|
346
|
-
},
|
|
347
|
-
contempt: {
|
|
348
|
-
type: "integer",
|
|
349
|
-
minimum: -100,
|
|
350
|
-
maximum: 100,
|
|
351
|
-
description: "Lc0 contempt bias. Signed 0-100 strength (same scale as the web UI's ContemptStrength slider — server multiplies by 8 to get the internal cp bias). 0 = objective (default). Positive favours White, negative favours Black. Typical: ±15 light nudge, ±30-60 real fighting play, ±80-100 maximum steer. Not applied to Stockfish. See engine_usage_primer for when to use.",
|
|
352
|
-
},
|
|
353
|
-
contract_id: { type: "string", description: "Which of your running rentals to use, from list_cloud_engines. Any engine shape works (combo, stockfish-only, lc0-only). Omit it when only one rental can serve the request; if several can, the error lists their contract ids." },
|
|
354
|
-
contract_ids: { type: "array", items: { type: "string" }, description: "Several running rentals to analyse the same FEN in parallel, from list_cloud_engines. Results merge by engine. Use it to get Stockfish and Lc0 on one position from separate rentals in one call." },
|
|
355
|
-
engines: {
|
|
356
|
-
type: "array",
|
|
357
|
-
items: { type: "string", enum: ["stockfish", "lc0"] },
|
|
358
|
-
description: "Which engines to run. Default = both. Use `[\"lc0\"]` to skip Stockfish (e.g. while a deep_analyse job is holding the SF slot on the same rental). Use `[\"stockfish\"]` when only the objective read matters. The skipped engine's field is omitted from the response.",
|
|
313
|
+
description: "Analyse up to 10 positions on the engines you choose (`stockfish`, `lc0`, `human`), each engine on whichever running rental provides it. Engines run in parallel per position. Returns, per position, each engine's depth, candidate lines and best move.\n\n" +
|
|
314
|
+
"Positions: `fens` (list), `lines` (list of SAN move sequences from `fen` or the start), `file_id` + `node_ids` (list), or a single `fen` / `moves` / `file_id`+`node_id`. Examples: analyse three FENs with Stockfish → `{fens: [a, b, c], engines: [\"stockfish\"]}`; with the human engine and Stockfish → `engines: [\"stockfish\", \"human\"]`. For a whole file or subtree use auto_evaluate instead.\n\n" +
|
|
315
|
+
"Scores: `scoreCp` and `mate` are White's point of view (+20 = White +0.20; mate +5 = White mates in 5). lc0 and human lines carry `wdl`: win/draw/loss percent, White's point of view. Read the response's `eval_scale` before calling a position equal.\n\n" +
|
|
316
|
+
"GROUNDING: every claim about a position must trace back to engine output from this session. Don't invent evaluations, best moves or variations. When you have no data for a position, run it or say so.\n\n" +
|
|
317
|
+
"Storing: with `file_id`, each result is merged into the `ceoEval` of every node that reaches that position (`stored_on` in the response), keeping other engines' stored reads. quote_engine_eval cites them later.\n\n" +
|
|
318
|
+
"Contempt (`contempt`, lc0 only): signed -100..100, positive favours White. Use it to find ideas (e.g. -20 makes Black play for a win). Never quote a contempt eval as objective.\n\n" +
|
|
319
|
+
"PVs are capped at 6 plies (`pv_truncated: true` when cut). To see further, analyse the position at the end of the line; raise `pv_max_plies` only to verify a forcing line. Don't paste PVs into add_line as prep.\n\n" +
|
|
320
|
+
"Needs running rentals for the engines you ask for (see ensure_engines). Costs money while rentals run; use get_position_stats for casual questions.",
|
|
321
|
+
inputSchema: {
|
|
322
|
+
type: "object",
|
|
323
|
+
properties: {
|
|
324
|
+
engines: { type: "array", items: { type: "string", enum: ["stockfish", "lc0", "human"] }, description: "Engines to run. Default [\"stockfish\"]. Each needs a running rental that provides it." },
|
|
325
|
+
fens: { type: "array", items: { type: "string" }, description: "Positions as FENs (up to 10 in total with the other inputs)." },
|
|
326
|
+
lines: { type: "array", items: { type: "string" }, description: "SAN move sequences, each applied from `fen` (or the start position), e.g. [\"e4 c5 Nf3\", \"e4 e5 Nf3\"]." },
|
|
327
|
+
file_id: { type: "string", description: "Prep file id. With node_ids/node_id the FEN comes from the tree; with any input, results are stored on matching nodes." },
|
|
328
|
+
node_ids: { type: "array", items: { type: "string" }, description: "Node ids inside `file_id`." },
|
|
329
|
+
node_id: { type: "string", description: "One node id inside `file_id`. Root is 'r'." },
|
|
330
|
+
fen: { type: "string", description: "One position, or the start position for `moves` / `lines`." },
|
|
331
|
+
moves: { type: "string", description: "SAN moves applied on top of `fen` (or the start position) for a single position." },
|
|
332
|
+
movetime_ms: { type: "integer", minimum: 100, maximum: 15000, description: "Think time per position in ms (default 2000)." },
|
|
333
|
+
multipv: {
|
|
334
|
+
type: "object",
|
|
335
|
+
properties: { stockfish: { type: "integer", minimum: 1, maximum: 10 }, lc0: { type: "integer", minimum: 1, maximum: 10 }, human: { type: "integer", minimum: 1, maximum: 10 } },
|
|
336
|
+
description: "Candidate lines per engine. Defaults: stockfish 2 (extra lines cost it strength), lc0 and human 8 (wide slate of ideas, no strength cost).",
|
|
359
337
|
},
|
|
360
|
-
|
|
361
|
-
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
description: "
|
|
338
|
+
contempt: { type: "integer", minimum: -100, maximum: 100, description: "lc0 only. Positive favours White, negative Black. 0 = objective (default)." },
|
|
339
|
+
rentals: {
|
|
340
|
+
type: "object",
|
|
341
|
+
properties: { stockfish: { type: "string" }, lc0: { type: "string" }, human: { type: "string" } },
|
|
342
|
+
description: "Only when two running rentals provide the same engine: engine → contract_id from list_cloud_engines.",
|
|
365
343
|
},
|
|
344
|
+
pv_max_plies: { type: "integer", minimum: 1, maximum: 40, description: "Cap each PV (default 6 plies). Raise only to verify a forcing line." },
|
|
366
345
|
},
|
|
367
346
|
},
|
|
368
347
|
},
|
|
@@ -723,25 +702,27 @@ export const TOOLS = [
|
|
|
723
702
|
},
|
|
724
703
|
{
|
|
725
704
|
name: "auto_evaluate",
|
|
726
|
-
description: "
|
|
727
|
-
"
|
|
728
|
-
"
|
|
729
|
-
"
|
|
730
|
-
"Costs real money — one cloud_analyse per node. A 200-node walk at default movetime is ~5 min of engine time (calls serialise on the per-engine semaphore in the backend).",
|
|
705
|
+
description: "Analyse a whole prep file, or the subtree under `node_id`, on the engines you choose and store the evals on every node. Example: \"analyse the whole tree with Stockfish and the human engine and save the evals\" → `{id, engines: [\"stockfish\", \"human\"]}`.\n\n" +
|
|
706
|
+
"Each node gets only the requested engines it is missing (`only_missing`, default true), so adding an engine later runs just that engine. Transpositions are analysed once and stamped on every node that reaches the position. Stored evals merge per engine; other engines' stored reads are kept. Results are saved after every batch of up to 10 positions.\n\n" +
|
|
707
|
+
"**Async job, returns immediately:** `{ job_id, target_count, already_complete, estimated_seconds }`. Poll auto_evaluate_status(job_id); cancel with auto_evaluate_cancel. Do other work between polls.\n\n" +
|
|
708
|
+
"Needs running rentals for every engine you ask for (ensure_engines). Wall time ≈ target_count × movetime (engines run in parallel per position). Does not set visible NAGs: the eval symbol is your call.",
|
|
731
709
|
inputSchema: {
|
|
732
710
|
type: "object",
|
|
733
|
-
properties: {
|
|
734
|
-
id: { type: "string" },
|
|
711
|
+
properties: {
|
|
712
|
+
id: { type: "string", description: "Prep file id." },
|
|
713
|
+
engines: { type: "array", items: { type: "string", enum: ["stockfish", "lc0", "human"] }, description: "Engines to run. Default [\"stockfish\"]." },
|
|
735
714
|
node_id: { type: "string", description: "Subtree root (default 'r' = whole file)." },
|
|
736
|
-
only_missing: { type: "boolean", description: "
|
|
737
|
-
movetime_ms: { type: "integer", minimum: 500, maximum:
|
|
715
|
+
only_missing: { type: "boolean", description: "Run only the engines a node has no stored eval for (default true). false re-analyses everything." },
|
|
716
|
+
movetime_ms: { type: "integer", minimum: 500, maximum: 15000, description: "Think time per position in ms (default 2000)." },
|
|
717
|
+
contempt: { type: "integer", minimum: -100, maximum: 100, description: "lc0 only. Stored evals then carry the bias; use for idea finding, not for the stored record." },
|
|
718
|
+
rentals: { type: "object", properties: { stockfish: { type: "string" }, lc0: { type: "string" }, human: { type: "string" } }, description: "Only when two running rentals provide the same engine: engine → contract_id." },
|
|
738
719
|
},
|
|
739
720
|
required: ["id"],
|
|
740
721
|
},
|
|
741
722
|
},
|
|
742
723
|
{
|
|
743
724
|
name: "auto_evaluate_status",
|
|
744
|
-
description: "Poll the status of an auto_evaluate job. Response: `{ status: 'running' | 'done' | 'cancelled' | 'error' | 'not_found',
|
|
725
|
+
description: "Poll the status of an auto_evaluate job. Response: `{ status: 'running' | 'done' | 'cancelled' | 'error' | 'not_found', engines, target_count, processed, remaining, evaluated: {engine: n}, failed: {engine: n}, failed_node_ids: {engine: [ids]}, last_error?, aborted_reason?, done }`. Retry failures with cloud_analyse({file_id, node_ids}) or a new auto_evaluate. When `status: 'not_found'` the job either expired (kept ~15 min after completion), never existed, or the MCP restarted since it was created — re-run auto_evaluate.\n\n" +
|
|
745
726
|
"Typical poll cadence: every 3-5 s for small walks, every 10-30 s for large ones. Don't hammer — status is a pure in-memory read but polling doesn't speed the engine up.",
|
|
746
727
|
inputSchema: {
|
|
747
728
|
type: "object",
|
|
@@ -753,7 +734,7 @@ export const TOOLS = [
|
|
|
753
734
|
},
|
|
754
735
|
{
|
|
755
736
|
name: "auto_evaluate_cancel",
|
|
756
|
-
description: "Ask a running auto_evaluate job to stop
|
|
737
|
+
description: "Ask a running auto_evaluate job to stop after the batch in flight. Every batch already analysed is saved in the file. Idempotent — cancelling an already-finished job is a no-op with a clear note in the response.",
|
|
757
738
|
inputSchema: {
|
|
758
739
|
type: "object",
|
|
759
740
|
properties: {
|
|
@@ -764,34 +745,27 @@ export const TOOLS = [
|
|
|
764
745
|
},
|
|
765
746
|
{
|
|
766
747
|
name: "deep_analyse",
|
|
767
|
-
description: "Start
|
|
768
|
-
"Use
|
|
769
|
-
"
|
|
748
|
+
description: "Start one long think (up to 5 min) on one position, on one engine (`engine`, default stockfish). Returns a `job_id` immediately; poll deep_analyse_status, cancel with deep_analyse_cancel. It holds only that engine, so cloud_analyse on the other engines keeps working meanwhile.\n\n" +
|
|
749
|
+
"Use it when a critical position deserves depth: a novelty candidate, a sharp tactic, a hard endgame. Typical movetime: 30_000-60_000 for a careful check, 120_000-300_000 to find the truth.\n\n" +
|
|
750
|
+
"With `file_id`+`node_id`, the result is merged into that node's ceoEval (other engines' stored reads are kept) and stamped on its transpositions.",
|
|
770
751
|
inputSchema: {
|
|
771
752
|
type: "object",
|
|
772
|
-
properties: {
|
|
773
|
-
|
|
753
|
+
properties: {
|
|
754
|
+
engine: { type: "string", enum: ["stockfish", "lc0", "human"], description: "Default stockfish." },
|
|
755
|
+
file_id: { type: "string", description: "Prep file id. With `node_id`, the FEN comes from the tree and the result is stored on the node." },
|
|
774
756
|
node_id: { type: "string", description: "Node id inside `file_id`. Root is 'r'. When set, overrides `fen`/`moves`." },
|
|
775
757
|
fen: { type: "string", description: "Position as FEN. Only used if `node_id` is not set." },
|
|
776
758
|
moves: { type: "string", description: "Optional SAN moves on top of `fen`. Only used if `node_id` is not set." },
|
|
777
|
-
movetime_ms: {
|
|
778
|
-
|
|
779
|
-
|
|
780
|
-
|
|
781
|
-
description: "Think time in ms. Default 60_000 (1 min). Max 300_000 (5 min).",
|
|
782
|
-
},
|
|
783
|
-
multipv: {
|
|
784
|
-
type: "integer",
|
|
785
|
-
minimum: 1,
|
|
786
|
-
maximum: 10,
|
|
787
|
-
description: "Number of candidate lines (default 2). Stockfish gets weaker as multipv grows — each extra PV steals search bandwidth from the top choice — so keep this low unless you specifically want to see several candidates ranked deep.",
|
|
788
|
-
},
|
|
759
|
+
movetime_ms: { type: "integer", minimum: 5_000, maximum: 300_000, description: "Think time in ms. Default 60_000 (1 min). Max 300_000 (5 min)." },
|
|
760
|
+
multipv: { type: "integer", minimum: 1, maximum: 10, description: "Candidate lines. Default 2 for stockfish (extra lines cost it strength), 8 for lc0/human." },
|
|
761
|
+
contempt: { type: "integer", minimum: -100, maximum: 100, description: "lc0 only." },
|
|
762
|
+
rental: { type: "string", description: "Only when two running rentals provide this engine: the contract_id to use." },
|
|
789
763
|
},
|
|
790
764
|
},
|
|
791
765
|
},
|
|
792
766
|
{
|
|
793
767
|
name: "deep_analyse_status",
|
|
794
|
-
description: "Poll a deep_analyse job. Response: `{ status: 'running' | 'done' | 'cancelled' | 'error' | 'not_found', elapsed_ms, movetime_ms, result?, error? }`. `result`
|
|
768
|
+
description: "Poll a deep_analyse job. Response: `{ status: 'running' | 'done' | 'cancelled' | 'error' | 'not_found', elapsed_ms, movetime_ms, result?, error? }`. `result` when done: that engine's block from a cloud_analyse response (`{ depth, bestMove, lines: [{rank, depth, scoreCp?, mate?, wdl?, pv}] }`, PVs in SAN, wdl in percent); `stored_on` lists the nodes it was saved to.\n\n" +
|
|
795
769
|
"Poll cadence: every ~15-30s for long thinks; there's no penalty for polling more often but the engine progresses at its own pace.",
|
|
796
770
|
inputSchema: {
|
|
797
771
|
type: "object",
|
|
@@ -864,7 +838,7 @@ export const TOOLS = [
|
|
|
864
838
|
{
|
|
865
839
|
name: "quote_engine_eval",
|
|
866
840
|
description: "Return the stored engine eval for a node, or null if that node was never analysed. **Call this before writing prose or NAGs that quote engine numbers** — if it returns null, you have no measurement to cite. Do NOT infer an eval for the node from siblings or children; either analyse it (cloud_analyse with node_id) or omit the number from your prose.\n\n" +
|
|
867
|
-
"Response: `{ ceoEval: { sf
|
|
841
|
+
"Response: `{ ceoEval: { sf?: {cp | mate, depth}, lc0?: {w, d, l, depth}, human?: {w, d, l, depth} } | null }`. `cp` is White-POV centipawns (+20 = +0.20); `mate` White-POV moves to mate; `w/d/l` win/draw/loss percent from White's point of view. An engine missing from the object was never run on this position. Files analysed before 0.50 may hold an lc0 `cp` instead of w/d/l.",
|
|
868
842
|
inputSchema: {
|
|
869
843
|
type: "object",
|
|
870
844
|
properties: {
|
package/docs/engine-usage.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Engine usage guide
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
chess.ceo offers three cloud engines: Stockfish, Lc0 and the human engine. Each runs on its own rental, and `cloud_analyse` runs any mix of them on the same positions in parallel. This doc explains what each engine is good for, how to read the numbers, and how to run the work end to end. The difference between "this line is a draw" and "this line is easy to draw" is central to real prep, and it takes more than one engine to see it.
|
|
4
4
|
|
|
5
5
|
## Grounding: don't invent, run the engine
|
|
6
6
|
|
|
@@ -10,13 +10,13 @@ When you call `cloud_analyse`, chess.ceo runs Stockfish and Lc0 in parallel on t
|
|
|
10
10
|
|
|
11
11
|
**When you're inside a prep file, call `cloud_analyse` with `file_id`+`node_id` — never with a hand-typed FEN.** The server derives the FEN from the tree node and, critically, **auto-stores the resulting eval on that node's `ceoEval`**. This is what makes engine attribution trustworthy end-to-end:
|
|
12
12
|
|
|
13
|
-
1. `cloud_analyse({
|
|
13
|
+
1. `cloud_analyse({ file_id, node_ids, engines })`: runs analysis on the nodes' exact positions and merges the result into each node's `ceoEval` (`sf` as cp/mate, `lc0` and `human` as win/draw/loss %). Engines you didn't run keep their stored read.
|
|
14
14
|
2. Later, before you write "engines say X on this position" in a comment, call `quote_engine_eval({ id, node_id })` — it returns the stored eval or `null`.
|
|
15
15
|
3. If `quote_engine_eval` returns `null`, you have no measurement to cite. **Do NOT infer an eval for a node from siblings, from children, or from a "position that looks similar."** Either analyse the node (`cloud_analyse`) or omit the number from your prose entirely.
|
|
16
16
|
|
|
17
17
|
Concrete failure this rule blocks: the LLM says *"9...Bb7: both engines 0.00"* after only calling `cloud_analyse` on the child positions (post-1.d4, post-castling). Both continuations really returned 0.00, but the Bb7 node was never analysed, and the claim reads to the user as a measurement. With this protocol, `quote_engine_eval(node_id=Bb7)` would return null and the LLM would either analyse it or reword to *"both continuations run to 0.00, so this position looks balanced"* (soft inference, honestly labelled).
|
|
18
18
|
|
|
19
|
-
**Transposition propagation.** `cloud_analyse({file_id, node_id})` also stamps the resulting `ceoEval` on every OTHER node in the same file that reaches the same position by a different move order (3-field FEN match: pieces + side-to-move + castling).
|
|
19
|
+
**Transposition propagation.** `cloud_analyse({file_id, node_id})` also stamps the resulting `ceoEval` on every OTHER node in the same file that reaches the same position by a different move order (3-field FEN match: pieces + side-to-move + castling). Each position in the response lists `stored_on: [id, ...]`, every node it was saved to, and a follow-up `quote_engine_eval` on any of those twin nodes returns the same measurement — no second analysis needed. `auto_evaluate` also analyses each position once and returns `skipped_transpositions`, so you can see how much engine time it saved. See `list_transpositions` / `list_nodes({filter: "transpositions"})` for auditing where duplication exists before you start.
|
|
20
20
|
|
|
21
21
|
**Concrete failure modes to avoid:**
|
|
22
22
|
|
|
@@ -73,8 +73,9 @@ Use it to:
|
|
|
73
73
|
|
|
74
74
|
### Rule of thumb
|
|
75
75
|
|
|
76
|
-
- Objective truth ("does this hold?", "is this a mate?") → **trust Stockfish
|
|
77
|
-
- Practical prep ("which side is easier?", "which candidate is best?") → **
|
|
76
|
+
- Objective truth ("does this hold?", "is this a mate?") → **trust Stockfish**. Run it on every position you evaluate.
|
|
77
|
+
- Practical prep ("which side is easier?", "which candidate is best?") → **Lc0**
|
|
78
|
+
- What a human will play, and how a human would judge the position → **the human engine**
|
|
78
79
|
- Both agree → high confidence, ship the recommendation
|
|
79
80
|
- They disagree → look at both scores together and reason about *why*:
|
|
80
81
|
- Stockfish sharply higher: probably a tactic Lc0 didn't calculate
|
|
@@ -167,14 +168,15 @@ Absolute values are useful on their own too — a large "King safety" term flags
|
|
|
167
168
|
|
|
168
169
|
Complements `cloud_analyse` (best move + PV): different questions, two angles on the same position.
|
|
169
170
|
|
|
170
|
-
## Deep
|
|
171
|
+
## Deep thinks: `deep_analyse`
|
|
171
172
|
|
|
172
|
-
`cloud_analyse`
|
|
173
|
+
`cloud_analyse` is built for short thinks (2 s default) because most opening-tree questions are answered in 2-3 s. For **one specific critical position** where you want depth 35+ instead of the usual depth 22, use `deep_analyse`:
|
|
173
174
|
|
|
174
|
-
-
|
|
175
|
+
- One engine (`engine`, default `stockfish`). Lc0 and the human engine gain little past a few seconds, so a deep think is almost always Stockfish.
|
|
175
176
|
- Movetime up to 5 min.
|
|
176
|
-
- Async: returns a `job_id` immediately, poll `deep_analyse_status(job_id)`, cancel with `deep_analyse_cancel(job_id)` if
|
|
177
|
-
- **Holds only
|
|
177
|
+
- Async: returns a `job_id` immediately, poll `deep_analyse_status(job_id)`, cancel with `deep_analyse_cancel(job_id)` if partial depth is enough.
|
|
178
|
+
- **Holds only that engine.** The other engines stay callable through `cloud_analyse` while the deep think runs. Use the wait to walk other branches.
|
|
179
|
+
- With `file_id`+`node_id`, the result is merged into the node's stored eval; the other engines' stored reads stay.
|
|
178
180
|
|
|
179
181
|
Typical movetimes:
|
|
180
182
|
- `30_000 – 60_000` (30–60 s) — careful check on a candidate move
|
|
@@ -184,53 +186,49 @@ Only use it when you actually need the depth. Regular `cloud_analyse` handles th
|
|
|
184
186
|
|
|
185
187
|
## Per-engine `multipv` defaults
|
|
186
188
|
|
|
187
|
-
`
|
|
189
|
+
`multipv: {stockfish, lc0, human}` sets the candidate-line count per engine. Defaults:
|
|
188
190
|
|
|
189
|
-
- **Stockfish:
|
|
190
|
-
- **Lc0:
|
|
191
|
+
- **Stockfish: 2.** SF gets weaker as multipv grows: each extra PV steals search bandwidth from the top choice. Raise only when you need SF's read on a wide range of candidates.
|
|
192
|
+
- **Lc0 and human: 8.** Neural-net engines don't lose strength with multipv. 8 candidates give a real slate of ideas (Lc0) or of the moves people actually play (human).
|
|
191
193
|
|
|
192
194
|
Mental model:
|
|
193
195
|
- **Stockfish = the checker.** "Is this line objectively good? What's actually best?" Narrow, deep, one answer.
|
|
194
|
-
- **Lc0 = the explorer.** "What's practically interesting? What
|
|
196
|
+
- **Lc0 = the explorer.** "What's practically interesting? What long-term ideas are there?"
|
|
197
|
+
- **Human = the opponent.** "What will a person play here, and how will they judge it?"
|
|
195
198
|
|
|
196
|
-
|
|
199
|
+
## Running the work: rent, analyse, stop
|
|
197
200
|
|
|
198
|
-
|
|
201
|
+
Every analysis needs a running rental for each engine you ask for. Each rental runs exactly one engine.
|
|
199
202
|
|
|
200
|
-
`
|
|
203
|
+
1. **Plan:** `ensure_engines({engines, positions})` shows which engines are already running, the cheapest SKU and price for each missing one, and a time and cost estimate. It starts nothing.
|
|
204
|
+
2. **Ask:** show the user the price of anything you'd start and get a yes. Rentals bill per minute until stopped.
|
|
205
|
+
3. **Start:** `start_cloud_engine({machine_type})` for each missing engine. The engine comes from the SKU. Static engines are ready at once; others take ~1-5 min (`list_cloud_engines` shows when).
|
|
206
|
+
4. **Analyse:**
|
|
207
|
+
- A few positions: `cloud_analyse` with `fens`, `lines`, or `file_id`+`node_ids` (up to 10 per call), e.g. `{fens: [a, b, c], engines: ["stockfish", "human"]}`.
|
|
208
|
+
- A whole file or subtree: `auto_evaluate({id, engines})`. It runs only the engines each node is missing, analyses transpositions once, and saves as it goes. Adding an engine later runs just that engine.
|
|
209
|
+
- One critical position deep: `deep_analyse`.
|
|
210
|
+
5. **Read back:** `quote_engine_eval` returns the stored evals per node.
|
|
211
|
+
6. **Stop:** `stop_cloud_engine` when the work is done, unless the user wants the rental kept.
|
|
201
212
|
|
|
202
|
-
|
|
203
|
-
- `engines: ["stockfish"]` — SF only, when only the objective read matters and you want to skip the Lc0 latency.
|
|
204
|
-
- Default (omitted) — both engines. This is what you want for real prep decisions.
|
|
213
|
+
If two running rentals provide the same engine, the call fails with both contract ids; pass `rentals: {engine: contract_id}` to choose.
|
|
205
214
|
|
|
206
|
-
|
|
215
|
+
### Recipe: finding ideas
|
|
207
216
|
|
|
208
|
-
|
|
217
|
+
1. Run Lc0 or the human engine with a wide multipv on the position. The human engine shows the moves people play; Lc0 with contempt (e.g. `-20` to make Black play for a win) shows ambitious tries.
|
|
218
|
+
2. Collect the interesting candidates and run Stockfish on the positions after each one (`lines` makes this one call).
|
|
219
|
+
3. A candidate Stockfish refutes only with a line no human would find is a practical idea. A candidate that loses to a simple refutation is not.
|
|
209
220
|
|
|
210
|
-
|
|
221
|
+
### Eval symbols are your call
|
|
211
222
|
|
|
212
|
-
|
|
213
|
-
- **`lc0`**: Lc0 only. Fits the practical, human-feel read, or running alongside a `deep_analyse` that holds the Stockfish side.
|
|
214
|
-
- **`combo`**: both engines in one container. The default, and the right choice for prep decisions, where both reads are needed on the same position.
|
|
215
|
-
|
|
216
|
-
Rule of thumb: start a `combo` rental unless the task only needs one engine. A single-engine rental gives one read per position, so a prep walk that needs both reads on a single-engine rental takes two passes.
|
|
217
|
-
|
|
218
|
-
`list_cloud_machine_options` returns the price for each SKU. Show the user the price and get their confirmation before starting a rental; every rental bills per second until `stop_cloud_engine`.
|
|
219
|
-
|
|
220
|
-
Routing to a rental:
|
|
221
|
-
|
|
222
|
-
- `cloud_analyse`, `auto_evaluate`, and `deep_analyse` each take an optional `contract_id` (from `list_cloud_engines`).
|
|
223
|
-
- Without `contract_id`, the call uses the only running rental that can serve the request. If several can, the error lists their contract ids, and you pick one.
|
|
224
|
-
- Any shape works. With `engines` set, the rental must provide those engines: asking for `engines: ["lc0"]` on a stockfish-only rental fails.
|
|
225
|
-
- Do not guess contract ids. Call `list_cloud_engines` first if more than one rental is running.
|
|
223
|
+
The symbol you put on a position (=, +=, ±, +-) is written for a human reader, and no tool sets it for you. Use every engine you ran: a Stockfish 0.00 can still be `+=` when Lc0 or the human engine clearly prefer one side, because the symbol describes the practical picture, not only the objective one. Don't put a symbol on a position nobody analysed.
|
|
226
224
|
|
|
227
225
|
## Worked example
|
|
228
226
|
|
|
229
227
|
User is preparing Black against a 2600 opponent who plays 1.e4 c5 2.Nf3 d6 3.d4 cxd4 4.Nxd4 Nf6 5.Nc3 a6 6.Be3 e5. You want to know if 7.Nb3 or 7.Nf3 is more testing.
|
|
230
228
|
|
|
231
229
|
Two calls:
|
|
232
|
-
1. `cloud_analyse(
|
|
233
|
-
2. If Stockfish scores them equal but Lc0
|
|
230
|
+
1. `cloud_analyse({lines: ["<moves to 6...e5> Nb3", "<moves to 6...e5> Nf3"], engines: ["stockfish", "lc0", "human"]})`: all three engines on both positions in one call.
|
|
231
|
+
2. If Stockfish scores them equal but Lc0 or the human engine clearly prefer one, that's your practical answer. The user will find that line harder to face.
|
|
234
232
|
|
|
235
233
|
If the user is specifically preparing to *play* the black side in a must-win, add `contempt=-30` (or up to `-60` for a harder steer) on a follow-up call to see which lines Lc0 finds most fighting for Black. Compare against Stockfish's objective read to make sure the fighting choice isn't just losing.
|
|
236
234
|
|
|
@@ -242,7 +240,7 @@ If the user is specifically preparing to *play* the black side in a must-win, ad
|
|
|
242
240
|
|
|
243
241
|
**To see further into a line, walk the tree.** Take the position at the tail of your truncated PV, run a fresh `cloud_analyse` on it. That call gives you the multipv candidate set at THAT position — the opponent's actual options — which is what you need to decide whether to branch. Don't raise `pv_max_plies` unless you're verifying a forcing sequence (a mate, an obligated recapture chain).
|
|
244
242
|
|
|
245
|
-
Practical: raise `
|
|
243
|
+
Practical: raise `multipv.lc0` or `multipv.human` (default 8) to see the candidate spread on the current position, NOT to see further down one PV. If Lc0 shows moves 1-3 within 0.15 of each other, that's a branching point — three responses need coverage, not one PV.
|
|
246
244
|
|
|
247
245
|
## What NOT to do
|
|
248
246
|
|
package/docs/pgn-authoring.md
CHANGED
|
@@ -40,12 +40,12 @@ All mutations **auto-save** with optimistic locking. Response includes the new `
|
|
|
40
40
|
|
|
41
41
|
## Typical build order
|
|
42
42
|
|
|
43
|
-
0a. **Cloud engine running?** Call `
|
|
43
|
+
0a. **Cloud engine running?** Call `ensure_engines` first. Every substantive step below needs engines: `cloud_analyse` at critical positions, `auto_evaluate` for the whole tree, engine-backed NAGs at endpoints, describe_position's Stockfish eval-terms breakdown. Stockfish at minimum; add Lc0 or the human engine for the practical read. If what you need isn't running: STOP, tell the user prep needs engines, show the price ensure_engines returns, get explicit confirmation (rentals bill per minute), then `start_cloud_engine`. Never silently write prep without engines — the file ends up with placeholder NAGs the user has no way to distinguish from real ones.
|
|
44
44
|
|
|
45
45
|
0b. **`read_docs({ docs: ["pgn-authoring", "examples/najdorf-6-f4-repertoire"] })`** — do this once per session, before writing any prose. Not optional. Log analysis shows most sessions skip the example files and produce documented anti-patterns (long PVs in prose, restating what the app renders, verbose citations). Reading the reference PGN once inoculates the LLM against those.
|
|
46
46
|
1. `read_prep_file` — see what's there. Every node has an `id` you'll pass to the mutation and engine/DB tools. Use `view: "compact"` (default) plus `node_id` + `max_depth` to scope; the full tree of a 500+-node file can blow the token limit.
|
|
47
47
|
2. `apply_mutations([...])` — one call with your whole intended build (a mix of `add_move` / `add_line` for structure, plus any `set_comment`/`set_annotations` you already know at author time, plus any `set_nags` where you already have a clear judgment — novelty `$146`, `!?` speculative sac, obvious `?` blunder in a sideline you're rejecting).
|
|
48
|
-
3. `auto_evaluate(id)` — spawns a background job that PERSISTS engine numbers on every node. Does not touch visible NAGs. Cheap way to get every position's
|
|
48
|
+
3. `auto_evaluate({id, engines})` — spawns a background job that PERSISTS engine numbers on every node, for the engines you name (e.g. `["stockfish", "human"]`), filling only what each node is missing. Does not touch visible NAGs. Cheap way to get every position's engine reads baked into the file for later reference. Grab the returned `job_id` and either (a) poll `auto_evaluate_status(job_id)` every ~10-30s until done, or (b) fire and do useful work meanwhile (write more of the tree, walk the opponent's repertoire) and check back later — it takes roughly `target_count × movetime_ms` in wall time (engines run in parallel per position).
|
|
49
49
|
4. Once the job reports `done:true`, re-read the file and add NAGs where they carry real signal (see NAG discipline below). Use individual mutation tools for surgical follow-ups (fix one comment, add one arrow, promote a specific variation, prune a branch).
|
|
50
50
|
|
|
51
51
|
The build-cost math: a 200-move file via individual `add_move` calls is 200 saves ≈ 100+ seconds of tool overhead. The same file via one `apply_mutations` call is one save ≈ 500ms. Use batch by default.
|
|
@@ -210,7 +210,7 @@ This is the single biggest quality problem in current LLM output on this system:
|
|
|
210
210
|
|
|
211
211
|
**Correct pattern for building a variation.** At every ply:
|
|
212
212
|
|
|
213
|
-
1. What are the plausible replies? `get_position_stats` (frequencies), `cloud_analyse` with
|
|
213
|
+
1. What are the plausible replies? `get_position_stats` (frequencies), `cloud_analyse` with Lc0 or the human engine (8 candidate lines by default), `predict_human_move` (what people actually play).
|
|
214
214
|
2. If ≥2 are plausible → branch. `add_move` each, then recurse on each branch or `add_line` for each of the several straightforward continuations.
|
|
215
215
|
3. If exactly 1 → continue linearly; note *why* it's the only move in a comment.
|
|
216
216
|
|
|
@@ -85,7 +85,7 @@ If yes to all three, ship it. If no, either add what's missing or CUT the branch
|
|
|
85
85
|
|
|
86
86
|
The same tools as any prep file, but different order and different density.
|
|
87
87
|
|
|
88
|
-
0. **Cloud engine running?** A summary looks light but requires SHARPER analysis than a big file — every endpoint NAG has to be right, and there are few enough of them that a wrong one stands out. Call `
|
|
88
|
+
0. **Cloud engine running?** A summary looks light but requires SHARPER analysis than a big file — every endpoint NAG has to be right, and there are few enough of them that a wrong one stands out. Call `ensure_engines` first. Nothing running → STOP: tell the user a summary needs engines, show the price it returns, get explicit confirmation (rentals bill per minute), then `start_cloud_engine`. Do NOT build a summary from cached / guessed evals — the point of the summary is trust, and a placeholder `$14` at an endpoint the reader will internalize is worse than no summary.
|
|
89
89
|
0.5. **Which colour is the reader playing?** Ask if you don't know. This decides everything downstream — the mainline is their moves, the branches are their opponent's replies, the endpoint NAGs are judged from their POV, the voice is written to them. See the "Summaries are color-oriented" section above.
|
|
90
90
|
1. `create_prep_file(collection_id, name)` — name it with the colour AND "Summary" in the Event tag (`"Modern Defence — Tiger ...a6/...b5 — Black (Summary)"`, `"Petroff 6.Bd3 Bd6 — White (Summary)"`) so the reader picks it out of a list immediately.
|
|
91
91
|
2. `set_comment(root, "framing")` — the file's thesis, first thing.
|
package/package.json
CHANGED