@chessceo/mcp 0.49.4 → 0.49.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/index.js CHANGED
@@ -114,6 +114,21 @@ const SUMMARY_AUTHORING_DOC = loadBundledDoc("summary-authoring.md", "Summary au
114
114
  // all intact) — not summarised into English.
115
115
  const EXAMPLE_OVERVIEW_PGN = loadBundledDoc("examples/italian-fried-liver.pgn", "Italian Fried Liver overview example");
116
116
  const EXAMPLE_REPERTOIRE_PGN = loadBundledDoc("examples/najdorf-6-f4-white.pgn", "Najdorf 6.f4 White repertoire example");
117
+ // Per-engine guidance surfaced in list_cloud_machine_options' RESPONSE only
118
+ // (an AUTHED_TOOLS-gated call, both transports — see authedRequest's
119
+ // missing-token check and isAuthedToolCall's 401+WWW-Authenticate
120
+ // pre-check). Deliberately NOT in this tool's static `description` in
121
+ // tools.ts, since tools/list itself is reachable without authentication —
122
+ // putting it there would surface this to any caller who can list tools,
123
+ // not just ones who actually hold a token.
124
+ const ENGINE_GUIDE = {
125
+ stockfish: "Objective calculation — the ground truth for whether a line is actually winning/drawn/losing with best play, whether a tactic is real, whether a defense holds. Gives 0.00 to a large share of normal positions even when one side is much harder to play for a human — that means 'drawn with best play,' not 'trivial' or 'nothing to look for.'",
126
+ lc0: "Neural-net engine trained on human-style play, roughly 3500-strength. Its evaluation approximates how a strong human judges the position (practical chances, initiative, structure) rather than pure objective truth — good for finding ideas and gauging which side is easier to play over the board. Has real weaknesses in unconventional/irregular positions outside its training distribution: evals there can be unstable or it can miss a deep concrete tactic. When it disagrees sharply with Stockfish, look for the tactical justification before trusting its read.",
127
+ combo: "Runs Stockfish and Lc0 together — use when you want both the objective read and the practical/human-feel read on the same position.",
128
+ };
129
+ // Modern engines give equality very often, so small numbers are not "equal".
130
+ // Rough guide, not exact thresholds. Shown alongside ENGINE_GUIDE above.
131
+ const ENGINE_EVAL_SCALE = "Small numbers are not equality. ±0.00 to ±0.10 is equal in practice (+0.10 is '=' or '+=' at most). +0.20 to +0.40 is real pressure, not 'a bit better'. +0.50 and up is a clear advantage. Stockfish 0.00 with a plus from Lc0 is a practical edge, not equality. Check both engines before calling a position equal.";
117
132
  // v0.48: consolidated the five `read_*_guide` / `read_example_prep_files`
118
133
  // tools into ONE `read_docs`. LLM lists what it wants; we return them
119
134
  // in a single response. Also cleaner: enumerating available docs in one
@@ -367,8 +382,15 @@ async function callToolInner(name, args) {
367
382
  case "list_player_live_tournaments":
368
383
  // Note: snake_case fide_id, unlike the prep endpoints. Documented quirk.
369
384
  return get("/api/chess/live/player", { fide_id: Number(args.fide_id) });
370
- case "list_cloud_machine_options":
371
- return authedRequest("GET", "/api/agent/cloud-engines/options");
385
+ case "list_cloud_machine_options": {
386
+ const raw = await authedRequest("GET", "/api/agent/cloud-engines/options");
387
+ const resp = raw;
388
+ return {
389
+ ...(resp && typeof resp === "object" ? resp : { options: [] }),
390
+ engine_guide: ENGINE_GUIDE,
391
+ eval_scale: ENGINE_EVAL_SCALE,
392
+ };
393
+ }
372
394
  case "start_cloud_engine":
373
395
  return authedRequest("POST", "/api/agent/cloud-engines", {
374
396
  machineType: String(args.machine_type),
@@ -427,6 +449,7 @@ async function callToolInner(name, args) {
427
449
  }
428
450
  }
429
451
  }
452
+ converted.eval_scale = ENGINE_EVAL_SCALE;
430
453
  return converted;
431
454
  }
432
455
  case "deep_analyse":
package/dist/tools.js CHANGED
@@ -258,7 +258,7 @@ export const TOOLS = [
258
258
  },
259
259
  {
260
260
  name: "list_cloud_machine_options",
261
- description: "Returns the catalog of cloud-engine machine types the user can start — every engine shape (stockfish-only, lc0-only, and combo) they're entitled to, same visibility rules as the chess.ceo app (SKU, which engine(s) it runs, human display name, cost per hour, availability). ALWAYS call this before start_cloud_engine — SKU strings like 'rtx-5090-64' do not match the display names ('Stockfish 32 CPUs + Lc0 1× RTX 5090') and are NOT guessable, and a given SKU only supports the engine(s) shown in its `engine` field (a stockfish-only SKU cannot be started as lc0 or combo). Present the user the display names + prices; pass the SKU to start_cloud_engine.",
261
+ description: "Returns the catalog of cloud-engine machine types the user can start — every engine shape (stockfish-only, lc0-only, and combo) they're entitled to, same visibility rules as the chess.ceo app (SKU, which engine(s) it runs, human display name, cost per hour, availability), plus an `engine_guide` explaining what each engine type is actually good/bad for so you can pick the right one for the task, not just the cheapest. ALWAYS call this before start_cloud_engine — SKU strings like 'rtx-5090-64' do not match the display names ('Stockfish 32 CPUs + Lc0 1× RTX 5090') and are NOT guessable, and a given SKU only supports the engine(s) shown in its `engine` field (a stockfish-only SKU cannot be started as lc0 or combo). Present the user the display names + prices; pass the SKU to start_cloud_engine.",
262
262
  inputSchema: { type: "object", properties: {} },
263
263
  },
264
264
  {
@@ -89,6 +89,17 @@ If you actually want to estimate PRACTICAL drawing chance, ask the tools that an
89
89
 
90
90
  **Never conflate the objective eval with the practical outcome.** A 0.00 middlegame with a piece imbalance, opposite-side attacks, or a rating gap is a fighting game; the number just told you nobody has a forced win.
91
91
 
92
+ ## Eval scale: small numbers are not equality
93
+
94
+ Modern engines give equality very often, so the scale is compressed. A number that looks small can still be a real edge. Rough guide (approximate, not exact thresholds):
95
+
96
+ - **±0.00 to ±0.10**: equal in practice. A Stockfish +0.10 is "=" or "+=" at most.
97
+ - **+0.20 to +0.40**: real pressure, not "a bit better". The side with the plus has a clearly easier game; the other side has to find accurate moves for a long time.
98
+ - **+0.50 and up**: clear advantage in engine terms, and usually a real practical winning chance.
99
+ - **Stockfish 0.00 with a plus from Lc0**: objectively balanced, but Lc0 expects the side with the plus to play more easily. Report it as "practical edge", not "equal" and not "winning".
100
+
101
+ Never describe a position as equal from a 0.00 or ±0.10 reading alone when the other engine shows a plus. Check both numbers first.
102
+
92
103
  ## Lc0 contempt
93
104
 
94
105
  Contempt is an Lc0 option that skews its evaluation and move choice toward one side. Passing `contempt` to `cloud_analyse` sets it on the Lc0 leg only; Stockfish always analyzes objectively.
@@ -326,6 +326,20 @@ If you only quote the Lc0 number, disclose the bias in the same sentence — `{L
326
326
 
327
327
  Contempt scale is signed 0-100 (same as the web UI's ContemptStrength slider). Typical: `±10-20` a light nudge, `±30-60` real fighting play, `±80-100` maximum steer.
328
328
 
329
+ ### Claims to verify before you write them
330
+
331
+ Prose errors that reached a real study file (2026-10, London starter, all caught by the owner, not by the model):
332
+
333
+ - **"Wins back a pawn" / "pawn up".** Count the captures on the line before you claim material. A pawn push that the other side simply captures (e.g. `...c3` answered by `bxc3`) does not win anything back. Say what the move is for: activity, a file, a threat.
334
+ - **"Protected" / "defends".** Only say a piece is protected or defended if the recapture or the defending move is actually legal in that position. Check the square, not the idea.
335
+ - **"Blocks" / "stops".** Name the exact line or piece that is blocked. A knight on d2 does not block a queen on b6 from b2 (that is a different file/diagonal). If you can't name the blocked line, don't write it.
336
+ - **"Undefended" / "hangs".** Say which piece left which defending square, and check the other pieces that could still reach that square.
337
+ - **"Gains space".** Only for a move that actually advances or controls more squares. If the pawn is already on the square, say "has more space" or "controls the centre", not "after c4 gains space".
338
+ - **"Forcing" / "only move".** Only if each move in the PV is a check, a capture, or a direct threat. Otherwise say "natural" or "principled".
339
+ - **Engine words.** A number is not equality. Read the eval scale in `engine-usage.md` before writing "equal", "balanced", or "slight plus".
340
+
341
+ If a claim needs a board check, `describe_position` first (below), and drop the claim if the check doesn't support it.
342
+
329
343
  ### Before you write any commentary: describe_position
330
344
 
331
345
  **This is the biggest lever for prose quality in the whole system.** Live audit of a recent session — 13 `describe_position` calls versus 50+ `set_comment` ops. The nodes where `describe_position` was called first produced comments that grounded specifically in the position (correct piece squares, real pawn structure, actual weak squares). The nodes where it wasn't produced generic prose that pattern-matched to similar-*looking* positions and confidently named pieces on wrong squares. This gap is why `set_comment` now emits a warning whenever a substantive comment (≥40 chars) lands on a node whose position was never grounded via `describe_position` this session.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@chessceo/mcp",
3
- "version": "0.49.4",
3
+ "version": "0.49.6",
4
4
  "description": "Model Context Protocol server for chess.ceo — 11.7M+ games, ~1.5M FIDE player profiles, opening preparation, live broadcasts.",
5
5
  "type": "module",
6
6
  "bin": {