@kolbo/mcp 1.86.4 → 1.87.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -23,48 +23,20 @@ Load this file when the user wants a **Seedance 2 / Seedance 2.0** (ByteDance) v
23
23
  Then Locked Intro, then `SHOT N — 0:00–0:02 — Size / camera` beats whose ranges **sum exactly to Xs**. Last line repeats `Total: Xs / N shots / AR`.
24
24
  - Example (15s / 6 shots): `6 connected cinematic shots, 15 seconds total, 16:9, Multishot ON` + `Total: 15s / 6 shots / 16:9`
25
25
  - UGC / phone vertical: `N connected phone shots, Xs total, 9:16, Multishot ON` (never the word "cinematic").
26
- - A MULTI-shot prompt with only shot body and no Total / Multishot header is a **failed turn** — rewrite before calling `generate_*`. A single-shot prompt uses the single-shot header instead and carries no `Multishot ON`; see "Shot count comes from the USER".
26
+ - A prompt with only shot body and no Total / Multishot header is a **failed turn** — rewrite before calling `generate_*`.
27
27
  - **MCP `duration` must match the Total line.** Pass `duration: X` (whole seconds) on `generate_video` / `generate_elements` / `generate_video_from_image` equal to the `Xs` in `Total: Xs / …`. Mismatch = wrong-length clip.
28
28
  - **Then the Locked Intro** — `[GLOBAL LOOK]` / `[CAST]` / `[LOCATION]` (+ LOCATION MAP / CONTINUITY / PHYSICS for multi-shot) — before any shot. A one-liner `same character throughout` is not a character lock.
29
29
  - **Order inside each shot**: Subject → Action → Camera → Constraints → (Audio/SFX if relevant). Do NOT restack GLOBAL LOOK style inside the shot.
30
30
  - **Prompt length**: simple single-idea pieces ~120–280 words. Locked-intro cinematic typically 400–900 words. Shorter than ~120 words = random output. The 10,000-char cap below always wins.
31
- - **Shot count is user-directed.** If the user asks for N shots, deliver exactly N in one prompt unless they ask to split — and if they ask for ONE shot, deliver one shot with no `Multishot ON`.
31
+ - **Shot count is user-directed.** If the user asks for N shots, deliver exactly N in one prompt unless they ask to split.
32
32
  - **Always describe at least one camera movement per shot.**
33
33
  - **Tell Seedance what the camera is NOT doing** (e.g. `no cuts, no zoom, natural head movement`) — this is what locks POV.
34
34
  - **Final prompt is always English**, wrapped in a copy-ready code block. Detect intent in any language and reply in the user's language, but the prompt itself is English.
35
35
  - **HARD CAP: 10,000 characters TOTAL for the ENTIRE prompt** — measured as one single string including all shots, boilerplate, SFX lines, and the Total lines. It is per PROMPT, not per shot. **Never** split into multiple prompts, code blocks, or "part 1 / part 2" to evade the cap. Count the final prompt before output; if over, trim (cut adjectives, collapse boilerplate, shorten SFX lists, merge or drop shots) and re-count until it fits.
36
36
 
37
+ ## Locked Intro (DEFAULT for any multi-shot cinematic — including Elements)
37
38
 
38
- ## Shot count comes from the USERdecide this FIRST
39
-
40
- A SHOT is one uninterrupted camera take. A CUT is what separates two shots. Count
41
- what the user asked for before choosing the output shape.
42
-
43
- **ONE shot requested → write ONE shot.**
44
- - No `SHOT N` labels, no per-shot beats, and **no `Multishot ON`**. That flag declares
45
- "this clip contains hard cuts" — on a single take it is false, and it additionally
46
- forces prompt enhancement on the wire, rewriting the prompt the user just approved.
47
- - Header: `Single continuous shot, Xs total, AR`. Closing line: `Total: Xs / 1 shot / AR`.
48
- - Describe the take as ONE unbroken movement; internal beats are timestamps inside it
49
- (`0:00–0:05 — the camera pushes in past the doorway…`), never numbered shots.
50
- - **Keep the Locked Intro blocks.** GLOBAL LOOK / CAST / LOCATION / PHYSICS are the
51
- consistency stack, not the multi-shot part — a long continuous move needs them most.
52
-
53
- These all mean ONE shot, however long it runs and however far the camera travels:
54
- "one shot", "single shot", "one continuous take", "a oner", "no cuts", "unbroken",
55
- "one continuous camera movement". **Writing "N connected shots … no visible cut" is a
56
- contradiction** — no cut means one shot. That exact output is what this rule exists to
57
- stop; never emit it.
58
-
59
- **N shots requested → deliver exactly N.** Never round up to a nicer-sounding number,
60
- never add shots the user did not ask for, and never invent a maximum.
61
-
62
- **No count given → pick the SIMPLEST structure the idea needs.** One continuous take is
63
- very often right. A montage is a deliberate choice, never a default.
64
-
65
- ## Locked Intro (DEFAULT for any cinematic piece — single-shot and Elements included)
66
-
67
- After the Total lines, every prompt with recurring people, a recurring place, or more than one shot opens with the locked blocks. A single continuous take keeps ALL of them — only the `SHOT N` beats and `Multishot ON` are multi-shot-only. Skip entirely for: a bare POV/orb with no cast, 3×3 grid-panel mode, or video-edit tasks.
39
+ After the Total lines, every multi-shot promptand any piece with recurring people or a recurring place — opens with the locked blocks. Skip only for: true single-shot POV/orb, 3×3 grid-panel mode, or video-edit tasks.
68
40
 
69
41
  ```
70
42
  N connected cinematic shots, Xs total, AR, Multishot ON
@@ -17,7 +17,7 @@ Creative generations bill against the user's Kolbo credit balance. **Billing uni
17
17
  | **Music** | per generation (flat) | 15–60 cr | Suno v5 = 15 cr; ElevenLabs Music = 60 cr |
18
18
  | **Speech (TTS)** | per 100 characters | 2–5 cr/100 chars | ElevenLabs (5) × 500 chars = 25 cr |
19
19
  | **Sound effects** | per generation (flat) | 4–7 cr | |
20
- | **3D model** | per model (flat, × toggle multipliers on Meshy V7) | 5–300 cr | Trellis = 5 cr; Trellis 2 = 60 cr; Meshy v5/v6 = 150 cr; Meshy V7 = 186 cr base (no-texture ~124; +rigging ~217; +rigging+animation ~236); Marble 1.1 = 300 cr |
20
+ | **3D model** | per model (flat) | 5–300 cr | Trellis = 5 cr; Meshy v6 = 150 cr; Marble 1.1 = 300 cr |
21
21
  | **Transcription (stt)** | per minute of audio | `model.credit × duration_minutes` | |
22
22
 
23
23
  ## Calculation Formulas
@@ -130,18 +130,18 @@ visual_dna_ids: ["vdna_abc", // dana
130
130
 
131
131
  The match is **literal and case-insensitive**, so:
132
132
  - The `@name` must equal the stored `name` field (e.g. if `name: "esther_model"` → write `@esther_model`, not `@Esther`, not `@אסתר`, not `@the model`).
133
- - Any-language characters are supported — if the DNA was created with `name: "אסתר"` you write `@אסתר`. Use the EXACT stored string.
133
+ - Any-language characters are supported — if the DNA was created with `name: "אסתר"` you write `@אסתר`. Use the EXACT stored string. Asset tags are identifiers and are exempt from English-only prompt/dialogue rules: never translate, transliterate, lowercase, strip diacritics, replace spaces, or slugify an existing name. Resolve the selected id with `get_visual_dna` if the current name is unknown; never guess from the user's language or a production-log alias.
134
134
  - Mentions terminate at punctuation (`.,!?`), double-spaces, another `@`, or end of string. So `@maya, wearing...` matches `maya`.
135
135
 
136
136
  This composes with `@image1` / `@image2` positional tags for plain reference/source images — see "Reference Tagging" below.
137
137
 
138
138
  ### ⚠️ Naming rule for `create_visual_dna` — NO SPACES (MANDATORY)
139
139
 
140
- The `name` you set MUST be a **single token, lowercase, no spaces, ASCII-safe** `esther_model`, `dana`, `tokyo_neon`, `brand_red`. Never `Sarah Johnson`, never `the red dress`.
140
+ For a NEW DNA, prefer a short single token in the user's chosen language, such as `אסתר`, `ليلى`, `小雨`, or `esther_model`. ASCII and lowercase are not required. This naming recommendation never permits rewriting an EXISTING stored name or its prompt tag.
141
141
 
142
142
  Reason: the prompt parser stops the `@<token>` match at the first space (and at `.,!?` punctuation). So `@Sarah Johnson` matches *only* `Sarah` — if no DNA named `Sarah` exists, the mention is silently dropped and the DNA never binds. A single-token name is the only way to guarantee inline `@name` works in any sentence, in any language, without forcing the user to write awkward punctuation around it.
143
143
 
144
- Use underscores for multi-word concepts (`old_town`, not `Old Town`). When the user proposes a name with spaces, accept the intent but collapse it into a single token before storing (`"Sarah Johnson"` `sarah_johnson`) and tell them once how you'll refer to it. Source of truth: [kolbo-docs / Visual DNA & @ References](https://docs.kolbo.ai/kolbo-code/visual-dna).
144
+ For new multi-word names, suggest underscores in the same language. Do not silently translate or rename a user-specified name. For existing names, preserve the full stored string and selected id, including spaces; if binding fails, report the limitation instead of inventing an alias or renaming the user's DNA.
145
145
 
146
146
  ## Reference Tagging — `@image1` / `@video1` / `@Audio1`
147
147
 
package/src/tools/chat.js CHANGED
@@ -19,11 +19,12 @@ function registerChatTools(server, client) {
19
19
  system_prompt: z.string().optional().describe('System prompt for the conversation. Only applied when creating a new session.'),
20
20
  web_search: z.boolean().optional().describe('Enable web search for this message. Default: false'),
21
21
  deep_think: z.boolean().optional().describe('Enable deep think (extended reasoning). Default: false'),
22
+ thinking_level: z.string().optional().describe('Thinking effort ID from list_models type="text" thinkingLevels. The server uses the model catalog default when omitted or invalid. Separate from legacy deep_think; safeguards take precedence.'),
22
23
  enhance_prompt: z.boolean().optional().describe('Enhance the prompt. Default: false — only pass true if the user explicitly asks to enhance/improve the prompt.'),
23
24
  media_urls: z.array(z.string()).optional().describe('Public URLs of images, videos, or audio files to analyze. The model auto-routes to a vision-capable model when media is present. For a local file, get a URL first via the LOCAL FILE route in this tool\'s description.'),
24
25
  project_id: projectIdField
25
26
  },
26
- async ({ message, model, session_id, system_prompt, web_search, deep_think, enhance_prompt = false, media_urls, project_id }) => {
27
+ async ({ message, model, session_id, system_prompt, web_search, deep_think, thinking_level, enhance_prompt = false, media_urls, project_id }) => {
27
28
  // Every generate_* tool resolves its model this way; chat was the one
28
29
  // `model` arg that went straight to the API, which has no fuzzy matching.
29
30
  // So the display names list_models hands back ("Claude Fable 5") came
@@ -38,6 +39,7 @@ function registerChatTools(server, client) {
38
39
  system_prompt,
39
40
  web_search,
40
41
  deep_think,
42
+ ...(thinking_level !== undefined ? { thinking_level } : {}),
41
43
  enhance_prompt,
42
44
  media_urls,
43
45
  project_id
@@ -192,6 +192,9 @@ function registerModelTools(server, client, options = {}) {
192
192
  // model says no, and absence means the API doesn't expose the field.
193
193
  const formatSpecs = m => {
194
194
  const parts = [];
195
+ if (m.haveThinking && Array.isArray(m.thinkingLevels) && m.thinkingLevels.length) {
196
+ parts.push(`thinking_level: ${m.thinkingLevels.map(level => level.id).join('/')} (default ${m.thinkingDefault})`);
197
+ }
195
198
  const types = Array.isArray(m.types) ? m.types : [];
196
199
  const isVideoType = types.some(t =>
197
200
  ['text_to_video', 'img_to_video', 'video_to_video', 'elements',
@@ -1,51 +0,0 @@
1
- # 3D Generation (`generate_3d`)
2
-
3
- Three model families, discoverable via `list_models` with types `3d_text_to_model`,
4
- `3d_image_to_model`, `3d_multi_image_to_model` (plus `3d_world` for world/splat generation).
5
- Settings are **family-scoped** — params for a different family than the selected model are
6
- silently ignored, so match the params to the model you pass.
7
-
8
- ## Families & when to pick each
9
-
10
- | Family | Identifiers | Pick when | Base credits |
11
- |---|---|---|---|
12
- | **Meshy V7** | `fal-ai/meshy/v7/image-to-3d`, `fal-ai/meshy/v7/multi-image-to-3d` | Game-ready assets, characters (rigging!), highest fidelity, PBR | 186 |
13
- | Meshy v5/v6 | `meshy/v5/multi-image-to-3d`, `meshy/v6-preview/*` (incl. the only **text**-to-3D) | Text mode, or legacy compatibility | 150 |
14
- | Trellis v1 | `trellis-image-to-3d`, `trellis-multi-image-to-3d` | Fast + cheap drafts | 5 |
15
- | Trellis 2 | `trellis-2-image-to-3d` | Better quality than v1, 4K textures, polygon control | 60 |
16
-
17
- Multi-image mode: 2–4 images of the SAME object from different angles (Trellis: up to 6).
18
- More angles = better reconstruction. Output formats: GLB (always, in-app preview), FBX/OBJ/USDZ
19
- and textures via the full package.
20
-
21
- ## Meshy V7 settings (the full-control family)
22
-
23
- - `topology` `"triangle"|"quad"`, `target_polycount` 100–300000 (default 30000), `symmetry_mode` `"off"|"auto"|"on"`.
24
- - `should_remesh` (default true) — false keeps the raw reconstructed mesh, ignores topology/polycount.
25
- - `should_texture` (default true) — **false is cheaper** (~0.67× base). Disables `enable_pbr`,
26
- `texture_prompt`, `texture_image_url`.
27
- - `texture_prompt` (≤600 chars) and/or `texture_image_url` — guide texturing.
28
- - `enable_tpose` — output an A/T-pose character.
29
- - **Rigging**: `enable_rigging` auto-rigs a humanoid (best with clear limbs) + basic walk/run
30
- animations; `rigging_height_meters` (default 1.7). **~1.17× credit surcharge.**
31
- - **Animation**: `enable_animation` (requires `enable_rigging`) applies one preset from Meshy's
32
- ~697-action library via `animation_action_id` 0–696 (default 92 "Idle"; 0=Idle, 1=Walk, 14=Run,
33
- 4=Attack, 22=Dance, 290=Wave — full catalog at docs.meshy.ai/en/api/animation-library).
34
- **Additional ~1.09× surcharge on top of rigging.** Returns `animation_glb`/`animation_fbx` plus
35
- the rigged character files.
36
-
37
- Credit math is toggle-multiplied off the base price (server-computed; the exact quote comes back
38
- in the generation response) — e.g. base 186 → +rigging 217 → +animation 236; no-texture 124.
39
-
40
- ## Trellis settings
41
-
42
- - v1: `texture_size` `"512"|"1024"|"2048"`, `ss_guidance_strength`/`ss_sampling_steps`,
43
- `slat_guidance_strength`/`slat_sampling_steps` (more steps = higher quality, slower),
44
- `mesh_simplify`, `multiimage_algo` `"stochastic"|"multidiffusion"` (multi mode), `seed`.
45
- - Trellis 2: `resolution` `"512"|"1024"|"1536"`, `t2_texture_size` `"1024"|"2048"|"4096"`,
46
- `decimation_target` (polygons), `remesh` (default true), `tex_sampling_steps`, `seed`.
47
-
48
- ## Text mode (Meshy v6-preview only)
49
-
50
- `prompt` required; `art_style` `"realistic"|"sculpture"` (sculpture disables PBR),
51
- `enable_prompt_expansion` for AI prompt enrichment.