@kolbo/mcp 1.93.2 → 1.93.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@kolbo/mcp",
3
- "version": "1.93.2",
3
+ "version": "1.93.3",
4
4
  "description": "Kolbo AI MCP Server - Generate images, videos, music, speech, and sound effects from Claude Code",
5
5
  "main": "src/index.js",
6
6
  "bin": {
package/skill/SKILL.md CHANGED
@@ -221,11 +221,25 @@ A URL from `generate_*`, `list_media`, `get_media`, or a prior `upload_media` is
221
221
  - `upload_media` is only for a **local disk path** or an **external** (non-Kolbo) URL — `files`/`source_images`/`image_url` reject unknown hosts with `400`; a Kolbo URL passes through as-is.
222
222
  - Same rule after compaction: pull the URL from `.kolbo/production.md` and reuse it. Never download-then-reupload.
223
223
 
224
- ## ⚠️ Assets Before Shots (HARD RULE)
224
+ ## ⚠️ Route video per use case (decide — do not cargo-cult)
225
225
 
226
- For any film / ad / scene / episode / campaign the order is **Map → Create → Confirm → Shoot** (the directing guide — load `references/workflows/production-planning.md` + `filmmaking.md` before creating anything). Crack the concept first. Then every character, location, and prop becomes a Visual DNA **from a sheet** (`list_presets` search → `generate_image` with that `preset_id` → `create_visual_dna`). Do **not** register a DNA from a single portrait and skip the sheet. Publish the session plan (`Cast` / `Locations` / `Scene NN — slug`). Get a GATE lock on the asset set. **Only then** video. A shot against an unapproved cast is waste.
226
+ Pick the cheapest route that actually controls what the brief needs. Do not invent a pipeline.
227
227
 
228
- Scene dialogue is **never** `generate_speech` or `generate_lipsync`. Seedance 2 / 2.5 performs quoted lines written into the shot beat itself — English only. Full flow: `references/workflows/production-planning.md`.
228
+ **1. Recurring identity (cast / product / location must match across shots)**
229
+ Map → Visual DNA sheets → Confirm → `generate_elements` (or DNA-locked Multishot). Asset sheets earn their cost here.
230
+
231
+ **2. Composition must be locked before motion** (deliberate framing, Pixar-like kids beats, specific staging, hero product plate, user-approved look)
232
+ Generate the needed keyframe still(s) first, then animate with `generate_video_from_image` / first-last / Elements **using those images as real inputs**. Stills without attaching them to the video call are waste.
233
+
234
+ **3. Pure text-to-video / Multishot Locked Intro — only when keyframes are 100% unnecessary**
235
+ Use for generic b-roll, ambient motion, simple stock-like scenes, or any brief where the video model inventing composition is fine and stills would not improve control. If you are not sure keyframes add nothing, prefer route 2.
236
+
237
+ **Anti-patterns (HARD)**
238
+ - Do not generate N stills and then run a Multishot T2V that never attaches them.
239
+ - Do not default every "make a video" to keyframes (generic b-roll does not need them).
240
+ - Do not default every narration-only brief to T2V when the user asked for tightly designed cute/controlled shots — those often need keyframes.
241
+
242
+ Scene dialogue is **never** `generate_speech` or `generate_lipsync` on a Seedance shoot. Seedance 2 / 2.5 perform quoted lines written into the shot beat — English or Latin transliteration of Hebrew (`"shalom"`), never Hebrew script. For native Hebrew speech, prefer Gemini Omni Flash 1.1 or Gemini Omni 1. Full flow: `references/workflows/production-planning.md`.
229
243
 
230
244
  ## ⚠️ Load the matching skill BEFORE generating (HARD RULE)
231
245
 
@@ -452,7 +466,7 @@ If at this point you still don't know which `references/` file to load, default
452
466
  ## Media selection preferences
453
467
  Honor explicit models, presets, budget and inputs. Choose only eligible catalog candidates with all required capabilities. Use requested presets; otherwise use fitting presets when useful. For video generation, editing and lip-sync, when the user has not explicitly selected an output resolution, use the cheapest supported output resolution from the live catalog and pass it explicitly; do not inherit an expensive provider default. Preserve explicit user-selected resolution/settings. Finish fully, cinematic, professional, final, production and available credits are NOT permission to increase resolution. Never infer output resolution from reference media or export settings. A budget is a ceiling, not a spending target. Do not upscale or regenerate at a higher tier without explicit user authorization. If pricing or supported resolutions cannot be verified, inspect the catalog before dispatch; never invent a tier. Models with fixed output resolution use their native output. Never treat a policy refusal as a technical failure or route around safeguards.
454
468
  Default images and edits: GPT Image 2.5 Flare/Sunburst; medium for value, high for ordinary maximum quality. Reserve xhigh/max for exceptional dense or difficult multilingual text after medium/high prove insufficient; do not automatically spend on retries. Nano Banana 2 is secondary. Seedream 5.0 Pro favors cinematic aesthetics over complex instruction fidelity; Wan 2.7 Pro is another creative alternative. Z Image/P Image for cheap tests. Midjourney for artistic concepts only, never editing. Soul V2 for realistic people/UGC concepts; derive character sheets before registering finished Visual DNA. Mirage Film 2 for environments and cinematic inspiration.
455
- Default video: Seedance 2.5. Kling specializes in controlled single-image and first/last-frame shots. Wan 3.0 specializes in motion graphics and Hebrew/dialogue work. MiniMax H3 offers higher resolution; H3 Max favors speed at lower resolution. Gemini Omni Flash is a secondary Hebrew option (up to 10 seconds); Grok Imagine 1.5 and Seedance 2.0 are alternatives. P Video/Draft for cheap fast tests. Use base, edit or extend variants only with their required inputs.
469
+ Default video: Seedance 2.5 for general cinematic work (not Hebrew speech). Kling specializes in controlled single-image and first/last-frame shots. Wan 3.0 specializes in motion graphics and animated typography — native Hebrew speech is poor; attached-audio lip-sync works well. MiniMax H3 offers higher resolution and strong attached-audio lip-sync; H3 Max favors speed at lower resolution with the same audio lip-sync strength. **Native Hebrew dialogue:** Gemini Omni Flash 1.1 or Gemini Omni 1 (best). Seedance 2 / 2.5 do not speak Hebrew — use Latin transliteration in quotes on Seedance, or switch to Gemini Omni. Grok Imagine 1.5 and Seedance 2.0 are non-Hebrew alternatives. P Video/Draft for cheap fast tests. Use base, edit or extend variants only with their required inputs.
456
470
  Existing-video lip-sync: Sync 3 for active-speaker handling; PixVerse for cartoons/2D and economical faster work. Portrait lip-sync: Veed Fabric or HeyGen Avatar; P Avatar for budget work. LTX Audio to Video for camera/environment motion with audio-driven performance.
457
471
  Default music: Suno v6. ElevenLabs Music is an alternative, especially for duration-directed scoring. Both accept custom duration requests; validate the selected tool schema and inspect actual output duration.
458
472
 
@@ -13,7 +13,7 @@ Load this file when the user wants a **Seedance 2 / Seedance 2.0** (ByteDance) v
13
13
 
14
14
  ## Creative direction takes precedence
15
15
 
16
- The current user brief overrides template defaults and illustrative examples. Keep the two-layer organization, but include only relevant locks. State concrete camera trajectory and visible action prominently; optics numbers, equipment names, repetition and word counts are not guarantees of fidelity. Preserve a continuous-shot exception even when other scenes are multishot. For one shot use `Single continuous shot`, `Total: Xs / 1 shot / AR`, one SHOT heading and `multi_shots: false`; for multiple shots use `Multishot ON` and matching counts. AR comes from the brief, never a copied example. Keep dialogue in the user's requested language or phonetic spelling; test pronunciation rather than claiming guaranteed support or impossibility. Narration reserved for post does not belong in the generation prompt.
16
+ The current user brief overrides template defaults and illustrative examples. Keep the two-layer organization, but include only relevant locks. State concrete camera trajectory and visible action prominently; optics numbers, equipment names, repetition and word counts are not guarantees of fidelity. Preserve a continuous-shot exception even when other scenes are multishot. For one shot use `Single continuous shot`, `Total: Xs / 1 shot / AR`, one SHOT heading and `multi_shots: false`; for multiple shots use `Multishot ON` and matching counts. AR comes from the brief, never a copied example. Seedance does not speak Hebrew — use Latin transliteration in quotes or route native Hebrew to Gemini Omni. Narration reserved for post does not belong in the generation prompt.
17
17
 
18
18
  ## Universal Rules (apply to EVERY Seedance / Elements prompt)
19
19
 
@@ -142,7 +142,7 @@ Appearance locks WHO. Persona locks HOW THEY BEHAVE — without it Seedance rend
142
142
  ## Dialogue & expression
143
143
 
144
144
  - **Dialogue is PERFORMED by the model, never by a TTS tool.** Quoted lines in the prompt come back as synced speech with lip movement and room tone, together with the SFX you name in AUDIO. Scene dialogue therefore never routes through `generate_speech` or `generate_lipsync` — write the line in quotes inside its shot beat and let Seedance act it.
145
- - **Preserve requested dialogue and its language.** If the user requests Hebrew in Latin letters, preserve that phonetic text as dialogue, not an English translation. Native pronunciation and lip-sync require actual output inspection. Do not promise success or claim the language is impossible without current evidence. Offer a separately authorized dubbing pass only when needed; keep narration reserved for post out of the prompt.
145
+ - **Hebrew (HARD):** Seedance 2 / 2.5 do **not** speak Hebrew. Never put Hebrew-script dialogue in the prompt. Use Latin transliteration in quotes (`שלום` → `"shalom"`), or recommend Gemini Omni Flash 1.1 / Gemini Omni 1 for native Hebrew. Do not promise Seedance Hebrew success. Attached-audio lip-sync on 2.5 is unreliable unless audio length matches the clip and the prompt has no Hebrew script. Keep narration reserved for post out of the prompt.
146
146
  - `list_models` reports `sound_generation_type: "none"` for Seedance 2 / 2.5 because there is no in-app sound toggle (`sound_baked_in: true`). That field does NOT mean the model is silent. Do not read it as a reason to add TTS.
147
147
  - For silent tension, deliver it as expression, not speech: `He does not speak. His expression clearly says: "…"`.
148
148
 
@@ -11,7 +11,7 @@ Load this file when the user wants a **Seedance 2.5** video (they said "2.5" / "
11
11
 
12
12
  **Audio:** Seedance 2.5 emits real synced audio. `list_models` shows `sound_generation_type: none` only because there is no in-app toggle (`sound_baked_in: true`) — it does NOT mean the model is silent, and it is never a reason to reach for TTS. Quoted dialogue is PERFORMED (synced voices, lip movement, room tone) alongside the SFX named in AUDIO, so scene dialogue never goes through `generate_speech` or `generate_lipsync`; write the lines in quotes inside their shot beats.
13
13
 
14
- **Dialogue language follows the user.** Preserve requested Hebrew or Hebrew-in-Latin transliteration; do not translate it into English or change the spoken content. Pronunciation and lip-sync must be inspected in the generated output, not guaranteed from the prompt. Keep post-production VO out of the generation prompt. Asset tags always retain their exact stored spelling, including `@אביב` / `#ישראל` literally.
14
+ **Hebrew (HARD):** Seedance 2.5 does **not** speak Hebrew. Never put Hebrew-script dialogue in the prompt. Use Latin transliteration in quotes per speaker (`שלום` → `"shalom"`), or route native Hebrew speech to Gemini Omni Flash 1.1 / Gemini Omni 1. Attached-audio lip-sync sometimes works when audio length matches the clip exactly and the prompt has no Hebrew script; Kolbo accepts native audio uploads (no black-video workaround). Asset tags always retain their exact stored spelling, including `@אביב` / `#ישראל` literally. Keep post-production VO out of the generation prompt.
15
15
 
16
16
  **Use the cheapest supported tier unless the user selected an output resolution.** Resolution is a credit MULTIPLIER, not a flat rate. Relative to 720p: 480p ×0.44, 1080p ×2.25. A 30s pass costs ~540cr at 480p against ~1230cr at 720p and ~2770cr at 1080p. When no output resolution was selected and 480p is the cheapest supported tier, block the film at 480p, get the user's sign-off on staging, performance and timing, then re-run only the approved cut at a higher delivery resolution if the user explicitly authorizes that resolution increase. Approval of the creative cut alone does not authorize a more expensive resolution. If no output resolution was selected, use the cheapest supported tier from the live catalog even for final work; pass it explicitly.
17
17
 
@@ -9,10 +9,13 @@ starts here, **before** a single video credit is spent. Most users do not know
9
9
  this flow exists; they ask for a film and expect a film. Walk them through it
10
10
  rather than jumping to a prompt.
11
11
 
12
- Skip it only for a genuine one-off: a single clip, no recurring subject, nothing
13
- that has to match anything else.
12
+ **Skip asset mapping** only when keyframes and DNA are **100% unnecessary**: generic b-roll,
13
+ ambient motion, simple stock-like scenes where the video model inventing composition is fine.
14
+ For tightly designed kids/educational beats, Pixar-like staging, or any brief where composition
15
+ must be locked, generate keyframes (or DNA) first and attach them to the video call — do not
16
+ invent stills and then run a Multishot T2V that ignores them.
14
17
 
15
- ## The order is not negotiable
18
+ ## The order is not negotiable (when cast/product identity must lock)
16
19
 
17
20
  1. **Map** every element the script needs — including the **session plan** (names).
18
21
  2. **Create** each one as an asset (Visual DNA), grouped into the planned sessions.
@@ -617,7 +617,7 @@ function registerGenerateTools(server, client, options = {}) {
617
617
  // retired textToVideoGeneration path and was stale.
618
618
  server.tool(
619
619
  'generate_video',
620
- 'Generate a video from a text prompt using Kolbo AI. For SEVERAL different videos, pass all their prompts in `prompts` in ONE call (one combined widget) — never a series of separate calls. For animating an existing still image into motion, use generate_video_from_image instead. For a coordinated multi-scene video campaign, use generate_creative_director with workflow_type="video". Supports reference images (for style/composition guidance) and Visual DNA for character consistency. Seedance 2/2.5 PERFORM quoted dialogue natively (synced voice, lip movement, room tone) — do not route scene dialogue to generate_speech or generate_lipsync; write it in ENGLISH (other languages, Hebrew included, do not perform reliably). Resolution is a credit MULTIPLIER (vs 720p: 480p x0.44, 1080p x2.25, 4k x4.95), so draft at 480p and re-run only the approved cut at delivery resolution. ROUTE BEFORE CALLING: when reference images anchor IDENTITY (specific characters, a specific product, a location that must match) — especially 2+ of them — that is generate_elements, not this tool; reference_images here are loose style/composition hints. Decide the right tool FIRST: a mis-routed call still starts a PAID generation, and switching tools afterwards without cancel_generation leaves the user paying for both. Returns the final video URL when complete.',
620
+ 'Generate a video from a text prompt using Kolbo AI. For SEVERAL different videos, pass all their prompts in `prompts` in ONE call (one combined widget) — never a series of separate calls. For animating an existing still image into motion, use generate_video_from_image instead. For a coordinated multi-scene video campaign, use generate_creative_director with workflow_type="video". Supports reference images (for style/composition guidance) and Visual DNA for character consistency. Seedance 2/2.5 PERFORM quoted dialogue natively (synced voice, lip movement, room tone) — do not route scene dialogue to generate_speech or generate_lipsync; write it in ENGLISH or Latin transliteration of Hebrew ("shalom"), never Hebrew script — Seedance 2/2.5 do not speak Hebrew; prefer Gemini Omni Flash 1.1 or Gemini Omni 1 for native Hebrew. Resolution is a credit MULTIPLIER (vs 720p: 480p x0.44, 1080p x2.25, 4k x4.95), so draft at 480p and re-run only the approved cut at delivery resolution. ROUTE BEFORE CALLING: when reference images anchor IDENTITY (specific characters, a specific product, a location that must match) — especially 2+ of them — that is generate_elements, not this tool; reference_images here are loose style/composition hints. Decide the right tool FIRST: a mis-routed call still starts a PAID generation, and switching tools afterwards without cancel_generation leaves the user paying for both. Returns the final video URL when complete.',
621
621
  {
622
622
  prompt: z.string().optional().describe('Text description of the video to generate. Required unless `prompts` is provided.'),
623
623
  prompts: promptsField('videos'),
@@ -1427,7 +1427,7 @@ function registerGenerateTools(server, client, options = {}) {
1427
1427
  // ─── generate_elements ─────────────────────────────────────
1428
1428
  server.tool(
1429
1429
  'generate_elements',
1430
- 'Generate a video from reference elements (images, videos, and/or audio) + a text prompt. Use when the user wants to animate specific uploaded/referenced assets — e.g. "animate this product", "put these 3 characters into a scene". PRIMARY ROUTE FOR A DNA-ANCHORED MULTI-SHOT FILM: one call can carry the whole sequence — seedance-2-5 takes 4-30s, up to 30 shots and 20 Visual DNAs in a SINGLE generation (seedance-2: 4-15s, 9 DNAs) — instead of a stack of separate clips. DIALOGUE IS PERFORMED NATIVELY: quoted dialogue in the prompt comes back as synced voices with lip movement, room tone and the SFX named in the AUDIO block — never route scene dialogue to generate_speech or generate_lipsync. Write dialogue in ENGLISH; other languages (Hebrew included) do not perform reliably. COST: resolution is a multiplier. When list_models publishes `video_input_credit` and this call carries videos, charge that rate against nominal input seconds + nominal output seconds; MP4 padding within 0.15s of an integer snaps to that integer and larger fractions round up. Otherwise use the normal output-second rate. PROMPT CONTRACT (Seedance / Elements): Locked Intro only — Total line, then [GLOBAL LOOK] / [CAST] / [LOCATION] / SHOT N. Do NOT write SCENE CONTEXT / OPTICS / ACTION department packs. Every Visual DNA in visual_dna_ids MUST also appear in the prompt as @ExactDNAName (e.g. "@Zohar walks…") — never "Zohar\'s" or "the man on the left" as a substitute. IMPORTANT: different models accept different numbers and durations of inputs — call list_models type="elements" and read elements_max_images / elements_max_videos / elements_max_audio plus min_video_duration / max_video_duration before generating. For text-only → video use generate_video instead. For animating a single still image use generate_video_from_image. Returns the final video URL when complete.',
1430
+ 'Generate a video from reference elements (images, videos, and/or audio) + a text prompt. Use when the user wants to animate specific uploaded/referenced assets — e.g. "animate this product", "put these 3 characters into a scene". PRIMARY ROUTE FOR A DNA-ANCHORED MULTI-SHOT FILM: one call can carry the whole sequence — seedance-2-5 takes 4-30s, up to 30 shots and 20 Visual DNAs in a SINGLE generation (seedance-2: 4-15s, 9 DNAs) — instead of a stack of separate clips. DIALOGUE IS PERFORMED NATIVELY: quoted dialogue in the prompt comes back as synced voices with lip movement, room tone and the SFX named in the AUDIO block — never route scene dialogue to generate_speech or generate_lipsync. Write dialogue in ENGLISH or Latin transliteration of Hebrew ("shalom"), never Hebrew script — Seedance does not speak Hebrew; prefer Gemini Omni Flash 1.1 or Gemini Omni 1 for native Hebrew. COST: resolution is a multiplier. When list_models publishes `video_input_credit` and this call carries videos, charge that rate against nominal input seconds + nominal output seconds; MP4 padding within 0.15s of an integer snaps to that integer and larger fractions round up. Otherwise use the normal output-second rate. PROMPT CONTRACT (Seedance / Elements): Locked Intro only — Total line, then [GLOBAL LOOK] / [CAST] / [LOCATION] / SHOT N. Do NOT write SCENE CONTEXT / OPTICS / ACTION department packs. Every Visual DNA in visual_dna_ids MUST also appear in the prompt as @ExactDNAName (e.g. "@Zohar walks…") — never "Zohar\'s" or "the man on the left" as a substitute. IMPORTANT: different models accept different numbers and durations of inputs — call list_models type="elements" and read elements_max_images / elements_max_videos / elements_max_audio plus min_video_duration / max_video_duration before generating. For text-only → video use generate_video instead. For animating a single still image use generate_video_from_image. Returns the final video URL when complete.',
1431
1431
  {
1432
1432
  prompt: z.string().describe('Locked Intro prompt (Seedance/Elements): Total line, [GLOBAL LOOK], [CAST] with @ExactDNAName for every visual_dna_ids entry, [LOCATION], then SHOT N. Not SCENE CONTEXT/OPTICS/ACTION packs. Never substitute "the left man" or "Zohar\'s" for @Name. EVERY attached reference must also be tagged by its 1-based array position — `@Image 1`/`@Image 2` (reference_images), `@Video 1` (reference_videos), `@Audio 1` (reference_audio_urls) — and its job stated ("@Image 1 defines the character\'s face and wardrobe", "@Video 1 defines the camera move"). An untagged attachment is ignored by the engine even though it was uploaded and billed.'),
1433
1433
  model: z.string().optional().describe('Model identifier. If the user already named a family (Grok / Kling / Veo / Seedance / …), pass THAT family — never default to Seedance because Elements often uses it. Use list_models type="elements" for exact ids and elements_max_* caps. Do NOT omit (omitting = Smart Select).'),