@kolbo/mcp 1.95.0 → 1.96.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -356,3 +356,7 @@ Both are optional — the local install logs in via the browser on first use.
356
356
  | `update_video_editor_session` | Atomic rename, settings, tracks, clips, trims, speed, captions and effects |
357
357
 
358
358
  `create_video_editor_session` also accepts optional advanced `session_data` instead of clips/audio/texts. Read the schema and saved revision before editing. Changes affect saved state; reload an already-open editor before manual editing. Existing create/export names and arguments remain supported.
359
+
360
+ ### URL downloads
361
+
362
+ `download_media_from_url` starts a server download for a public media page URL. Use `get_download_status` with the returned `job_id` (not `get_generation_status`) and wait for `completed` before using `resultUrl`. `cancel_download` cancels an active job. This produces a cloud file and saves it to the account library when library sync succeeds; it does not write to the caller's computer. Video returns a verified MP4, audio returns MP3; maximum 500 MB. Existing `upload_media` still handles local/direct-file uploads. SDK routes reuse the authenticated utility pipeline at `POST /v1/downloads`, `GET /v1/downloads/:jobId`, and `DELETE /v1/downloads/:jobId`.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@kolbo/mcp",
3
- "version": "1.95.0",
3
+ "version": "1.96.2",
4
4
  "description": "Kolbo AI MCP Server - Generate images, videos, music, speech, and sound effects from Claude Code",
5
5
  "main": "src/index.js",
6
6
  "bin": {
@@ -1,6 +1,6 @@
1
1
  # AUTO-GENERATED — do not edit
2
2
 
3
- This tree is mirrored from kolbo-code@d6af561, the single source of truth.
3
+ This tree is mirrored from kolbo-code@8d24f6e, the single source of truth.
4
4
  Canonical source: packages/opencode/skills/kolbo/
5
5
  Distribution: .github/workflows/sync-skill-to-plugin.yml
6
6
 
package/skill/SKILL.md CHANGED
@@ -142,6 +142,7 @@ Font tools (when exposed by the installed MCP): `list_fonts`, `get_font`, `uploa
142
142
  | Tool | Purpose |
143
143
  |------|---------|
144
144
  | `list_models` / `list_voices` / `check_credits` / `show_plans` / `get_generation_status` / `cancel_generation` / `get_session_usage` | Discovery + status. `list_models` with no args returns the recommended shortlist out of ~428 — pass `type` for a full category with per-model caps. `cancel_generation` stops an in-flight job and refunds what it can: use it when the user changes their mind mid-generation instead of letting it run. `show_plans` renders the balance + upgrade card for pricing/plan/upgrade questions. |
145
+ | `download_media_from_url` / `get_download_status` / `cancel_download` | Download a public social/video page URL to Kolbo storage. Async job; use download status, not generation status. See `workflows/media-library.md`. |
145
146
  | `upload_media` / `create_upload_ticket` / `list_media` / `get_media` / `get_media_stats` / `favorite_media` / `unfavorite_media` / `delete_media` / `restore_media` / `permanently_delete_media` / `move_media` / `bulk_*_media` / `*_media_folder` | Media library — see `workflows/media-library.md`. Getting a LOCAL file in depends on where the server runs: `upload_media` with a path only works on a local (stdio) install; over a remote connector use `create_upload_ticket` and POST the file yourself. |
146
147
  | `create_visual_dna` / `update_visual_dna` / `generate_character_sheet` / `list_visual_dnas` / `get_visual_dna` / `delete_visual_dna` / `*_visual_dna_folder` (5 folder tools) | Visual DNA (+ character sheet, character folders) — see `workflows/visual-dna.md`. Edit with `update_visual_dna`; never delete+recreate. |
147
148
  | `list_moodboards` / `get_moodboard` / `list_presets` / `list_cinematic_presets` | Style overlays + presets. `list_presets` spans FOUR distinct catalogs (`image`, `image_edit`, `video`, `music`; `text_to_video` is an alias for `video`, `shorts` is empty) — the `video` one holds 200+ Seedance shot recipes. `list_cinematic_presets` is a separate tool feeding the `cinematic` arg, never `preset_id`. Full doctrine + intent→catalog map: `references/workflows/presets.md`. Never omit `preset_id` after claiming a preset was used. |
@@ -124,3 +124,12 @@ The `*_music_library` tools (`search_music_library` / `browse_music_library` / `
124
124
  - `import_music_track_to_library` charges the same way AND also copies the clean track into the media library.
125
125
  - `analyze_script_for_music` turns a script into search terms for `search_music_library`.
126
126
  - Use this family when the user needs music cleared for commercial use. When free stock will do, use `search_stock_media` with `mediaType: "music"` instead.
127
+
128
+
129
+ ## Download a public video or audio page
130
+
131
+ For a YouTube, Instagram, TikTok, Facebook, X or other supported public media page, call `download_media_from_url`. Use `output_type: "video"` (default) for MP4 or `"audio"` for MP3; choose an optional resolution through `quality`. This is a download, not AI generation. Maximum output size is 500 MB.
132
+
133
+ Keep the returned `job_id`. Check `get_download_status` at reasonable intervals (at least two seconds); never use `get_generation_status` for these jobs or start another download while the first is pending. Only `completed` means the returned `resultUrl` is ready. Use `cancel_download` if the user cancels. Do not repeatedly retry unavailable/private media.
134
+
135
+ The result is hosted by Kolbo, not saved to the user's local folder automatically. If a local file was requested and your host supports filesystem downloads, save the completed result URL there. Never claim local delivery from a cloud URL alone. For an existing local file or direct file upload, keep using `upload_media` instead.
package/src/apps/theme.js CHANGED
@@ -85,7 +85,7 @@ body {
85
85
  .k-prompt { color: var(--text-muted); font-size: 12.5px; margin-bottom: 4px; word-break: break-word;
86
86
  display: -webkit-box; -webkit-line-clamp: 2; -webkit-box-orient: vertical; overflow: hidden;
87
87
  user-select: text; -webkit-user-select: text; cursor: text; }
88
- .k-prompt.expanded { -webkit-line-clamp: unset; }
88
+ .k-prompt.expanded { -webkit-line-clamp: unset; white-space: pre-wrap; }
89
89
  .k-text-tools { display: flex; gap: 2px; justify-content: flex-end; margin: 0 0 10px; }
90
90
  .k-text-btn {
91
91
  display: inline-flex; align-items: center; gap: 4px;
@@ -135,6 +135,8 @@ function makeExpandable(node, raw) {
135
135
  // Synchronous layout read — rAF would never fire in a hidden/backgrounded
136
136
  // iframe, leaving long prompts stuck without the expand affordance.
137
137
  var overflow = node.scrollHeight > node.clientHeight + 2 || node.scrollWidth > node.clientWidth + 2;
138
+ // Hidden host iframes report zero dimensions. Long prompts still need an expand button.
139
+ if (node.id === 'prompt' && text.length > 120) overflow = true;
138
140
  if (overflow) node.classList.add('k-clamped');
139
141
  var tools = document.createElement('div');
140
142
  tools.className = 'k-text-tools';
@@ -282,6 +284,15 @@ function displayKind(sc) {
282
284
  return kind || 'image';
283
285
  }
284
286
 
287
+ function resolutionLabel(sc) {
288
+ var s = sc.settings || {};
289
+ var resolution = String(s.resolution || '');
290
+ var model = String(sc.model || '');
291
+ var draft = /-draft$/i.test(resolution) || /-draft$/i.test(model) || s.is_draft === true || sc.is_draft === true;
292
+ var pixels = resolution.replace(/-draft$/i, '');
293
+ return draft ? (pixels ? pixels + ' Draft' : 'Draft') : resolution;
294
+ }
295
+
285
296
  function renderChips(sc) {
286
297
  var h = modelChipHTML(modelLabel(sc), sc.model_icon);
287
298
  var s = sc.settings || {};
@@ -293,7 +304,8 @@ function renderChips(sc) {
293
304
  var shotLabel = s.shots > 1 ? (s.shots + ' shots') : (s.multi_shot ? 'multishot' : '');
294
305
  if (s.duration) h += chip(ICONS.clock + ' ' + fmtDur(s.duration) + (shotLabel ? ' · ' + shotLabel : ''));
295
306
  else if (shotLabel) h += chip(shotLabel);
296
- if (s.resolution) h += chip(esc(s.resolution));
307
+ var resolutionText = resolutionLabel(sc);
308
+ if (resolutionText) h += chip(esc(resolutionText));
297
309
  if (s.aspect_ratio) h += chip(esc(s.aspect_ratio));
298
310
  if (s.quality) h += chip(esc(s.quality) + ' quality');
299
311
  if (s.enhance_prompt) h += chip(ICONS.sparkle + ' enhanced');
@@ -1343,7 +1355,7 @@ function openPromptRow(placeholder, onSend) {
1343
1355
  var PRE_IMAGE_KEYS = ['source_images', 'reference_images', 'image_url', 'mask_image_url',
1344
1356
  'additional_images', 'first_frame', 'last_frame', 'seed_reference_image_url',
1345
1357
  'elements', 'files', 'keyframes', 'source'];
1346
- var PRE_VIDEO_KEYS = ['source_video', 'reference_videos'];
1358
+ var PRE_VIDEO_KEYS = ['source_video', 'video_url', 'reference_videos'];
1347
1359
  var PRE_AUDIO_KEYS = ['audio', 'audio_url', 'reference_audio_urls', 'seed_reference_audio_urls'];
1348
1360
 
1349
1361
  // The card mounts the moment the tool is CALLED, so the only thing it knows is
@@ -1375,6 +1387,7 @@ function preRefSc(toolName, a) {
1375
1387
  PRE_AUDIO_KEYS.forEach(function (k) { take(a[k], aud); });
1376
1388
  return {
1377
1389
  tool: toolName,
1390
+ is_draft: /-draft$/i.test(String(a.model || '')),
1378
1391
  kind: kindFromTool(toolName, null),
1379
1392
  count: a.num_images || (Array.isArray(a.prompts) ? a.prompts.length : 1),
1380
1393
  reference_images: img,
@@ -9,6 +9,7 @@
9
9
  */
10
10
 
11
11
  const READ_ONLY = [
12
+ 'get_download_status',
12
13
  'get_video_editor_schema', 'list_video_editor_sessions', 'get_video_editor_session',
13
14
  'list_fonts', 'get_font', 'get_font_upload_status',
14
15
  'get_creative_director_status', 'get_generation_status', 'list_models',
@@ -70,6 +71,7 @@ const PRIVATE_WRITE = [
70
71
  ];
71
72
 
72
73
  const DESTRUCTIVE_WRITE = [
74
+ 'cancel_download',
73
75
  'extend_music', 'cover_music',
74
76
  'update_video_editor_session',
75
77
  'delete_font',
@@ -102,6 +104,7 @@ const DESTRUCTIVE_WRITE = [
102
104
  ];
103
105
 
104
106
  const OPEN_WORLD_WRITE = [
107
+ 'download_media_from_url',
105
108
  'import_music_audio',
106
109
  'publish_html_artifact', 'create_review_share_link', 'blender_capture_viewport',
107
110
  // Adobe edits add bins, clips, sequences or caption tracks; none delete or overwrite.
@@ -617,7 +617,7 @@ function registerGenerateTools(server, client, options = {}) {
617
617
  // retired textToVideoGeneration path and was stale.
618
618
  server.tool(
619
619
  'generate_video',
620
- 'Generate a video from a text prompt using Kolbo AI. For SEVERAL different videos, pass all their prompts in `prompts` in ONE call (one combined widget) — never a series of separate calls. For animating an existing still image into motion, use generate_video_from_image instead. For a coordinated multi-scene video campaign, use generate_creative_director with workflow_type="video". Supports reference images (for style/composition guidance) and Visual DNA for character consistency. Seedance 2/2.5 PERFORM quoted dialogue natively (synced voice, lip movement, room tone) — do not route scene dialogue to generate_speech or generate_lipsync; write it in ENGLISH or Latin transliteration of Hebrew ("shalom"), never Hebrew script — Seedance 2/2.5 do not speak Hebrew; prefer Gemini Omni Flash 1.1 or Gemini Omni 1 for native Hebrew. Resolution is a credit MULTIPLIER (vs 720p: 480p x0.44, 1080p x2.25, 4k x4.95), so draft at 480p and re-run only the approved cut at delivery resolution. ROUTE BEFORE CALLING: when reference images anchor IDENTITY (specific characters, a specific product, a location that must match) — especially 2+ of them — that is generate_elements, not this tool; reference_images here are loose style/composition hints. Decide the right tool FIRST: a mis-routed call still starts a PAID generation, and switching tools afterwards without cancel_generation leaves the user paying for both. Returns the final video URL when complete.',
620
+ 'Generate a video from a text prompt using Kolbo AI. For SEVERAL different videos, pass all their prompts in `prompts` in ONE call (one combined widget) — never a series of separate calls. For animating an existing still image into motion, use generate_video_from_image instead. For a coordinated multi-scene video campaign, use generate_creative_director with workflow_type="video". Supports reference images (for style/composition guidance) and Visual DNA for character consistency. Seedance 2/2.5 PERFORM quoted dialogue natively (synced voice, lip movement, room tone) — do not route scene dialogue to generate_speech or generate_lipsync; write it in ENGLISH or Latin transliteration of Hebrew ("shalom"), never Hebrew script — Seedance 2/2.5 do not speak Hebrew; prefer Gemini Omni Flash 1.1 or Gemini Omni 1 for native Hebrew. Resolution is a credit MULTIPLIER (vs 720p: 480p x0.44, 1080p x2.25, 4k x4.95), so use a lower resolution for previews. For actual Seedance 2.5 Draft, explicitly pass resolution="480p-draft" with model="seedance-2-5", then use edit_video draft_quote/draft_enhance for paid finalization. Plain "480p" is a regular generation, not Draft. ROUTE BEFORE CALLING: when reference images anchor IDENTITY (specific characters, a specific product, a location that must match) — especially 2+ of them — that is generate_elements, not this tool; reference_images here are loose style/composition hints. Decide the right tool FIRST: a mis-routed call still starts a PAID generation, and switching tools afterwards without cancel_generation leaves the user paying for both. Returns the final video URL when complete.',
621
621
  {
622
622
  prompt: z.string().optional().describe('Text description of the video to generate. Required unless `prompts` is provided.'),
623
623
  prompts: promptsField('videos'),
@@ -1442,7 +1442,7 @@ function registerGenerateTools(server, client, options = {}) {
1442
1442
  preset_id: z.string().optional().describe('Preset ID from list_presets type="video" (optional)'),
1443
1443
  enhance_prompt: z.boolean().optional().describe('Enhance the prompt. Default: false — only pass true if the user explicitly asks to enhance/improve the prompt.'),
1444
1444
  visual_dna_ids: z.array(z.string()).optional().describe('Array of Visual DNA profile IDs to apply for character/style consistency across outputs. **Cap: pass at most `max_visual_dna` IDs from list_models for the chosen model.**'),
1445
- resolution: z.string().optional().describe('Video resolution tier (vertical pixels): "720p" / "1080p" / "1440p" / "2160p". Model-dependent — call list_models and read supported_resolutions.'),
1445
+ resolution: z.string().optional().describe('Video resolution or named tier: read supported_resolutions from list_models. For Seedance 2.5 Draft explicitly use model="seedance-2-5", resolution="480p-draft". Plain "480p" generates regular video, NOT Draft. Finalize an actual draft with edit_video operation="draft_quote" then "draft_enhance".'),
1446
1446
  sound_enabled: z.boolean().optional().describe('Enable (`true`) or disable (`false`) AI-generated synced audio on the output video. Honored by `sound_generation_type: "native"` models (Kling O3/V3, Veo 3.1, PixVerse V6). Omit to use `sound_enabled_by_default`. Enabling sound may apply `sound_credit_multiplier` to cost.'),
1447
1447
  keyframes: z.array(z.object({
1448
1448
  image_url: z.string().describe('Public URL of the keyframe image'),
@@ -1639,7 +1639,7 @@ function registerGenerateTools(server, client, options = {}) {
1639
1639
  aspect_ratio: z.string().optional().describe(aspectRatioDescribe('16:9') + ' Auto-detected from the first frame if omitted.'),
1640
1640
  enhance_prompt: z.boolean().optional().describe('Enhance the prompt. Default: false — only pass true if the user explicitly asks to enhance/improve the prompt.'),
1641
1641
  visual_dna_ids: z.array(z.string()).optional().describe('Array of Visual DNA profile IDs to apply. **Cap: pass at most `max_visual_dna` IDs from list_models for the chosen model; if `supports_visual_dna: false`, DNA is silently ignored.**'),
1642
- resolution: z.string().optional().describe('Video resolution tier (vertical pixels): "720p" / "1080p" / "1440p" / "2160p". Model-dependent — call list_models and read supported_resolutions.'),
1642
+ resolution: z.string().optional().describe('Video resolution or named tier: read supported_resolutions from list_models. For Seedance 2.5 Draft explicitly use model="seedance-2-5", resolution="480p-draft". Plain "480p" generates regular video, NOT Draft. Finalize an actual draft with edit_video operation="draft_quote" then "draft_enhance".'),
1643
1643
  sound_enabled: z.boolean().optional().describe('Enable (`true`) or disable (`false`) AI-generated synced audio on the output video. Honored by `sound_generation_type: "native"` models (Kling O3/V3, Veo 3.1, PixVerse V6). Omit to use `sound_enabled_by_default`. Enabling sound may apply `sound_credit_multiplier` to cost.'),
1644
1644
  project_id: projectIdField,
1645
1645
  session_id: sessionIdField
@@ -1850,7 +1850,7 @@ function registerGenerateTools(server, client, options = {}) {
1850
1850
  duration: z.number().optional().describe('Output duration in seconds. Must be in `supported_durations` from list_models, OR within `min_output_duration`-`max_output_duration`. Default: matches source'),
1851
1851
  enhance_prompt: z.boolean().optional().describe('Enhance the prompt. Default: false — only pass true if the user explicitly asks to enhance/improve the prompt.'),
1852
1852
  visual_dna_ids: z.array(z.string()).optional().describe('Array of Visual DNA profile IDs to apply for character/style consistency. **Cap: pass at most `max_visual_dna` IDs from list_models for the chosen model; if `supports_visual_dna: false`, DNA is silently ignored.**'),
1853
- resolution: z.string().optional().describe('Video resolution tier (vertical pixels): "720p" / "1080p" / "1440p" / "2160p". Model-dependent — call list_models and read supported_resolutions.'),
1853
+ resolution: z.string().optional().describe('Video resolution or named tier: read supported_resolutions from list_models. For Seedance 2.5 Draft explicitly use model="seedance-2-5", resolution="480p-draft". Plain "480p" generates regular video, NOT Draft. Finalize an actual draft with edit_video operation="draft_quote" then "draft_enhance".'),
1854
1854
  reference_images: z.array(z.string()).optional().describe('Array of reference image URLs for models that support additional image inputs. **Cap: pass at most `max_images` URLs from list_models — if `max_images === 0` the model does not accept image refs.** Examples: character reference images for Kling O1/O3, style reference for Aleph/gen4_aleph, character image for WAN VACE video-edit.'),
1855
1855
  reference_videos: z.array(z.string()).optional().describe('Array of additional reference video URLs for models that support multiple video inputs. **Cap: pass at most `max_videos` URLs from list_models — if `max_videos <= 1` only the source_video is accepted.** Example: WAN 2.6 reference-to-video accepts 1–3 reference videos.'),
1856
1856
  elements: z.array(z.string()).optional().describe('Array of element image URLs. **Cap: pass at most `max_elements` URLs from list_models — if `max_elements === 0` the model does not accept elements.** Elements are style or character reference assets alongside the main video.'),
@@ -64,6 +64,23 @@ function uploadTicketPayload(ticket) {
64
64
  }
65
65
 
66
66
  function registerMediaTools(server, client, options = {}) {
67
+
68
+ server.tool('download_media_from_url',
69
+ 'Download a public YouTube, Instagram, TikTok, Facebook, X or other supported media page through Kolbo servers. Use this for social page URLs; upload_media is for existing files or direct file URLs. Returns an asynchronous download job, NOT a finished file. Call get_download_status with job_id until terminal; do not use get_generation_status or resubmit while pending. Completed video output is one MP4 with its expected audio, audio output is MP3, maximum 500 MB. Returns a cloud URL, not a local filesystem path. Private/login-restricted or unavailable media may fail.',
70
+ { url: z.string().url(), quality: z.enum(['best','2160','1440','1080','720','480','360']).optional(), output_type: z.enum(['video','audio']).optional() },
71
+ async ({ url, quality = 'best', output_type = 'video' }) => {
72
+ const result = await client.post('/v1/downloads', { url, quality, outputType: output_type });
73
+ return { content: [{ type: 'text', text: JSON.stringify({ ...result, job_id: result.jobId, next_tool: 'get_download_status' }) }] };
74
+ });
75
+ server.tool('get_download_status',
76
+ 'Read a URL-download job owned by this account. Use the job_id returned by download_media_from_url. pending/processing means wait before checking again; completed includes resultUrl/url; failed/cancelled is terminal. Do not report a file ready until completed. This is separate from AI generation status.',
77
+ { job_id: z.string().min(1).max(100) },
78
+ async ({ job_id }) => ({ content: [{ type: 'text', text: JSON.stringify(await client.get(`/v1/downloads/${encodeURIComponent(job_id)}`)) }] }));
79
+ server.tool('cancel_download',
80
+ "Cancel this account's pending or processing URL-download job. Does not delete a previously completed file.",
81
+ { job_id: z.string().min(1).max(100) },
82
+ async ({ job_id }) => ({ content: [{ type: 'text', text: JSON.stringify(await client.delete(`/v1/downloads/${encodeURIComponent(job_id)}`)) }] }));
83
+
67
84
  // `opts.apps` is set only by kolbo-api's per-request server (see createServer
68
85
  // in ../index.js), which makes it a TRANSPORT signal — deliberately not
69
86
  // `appsEnabled()`, which also returns true for stdio hosts that advertise UI.