@kolbo/mcp 1.28.0 → 1.30.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,6 +4,8 @@ Use [Kolbo AI](https://kolbo.ai) as native tools in Claude Code and Claude Deskt
4
4
 
5
5
  Generate images, videos, music, speech, sound effects, multi-scene campaigns, and conversational chat — all from natural language in your coding environment. 100+ AI models behind Smart Select routing, with reusable Visual DNA profiles for character/style consistency.
6
6
 
7
+ **✨ Interactive widgets (v1.30+):** in claude.ai and Claude Desktop, generations render as live Kolbo cards — real-time progress with model + settings chips, an inline result gallery / video player, and one-click **Animate · Edit · Recreate · Download** actions. Library and model searches render as browsable grids with audio preview. Text-only clients (Claude Code, Cursor) keep the classic text responses.
8
+
7
9
  ## Set up — paste one prompt, or one config block (keyless, no API key)
8
10
 
9
11
  ### Easiest: paste this prompt to your AI
@@ -111,7 +113,7 @@ Just ask your agent naturally:
111
113
 
112
114
  Without the optional skill, the config block alone already exposes every tool — you just describe what you want. With the skill installed, each of these is also routed to the right MCP tool with the right defaults — UGC mode picks 9:16 + sound-off + no-captions, marketplace mode enforces compliance (pure white bg, no text, no props), product photoshoot mode uses the right aspect for the platform (2:3 Pinterest, 16:9 hero banner, 1:1 IG feed), etc. The routing logic is shared with [Kolbo Code](https://github.com/Zoharvan12/kolbo-code), so the behavior is identical however you connect.
113
115
 
114
- ## Available Tools (52)
116
+ ## Available Tools (86)
115
117
 
116
118
  **Generation**
117
119
  | Tool | Description |
@@ -211,6 +213,17 @@ Every generation tool also accepts an optional `project_id` arg that routes the
211
213
  | `analyze_script_for_stock` | AI: turn a script into b-roll search terms (`queries[]`, `mediaType`, `keywords`). |
212
214
  | `import_stock_asset` | Copy a stock asset into the media library (CDN copy, stable URL). Free. |
213
215
 
216
+ **Shorts Creator** (long video → viral vertical shorts, two-phase)
217
+ | Tool | Description |
218
+ |------|-------------|
219
+ | `shorts_analyze` | Phase 1: analyze a long video (Kolbo media-library URL, ≤30 min) → AI-picked best moments with titles, hooks, scores, accent beats. Flat 15 credits. Polls until moments are ready (~1-3 min). |
220
+ | `shorts_list_presets` | List restyle presets (identifier, name, preview video, default mode/subtitle style). |
221
+ | `shorts_get_transcript` | Word-level Scribe transcript of the source video (`words`, `language`, `sourceDuration`) — the base for the Review & Edit workflow (build `delete_ranges` cuts + edited `srt_content`). |
222
+ | `shorts_estimate` | Price a selection before rendering — free. Per-short credits + chunk counts. `delete_ranges` cuts shorten the effective duration (cheaper). |
223
+ | `shorts_render` | Phase 2: render up to 5 shorts (15-90s each) from picked moments — `accents` mode (restyle strongest beats, cheaper) or `full` (restyle everything), optional burned-in subtitles. Per short: optional `delete_ranges` (cut dead air, absolute source seconds, ≥8s must remain) and `srt_content` (user-edited SRT, cut-timeline times, ≤200KB). Polls until done (~5-20 min), returns final URLs. Failed shorts auto-refund. |
224
+ | `shorts_status` | One-shot job state read (moments / shorts / phase) — resume after a timeout. |
225
+ | `shorts_cancel` | Cancel a job and refund unused credits. |
226
+
214
227
  **Discovery & Account**
215
228
  | Tool | Description |
216
229
  |------|-------------|
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@kolbo/mcp",
3
- "version": "1.28.0",
3
+ "version": "1.30.0",
4
4
  "description": "Kolbo AI MCP Server - Generate images, videos, music, speech, and sound effects from Claude Code",
5
5
  "main": "src/index.js",
6
6
  "bin": {
@@ -41,6 +41,7 @@
41
41
  "README.md"
42
42
  ],
43
43
  "dependencies": {
44
+ "@modelcontextprotocol/ext-apps": "^1.7.4",
44
45
  "@modelcontextprotocol/sdk": "1.29.0",
45
46
  "form-data": "^4.0.5",
46
47
  "zod": "^3.25.0"
@@ -1,7 +1,7 @@
1
1
  # AUTO-GENERATED — do not edit
2
2
 
3
3
  This skill/ tree is mirrored from kolbo-code (the single source of truth)
4
- by .github/workflows/sync-skill-to-plugin.yml — synced from kolbo-code@2d43295.
4
+ by .github/workflows/sync-skill-to-plugin.yml — synced from kolbo-code@9cf6c35.
5
5
 
6
6
  It is the skill that 'npx @kolbo/mcp install' deploys into the user's agent.
7
7
  To change it, edit packages/opencode/skills/kolbo/ in kolbo-code and push;
package/skill/SKILL.md CHANGED
@@ -1,5 +1,5 @@
1
1
  ---
2
- version: 0.4.0
2
+ version: 0.5.0
3
3
  name: kolbo
4
4
  description: |
5
5
  Generate, edit, or analyze creative media via the Kolbo AI MCP server:
@@ -39,6 +39,25 @@ Once per conversation, before any other Kolbo tool call:
39
39
 
40
40
  If the user is on a whitelabel build (`sapir`, etc.), they must use their branded command — not `kolbo`. See `references/workflows/troubleshooting.md`.
41
41
 
42
+ ## 🎬 Confirm the Creative Brief BEFORE Generating (CRITICAL — read first)
43
+
44
+ Never fire a paid generation the moment the user says "make X". First **present the brief back as a confirmation the user can change** — this is the single most important interaction. It gives the user control over what gets created and what it costs, instead of silently spending credits on defaults.
45
+
46
+ **Before ANY paid image / video / music / speech / 3D generation**, unless the user has *explicitly* dictated every key parameter in this message, ask ONE labeled question (the UI renders it as an options card) confirming:
47
+
48
+ - **Model** — your recommended pick as the default option, plus 1–2 alternatives (with their credit cost).
49
+ - **Aspect ratio** — e.g. `1:1 / 9:16 / 16:9` (offer the sensible default first).
50
+ - **Count** — how many (1 / 4 / …).
51
+ - **Resolution / quality / duration** — where the model supports it.
52
+ - **Creative direction** — style / mood / scene, when the user was vague ("4 cats" → offer style options: photoreal / illustrated / cinematic / surprise-me).
53
+ - **Credit cost** — state the total (`✦ N credits`) right in the question so cost is never a surprise.
54
+
55
+ Then generate **only** with the confirmed parameters. If the user changes an option, use the change. This mirrors the approval-card flow: propose → let them adjust → confirm → generate.
56
+
57
+ **Only skip the brief confirmation when** the user's message already pins model + aspect + count + creative direction (e.g. "generate 4 photoreal tabby cats, 1:1, z-image/turbo") — then just state the cost one-liner and fire. A low credit cost is **not** a reason to skip: cheap ≠ no-confirmation. What matters is whether the user actually chose the parameters.
58
+
59
+ For multi-scene / batch work this pairs with `generate_creative_director` (see below) — still confirm the brief first.
60
+
42
61
  ## Routing Index — Read These Files on Demand
43
62
 
44
63
  | If the user wants to… | Read first |
@@ -49,7 +68,6 @@ If the user is on a whitelabel build (`sapir`, etc.), they must use their brande
49
68
  | Generate a **Veo 3 / 3.1** video | `references/models/veo.md` |
50
69
  | Build a **multi-scene set** (Creative Director, storyboard, campaign batch, 4+ angles) | `references/models/creative-director.md` |
51
70
  | Generate **music** (Suno, song, lyrics, jingle, score) | `references/models/music.md` |
52
- | Find an **existing / stock / library / royalty-free track** to score a video, ad, or voiceover | `references/workflows/music-library.md` |
53
71
  | Build an **HTML presentation / slide deck** | `references/models/html-presentation.md` |
54
72
  | Build a **landing page / marketing site** | `references/models/landing-page.md` |
55
73
  | Build a **dashboard / data viz / interactive widget / mini-game / UI mockup** | `references/models/visual-code.md` |
@@ -82,7 +100,7 @@ Each `references/models/*.md` mirrors the matching skill prompt in `kolbo-api/sr
82
100
  | `generate_video_from_video` | Restyle/transform an existing video. Keeps original motion. |
83
101
  | `generate_elements` | Reference-driven video. **Primary route for DNA → video.** |
84
102
  | `generate_first_last_frame` | Keyframe interpolation between two frames. |
85
- | `generate_lipsync` | Lipsync audio to an image or video face. Sync-3 adds multi-person speaker selection (`active_speaker_detection`), `emotion`, `model_mode`, `temperature`. |
103
+ | `generate_lipsync` | Lipsync audio to an image or video face. |
86
104
  | `generate_music` | Music generation (Suno + variants). |
87
105
  | `generate_speech` | TTS. Use `list_voices` to pick a voice. |
88
106
  | `generate_sound` | Sound effects. |
@@ -96,7 +114,6 @@ Each `references/models/*.md` mirrors the matching skill prompt in `kolbo-api/sr
96
114
  | `create_visual_dna` / `list_visual_dnas` / `get_visual_dna` / `delete_visual_dna` | Visual DNA — see `workflows/visual-dna.md` |
97
115
  | `list_moodboards` / `get_moodboard` / `list_presets` | Style overlays |
98
116
  | `chat_send_message` / `chat_list_conversations` / `chat_get_messages` | Kolbo chat with optional `media_urls` (up to 10 per call) |
99
- | `search_music_library` / `analyze_script_for_music` / `browse_music_library` / `get_music_library_facets` / `get_music_track_audio` / `get_music_track_related` / `get_music_track_lyrics` | **Stock / production music library** — find a licensed ready-made track (NOT `generate_music`, which composes a new song). See `workflows/music-library.md` |
100
117
  | `app_builder_*` (9 tools) | Full React app generation — see `workflows/app-builder.md` |
101
118
  | `publish_html_artifact` | Publish HTML / SVG / Mermaid to `sites.kolbo.ai`. Server dedupes by content hash. Strict CSP. |
102
119
 
@@ -145,8 +162,8 @@ Model types for `list_models`: `text_to_img`, `image_editing`, `text_to_video`,
145
162
 
146
163
  Full tables + formulas in `references/workflows/cost-and-validation.md`. Quick rules:
147
164
 
148
- - **Skip cost confirmation** when the user already specified model + count + duration, OR when a single generation costs < 5 credits.
149
- - **Required cost confirmation** otherwise: one-line summary, suggest cheaper alternative if available, wait for confirm.
165
+ - **Skip the brief/cost confirmation ONLY** when the user's message already pins model + count + aspect + creative direction (see "Confirm the Creative Brief" above). Low cost alone is **not** a reason to skip cheap generations still get the one labeled confirmation unless the user chose the parameters.
166
+ - **Otherwise confirm** via the labeled-question card: the parameters + the credit cost, suggest a cheaper alternative if one fits, wait for the user's pick. Never fire on defaults the user didn't choose.
150
167
  - **Batch totalling 100+ credits**: run `check_credits` first.
151
168
  - **Quote real cost**: after firing, log `credits_used` (from the tool result) to `.kolbo/production.md` — never `base × count`.
152
169
 
package/skill/VERSION CHANGED
@@ -1 +1 @@
1
- 0.4.0
1
+ 0.5.0
@@ -13,18 +13,18 @@ Load this file when the user wants a **Seedance 2 / Seedance 2.0** (ByteDance) v
13
13
 
14
14
  - **First line ALWAYS declares shot structure**: total duration, shot count, aspect ratio. Example: `Total: 15s / 6 shots / 16:9`. Put it at the BOTTOM of the prompt too.
15
15
  - **Order inside each shot**: Subject → Action → Camera → Style → Constraints → (Audio/SFX if relevant).
16
- - **Prompt length**: aim for ~120–280 words TOTAL across all shots combined (not per shot). Shorter than ~120 words = random output. Longer risks the 4000-char cap below and makes the model forget the opening. For 6-shot prompts, keep each shot 1–2 tight sentences.
16
+ - **Prompt length**: aim for ~120–280 words TOTAL across all shots combined (not per shot). Shorter than ~120 words = random output. Longer risks the 8000-char cap below and makes the model forget the opening. For 6-shot prompts, keep each shot 1–2 tight sentences.
17
17
  - **Character lock**: if a character recurs, open with `same character throughout all shots` to stop identity drift.
18
18
  - **Max 3 shots per single-shot prompt; max 6 shots in a multi-shot montage.** More causes drift.
19
19
  - **Always describe at least one camera movement per shot.**
20
20
  - **Tell Seedance what the camera is NOT doing** (e.g. `no cuts, no zoom, natural head movement`) — this is what locks POV.
21
21
  - **Final prompt is always English**, wrapped in a copy-ready code block. Detect intent in any language and reply in the user's language, but the prompt itself is English.
22
- - **HARD CAP: 4000 characters TOTAL for the ENTIRE prompt** — measured as one single string, including ALL shots, ALL boilerplate, ALL SFX lines, the opening style block, the closing `Total: …` line, every newline, every space, every punctuation mark. This is non-negotiable.
23
- - Applies to ANY prompt: 1 shot or 6 shots, single POV or full montage — the WHOLE thing must fit under 4000 chars combined.
24
- - It is NOT 4000 chars per shot. It is 4000 chars per prompt.
25
- - If your draft exceeds 4000 chars, trim aggressively in this order: (1) cut redundant adjectives, (2) collapse the opening cinematic boilerplate, (3) shorten SFX lists, (4) merge or drop shots — keep escalation beats and cut filler beats, (5) tighten action descriptions to verb-led essentials.
22
+ - **HARD CAP: 8000 characters TOTAL for the ENTIRE prompt** — measured as one single string, including ALL shots, ALL boilerplate, ALL SFX lines, the opening style block, the closing `Total: …` line, every newline, every space, every punctuation mark. This is non-negotiable.
23
+ - Applies to ANY prompt: 1 shot or 6 shots, single POV or full montage — the WHOLE thing must fit under 8000 chars combined.
24
+ - It is NOT 8000 chars per shot. It is 8000 chars per prompt.
25
+ - If your draft exceeds 8000 chars, trim aggressively in this order: (1) cut redundant adjectives, (2) collapse the opening cinematic boilerplate, (3) shorten SFX lists, (4) merge or drop shots — keep escalation beats and cut filler beats, (5) tighten action descriptions to verb-led essentials.
26
26
  - **Never** split into multiple prompts, multiple code blocks, or "part 1 / part 2" to evade the cap.
27
- - Before outputting, internally count the characters of the final prompt as a single string. If > 4000, rewrite tighter and re-count. Repeat until ≤ 4000. Only then show the user.
27
+ - Before outputting, internally count the characters of the final prompt as a single string. If > 8000, rewrite tighter and re-count. Repeat until ≤ 8000. Only then show the user.
28
28
 
29
29
  ## The 5 Formats
30
30
 
@@ -83,7 +83,7 @@ When the user uploads a 3×3 grid image and asks for Seedance prompts, switch to
83
83
  - Final prompt(s) ALWAYS in a fenced code block ready to paste into the Seedance `prompt` field (or pass as `prompt` on `generate_video` / `generate_elements`).
84
84
  - After the code block, give a 1-line "why this works" note (camera/escalation/physics choice).
85
85
  - If user asked in any language other than English, write your explanation in their language but keep the prompt itself English.
86
- - **Never exceed 4000 characters TOTAL** for the entire prompt as one string — that is the WHOLE prompt including every shot, every line of boilerplate, every SFX list, every newline. NOT 4000 per shot — 4000 for the prompt as one combined unit. Count before output. If over, rewrite tighter (cut adjectives, collapse boilerplate, merge or drop shots). NEVER split into multiple prompts / multiple code blocks / "part 1 / part 2" to work around the limit.
86
+ - **Never exceed 8000 characters TOTAL** for the entire prompt as one string — that is the WHOLE prompt including every shot, every line of boilerplate, every SFX list, every newline. NOT 8000 per shot — 8000 for the prompt as one combined unit. Count before output. If over, rewrite tighter (cut adjectives, collapse boilerplate, merge or drop shots). NEVER split into multiple prompts / multiple code blocks / "part 1 / part 2" to work around the limit.
87
87
 
88
88
  ## Seedance + Visual DNA / References
89
89
 
@@ -119,6 +119,68 @@ For ads that feature a specific product:
119
119
 
120
120
  If the user gives a **product URL** instead of a photo, see `workflows/research-first.md` — scrape, extract images, re-host via `upload_media`, persist as a brand kit at `.kolbo/brand-kits/<slug>.md`.
121
121
 
122
+ <!-- SKILL-ONLY: no server parity — UGC output routes through the creative-director / veo / seedance enhancers, so any server-side rule lives in those parity files, not here. -->
123
+
124
+ ## Multi-Slot Board Method (structured shot specs + character consistency)
125
+
126
+ For any multi-shot UGC / review / how-to where the SAME presenter must stay identical across shots, compose the prompt as explicit **slots** and lock identity with a **board-first** pass. This is a prompt-only convention — no special MCP mode; it uses `generate_image` (board) + `generate_elements` / `generate_video_from_image` (per-slot animate) that already exist.
127
+
128
+ ### 1. Structured input slots
129
+
130
+ Define each shot as one row. Fill every column before generating — blanks are where identity/quality drift creeps in.
131
+
132
+ | Slot | Arc role | POV / framing | Presenter action | Product visibility | Aspect | Audio |
133
+ |---|---|---|---|:-:|:-:|:-:|
134
+ | 1 | hook | selfie arm, chest-up, eye contact | states the problem / grabs attention | held up to camera | 9:16 | monologue seg 1 |
135
+ | 2 | demo | slightly wider, hands in frame | uses / demonstrates the product | in active use | 9:16 | monologue seg 2 |
136
+ | 3 | payoff | back to selfie framing | reaction + soft CTA | resting in hand / on surface | 9:16 | monologue seg 3 |
137
+
138
+ Scale to 2–6 slots. Keep `hook → demo → payoff` as the minimum arc; add `tension` / `proof` slots between demo and payoff for longer reviews.
139
+
140
+ ### 2. Rendering rules (hard invariants — apply to EVERY slot)
141
+
142
+ - One aspect ratio across all slots (UGC = `9:16`). Never mix.
143
+ - **No on-image text**, captions, subtitles, watermarks, or lower-thirds (users add captions in post).
144
+ - **Identity lock**: same presenter, same wardrobe, same lighting environment across all slots — open the prompt with `same character throughout all shots`.
145
+ - Hands and product must read cleanly — no deformed hands, no floating / clipping product, product logo legible when held.
146
+ - Phone-shot aesthetic (handheld sway, window/screen key) unless the mode is polished (`tv_spot`, `product_showcase`).
147
+
148
+ ### 3. Board-first consistency (the grid technique)
149
+
150
+ Before animating, generate ONE composite board image that locks the presenter's identity, then animate each panel:
151
+
152
+ 1. `generate_image` a labeled N-panel grid (2×2 or 1×N) of the presenter across the slot poses — front hook pose, hands-on-product demo pose, reaction pose — locked to `visual_dna_ids` (the presenter's Visual DNA). Aspect `16:9` for the board sheet.
153
+ 2. Treat that board image's CDN URL as the **`board_media_id`** — the single source of truth for identity.
154
+ 3. Animate each slot with `generate_video_from_image` / `generate_elements`, passing the board panel (and product) as `reference_images` and tagging `@image1`, so every clip inherits the same face/wardrobe.
155
+
156
+ This mirrors how the best UGC pipelines keep a character consistent: lock once as a board, then move each shot — not N independent generations that drift.
157
+
158
+ ### 4. Structured parameters (what to carry per generation)
159
+
160
+ Track these so each slot's call is reproducible and the arc stays coherent:
161
+
162
+ | Param | Meaning | Maps to |
163
+ |---|---|---|
164
+ | `arc_role` | hook / tension / demo / proof / payoff | prompt framing + shot order |
165
+ | `board_media_id` | the locked board image URL | `reference_images` (`@image1`) |
166
+ | `character_media_id` / `visual_dna_id` | presenter identity | `visual_dna_ids` (`@<dna-name>`) |
167
+ | `product_media_id` | product photo URL | `reference_images` (`@image2`) |
168
+ | `input_tier` | `draft` (fast preview) vs `hero` (final) | model + resolution choice |
169
+ | `monologue_segment` | the spoken line for this slot | prompt audio/dialogue line |
170
+ | `aspect_ratio` / `duration` / `sound_enabled` | per UGC Family Defaults above | MCP call args |
171
+
172
+ ### 5. Worked example (brief → slots → board → clips)
173
+
174
+ Brief: *"15s UGC review of a skincare serum, tech-savvy woman creator."*
175
+
176
+ 1. Ensure/create presenter Visual DNA (tech-savvy woman) → `visual_dna_id`.
177
+ 2. Board: `generate_image` a 3-panel `16:9` sheet — (a) chest-up hook holding the serum, (b) hands applying it, (c) thumbs-up reaction — `same character throughout all shots`, locked to the DNA. → `board_media_id`.
178
+ 3. Slots (each `9:16`, ~5s, sound OFF, animate from the matching board panel + product `@image2`):
179
+ - Slot 1 (hook): "Before this serum my routine was five products…" holding it to camera.
180
+ - Slot 2 (demo): hands applying, product in active use.
181
+ - Slot 3 (payoff): reaction + "…now it's just one step." soft CTA.
182
+ 4. Deliver as a structured message (setup + monologue + media), humanized refs (names/thumbnails, not raw IDs).
183
+
122
184
  ## UX Rules
123
185
 
124
186
  1. **Always pick a mode explicitly.** Don't auto-pick from one ambiguous word. If the user said "make me an ad" with no other signal, offer labeled options: `[UGC / TV Spot / Product Showcase / Surprise me]`.
@@ -0,0 +1,124 @@
1
+ 'use strict';
2
+
3
+ /**
4
+ * Minimal MCP Apps iframe bridge (io.modelcontextprotocol/ui, protocol 2026-01-26).
5
+ *
6
+ * Injected as an inline <script> into every ui://kolbo/* widget. Hand-rolled instead
7
+ * of shipping the 337KB @modelcontextprotocol/ext-apps browser bundle — implements the
8
+ * same JSON-RPC-over-postMessage handshake the official App class performs:
9
+ *
10
+ * widget → host request ui/initialize { appInfo, appCapabilities, protocolVersion }
11
+ * widget → host notification ui/notifications/initialized
12
+ * host → widget notification ui/notifications/tool-result | tool-input | host-context-changed
13
+ * widget → host request tools/call | ui/message | ui/open-link
14
+ * widget → host notification ui/notifications/size-changed
15
+ *
16
+ * Host-bound calls made before the handshake completes are queued (avoids the
17
+ * hidden-iframe race documented in claude-ai-mcp#61/#149).
18
+ *
19
+ * Exposed global: window.kolbo
20
+ * .ready(fn) — fn(hostContext) after handshake
21
+ * .onToolResult(fn) — fn(result) for ui/notifications/tool-result
22
+ * .onThemeChange(fn) — fn(hostContext) on host-context-changed
23
+ * .callTool(name, args) — Promise<CallToolResult>
24
+ * .sendMessage(text) — append a user chat message (returns Promise)
25
+ * .openLink(url) — open external URL
26
+ * .notifySize() — report content size to host
27
+ */
28
+
29
+ const BRIDGE_JS = `
30
+ (function () {
31
+ var nextId = 1;
32
+ var pending = {}; // id -> {resolve, reject}
33
+ var initialized = false;
34
+ var queue = []; // deferred host-bound sends until initialized
35
+ var hostContext = null;
36
+ var readyFns = [], toolResultFns = [], themeFns = [];
37
+
38
+ function post(msg) { window.parent.postMessage(msg, '*'); }
39
+
40
+ function request(method, params) {
41
+ return new Promise(function (resolve, reject) {
42
+ var id = nextId++;
43
+ pending[id] = { resolve: resolve, reject: reject };
44
+ var msg = { jsonrpc: '2.0', id: id, method: method, params: params || {} };
45
+ if (initialized || method === 'ui/initialize') post(msg);
46
+ else queue.push(msg);
47
+ });
48
+ }
49
+
50
+ function notify(method, params) {
51
+ var msg = { jsonrpc: '2.0', method: method, params: params || {} };
52
+ if (initialized || method === 'ui/notifications/initialized') post(msg);
53
+ else queue.push(msg);
54
+ }
55
+
56
+ window.addEventListener('message', function (ev) {
57
+ var m = ev.data;
58
+ if (!m || m.jsonrpc !== '2.0') return;
59
+ if (m.id != null && (m.result !== undefined || m.error !== undefined)) {
60
+ var p = pending[m.id];
61
+ if (!p) return;
62
+ delete pending[m.id];
63
+ if (m.error) p.reject(new Error(m.error.message || 'host error'));
64
+ else p.resolve(m.result);
65
+ return;
66
+ }
67
+ if (m.method === 'ui/notifications/tool-result') {
68
+ toolResultFns.forEach(function (f) { try { f(m.params || {}); } catch (e) {} });
69
+ } else if (m.method === 'ui/notifications/host-context-changed') {
70
+ hostContext = (m.params && m.params.hostContext) || m.params || hostContext;
71
+ themeFns.forEach(function (f) { try { f(hostContext); } catch (e) {} });
72
+ } else if (m.method === 'ui/resource-teardown' && m.id != null) {
73
+ post({ jsonrpc: '2.0', id: m.id, result: {} });
74
+ } else if (m.id != null) {
75
+ // Unknown host request — respond empty so the host isn't left hanging.
76
+ post({ jsonrpc: '2.0', id: m.id, result: {} });
77
+ }
78
+ });
79
+
80
+ request('ui/initialize', {
81
+ protocolVersion: '2026-01-26',
82
+ appInfo: { name: 'kolbo-widget', version: '1.0.0' },
83
+ appCapabilities: {}
84
+ }).then(function (res) {
85
+ hostContext = (res && res.hostContext) || null;
86
+ notify('ui/notifications/initialized');
87
+ initialized = true;
88
+ queue.forEach(post);
89
+ queue = [];
90
+ readyFns.forEach(function (f) { try { f(hostContext); } catch (e) {} });
91
+ }).catch(function () { /* host without apps support — widget stays static */ });
92
+
93
+ function notifySize() {
94
+ var el = document.documentElement;
95
+ notify('ui/notifications/size-changed', {
96
+ width: el.scrollWidth, height: el.scrollHeight
97
+ });
98
+ }
99
+
100
+ var sizeTimer = null;
101
+ new MutationObserver(function () {
102
+ clearTimeout(sizeTimer);
103
+ sizeTimer = setTimeout(notifySize, 120);
104
+ }).observe(document.documentElement, { childList: true, subtree: true, attributes: true });
105
+
106
+ window.kolbo = {
107
+ ready: function (f) { if (initialized) f(hostContext); else readyFns.push(f); },
108
+ onToolResult: function (f) { toolResultFns.push(f); },
109
+ onThemeChange: function (f) { themeFns.push(f); },
110
+ callTool: function (name, args) { return request('tools/call', { name: name, arguments: args || {} }); },
111
+ sendMessage: function (text) {
112
+ return request('ui/message', { role: 'user', content: [{ type: 'text', text: text }] });
113
+ },
114
+ openLink: function (url) { return request('ui/open-link', { url: url }); },
115
+ updateModelContext: function (text) {
116
+ return request('ui/update-model-context', { content: [{ type: 'text', text: text }] });
117
+ },
118
+ notifySize: notifySize,
119
+ hostContext: function () { return hostContext; }
120
+ };
121
+ })();
122
+ `;
123
+
124
+ module.exports = { BRIDGE_JS };
@@ -0,0 +1,73 @@
1
+ 'use strict';
2
+
3
+ const { BRIDGE_JS } = require('./bridge');
4
+ const { KOLBO_CSS, KOLBO_LOGO_SVG } = require('./theme');
5
+
6
+ /**
7
+ * Assemble a self-contained widget page. No build step — each widget module
8
+ * provides a body skeleton + its script; we wrap with theme, bridge, and the
9
+ * shared runtime helpers (theme sync, escaping, chips, formatting).
10
+ */
11
+ function widgetPage({ title, body, script }) {
12
+ return `<!DOCTYPE html>
13
+ <html lang="en">
14
+ <head>
15
+ <meta charset="UTF-8">
16
+ <meta name="viewport" content="width=device-width, initial-scale=1">
17
+ <title>${title}</title>
18
+ <link rel="preconnect" href="https://fonts.googleapis.com">
19
+ <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
20
+ <link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700&family=JetBrains+Mono:wght@400;500&display=swap" rel="stylesheet">
21
+ <style>${KOLBO_CSS}</style>
22
+ </head>
23
+ <body>
24
+ ${body}
25
+ <script>${BRIDGE_JS}</script>
26
+ <script>
27
+ // ---- shared widget runtime ----
28
+ var KOLBO_LOGO = ${JSON.stringify(KOLBO_LOGO_SVG)};
29
+ function esc(s) {
30
+ return String(s == null ? '' : s).replace(/[&<>"']/g, function (c) {
31
+ return { '&': '&amp;', '<': '&lt;', '>': '&gt;', '"': '&quot;', "'": '&#39;' }[c];
32
+ });
33
+ }
34
+ function el(id) { return document.getElementById(id); }
35
+ function fmtCredits(n) { return (n == null) ? '' : (Math.round(n * 100) / 100) + ' cr'; }
36
+ function fmtDur(s) { if (s == null) return ''; s = Math.round(s); return s >= 60 ? Math.floor(s/60) + 'm ' + (s%60) + 's' : s + 's'; }
37
+ function applyTheme(ctx) {
38
+ try {
39
+ var theme = ctx && (ctx.theme || (ctx.styles && ctx.styles.theme));
40
+ if (theme === 'light') document.documentElement.setAttribute('data-theme', 'light');
41
+ else if (theme === 'dark') document.documentElement.removeAttribute('data-theme');
42
+ } catch (e) {}
43
+ }
44
+ window.kolbo.ready(function (ctx) { applyTheme(ctx); window.kolbo.notifySize(); });
45
+ window.kolbo.onThemeChange(applyTheme);
46
+
47
+ // Model chip: real icon when the API provided one, brand monogram fallback.
48
+ function modelChipHTML(name, iconUrl) {
49
+ if (!name) return '';
50
+ var inner = iconUrl
51
+ ? '<img src="' + esc(iconUrl) + '" onerror="this.outerHTML=monogram(\\'' + esc(name).replace(/'/g, '') + '\\')" alt="">'
52
+ : monogram(name);
53
+ return '<span class="k-chip brand">' + inner + esc(name) + '</span>';
54
+ }
55
+ function monogram(name) {
56
+ return '<span class="k-mono-icon">' + esc(String(name).trim().charAt(0).toUpperCase()) + '</span>';
57
+ }
58
+ // Pull structuredContent out of a tools/call result (host bridge shape).
59
+ function structured(res) {
60
+ if (!res) return null;
61
+ if (res.structuredContent) return res.structuredContent;
62
+ try {
63
+ var t = (res.content || []).filter(function (c) { return c.type === 'text'; })[0];
64
+ return t ? JSON.parse(t.text) : null;
65
+ } catch (e) { return null; }
66
+ }
67
+ </script>
68
+ <script>${script}</script>
69
+ </body>
70
+ </html>`;
71
+ }
72
+
73
+ module.exports = { widgetPage };
@@ -0,0 +1,142 @@
1
+ 'use strict';
2
+
3
+ /**
4
+ * MCP Apps integration (io.modelcontextprotocol/ui) — Kolbo interactive widgets.
5
+ *
6
+ * Registers the ui://kolbo/* HTML resources and provides the helpers tool files
7
+ * use to attach widgets to results. Everything here is ADDITIVE: text-only hosts
8
+ * (Claude Code, Cursor, old clients) ignore `_meta` + `structuredContent` and see
9
+ * exactly the same text responses as before.
10
+ */
11
+
12
+ const {
13
+ registerAppResource,
14
+ getUiCapability,
15
+ RESOURCE_MIME_TYPE,
16
+ RESOURCE_URI_META_KEY,
17
+ } = require('@modelcontextprotocol/ext-apps/server');
18
+
19
+ const { generationWidgetHtml } = require('./widgets/generation');
20
+ const { mediaGridWidgetHtml } = require('./widgets/mediaGrid');
21
+ const { catalogWidgetHtml } = require('./widgets/catalog');
22
+ const { transcriptWidgetHtml } = require('./widgets/transcript');
23
+
24
+ const UI = {
25
+ generation: 'ui://kolbo/generation.html',
26
+ mediaGrid: 'ui://kolbo/media-grid.html',
27
+ catalog: 'ui://kolbo/catalog.html',
28
+ transcript: 'ui://kolbo/transcript.html',
29
+ };
30
+
31
+ const WIDGET_BUILDERS = {
32
+ [UI.generation]: generationWidgetHtml,
33
+ [UI.mediaGrid]: mediaGridWidgetHtml,
34
+ [UI.catalog]: catalogWidgetHtml,
35
+ [UI.transcript]: transcriptWidgetHtml,
36
+ };
37
+
38
+ // Widgets are pure functions of source — build once per process.
39
+ const htmlCache = new Map();
40
+ function widgetHtml(uri) {
41
+ if (!htmlCache.has(uri)) htmlCache.set(uri, WIDGET_BUILDERS[uri]());
42
+ return htmlCache.get(uri);
43
+ }
44
+
45
+ /** Register all Kolbo widget resources on an McpServer. */
46
+ function registerApps(server) {
47
+ for (const [uri, name] of [
48
+ [UI.generation, 'Kolbo Generation Widget'],
49
+ [UI.mediaGrid, 'Kolbo Library Widget'],
50
+ [UI.catalog, 'Kolbo Model Catalog Widget'],
51
+ [UI.transcript, 'Kolbo Transcription Widget'],
52
+ ]) {
53
+ registerAppResource(server, name, uri, { mimeType: RESOURCE_MIME_TYPE }, async () => ({
54
+ contents: [{ uri, mimeType: RESOURCE_MIME_TYPE, text: widgetHtml(uri) }],
55
+ }));
56
+ }
57
+ }
58
+
59
+ /** `_meta` for a tool RESULT (and optionally for tool registration). */
60
+ function uiMeta(uri) {
61
+ return { [RESOURCE_URI_META_KEY]: uri, ui: { resourceUri: uri } };
62
+ }
63
+
64
+ /**
65
+ * Should this server instance produce widget results?
66
+ * - `opts.apps === true` — set by the kolbo-api remote connector (claude.ai),
67
+ * where the stateless transport makes client capabilities unavailable per-call.
68
+ * - stdio hosts (Claude Desktop) — detected from the initialize handshake.
69
+ * - `KOLBO_MCP_APPS=1|0` env — manual override for local testing.
70
+ */
71
+ function appsEnabled(server, opts = {}) {
72
+ if (process.env.KOLBO_MCP_APPS === '0') return false;
73
+ if (opts.apps === true || process.env.KOLBO_MCP_APPS === '1') return true;
74
+ try {
75
+ const caps = server?.server?.getClientCapabilities?.();
76
+ return getUiCapability(caps) !== undefined;
77
+ } catch (_) {
78
+ return false;
79
+ }
80
+ }
81
+
82
+ /**
83
+ * Build a widget-carrying tool result. `text` stays the LLM-facing source of
84
+ * truth; `structured` goes to the widget only.
85
+ */
86
+ function uiResult(uri, text, structured) {
87
+ return {
88
+ content: [{ type: 'text', text }],
89
+ structuredContent: structured,
90
+ _meta: uiMeta(uri),
91
+ };
92
+ }
93
+
94
+ /* ------------------------------------------------------------------ */
95
+ /* Model icon lookup (name/identifier → absolute avatar URL) */
96
+ /* ------------------------------------------------------------------ */
97
+
98
+ const ICON_TTL_MS = 10 * 60 * 1000;
99
+ const iconCache = new Map(); // apiBase → { at, byKey: Map<lowername, url> }
100
+
101
+ async function modelIconMap(client) {
102
+ const cacheKey = client.apiBase || 'default';
103
+ const hit = iconCache.get(cacheKey);
104
+ if (hit && Date.now() - hit.at < ICON_TTL_MS) return hit.byKey;
105
+ const byKey = new Map();
106
+ try {
107
+ const res = await client.request('GET', '/v1/models');
108
+ const models = res?.models || res?.data?.models || [];
109
+ for (const m of models) {
110
+ if (!m || !m.avatar) continue;
111
+ // The API usually resolves avatars to absolute URLs; bare filenames (older
112
+ // deployments / internal calls) resolve against the app's public icon dir.
113
+ const url = /^https?:\/\//i.test(m.avatar)
114
+ ? m.avatar
115
+ : `https://app.kolbo.ai/models_icons/${encodeURIComponent(m.avatar)}`;
116
+ if (m.name) byKey.set(String(m.name).toLowerCase(), url);
117
+ if (m.identifier) byKey.set(String(m.identifier).toLowerCase(), url);
118
+ }
119
+ } catch (_) {
120
+ /* fail open — widgets fall back to monogram chips */
121
+ }
122
+ iconCache.set(cacheKey, { at: Date.now(), byKey });
123
+ return byKey;
124
+ }
125
+
126
+ /** Resolve one model's icon URL; null → widget renders a monogram. */
127
+ async function modelIcon(client, modelName) {
128
+ if (!modelName) return null;
129
+ const map = await modelIconMap(client);
130
+ return map.get(String(modelName).toLowerCase()) || null;
131
+ }
132
+
133
+ module.exports = {
134
+ UI,
135
+ registerApps,
136
+ uiMeta,
137
+ uiResult,
138
+ appsEnabled,
139
+ modelIcon,
140
+ modelIconMap,
141
+ widgetHtml, // exported for smoke tests
142
+ };