@kolbo/mcp 1.21.0 → 1.22.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,90 @@
1
+ <!-- PARITY: this file mirrors getSeedancePromptSystemPrompt() in
2
+ kolbo-api/src/config/systemPrompt.js (lines ~775–855).
3
+ When that function changes, update this file in the same session.
4
+ See packages/opencode/CLAUDE.md "MCP & Skill Sync Rule". -->
5
+
6
+ # Seedance 2 — Prompt Rules
7
+
8
+ Load this file when the user wants a **Seedance 2 / Seedance 2.0** (ByteDance) video. For any other video model, see `models/veo.md`, `models/prompt-copilot.md`, or generic video rules in `SKILL.md`.
9
+
10
+ **Kolbo MCP routing:** Seedance is a video model — call `generate_video` (text-to-video) or `generate_elements` (when video references / Visual DNA / first-last frames are involved). Run `list_models({ type: "text_to_video" })` and pick a Seedance variant by name.
11
+
12
+ ## Universal Rules (apply to EVERY Seedance prompt)
13
+
14
+ - **First line ALWAYS declares shot structure**: total duration, shot count, aspect ratio. Example: `Total: 15s / 6 shots / 16:9`. Put it at the BOTTOM of the prompt too.
15
+ - **Order inside each shot**: Subject → Action → Camera → Style → Constraints → (Audio/SFX if relevant).
16
+ - **Prompt length**: aim for ~120–280 words TOTAL across all shots combined (not per shot). Shorter than ~120 words = random output. Longer risks the 4000-char cap below and makes the model forget the opening. For 6-shot prompts, keep each shot 1–2 tight sentences.
17
+ - **Character lock**: if a character recurs, open with `same character throughout all shots` to stop identity drift.
18
+ - **Max 3 shots per single-shot prompt; max 6 shots in a multi-shot montage.** More causes drift.
19
+ - **Always describe at least one camera movement per shot.**
20
+ - **Tell Seedance what the camera is NOT doing** (e.g. `no cuts, no zoom, natural head movement`) — this is what locks POV.
21
+ - **Final prompt is always English**, wrapped in a copy-ready code block. Detect intent in any language and reply in the user's language, but the prompt itself is English.
22
+ - **HARD CAP: 4000 characters TOTAL for the ENTIRE prompt** — measured as one single string, including ALL shots, ALL boilerplate, ALL SFX lines, the opening style block, the closing `Total: …` line, every newline, every space, every punctuation mark. This is non-negotiable.
23
+ - Applies to ANY prompt: 1 shot or 6 shots, single POV or full montage — the WHOLE thing must fit under 4000 chars combined.
24
+ - It is NOT 4000 chars per shot. It is 4000 chars per prompt.
25
+ - If your draft exceeds 4000 chars, trim aggressively in this order: (1) cut redundant adjectives, (2) collapse the opening cinematic boilerplate, (3) shorten SFX lists, (4) merge or drop shots — keep escalation beats and cut filler beats, (5) tighten action descriptions to verb-led essentials.
26
+ - **Never** split into multiple prompts, multiple code blocks, or "part 1 / part 2" to evade the cap.
27
+ - Before outputting, internally count the characters of the final prompt as a single string. If > 4000, rewrite tighter and re-count. Repeat until ≤ 4000. Only then show the user.
28
+
29
+ ## The 5 Formats
30
+
31
+ ### 1. Transformations (highest-performing format)
32
+ - Numbered shots, beat by beat.
33
+ - Escalation arc: **calm → threat → transformation → aftermath**.
34
+ - 6 shots / 15s / 16:9 is the proven structure.
35
+ - Opening boilerplate: `Montage, multi-shot action Hollywood movie, don't use one camera angle or single cut, cinematic lighting, photorealistic, 35mm film, professional color grading, sharp focus, high detail texture, film grain, depth of field mastery, ARRI ALEXA aesthetic`.
36
+ - **Realism trick**: for monsters/creatures, append `no 3D, no cartoon, no VFX` to force ultra-realism.
37
+ - **Comedy trick**: append `add a visual gag in the background` and Seedance invents one.
38
+
39
+ ### 2. Orbs (single continuous POV with powers)
40
+ - **One shot only**, first-person, 15 seconds, hands always visible in frame.
41
+ - Boilerplate: `Single continuous shot, first-person POV perspective, the camera IS her eyes, hyper-chaotic handheld motion, completely unstabilized, violent raw human movement, constant micro-jitters, aggressive head swings, abrupt jerks, frequent over-rotation and harsh correction, moments of near motion blur loss, no smoothness at all, no stabilization, wide-angle lens (strong distortion), subtle chromatic aberration near frame edges, her hands always visible in frame, no music only raw SFX, cinematic lighting, photorealistic, grounded realism, strong 35mm film look, heavy film grain, sharp but imperfect focus, noticeable focus breathing, motion blur on fast actions, halation on highlights, soft highlight rolloff, slightly desaturated tones, ARRI ALEXA aesthetic, practical VFX feel, minimal CGI look, natural imperfections`.
42
+ - **Inline VFX syntax**: describe powers with bracketed VFX tags inside the action, e.g. `[VFX: branching electric circuits pulsing with white-blue current, sparks jumping between fingers]`.
43
+ - **Always include a slow-motion ramp + snap-back**: `RAMPS TO SLOW MOTION as ... — SNAPS BACK ...`.
44
+ - **End with an explicit SFX list line** (electric crackle, energy burst, slow-mo hum stretch, snap impact, etc).
45
+
46
+ ### 3. POVs (locked first-person, no powers)
47
+ - One continuous shot, POV perspective. Always state what the camera is NOT doing: `no cuts, no zoom, natural head movement`.
48
+ - Describe ambient environment density (other actors, dust, sunlight, debris).
49
+ - Short prompts can hit hard — don't pad if the concept is tight.
50
+
51
+ ### 4. Fights
52
+ - Always supply: **clear location, clear power mismatch, defined escalation arc**.
53
+ - Describe choreography beat by beat — Seedance executes what you write.
54
+ - Single continuous shot 15s works for two-fighter scenes; describe camera moves between beats (`crests rooftop edge`, `full 360 orbit`, `pulls back to wide`, `descends with them`).
55
+ - Use `Guy Ritchie speed-ramping with Snyder impact slow-motion` as the style anchor when comedic/stylized.
56
+
57
+ ### 5. Animation (3D stylized)
58
+ - Break the 15s into **timed segments** (`0–3s`, `3–6s`, `6–9s`, `9–12s`, `12–15s`) and describe each explicitly.
59
+ - Reference the input image as `@image is the first keyframe and style reference.`
60
+ - Style anchor: `Cinematic stylized 3D animation, photorealistic <env>, stylized characters`.
61
+ - Describe physics as precisely as character actions (particle simulation, volumetric dust, sand displacement, energy VFX).
62
+
63
+ ## Grid Storyboard Mode (3×3 grid input)
64
+
65
+ When the user uploads a 3×3 grid image and asks for Seedance prompts, switch to this mode:
66
+
67
+ 1. **Analyze all 9 panels.** Summarize what you see in each row (2–3 sentences per row).
68
+ 2. **Confirm parameters if missing** (one short clarifying question max):
69
+ - Duration per video (default: 10s)
70
+ - Output type: `9 separate full-screen videos` (default) OR `single animated grid video`
71
+ - Motion intensity (default: 70–80)
72
+ - Style (slow-mo, dramatic, epic, realistic physics, etc.)
73
+ 3. **Default behavior: 9 separate full-screen 16:9 prompts**, each panel expanded to full frame. Never animate the whole grid unless explicitly asked.
74
+ 4. **Each prompt must include** camera, lighting, physics, emotion, particle effects, character consistency (lock the recurring subject in line 1).
75
+ 5. **Never invent actions not present in the source panel.**
76
+ 6. **Output format**:
77
+ - First: short panel-by-panel analysis (row 1 / row 2 / row 3).
78
+ - Then: a clean JSON object with 9 prompts keyed `panel_1` … `panel_9`.
79
+ - Finally: 1–2 sentences on motion strategy + improvement suggestions.
80
+
81
+ ## Output Discipline
82
+
83
+ - Final prompt(s) ALWAYS in a fenced code block ready to paste into the Seedance `prompt` field (or pass as `prompt` on `generate_video` / `generate_elements`).
84
+ - After the code block, give a 1-line "why this works" note (camera/escalation/physics choice).
85
+ - If user asked in any language other than English, write your explanation in their language but keep the prompt itself English.
86
+ - **Never exceed 4000 characters TOTAL** for the entire prompt as one string — that is the WHOLE prompt including every shot, every line of boilerplate, every SFX list, every newline. NOT 4000 per shot — 4000 for the prompt as one combined unit. Count before output. If over, rewrite tighter (cut adjectives, collapse boilerplate, merge or drop shots). NEVER split into multiple prompts / multiple code blocks / "part 1 / part 2" to work around the limit.
87
+
88
+ ## Seedance + Visual DNA / References
89
+
90
+ When a character must stay consistent, pair Seedance with Visual DNA via `generate_elements` (NOT `generate_video` — text-to-video silently drops `visual_dna_ids`). Tag the DNA inside the prompt with `@<dna-name>` — see `workflows/visual-dna.md`. For grid/storyboard inputs, the source frame is `@image1`.
@@ -0,0 +1,110 @@
1
+ <!-- PARITY: this file mirrors getVeoPromptSystemPrompt() in
2
+ kolbo-api/src/config/systemPrompt.js (lines ~1156–1256).
3
+ When that function changes, update this file in the same session. -->
4
+
5
+ # Veo 3 / 3.1 — Prompt Rules
6
+
7
+ Load this file when the user wants a **Veo 3 / Veo 3.1** (Google) video. For other video models see `models/seedance.md`, `models/prompt-copilot.md`, or generic video rules in `SKILL.md`.
8
+
9
+ **Kolbo MCP routing:**
10
+ - Text-to-video → `generate_video` with `model: "veo-3.1"` (or via `list_models({ type: "text_to_video" })`).
11
+ - Image-to-video → `generate_video_from_image`.
12
+ - First-and-last frame → `generate_first_last_frame`.
13
+ - Ingredients-to-video (multi-reference) → `generate_elements` with `reference_images` and/or `visual_dna_ids`.
14
+
15
+ ## CRITICAL Kolbo Platform Rules
16
+
17
+ - **Aspect ratio, resolution, and clip length are MCP-tool params** (`aspect_ratio`, `resolution`, `duration`). **NEVER include "16:9", "9:16", "720p", "1080p", "4 seconds", "8s", or any duration / aspect / resolution string inside the prompt body.**
18
+ - Pass `sound_enabled: true/false` as a separate param when the user mentions audio — see SKILL.md "Sound on/off".
19
+ - Don't write Python / Vertex AI / API call syntax. The user is generating through Kolbo's MCP tools.
20
+
21
+ ## Model Capabilities (informs recommendations, never in the prompt body)
22
+ - Resolution: 720p or 1080p (`resolution` param)
23
+ - Aspect: 16:9 or 9:16 (`aspect_ratio` param)
24
+ - Clip length: 4s, 6s, or 8s (`duration` param)
25
+ - Synchronous audio: dialogue, SFX, ambient, music — all guided by prompt text. Veo 3.1 has `sound_generation_type: "native"` and `sound_enabled_by_default: true` — if the user said "no sound", you MUST pass `sound_enabled: false`.
26
+ - Image-to-video, first-and-last frame, ingredients-to-video (up to multiple reference images)
27
+ - Add/remove object (uses Veo 2 under the hood; no audio for that mode)
28
+ - All output watermarked with SynthID
29
+
30
+ ## The Veo Prompt Formula (use for EVERY prompt)
31
+
32
+ `[Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]`
33
+
34
+ - **Cinematography** — camera work and shot composition (the most powerful tone-control lever)
35
+ - **Subject** — main character or focal point
36
+ - **Action** — what the subject is doing (strong verbs)
37
+ - **Context** — environment, background, time of day
38
+ - **Style & Ambiance** — overall aesthetic, mood, lighting, film stock
39
+
40
+ Example shape: `Medium shot, a tired corporate worker, rubbing his temples in exhaustion, in front of a bulky 1980s computer in a cluttered office late at night. The scene is lit by harsh fluorescent overhead lights and the green glow of the monochrome monitor. Retro aesthetic, shot as if on 1980s color film, slightly grainy.`
41
+
42
+ ## The Language of Cinematography (Veo's strongest lever)
43
+
44
+ - **Camera movement**: `dolly shot`, `tracking shot`, `crane shot`, `aerial view`, `slow pan`, `POV shot`, `arc shot`, `whip pan`, `handheld`, `static`. Always name at least one.
45
+ - **Composition**: `wide shot`, `close-up`, `extreme close-up`, `low angle`, `high angle`, `two-shot`, `over-the-shoulder`.
46
+ - **Lens & focus**: `shallow depth of field`, `wide-angle lens`, `soft focus`, `macro lens`, `deep focus`, `anamorphic 2.39:1`.
47
+
48
+ ## Directing the Soundstage (Veo 3.1 strength)
49
+
50
+ Veo bakes audio directly from prompt instructions. Use these conventions:
51
+ - **Dialogue**: put speech in **quotation marks** with speaker attribution.
52
+ `A woman says, "We have to leave now."`
53
+ `The detective replies in a weary voice, "Of all the offices in this town, you had to walk into mine."`
54
+ - **Sound effects**: prefix with `SFX:`. Example: `SFX: thunder cracks in the distance, rain hits the window`.
55
+ - **Ambient noise**: prefix with `Ambient noise:` or `Ambient:`. Example: `Ambient noise: the quiet hum of a starship bridge`.
56
+ - **Music**: describe inline. Example: `A swelling orchestral score begins to play.`
57
+
58
+ ## Negative / Exclusion Prompts (Veo prefers positive framing)
59
+
60
+ - Describe what you WANT, not what you don't want.
61
+ - ❌ "no buildings, no roads"
62
+ - ✅ "a desolate, untouched landscape with bare earth and scrub grass"
63
+
64
+ ## Advanced Workflows
65
+
66
+ ### 1. First-and-Last-Frame Transition (`generate_first_last_frame`)
67
+ The user provides two images (`first_frame_url` + `last_frame_url`). The prompt describes ONLY the transition between them.
68
+ - Describe the **camera move** that bridges the two frames (`smooth 180-degree arc`, `slow dolly through`, `whip pan reveal`, `time-lapse fade`).
69
+ - Include any audio (dialogue / SFX / score) that plays during the transition.
70
+ - Don't re-describe either frame — Veo can see them.
71
+ Example: `The camera performs a smooth 180-degree arc shot, starting with the front-facing view of the singer and circling around her to seamlessly end on the POV shot from behind her. She sings, "When you look me in the eyes, I can see a million stars."`
72
+
73
+ ### 2. Ingredients-to-Video (`generate_elements`, multi-reference consistency)
74
+ The user provides reference images for characters / objects / setting via `reference_images` (and/or `visual_dna_ids`). The prompt references each one and describes the scene.
75
+ - Open with: `Using @image1 for the <character A>, @image2 for the <character B>, and @image3 for the <setting>, create...` — see `workflows/visual-dna.md` for tag rules.
76
+ - Then describe shot type + action + dialogue + audio.
77
+ - Great for dialogue scenes, multi-character shots, character-locked sequences.
78
+
79
+ ### 3. Timestamp Prompting (multi-shot single generation)
80
+ Direct a multi-shot sequence with precise pacing inside one prompt by tagging each segment with a time range.
81
+ Format:
82
+ `[00:00-00:02] <shot 1 — cinematography + subject + action + audio>`
83
+ `[00:02-00:04] <shot 2 — ...>`
84
+ `[00:04-00:06] <shot 3 — ...>`
85
+ - Use for 4s / 6s / 8s clips, sized to whatever `duration` param is set to.
86
+ - Each segment should change at least one of: angle, framing, subject, or location.
87
+ - Add `SFX:`, dialogue in quotes, and emotion cues inside each segment.
88
+
89
+ ### 4. Image-to-Video (`generate_video_from_image`)
90
+ Veo can animate a source image with strong prompt adherence.
91
+ - The model can see the image — describe **what happens**, not what's already there.
92
+ - Always name a camera move + at least one audio element.
93
+ - Concise. Action-led.
94
+
95
+ ## Negative Prompts (when you must specify exclusions)
96
+
97
+ If a tool exposes a separate negative-prompt field, write a short positive description of what to AVOID — e.g. `cartoon, animated, low resolution, watermark, text overlay`. Most of the time, positive prompting is better.
98
+
99
+ ## Output Discipline
100
+
101
+ - Pass the prompt as the `prompt` field on the chosen tool.
102
+ - **NEVER** include aspect ratio, resolution, or duration inside the prompt body.
103
+ - When summarizing the call to the user, state separately:
104
+ - **Aspect:** 16:9 or 9:16 — one-line why
105
+ - **Resolution:** 720p or 1080p — one-line why (1080p for hero shots, 720p for drafts / cost-sensitive)
106
+ - **Duration:** 4s / 6s / 8s — one-line why (match it to the action density)
107
+ - **Sound:** `sound_enabled: true/false` — explicit if the user mentioned audio
108
+ - **Workflow:** text-to-video / image-to-video / first-and-last-frame / ingredients-to-video / timestamp — which Kolbo MCP tool you'll call
109
+ - **Why this works:** 1 line on the key cinematography / audio choice
110
+ - If the user asks in any language other than English, write explanations in their language but keep the prompt itself English (Veo handles English best for cinematography vocab; dialogue inside quotes can be in any language).
@@ -0,0 +1,80 @@
1
+ <!-- PARITY: this file mirrors getVisualCodeSystemPrompt() + HTML_ARTIFACT_BOILERPLATE
2
+ in kolbo-api/src/config/systemPrompt.js (lines ~1625–1683).
3
+ When that function changes, update this file in the same session. -->
4
+
5
+ # Visual Code — Interactive HTML Artifact Rules
6
+
7
+ Load this file when the user wants to **build an interactive HTML artifact where the visual rendered result matters as much as the logic** — dashboards, data visualizations, interactive widgets, animated components, mini-games, UI mockups, charts, tools, demos.
8
+
9
+ If the user asks for a **presentation** → see `models/html-presentation.md`. If they ask for a **landing page** → see `models/landing-page.md`. Everything else visual-and-interactive is here.
10
+
11
+ **Kolbo Code routing:** write the artifact as a single HTML block in your reply. Kolbo Code's panel renders it as a previewable artifact card. Call `publish_html_artifact({ title, content })` to publish to `sites.kolbo.ai` after approval.
12
+
13
+ ## What This Skill Is For
14
+
15
+ - **Dashboards** — KPI cards, tables, filterable views, charts (Chart.js / D3).
16
+ - **Data visualizations** — bar / line / pie / scatter, network graphs, heatmaps, geo maps.
17
+ - **Interactive widgets** — calculators, configurators, color pickers, gradient generators, font playgrounds, regex testers.
18
+ - **Mini-games** — snake, tetris, breakout, memory match, typing trainer, anything that fits in <1000 lines of vanilla JS or Canvas API.
19
+ - **Animated components** — splash screens, hero animations, scroll-driven effects, loading states, transition demos.
20
+ - **UI mockups** — settings pages, onboarding flows, chat UIs, e-commerce product pages — fully interactive even if data is mocked.
21
+ - **Tools** — JSON formatter, base64 encoder, color contrast checker, lorem ipsum generator (the irony noted).
22
+
23
+ ## Picking the Tech Stack
24
+
25
+ - **Vanilla HTML + CSS + JS + Tailwind** is the default. Reach for it first.
26
+ - **Chart.js** for standard charts (bar, line, pie, doughnut, radar). Easy and good-looking.
27
+ - **D3.js** for custom / complex visualizations (network graphs, force layouts, custom interactions).
28
+ - **Three.js** for 3D scenes, WebGL, generative art.
29
+ - **Canvas API** for mini-games, particle systems, animations not suited to DOM.
30
+ - **GSAP** for serious animation timelines / scroll-triggered sequences.
31
+ - **Framer Motion** for animations on a React app.
32
+ - **React 18 + Babel standalone** for genuinely component-driven apps (state-heavy UIs). Don't reach for React for static widgets.
33
+ - **Lucide icons** via CDN for any iconography. Stop using emoji where icons fit better.
34
+
35
+ ## Architecture Patterns
36
+
37
+ - For widgets with state: keep state in one object `const state = { ... }` and a single `render()` function that reads from it. Mutate state, call render. Easy to reason about, fast to iterate.
38
+ - For data viz: separate `prepareData()` from `renderChart()`. Don't tangle the two.
39
+ - For games: classic game loop — `requestAnimationFrame(tick)` → update → render. Keep entity objects in arrays.
40
+ - For React apps: use hooks (`useState`, `useEffect`, `useMemo`). Don't pull in Redux for a toy app.
41
+
42
+ ## Quality Bar
43
+
44
+ - **Real data when the user provides it.** Don't paraphrase numbers — render them verbatim.
45
+ - **Empty / loading / error states** all handled.
46
+ - **Keyboard accessibility** for anything interactive. Tab order makes sense, focus rings visible, Enter / Space activate buttons.
47
+ - **Hover and active states** on every interactive element. Cursor: pointer where appropriate.
48
+ - **Mobile-responsive** unless it's fundamentally desktop-only (complex dashboard) — in which case say so in the lead-in.
49
+ - **Animations under 400ms** for micro-interactions, custom easing not linear. Include `@media (prefers-reduced-motion: reduce)`.
50
+ - **Don't ship broken JS.** Mentally verify every `addEventListener`, every `querySelector` matches a real element.
51
+
52
+ ## Anti-AI-Slop (same principles as the landing-page skill, applied lightly)
53
+
54
+ - ❌ NEVER use `Inter` / `Roboto` / `Arial` / system fonts as default. Pick distinctive Google Fonts or Fontshare.
55
+ - ❌ NEVER default to purple-violet gradient on white.
56
+ - ❌ NEVER default to `Space Grotesk` everywhere — pick something else most of the time.
57
+ - Pick a deliberate palette tied to the artifact's mood, not Tailwind defaults.
58
+ - For dashboards: use a single dominant brand color + neutral grays + one accent for emphasis. Avoid the "rainbow chart with 8 colors" look — limit each chart to 1–3 colors.
59
+ - Hover / focus states on every interactive element. Cursor: pointer where appropriate.
60
+
61
+ ## RTL / Multilingual
62
+
63
+ - Set `<html lang dir>` correctly when content is in an RTL language.
64
+ - For mixed-language UIs (e.g. RTL text inside an LTR dashboard), use `dir="auto"` or explicit `dir` per element.
65
+
66
+ ## Output Discipline — HTML Artifact (NON-NEGOTIABLE)
67
+
68
+ - Reply MUST contain exactly ONE ` ```html ... ``` ` fenced code block with a COMPLETE, self-contained HTML document.
69
+ - Document must start with `<!DOCTYPE html>` and include `<html>`, `<head>` (with `<meta charset="UTF-8">` + `<meta name="viewport" content="width=device-width, initial-scale=1">`), and `<body>`.
70
+ - Embed ALL CSS inside `<style>` and ALL JavaScript inside `<script>`. No external CSS files, no relative asset paths. CDN URLs are fine.
71
+ - Approved CDN libraries: Tailwind, GSAP, Chart.js, D3.js, Three.js, Lucide Icons, Framer Motion, React 18 + Babel standalone, Vue 3, date-fns.
72
+ - Outside the html block: one-line lead-in and a short note about how to iterate. Nothing else.
73
+
74
+ ## Media Integration
75
+
76
+ If the conversation contains generated Kolbo media URLs (images, videos, audio), USE the actual URLs inside `<img>` / `<video>` / `<audio>` tags. Never substitute placeholders when real assets are available.
77
+
78
+ ## Publishing
79
+
80
+ After approval, call `publish_html_artifact({ title, content })` to publish to `sites.kolbo.ai` with strict CSP (`connect-src 'none'`, `form-action 'none'`). The page can't exfiltrate data; CDN libraries still load.
@@ -0,0 +1,41 @@
1
+ # App Builder
2
+
3
+ Load this file when the user wants to build / edit / iterate on a React app via Kolbo's App Builder ("build me a todo app", "add dark mode to my app", "give me the GitHub repo").
4
+
5
+ Use the App Builder tools to generate and iterate on full React apps from a text prompt. The backend auto-provisions a GitHub repo, Supabase database (when the app needs storage), and a live hosted deployment — all in one flow.
6
+
7
+ ## Standard Workflow
8
+
9
+ 1. **Find project ID**: `app_builder_list_projects` → pick the right project
10
+ 2. **Create session**: `app_builder_create_session` with `project_id`
11
+ 3. **Generate app**: `app_builder_generate_app` with `session_id` + `prompt`
12
+ - Fires the build in the background, polls until `build_status === "deployed"` (up to 5 min)
13
+ - Always surface the `deployment_url` to the user: **"Your app is live at: [url]"**
14
+ 4. **Iterate**: `app_builder_list_generations` → get `generation_id` → `app_builder_edit_app` with natural language instruction
15
+
16
+ No manual polling needed — `generate_app` and `edit_app` block until the build completes.
17
+
18
+ ## Local Dev Workflow
19
+
20
+ If the user wants to run the app locally or connect to the database directly:
21
+ ```
22
+ app_builder_get_session(session_id) → returns:
23
+ github_repo_url → git clone <url> && npm install && npm run dev
24
+ supabase_url → paste into .env as NEXT_PUBLIC_SUPABASE_URL
25
+ supabase_anon_key → paste into .env as NEXT_PUBLIC_SUPABASE_ANON_KEY
26
+ ```
27
+
28
+ ## ⚠️ Rules
29
+
30
+ - **Always confirm before `app_builder_delete_session`** — permanently deletes the GitHub repo, Supabase DB (unless user-connected), deployed files, and history. IRREVERSIBLE.
31
+ - **On build timeout** (rare): use `app_builder_get_build_status` to check manually, then continue or report.
32
+
33
+ Whitelabel works automatically — the MCP client routes App Builder calls through whitelabel API endpoints.
34
+
35
+ ## Routing examples
36
+
37
+ | User says | Sequence |
38
+ |---|---|
39
+ | "Build me a todo app" / "Make a landing page with waitlist" | `app_builder_list_projects` → `app_builder_create_session` → `app_builder_generate_app` → show `deployment_url` |
40
+ | "Add dark mode to my app" / "Add a contact form" | `app_builder_list_generations` → `app_builder_edit_app` |
41
+ | "Give me the GitHub repo" / "Supabase credentials" | `app_builder_get_session` → return `github_repo_url` + `supabase_url` + `supabase_anon_key` |
@@ -0,0 +1,138 @@
1
+ # Cost Awareness, Validation & Constraints
2
+
3
+ Load this file when you need to: confirm cost before firing a generation, validate input params against a model's caps, or quote real cost after a generation completes.
4
+
5
+ ## Billing Units by Type
6
+
7
+ Creative generations bill against the user's Kolbo credit balance. **Billing units differ by type** — apply the correct formula before generating.
8
+
9
+ | Type | Billing unit | Credit range | Example |
10
+ |------|-------------|-------------|---------|
11
+ | **Image** | per image (flat) | 1–30 cr | Flux.1 Fast = 1 cr, Midjourney = 4 cr. If `resolution` is set, check `resolutionMultipliers` — some families multiply cost significantly at higher tiers. |
12
+ | **Image edit** | per image (flat) | 2–20 cr | |
13
+ | **Video** | **cr/s × duration** | 2–30 cr/s | Kandinsky 5 Fast × 5s = 10 cr; Seedance 2.0 × 10s = 300 cr. Check `resolutionMultipliers` + `soundCreditMultiplier`. |
14
+ | **Video from image** | **cr/s × duration** | 4–30 cr/s | Same per-second rule. |
15
+ | **Elements (ref-to-video)** | **cr/s × duration** | 4–30 cr/s | Check `credit` and multipliers in `list_models type="elements"`. |
16
+ | **Lipsync** | **cr/s × duration** | 5–20 cr/s | |
17
+ | **Music** | per generation (flat) | 15–60 cr | Suno v5 = 15 cr; ElevenLabs Music = 60 cr |
18
+ | **Speech (TTS)** | per 100 characters | 2–5 cr/100 chars | ElevenLabs (5) × 500 chars = 25 cr |
19
+ | **Sound effects** | per generation (flat) | 4–7 cr | |
20
+ | **3D model** | per model (flat) | 5–300 cr | Trellis = 5 cr; Meshy v6 = 150 cr; Marble 1.1 = 300 cr |
21
+ | **Transcription (stt)** | per minute of audio | `model.credit × duration_minutes` | |
22
+
23
+ ## Calculation Formulas
24
+
25
+ Apply when confirming cost before firing:
26
+
27
+ - **Video / Lipsync**: `total = model_credit_per_second × duration_seconds`. Never assume the credit shown is a flat per-generation cost for these types.
28
+ - **Music**: flat per generation — `total = model_credit` (duration does not change cost).
29
+ - **TTS**: `total = model_credit × ceil(character_count / 100)`. Count actual characters first. 1000 chars with ElevenLabs = 50 credits.
30
+ - **Images / 3D / Sound effects**: `total = model_credit × quantity`.
31
+ - **Resolution / audio multipliers**: if `resolution` is set or model has native audio, read `resolutionMultipliers[tier]` and `soundCreditMultiplier`. Formula: `final = base × resolutionMult × (sound ? soundMult : 1) × durationSeconds`.
32
+
33
+ ### Tier label → pixel mapping (rough)
34
+
35
+ - Images: `"1K"` ≈ 1024px, `"2K"` ≈ Full HD (1920×1080), `"3K"` ≈ QHD (2560×1440), `"4K"` ≈ UHD (3840×2160). Picker shows only tiers the model supports (per `supported_resolutions`).
36
+ - Videos: `"720p"` / `"1080p"` / `"1440p"` / `"2160p"` = vertical pixels. Some models use model-specific labels like `"512P"` / `"1024P"` (Hailuo).
37
+
38
+ ## When to Confirm Cost
39
+
40
+ **Skip cost confirmation when:**
41
+ - The user already specified model + count + duration ("make 5 videos, seedance 2 fast, 15s" IS the confirmation).
42
+ - A single generation costs under 5 credits.
43
+
44
+ **Required cost confirmation when:**
45
+ - Anything else — present a one-line summary: "8 videos × 5s × [model] @ X cr/s = **Y credits**. Proceed?"
46
+ - Suggest a cheaper alternative if one exists.
47
+ - Wait for the user's confirm before firing.
48
+
49
+ **Batch totalling 100+ credits:** run `check_credits` first and include the available balance in the summary.
50
+
51
+ ## ⚠️ Quote Real Cost, Never Estimates (CRITICAL)
52
+
53
+ Pre-flight formulas above are for **preview only**. After firing, every generation returns `credits_used` (multiplier-adjusted total) and `credits_breakdown` (per-model attribution).
54
+
55
+ ```json
56
+ {
57
+ "credits_used": 12,
58
+ "credits_breakdown": [
59
+ { "model": "nano-banana-2", "base": 8, "final": 12, ... }
60
+ ],
61
+ "urls": [...]
62
+ }
63
+ ```
64
+
65
+ **Log `credits_used` to `.kolbo/production.md`**, not `base × count`. The multiplier-adjusted number is the only truth.
66
+
67
+ When the user asks "how much did I spend?" → call `get_session_usage` for the real, multiplier-adjusted session total + per-tool + per-model breakdowns (same numbers as the desktop bottom-bar counter).
68
+
69
+ ## Validation Pattern — Every Generation
70
+
71
+ Before submitting:
72
+
73
+ 1. Call `list_models type=<tool-type>` (text mode is enough for picking; `format: "json"` for programmatic comparison).
74
+ 2. For each input array (refs / DNAs / elements) — check `length <= <cap>` from the canonical field reference below. If over, drop the lowest-priority entries OR ask the user.
75
+ 3. For each enumerated value (`aspect_ratio` / `resolution` / `duration`) — check it's in `supported_*`. If not, **do not silently substitute**; show the user the allowed set and ask.
76
+ 4. For each duration-bearing file (source_video for lipsync/v2v, audio for lipsync/elements) — pre-check duration against the min/max range. Use ffmpeg if needed.
77
+ 5. For uploads — pre-check size against `max_file_size`.
78
+
79
+ The MCP tool descriptions also embed the cap field name on the relevant parameter (e.g. `reference_images: "...Cap: pass at most max_reference_images..."`) — use those as inline reminders.
80
+
81
+ ## Canonical Field Reference — Which `list_models` Field Controls Which Input
82
+
83
+ The same conceptual slot (e.g. "max reference images") lives under **different field names per model family**. Read the row for your tool, not the model name.
84
+
85
+ | Your input | Tool(s) | Field on the model | What `0` / `null` means |
86
+ |---|---|---|---|
87
+ | `reference_images` | `generate_image`, `generate_image_edit` (uses `source_images`), `generate_creative_director`, `generate_video` | `max_reference_images` | `0` = no refs |
88
+ | `reference_images` | `generate_elements` | `elements_max_images` | `0` = no image refs |
89
+ | `reference_images` | `generate_video_from_video` | `max_images` | `0` = no secondary image input |
90
+ | `reference_videos` | `generate_elements` | `elements_max_videos` | `0` = no video refs |
91
+ | `reference_videos` | `generate_video_from_video` | `max_videos` | `<= 1` = only the source_video |
92
+ | `elements` | `generate_video_from_video` | `max_elements` | `0` = no elements |
93
+ | `audio_url` | `generate_elements` | `elements_max_audio` (+ `max_audio_duration` for the file) | `0` = no audio ref |
94
+ | `visual_dna_ids` | every DNA-aware tool | `max_visual_dna` (+ `supports_visual_dna` boolean) | `null` / `0` / `false` = model rejects DNA |
95
+ | `aspect_ratio` | any | `supported_aspect_ratios` (or `_by_type[<type>]` when multimodal) | empty → `default_aspect_ratio` if set |
96
+ | `resolution` | any | `supported_resolutions` (+ `resolution_multipliers` for cost) | empty → no resolution tiering |
97
+ | `duration` (video output) | video tools | `supported_durations`, else `min_output_duration`–`max_output_duration` | both null → omit and let server default |
98
+ | **input** video duration | `lipsync-video`, `generate_video_from_video` | `min_video_duration` – `max_video_duration` | outside range → reject |
99
+ | input audio duration | `generate_lipsync`, `generate_elements` audio | `min_audio_duration` – `max_audio_duration` (+ `audio_max_follows_video_duration` for lipsync) | outside range → reject |
100
+ | audio file format | any audio input | `supported_audio_formats` (e.g. `["mp3","wav","m4a"]`; empty = all) | pre-validate before upload |
101
+ | recording duration | `text_to_speech` recording UX | `min_recording_duration` – `max_recording_duration` | usually null for plain TTS |
102
+ | upload file size | every file upload | `max_file_size` (bytes) | null → use platform default |
103
+ | `num_images` | image tools | `images_per_request` overrides for fixed-output models (Midjourney returns 4) | null → `num_images` honored as-is |
104
+ | `prompt` | every tool | `requires_prompt`, `min_prompt_length`, `max_prompt_length` | null → unconstrained |
105
+ | sound on/off | video tools | `sound_generation_type` (`"native"` vs `"none"`), `sound_enabled_by_default`, `sound_credit_multiplier` | not `"native"` → can't emit synced audio |
106
+ | capability gate | route decision | `supports_visual_dna`, `supports_first_last_frame`, `supports_audio_input` | `false` → the controller silently drops that param |
107
+
108
+ Cost formula: `final_cost = credit × resolution_multipliers[resolution] × (sound_enabled ? sound_credit_multiplier : 1)`, multiplied by `num_images` / `scene_count` as applicable.
109
+
110
+ ## Decision Rule for Resolution
111
+
112
+ 1. **User specified resolution explicitly** ("4K", "1080p", "480p") → ALWAYS verify in `supported_resolutions` BEFORE firing. If not supported:
113
+ - ❌ Do **NOT** silently substitute. The user asked for 480p; sending 720p without consent burns 1.5–2× the credits they expected.
114
+ - ✅ Show them the supported set in one line and ask:
115
+ > "Seedance 2 elements supports `[720p, 1080p, 1440p, 2160p]` — 480p isn't available. Closest cheap option is 720p (~+0 credits over your intent). Want 720p, or pick another?"
116
+ - Only fire after they reply.
117
+ 2. **User specified quality intent without numbers** ("draft", "quick test", "final delivery", "for client", "production"):
118
+ - draft / quick / preview → cheapest in `supported_resolutions` (1K / 720p)
119
+ - normal / standard → middle tier (typically 2K / 1080p)
120
+ - final / production / hero → highest the user's budget allows (3K-4K / 1440p-2160p)
121
+ 3. **No quality signal AND cost difference >2×** OR total batch ≥4 outputs → **ask the user once** with a one-line cost comparison, then default to standard if they don't reply.
122
+ 4. **No quality signal AND cost difference ≤1.5×** → quietly use the cheapest supported, no need to interrupt.
123
+ 5. **Sound on a video model with `sound_credit_multiplier > 1`** → if user didn't ask for sound, leave it off. If user said "with sound" / "with music", enable it.
124
+
125
+ ## Defaults When Nothing Is Specified
126
+
127
+ - **Image**: `1K` (or the cheapest in `supported_resolutions`).
128
+ - **Video**: `720p` (or the cheapest), with `default_duration` (or shortest in `supported_durations`).
129
+ - **Sound**: respect `sound_enabled_by_default`; if false, leave off.
130
+
131
+ ## Always Log the Resolution / Duration / Sound Choices
132
+
133
+ Production-log entries should include the resolution and (for video) duration + sound state alongside the URL, so the user can see what they paid for:
134
+
135
+ ```md
136
+ - still: https://...01-coffee.png (flux-2-pro · 1K, 2026-05-14)
137
+ - video: https://...02-rain.mp4 (kling-2 · 1080p · 5s · sound-off, 2026-05-14)
138
+ ```
@@ -0,0 +1,126 @@
1
+ # DTC Ads — Composed Brand Image Workflow
2
+
3
+ Load this file when the user wants a **DTC ad image** composed from brand identity + ad format + optional avatar/product/reference media. For ad **video** see `workflows/marketing-studio.md`. For brand **product imagery** (Pinterest pin, hero banner, ad pack) see `workflows/product-photoshoot.md`. For marketplace listings see `workflows/marketplace-cards.md`.
4
+
5
+ ## What This Is
6
+
7
+ A DTC ad is built from **5 composable blocks**:
8
+
9
+ 1. **Ad format** — the structural template (headline-driven, bullet-points, us-vs-them, before-after, founder-statement, etc.). Defines the layout shape.
10
+ 2. **Brand kit** *(optional)* — palette, fonts, logo, tone, voice. Keeps every ad in a campaign visually consistent.
11
+ 3. **Avatar** *(optional)* — a presenter face (curated character or trained Visual DNA). Use when the brand has a specific founder, model, or recurring presenter.
12
+ 4. **Product** *(optional)* — the item being sold. One product image, or a product brief from a URL.
13
+ 5. **Reference media** *(optional)* — up to ~14 reference images to anchor style / composition / setting.
14
+
15
+ You don't need all 5. The minimum is: a **prompt** + an **ad format**. Everything else is opt-in based on what the user provides.
16
+
17
+ ## End-to-End Flow
18
+
19
+ ```
20
+ 1. Pick an ad format → ask user (labeled options, never auto-pick)
21
+ 2. Pick / build brand kit → workflows/research-first.md persists to .kolbo/brand-kits/<slug>.md
22
+ 3. Attach avatar → workflows/visual-dna.md ("character" type DNA)
23
+ 4. Attach product → upload_media → reference_images
24
+ 5. Attach reference media → upload_media → reference_images (up to ~14 total)
25
+ 6. Generate → generate_creative_director (multi-variant) or generate_image (single)
26
+ 7. Deliver → image URLs + brief one-line summary
27
+ ```
28
+
29
+ ## Ad Format — Always Ask Explicitly
30
+
31
+ Picking an ad format is **mandatory and creative** — don't auto-pick from the user's phrasing. The catalogue is small and the choice changes the layout shape dramatically. Always present labeled options:
32
+
33
+ | Format type | Examples |
34
+ |---|---|
35
+ | **Headline-driven** | Big hero phrase + small product. "Hero word" style. |
36
+ | **Bullet points** | 3–5 benefit bullets + product hero. SaaS / DTC standard. |
37
+ | **Us vs Them** | Side-by-side comparison column. Competitor takedown style. |
38
+ | **Before / After** | Split frame showing transformation. Great for skincare, fitness, home. |
39
+ | **Founder statement** | Founder portrait + quote + product. Trust-builder. |
40
+ | **Lifestyle hero** | Product in-use in an aspirational scene. No copy hero. |
41
+ | **Pure product** | Clean studio product shot with brand framing. |
42
+ | **Testimonial** | Customer quote + face + product. Social proof. |
43
+ | **Pattern interrupt** | Bold color block / typographic shock / surreal composition. Scroll-stopper. |
44
+
45
+ When the user says "make me an ad" without naming a format, offer 3 of these in a labeled question (don't dump all 9). Pick the 3 that best fit the product / brand / phase the user mentioned.
46
+
47
+ ## Brand Kit Reuse
48
+
49
+ If `.kolbo/brand-kits/<slug>.md` exists for the brand (see `workflows/research-first.md`), **Read it first** and pull `primary_color`, `accent_color`, `text_color`, `bg_color`, `fonts`, `tone`, `target_user`, `logo_url`. Bake these into the prompt:
50
+
51
+ - Exact hex codes for every color (`#FF4D2E` not "orange")
52
+ - Named fonts (`Inter Bold for headline, Inter Regular for body`)
53
+ - Tone descriptors from `### Voice & Audience`
54
+ - Logo as `reference_images[0]` with `@image1` reference in the prompt ("place logo from `@image1` top-left at 8% width, no recolor")
55
+
56
+ If no brand kit exists and the user gives a brand URL, run `workflows/research-first.md` to build one. Then come back here.
57
+
58
+ ## Avatar Workflow
59
+
60
+ For ads featuring a specific presenter (founder, recurring model, character):
61
+
62
+ 1. **Check if a Visual DNA exists** — `list_visual_dnas`. Match by name or recent use.
63
+ 2. **If yes** — pass `visual_dna_ids: ["<id>"]` and reference as `@<dna-name>` in the prompt.
64
+ 3. **If no** and the user wants a specific person — create one per `workflows/visual-dna.md` (always generate 2 reference images first; lock single-token lowercase name).
65
+ 4. **If no** and the brief doesn't need a specific face — skip the avatar entirely; the model will synthesize a plausible presenter.
66
+
67
+ ## Product Workflow
68
+
69
+ For ads featuring a specific product:
70
+
71
+ | User provides | Do |
72
+ |---|---|
73
+ | **Product photo** (local file or URL) | `upload_media({ source })` → tag as `@image1` in prompt → log to `.kolbo/production.md` under `### Products` |
74
+ | **Product URL only** (no photo) | Run `workflows/research-first.md` first to scrape hero images + brand palette; re-host via `upload_media` → use Kolbo CDN URL |
75
+ | **Multiple angles** | Upload all in parallel (one `upload_media` call each) → pass all in `reference_images` → tag `@image1`, `@image2`, … per `workflows/visual-dna.md` reference-tagging rules |
76
+ | **Nothing — text only** | Ask once: "Do you have a product photo? It dramatically improves fidelity." If they say no, proceed text-only but warn quality may be lower |
77
+
78
+ **Always log products in `.kolbo/production.md`** so subsequent ads in the same workspace reuse the same CDN URL without re-uploading.
79
+
80
+ ## Reference Media Cap
81
+
82
+ Up to **~14 reference images per call**. Higher = the model gets confused about which reference plays which role. Use **`@image1` / `@image2` / …** tags to bind each reference to a role:
83
+
84
+ ```
85
+ Headline ad with @maya (the founder) holding @image1 (the product),
86
+ shot in the style of @image2 (lifestyle reference).
87
+ Match the palette from the brand kit (#FF4D2E primary, #1A1A1A text).
88
+ ```
89
+
90
+ See `workflows/visual-dna.md` for the full tagging system.
91
+
92
+ ## Generate
93
+
94
+ **Pick the right Kolbo MCP tool based on output count:**
95
+
96
+ - **Single ad image** → `generate_image` with `model: "<from list_models>"`. Use Nano Banana 2 for character/lifestyle, GPT Image 2 for layouts with dense on-image text or infographics, Nano Banana Pro for hero/brand-final assets.
97
+ - **Multi-variant set** (3–8 variants of the same ad concept with different palettes / angles / hooks) → `generate_creative_director` with `scene_count`. The director plans each variant's prompt internally.
98
+ - **Identical prompt, just different seeds** (rare for ads — usually you want varied direction) → `generate_image` with `num_images: 1–4`.
99
+
100
+ ## Output Settings — Always Confirm
101
+
102
+ These materially change output and cost. Ask once, labeled options, before firing:
103
+
104
+ | Setting | Common options for ads |
105
+ |---|---|
106
+ | `aspect_ratio` | `1:1` (IG feed) / `9:16` (Reels / TikTok / Stories) / `4:5` (IG portrait) / `16:9` (YouTube, banners) / `1.91:1` (Facebook feed) |
107
+ | `resolution` | `1K` (drafts, fast iteration) / `2K` (standard delivery) / `4K` (hero / print) |
108
+ | Quantity | `1` (test) / `3–4` (variant exploration) / `8` (full ad pack via Creative Director) |
109
+
110
+ Default-to-cheapest when the user hasn't expressed a quality intent and the difference is ≤ 2× cost.
111
+
112
+ ## Failure Handling
113
+
114
+ - **Content-policy refusal** → don't retry the same prompt. Suggest less-explicit phrasing or a different product framing.
115
+ - **Brand asset not loading** (logo URL 404, hex code typo) → fix the brand kit file, then retry.
116
+ - **Watermarks / extra text appearing uninvited** → add explicit prompt constraints: "NO captions, NO subtitles, NO watermarks, NO extra text beyond what's specified." This is the most common DTC ad failure mode — models love to invent copy.
117
+ - **Generic 5xx / rate-limit** → retry ONCE with the same payload after a short pause. See SKILL.md "Detecting failed generations".
118
+
119
+ ## UX Rules
120
+
121
+ 1. **Always pick an ad format explicitly** with the user — never auto-pick.
122
+ 2. **Always confirm aspect ratio + resolution + quantity** before firing.
123
+ 3. **Always check for a brand kit** before scraping fresh — `Read .kolbo/brand-kits/<slug>.md` first.
124
+ 4. **Always log products + brand kits in `.kolbo/production.md`** so future ads reuse instead of re-uploading / re-scraping.
125
+ 5. **No auto-retry on failure** — surface the reason and let the user adjust.
126
+ 6. **Strict NO uninvited additions** in every ad prompt: "NO captions, NO subtitles, NO watermarks, NO extra text beyond what's specified."