@slatesvideo/shared 0.6.8 → 0.6.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/manual/content.d.ts
CHANGED
|
@@ -1,2 +1,2 @@
|
|
|
1
|
-
export declare const APP_MANUAL = "# Slates \u2014 Complete Reference for AI Assistants\r\n\r\n<!-- The heading above must stay first: shipped app builds reject this file if it\r\n does not begin with one. See slate/CLAUDE.md; check:llm-docs enforces it. -->\r\n\r\n<system_role>\r\nYou are a support assistant for the Slates desktop application. Answer user questions using ONLY the information in the <slates_reference> below. Be concise and direct. Use numbered steps for procedures. Use bullet points for explanations when helpful.\r\n</system_role>\r\n\r\n<rules>\r\n- If the answer cannot be found in the <slates_reference>, say: \"That isn't covered in the Slates reference.\" Do not guess or invent features.\r\n- If the user asks how to do something, give step-by-step instructions from the workflows and features described here.\r\n- If the user reports an error, check the TROUBLESHOOTING section first.\r\n- Refer to the FEATURES NOT IN SLATES section before answering questions about capabilities that might not exist.\r\n- Quote the exact error message when referencing troubleshooting entries.\r\n</rules>\r\n\r\n<slates_reference>\r\n\r\n<!-- BEGIN:GENERATED header -->\n# SLATES v1.5.6 \u2014 Complete Reference\n\n> **Freshness.** Generated from the Slates source of truth for app version **1.5.6**, last changed **2026-09-09**. The canonical copy of this file is <https://slates.video/slates-reference.md>. If a model, price or feature the user mentions is missing below, this copy is out of date: re-fetch that URL before answering, and say so.\n<!-- END:GENERATED header -->\r\n\r\n**What is Slates?** Desktop app (Windows 10/11, macOS 12+) for AI image and video creation. One-time purchase, no subscription. Every license includes 1,000 free credits, and Slates Pro starts with 3,000 credits. Every generation runs on Slates Credits \u2014 there are no API keys to set up, and credits never expire.\r\n\r\n---\r\n\r\n## INTENDED WORKFLOW\r\n\r\nThe designed start-to-finish flow:\r\n\r\n1. **Create project** \u2014 New project with name/description. Creates folder on your disk.\r\n2. **Build visual assets** \u2014 Generate images, create characters (with character sheets for consistency), environments (with environment grids), and styles. This is your visual library.\r\n3. **Create storyboard** \u2014 Add scenes. Every picture you drop in becomes a **Shot**: one beat of the piece, holding its references, its model, its settings and (when you want them) its words.\r\n4. **Write the piece** \u2014 Switch the storyboard to **Document** and write. What is said, what happens, the framing, the prompt each beat will send. Nothing is required and nothing is asked for; a visuals-only piece is finished as it stands.\r\n5. **Read it before you pay for it** \u2014 The header states how many generations, how many cuts, how long it runs and what it will cost. Press play for a rough cut at the real timing. Re-chop with split and merge and watch the price move.\r\n6. **Generate** \u2014 Select the beats you want and fire them in one approved batch. Each result lands under the row that made it.\r\n7. **Organize** \u2014 Switch to **Board** to re-order. Drag a beat and the whole thing moves with it.\r\n8. **Export to timeline** \u2014 Send clips to the built-in multi-track video editor.\r\n9. **Edit** \u2014 Trim, reorder, add markers, adjust timing.\r\n10. **Final export** \u2014 Export to MP4 directly, or export DaVinci Resolve XML for professional color grading.\r\n\r\n---\r\n\r\n## MODEL REFERENCE TABLE\r\n\r\n**Generating 4K video is a Slates Pro feature** \u2014 every tier generates video up to 1080p, and 4K images are open to everyone. Exporting your finished timeline at 4K is available on every tier.\r\n\r\nEvery model runs on Slates Credits. **The exact credit cost appears on the Generate button before anything fires.** The tables below are generated from the app's own model registry and rate tables, so they describe exactly what the model picker offers in this version: aspect ratios, resolutions, durations, reference-image limits, and the credit price of each. Bigger credit packs lower your per-credit cost, and Slates Pro gets the best pack rate on every purchase.\r\n\r\n<!-- BEGIN:GENERATED model-tables -->\n### Image Models\n\n| Model | Aspect Ratios | Resolutions | Max Refs | Credits per image |\n|-------|--------------|-------------|----------|-------------------|\n| **GPT Image 2.5 Flare** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 \u00B7 3K 3 \u00B7 4K 5 (default quality; at max: 2K 8 \u00B7 3K 11 \u00B7 4K 20) |\n| **GPT Image 2.5 Sunburst** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 \u00B7 3K 3 \u00B7 4K 5 (default quality; at max: 2K 8 \u00B7 3K 11 \u00B7 4K 20) |\n| **Nano Banana 2** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 4 \u00B7 2K 6 \u00B7 4K 8 |\n| **NB2 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K | 4 | 1K 2 |\n| **Nano Banana Pro** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 8 \u00B7 2K 8 \u00B7 4K 15 |\n| **FLUX.2 Max** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 4 | 1K 4 \u00B7 2K 5 \u00B7 4K 8 |\n| **Seedream 5 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 2K / 3K / 4K | 10 | 2K 2 \u00B7 3K 2 \u00B7 4K 2 |\n\n### Video Models\n\n| Model | Duration | Aspect Ratios | Resolutions | Max Refs | Audio | Credits per second |\n|-------|----------|--------------|-------------|----------|-------|--------------------|\n| **Seedance 2.0** | 4-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p / 4K | 9 | Included | 480p 3.5 \u00B7 720p 7.5 \u00B7 1080p 18.5 \u00B7 4K 39 |\n| **Seedance 2.5** | 4-30s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 30 | Included | 480p 5.1 \u00B7 720p 11.6 \u00B7 1080p 20.5 |\n| **Seedance 2.5 Edit** | Follows the source clip (4-30s) | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 0 | Included | 480p 6.2 \u00B7 720p 13.9 \u00B7 1080p 24.6 |\n| **Kling V3.0 Standard** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (6.3 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (8.4 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Omni** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (5.6 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Omni Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (7 with audio) \u00B7 4K 21 |\n| **Kling O3 Edit** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 6.3 |\n| **Kling O3 Edit Pro** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 8.4 |\n| **MiniMax H3** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 2K / 4K | 9 | Included | 480p 2.5 \u00B7 768p 3 \u00B7 2K 6.5 \u00B7 4K 8 |\n| **MiniMax H3 Max** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 1080p | 9 | Included | 480p 2.5 \u00B7 768p 4 \u00B7 1080p 8 |\n| **Gemini Omni Flash** | 3-10s | 16:9, 9:16 | 720p | 7 | Included | 720p 6.4 |\n| **Omni Flash Edit** | Follows the source clip (3-10s) | 16:9, 9:16 | 720p | 0 | Included | 720p 6.4 |\n| **LTX-2.5** | 6/8/10/12/14/16/18/20s at 720p/1080p; 6/8/10s at 1440p/4K | 16:9, 9:16 | 720p / 1080p / 1440p / 4K | 0 | Included | 720p 4.5 \u00B7 1080p 6.5 \u00B7 1440p 9.5 \u00B7 4K 15 |\n| **LTX-2.5 Pro** | 6, 8, 10s | 16:9, 9:16 | 720p / 1080p | 0 | Included | 720p 6 \u00B7 1080p 8.5 |\n| **Veo 3.1 Fast** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 5 (7.5 with audio) \u00B7 1080p 5 (7.5 with audio) \u00B7 4K 15 (17.5 with audio) |\n| **Veo 3.1 Standard** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 10 (20 with audio) \u00B7 1080p 10 (20 with audio) \u00B7 4K 20 (30 with audio) |\n\n**Seedance 2.0 \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 27.4 credits per second at 1080p.\n**Seedance 2.5 \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 16.3 credits per second at 720p.\n**Seedance 2.5 Edit \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 19.8 credits per second at 720p.\n\n### Audio Models\n\n| Model | Length | Credits |\n|-------|--------|---------|\n| **Seed Audio 1.0** | 3-120s | 1 at 3s \u00B7 3 at 15s \u00B7 19 at 120s |\n| **Inworld TTS-2** | up to 2,000 characters of text per take | 1 at 250 characters \u00B7 2 at 2,000 characters (billed per 250) |\n| **Sound Effects** | 1-22s | 1 at 1s \u00B7 1 at 4s \u00B7 3 at 22s |\n\n### Tools (Lip Sync, Motion Transfer)\n\nThese are real Kling endpoints that take a clip or a still as their subject, not models you prompt from scratch. Both bill in 5-second blocks.\n\n| Tool | Input | Billed in | Credits per block |\n|------|-------|-----------|-------------------|\n| **Kling Lip Sync** | Video source | 5s block | 4 |\n| **Kling Lip Sync (Avatar v2 Standard)** | Still-image source | 5s block | 14 |\n| **Kling Lip Sync (Avatar v2 Pro)** | Still-image source | 5s block | 29 |\n| **Kling Motion Control Standard** | Still image + reference video | 5s block | 32 |\n| **Kling Motion Control Pro** | Still image + reference video | 5s block | 42 |\n<!-- END:GENERATED model-tables -->\r\n\r\n---\r\n\r\n## WHICH MODEL TO USE\r\n\r\n### Images\r\n\r\n**Nano Banana 2 is the default image model.** It is the best all-round image model in the app: the most reference images of any image model, every aspect ratio, and output up to 4K. Brief it like a creative director rather than with tag soup. It is also the only model that supports the 2x2 / 3x3 grid exploration wrapper.\r\n\r\n- **NB2 Lite** is the fast, cheap draft seat in the same Nano Banana family. Roughly half the price of NB2 full and noticeably faster, 1K output only. Iterate here, finish on NB2.\r\n- **Nano Banana Pro** is the hero-frame and typography tier. Reach for it when spatial composition, cinematic lighting and skin, or fine in-image type have to be perfect. NB2 gets you most of the way there, so this is a deliberate step up, never a default.\r\n- **GPT Image 2.5** is the strongest image model in the app: it follows a long instruction more faithfully than anything else here, and it is the one to pick when the picture simply has to be right. It is also the sharp-text model, which is what makes it the choice for character sheets, shot grids, ordered panels and anything with words in the picture. It comes in two seats that cost exactly the same, and the difference is speed against quality. **Flare** is the fast one: OpenAI describes its quality as comparable to the older GPT Image 2, at roughly half the wait. **Sunburst** is OpenAI's most capable image model, better than GPT Image 2, and deliberately slower. Use Flare while you are still exploring, then re-run the shot you like on Sunburst for the final \u2014 and reach for Sunburst directly when several reference images all have to survive into one frame, or when an edit must change one region and leave identity, geometry and lighting untouched.\n Its quality knob has **five** settings \u2014 `low`, `medium`, `high`, `xhigh`, `max` \u2014 spanning about 36\u00D7 from cheapest to dearest, which makes it the biggest cost lever on the model. The steps are uneven rather than a constant multiplier: `max` is four times `high`, but `xhigh` is only about 1.8 times it. **`high` is the default and the everyday setting.** `medium` is for drafts and is cheap enough to iterate on freely. `max` is the ceiling, for finished frames and exact character-level text; `xhigh` sits just under it for about half the price and is worth trying first. Go past `high` deliberately, not by habit.\n 4K is worth it only once your references and prompt are already settled: prove the shot at 3K, then re-run the finished prompt at 4K. Iterating at 4K is the most common way to waste credits on this model. Its 4K tier is API-only, so even a paid ChatGPT account cannot render it. Note that its resolution tiers are **not** a price ladder: the pixel classes are token-priced by OpenAI, so the cheapest seat is not the smallest one. Read the prices in the table above rather than assuming.\n **If you have used GPT Image 2 before, the quality names all shifted by one.** What it called `medium` is now called `high`, and what it called `high` is now `max` \u2014 the same pictures at the same prices, renamed. A remembered setting will quietly buy you a cheaper tier than it used to.\r\n **It is the only image model that can give you a transparent background.** Set **Background** to *Transparent* on the prompt bar and you get a real alpha channel \u2014 a cut-out for a logo, sticker or overlay \u2014 rather than a painted-in backdrop. *Auto* is the default and lets the model decide from your prompt; *Opaque* forces a filled background. It costs nothing either way. Slates always saves PNG, which is what carries the transparency, so there is nothing else to set.\r\n **Square and 4:3 frames cost more than 16:9 on this model, and only on this model.** OpenAI charges by image tokens rather than by pixels, and a square frame uses about 1.8 times the tokens of a 16:9 frame the same size, and 4:3 or 3:4 about 1.37 times. The credit prices in the table above are the 16:9 numbers; pick 1:1 or 4:3 and the price on the Generate button goes up to match. 9:16 costs the same as 16:9. Every other image model charges the same whatever the shape.\r\n- **FLUX.2 Max** and **Seedream 5 Lite** are the less content-restricted options. Seedream is flat-priced at every resolution it offers, so there is no reason to pick a lower one. Both auto-route to their edit endpoint when you attach reference images.\r\n\r\n### Video\r\n\r\n**Seedance 2.0 is the default video model.** Reach for it the moment physics, effects, destruction or scale matter, and for hero shots. It takes many reference images, generates native audio at no extra cost, and is the only Seedance seat that reaches 4K (generating 4K video needs Slates Pro; timeline export at 4K does not). It is also the cheaper of the two seats at every resolution they share. A face in a reference image routes it to a different provider, which is why **Seedance 2.0 \u00B7 Face** is its own row in the model picker at its own price.\r\n\r\n- **Seedance 2.5 is a second seat, not an upgrade.** It buys much longer single takes and far more reference images, plus better prompt adherence. What it gives up is 4K, and it costs more than 2.0 at every resolution the two share \u2014 so 2.0 stays the model for 4K, and for the same resolution at a lower price. Because 2.5 runs longer, a long clip on 2.5 can cost more than a shorter, higher-resolution one on 2.0 \u2014 read the Generate button, not the resolution. **Seedance 2.5 Edit** is its clip-editing row: attach a clip, describe the change, and the output length follows the source.\r\n- **Kling** is the cost-effective workhorse and the most flexible family: strong start-frame adherence for identity, layout and text, acting, dialogue, multi-shot (up to 6 cuts), and the widest range of clip lengths. **Kling V3.0 Omni** adds multi-character dialogue in English, Chinese, Japanese, Korean and Spanish. Standard and Pro are the same model at two fidelity and price tiers. **Kling O3 Edit** takes an existing clip and changes what you describe, with subject and style reference images, while the original audio is preserved verbatim. Kling is also the only engine behind the Lip Sync and Motion Control tools.\r\n- **MiniMax H3** is the seat to pick when the SOUND is part of what you are writing. Every other video model treats audio as a switch; H3 takes it as three separate instructions in one prompt \u2014 the lines and action sounds tied to a moment, the ambience running underneath, and a score only the audience hears \u2014 and generates all of it with the picture in a single pass. It is also the only model where you say how much of a reference should survive, including moving one subject's characteristic onto a different subject. It runs 5-15 seconds at 480p, 768p, 2K or 4K, and takes up to nine reference images plus reference video and audio. Two things to watch: **the first five reference images are free and every one after that costs extra**, so attach what the shot needs rather than the maximum; and 2K and 4K are upscales of a 768p render rather than larger generations \u2014 in our own testing the 2K pass showed more artifacting than the 768p original it was built from, at more than twice the price. Generate and judge at 768p; step up only when a delivery spec demands the pixels.\r\n- **MiniMax H3 Max** is the same model post-trained by fal for SPEED, and it is the more expensive seat, not the cheaper one. It is dramatically faster: on the same 5-second 768p prompt it finished in about 5 seconds against about 57 seconds for H3 \u2014 roughly 12x (measured 2026-08-27). It stops at 768p and costs more per second than H3 at the resolution they share. It still animates a start frame and an end frame, so image-to-video works normally; what it does not have is the reference set \u2014 the extra identity, style and environment images plus reference video and audio that base H3 reads. Pick it when a fast turnaround on a text-to-video or start-frame shot is worth paying for; pick H3 for resolution, references, or the same tier at a lower price.\r\n- **LTX-2.5** is the VOLUME seat \u2014 the cheapest native 1080p second in the catalogue, with synchronised audio included free at every resolution, so it is the model to reach for when the job is many takes rather than one hero shot. Two things are unique to it. It makes the LONGEST clips of anything here, up to 20 seconds, and it is the only model that reaches 1440p. It is also the only one with native MULTISHOT: a single generation can carry two to four connected shots that hold the character, lighting and voice across the cuts, which everywhere else means generating separate clips and watching identity drift between them. Its constraints are unusually sharp, though. Durations are EVEN NUMBERS ONLY starting at six \u2014 6, 8, 10, 12, 14, 16, 18, 20, with no 5-second or 7-second clip \u2014 and above 1080p that ceiling drops to 10 seconds. Aspect ratios are 16:9 and 9:16 only. And it takes FRAMES, not references: a start frame and an optional end frame that generates a transition between them, but no identity, style or environment reference images at all, so cross-shot character consistency belongs on MiniMax H3 or Kling. Because sound is generated in the same pass, write the audio into the prompt and anchor every cue to something visible \u2014 anything unanchored gets invented.\r\n- **LTX-2.5 Pro** is the fidelity seat of that pair, and it is NOT simply a better LTX. It renders the picture with more compute on busy frames, but on a narrower envelope than the base row: 720p and 1080p only (no 1440p, no 4K) and 6, 8 or 10 seconds only, for about a third more per second. Reaching for it because the name says Pro costs more AND takes away the reach. Pick it when a specific shot needs the extra fidelity and fits inside 1080p and ten seconds; pick base LTX for length, resolution and volume.\r\n- **Gemini Omni Flash** is the cheap 720p seat with native synced audio included in one pass. **Omni Flash Edit** is the prompt-only clip editor: no reference images, one short instruction plus \"Keep everything else the same.\" Long descriptive prompts destroy it.\r\n- **Veo 3.1** is niche and is never a default. Pick it only when you specifically want Google's audio pass. It has the fewest aspect ratios and reference slots of any video model, fixed durations, and the highest per-clip cost.\r\n\r\nBoth edit models take an existing clip as their canvas, so their output length follows the source clip rather than a duration you choose.\r\n\r\n### Audio\r\n\r\nAudio is a third media type alongside images and video \u2014 generated as its own asset, shown in the gallery's **Audio** tab, and dragged onto an audio track in the timeline. This is separate from the audio some VIDEO models generate *inside* a clip (see AUDIO IN GENERATION below): use a video model when the sound must be locked to what is on screen, and these when you need audio you can move, trim, re-use, or layer.\r\n\r\n**Seed Audio 1.0 is the default.** A room with dialogue *and* clatter *and* ambience is one generation, not three layered ones, and because it is cheap you can run five takes and keep the best. It makes a whole audio SCENE from one plain sentence. **It has no length setting of its own** \u2014 Slates writes your chosen duration into the prompt, and that is what you are charged. You describe the voice in words; there is no voice list to pick from. Set **Languages** to Mixed if one scene needs more than one language (it costs the same).\r\n\r\n**Sound Effects** makes one effect, or a seamless loop. It is the only surface with an exact duration, so an effect can land on a specific frame. Describe the physical cause (\"heavy oak door slams shut in a stone hallway\"), not the label (\"door sound\"). **Loop** makes it seamless for beds; **Wording** controls how literally your description is followed. Seed Audio is actually the better tool for *long* ambience beds, so the two are not redundant in the direction you would expect.\r\n\r\nKling's `SFX:` / `Ambient noise:` prompt syntax belongs to video prompts and makes Seed Audio results *worse* \u2014 write plain sentences there instead.\r\n\r\n**Inworld TTS-2 is the voice seat** \u2014 type the words, pick a voice, press Generate. In the prompt box's Audio lane pick **Voice** in the model picker; the prompt is the exact text that gets spoken (nothing is added or rewritten \u2014 open \"See what gets sent\" to confirm), and the **Voice** control on the bar opens the voice picker: **Presets** (ready-made voices with gender, accent and age filters \u2014 every one plays the same audition line, so you compare voices rather than scripts), **Clips** (any character's voice, or any audio clip in the project, cloned for the take), or **Describe** (a voice in words). The character counter beside the bar is the bill: the generated audio table above gives the text cap and billing buckets, and the Generate button shows the price. Direction goes in square brackets (`[whispering] \u2026`) \u2014 anything in parentheses is read aloud. Cloning a real person's voice needs their permission. Right-click any audio clip in the Audio tab \u2192 **Use as voice** lands you on the Voice lane with that clip as the voice. Studio Agent and the MCP/CLI do the same through `slates_generate_audio` (a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId`, or a `voiceDescription`).\r\n\r\n**There is no music generation.** For a song, use an external tool and import the audio (see PROJECTS \u2192 Supported File Formats). For spoken lines inside a scene, let Seed Audio perform them, put them in the video prompt on a model with native audio (Seedance, Kling Omni, Omni Flash, Veo), or use Kling Lip-Sync's text-to-speech against a shot.\r\n\r\n### Tools\r\n\r\n**Tools** is not a model family. It is two real Kling endpoints that take a clip or a still as their subject: **Kling Lip Sync** and **Kling Motion Control**. Both bill in 5-second blocks and are described under GENERATION MODES below.\r\n\r\n**How pricing works:** every generation is priced in Slates Credits, and the exact cost is shown on the Generate button before you commit. Bigger credit packs give more credits per dollar; Slates Pro locks in the best pack rate on every purchase, forever.\r\n\r\n---\r\n\r\n## GENERATION MODES\r\n\r\n### Create Image\r\nPrompt \u2192 select image model \u2192 set aspect ratio + resolution \u2192 generate. Batch grids available for quick iteration.\r\n\r\n### Text-to-Video\r\nPrompt \u2192 select video model \u2192 set duration + aspect ratio + resolution \u2192 generate. Output: MP4.\r\n\r\n### Image-to-Video (I2V)\r\nAttach start image + prompt \u2192 select model \u2192 generate video from that image. Optional: attach end image (Veo) for guided transitions.\r\n\r\n### Ingredients / References\r\nUse @character_name, @environment_name, or #style_name in prompt to attach reference images for visual consistency. Kling: up to 4 total references. Veo: up to 3. Nano Banana 2: up to 14. The @mentions auto-complete from your project's characters/environments/styles.\r\n\r\n### Lip Sync\r\n**Kling only.** Pick **Kling Lip Sync** under the Tools family in the model picker. Source: video or still image.\r\n\r\n- **Audio source** \u2014 Text to speech (type the line; six English/UK voices plus a storyteller, with a speed control) OR Upload audio (bring your own recording, max 5MB).\r\n- **Avatar tier** \u2014 only appears for a still-image source: Avatar v2 Standard (the value tier) or Avatar v2 Pro (higher fidelity, higher rate).\r\n- 5s output blocks.\r\n\r\nWorks very well with human-like characters. Less reliable with animals or non-human characters.\r\n\r\n### Motion Transfer\r\n**Kling only.** Pick **Kling Motion Control** under the Tools family. Target: still image (your character). Source: reference video (the motion).\r\n\r\n- **Engine** \u2014 Kling MC Standard (value tier) or Kling MC Pro (higher fidelity, higher rate).\r\n- **Orientation** \u2014 *Match video* copies skeleton and depth from the clip (best for dancing, walking, full-body action; driving clips up to 30s). *Match image* keeps your character's pose and angle and uses the video only as motion hints (best for close-ups; up to 10s).\r\n- 5s output.\r\n\r\n> **Note:** these two tools used to offer a second \"Seedance 2.0\" engine. It was not a separate engine \u2014 picking it made Slates write a sentence into your prompt that you never saw, which is no longer allowed anywhere in the app (see \"What gets sent\" below). Both tools are now Kling endpoints only.\r\n\r\n### Edit Image\r\nRight-click any image asset \u2192 open viewer \u2192 switch to Edit Mode. Enter an edit prompt describing the changes you want. Select edit model: Nano Banana 2 (supports up to 14 reference images), FLUX.2 Max, or Seedream 5 Lite. Choose resolution and aspect ratio. The result saves as a new asset with the original preserved. Useful for refining generated images without starting from scratch.\r\n\r\n### Edit Video (Kling O3 Edit / Omni Flash Edit)\r\nRight-click any video clip (gallery or timeline) \u2192 \"Edit with AI\". The clip attaches to the prompt box as the source; describe the CHANGE, not the whole scene (\"replace the man with @marcus\", \"make it a rainy night, keep everything else\"). Two engines in the model picker:\r\n- **Kling O3 Edit (default):** attach subject images (role: Subject) to swap someone in, or style images (role: Style) for a look \u2014 max 4 combined refs. Clips 3-15s. Original audio preserved.\r\n- **Omni Flash Edit (cheapest):** prompt only \u2014 no reference images; keep instructions simple and add \"Keep everything else the same.\" Clips 3-10s, 720p output.\r\n\r\nOutput length follows the source clip; the credit cost (clip seconds, rounded up, at the per-second rate) shows on the Generate button. The edited clip saves as a NEW asset linked to the original \u2014 chain edits freely. Trim longer clips on the timeline first.\r\n\r\n### Multi-Shot (Kling V3.0/Omni)\r\nEnable multi-shot toggle \u2192 multiple scene prompts in one generation, each with different framing. 6-axis camera controls per shot. Results can be hit-or-miss, but worth trying for quick multi-cut sequences. For more reliable results, most users prefer generating multiple short 5s clips separately using Kling V3.0 Omni in ingredients mode and assembling them on the timeline.\r\n\r\n---\r\n\r\n### Generate Audio\r\n\r\nSwitch the prompt box's lane pill from Image/Video to **Audio**, pick a surface, and generate. The prompt box offers three: **Seed Audio 1.0**, **Voice** (Inworld TTS-2) and **Sound Effects**. The result lands in the gallery's Audio tab as its own asset with a waveform and an inline player, and can be dragged onto an audio track in the timeline.\r\n\r\n- **Scene (Seed Audio 1.0)** \u2014 one plain sentence describing the moment. Set **Length**; Slates writes it into the prompt for you and that is exactly what you're billed for (open \"See what gets sent\" under the prompt box to read the appended text). **Say the crowd/room size out loud** \u2014 \"applause\" returns a full auditorium when you meant three people at an open mic. Ask for a few seconds more than the clip needs so the edit has fade handles. Describe the voice you want in the sentence itself (\"a weary dock foreman in his fifties, gravel in his voice\") \u2014 there is no voice picker.\r\n- **Voice (Inworld TTS-2)** \u2014 the prompt is the words to be spoken, verbatim. Pick the voice with the **Voice** control on the bar (presets you can play first, any clip in the project, a character's voice, or a description); the character counter is the bill. See MODELS \u2192 Inworld TTS-2 for direction tags and the cloning rules.\r\n- **Sound Effect** \u2014 describe the physical cause and set the length to roughly the event (\u22481s for an impact, 2\u20134s for a whoosh, 8\u201322s + **Loop** for a bed). **Wording** sets how literally the description is followed: Interpretive, Balanced (default), or Literal.\n\n#### Use your own voice recording\n\nIn the bottom prompt box, choose **Audio**, then **Inworld TTS-2** in the model picker. Open **Voice \u2192 Clips \u2192 Import voice clip** and select your recording. The import adds an audio asset to this project without generating anything. Click its play button to audition it, then click the recording's name to choose it. Type the words you want spoken in the prompt box and press **Generate**, which shows the price. The new take appears in **Gallery \u2192 Audio**.\n\nUse a clean recording of one speaker whose voice you have permission to use. Slates clones the recording for each take; there is no separate training wizard or persistent vendor voice to manage. **Presets** lets you audition ready-made voices; **Describe** lets you write a voice description and choose **Use this description**. Choosing in the prompt box sets up the next take; only Generate spends credits.\n\nTo attach your recording or a generated take to a character, open **Gallery \u2192 Characters**, then **Add voice** (or **Change voice**) on that character's card. Choose **Clips** and click the clip's name. Attaching an existing clip is free. The card displays its waveform and player. That character's voice is also listed under **Voice \u2192 Clips \u2192 Characters** in the prompt box. Selecting a preset or description from a character card generates and attaches a take; read the cost shown in that picker before choosing.\n\nFor **Inworld TTS-2**, choose the character under **Voice \u2192 Clips \u2192 Characters**. Typing `@name` does not choose a TTS voice: the prompt is spoken verbatim. Automatic voice attachment from a character mention belongs to **Seed Audio** only.\n\r\n---\r\n\r\n## AUDIO IN GENERATION\r\n\r\nThis section is about audio generated **inside a video clip**. For audio as its own asset, see Generate Audio above.\r\n\r\n**Veo 3.1 native audio:** Generates audio WITH video. Prompt syntax: `\"Hello!\"` for dialogue, `SFX: [sound]` for effects, `Ambient noise: [description]` for ambience. Max 10s dialogue. Add `(no subtitles)` to suppress text overlays.\r\n\r\n**Kling V3.0 Omni dialogue:** Multi-character dialogue with distinct voices. Languages: EN, ZH, JA, KO, ES. `Background music: [description]` for music. Max 10s dialogue.\r\n\r\n**Kling V3.0 sound co-generation:** Synchronized sound effects generated with video.\r\n\r\n\u26A0\uFE0F **This prompt syntax is video-only.** `SFX:`, `Ambient noise:` and `Background music:` are Kling/Veo conventions \u2014 the audio models above have no parser for them and will treat them as words in the scene.\r\n\r\n---\r\n\r\n## PROMPT SYSTEM\r\n\r\n### Unified Create Surface (roles + model-on-button)\r\nThe old mode tabs (text-to-video / frames-to-video / ingredients / create-image) are ONE \"Create\" surface. An **Image | Video | Audio pill** on the prompt bar switches your output lane \u2014 it remembers and restores the last model you used in each lane (pick Seedance once and the Video lane stays Seedance until you change it). Only the active lane shows its name; the other two are icons. Every attachment in the reference tray carries a tappable ROLE badge \u2014 Reference / First frame / Last frame / Subject / Style \u2014 you say what each attachment is; nothing is inferred. Your model choice sticks across generations and workflow actions (\"use as first frame\" keeps your chosen video model). Attaching a video via \"Edit with AI\" flips the surface into Edit Video mode. Lip Sync and Motion Transfer live under the **Tools** family in the model picker.\r\n\r\n### Floating Prompt Box \u2014 the bar holds everything\r\nPersistent across all pages. **There is no settings panel and no gear button.** Everything sits on one bottom bar, left to right:\r\n\r\n1. **Media toggle** \u2014 Image / Video / Audio.\r\n2. **Model picker** \u2014 a searchable menu plus a detached submenu. The main menu lists model families with a vendor glyph tile each; picking one opens that family's models beside it, every row carrying capability chips (resolution, clip length, audio, references) and its per-unit rate. Type to search across every model. The submenu is anchored to the row you opened it from, so it never travels. The trigger on the bar shows the model name and nothing else \u2014 no chevron, no resolution appended. Everything listed is a real model or endpoint.\r\n3. **Parameter controls** \u2014 one per setting the chosen model actually has (resolution, aspect, duration, length, quality, count, grid, face-in-reference, audio, loop, and so on). The trigger shows the current value; the explanation lives *inside* the menu as a subtitle under each option, along with what that option costs. A setting with only one possible value still shows, muted and non-interactive, so the row never changes shape.\r\n4. **Sliders for ranges.** A setting with a long list of steps (video duration, audio length) opens a ruler instead of a many-row menu. The handle moves between the values the model actually declares, so it cannot land on one the model will not accept, and the price for the selected value is shown on the ruler.\r\n5. **`More \u25BE`** \u2014 if the model has more parameters than fit the current window width, the extras are collected into a generated `More` dropdown automatically. Widen the window (or close the Studio Agent panel) and they move back onto the bar.\r\n6. **Generate** \u2014 reads `Generate \u00B7 <cost>`. **Cost only** \u2014 the model name is not repeated on the button (it's in the model picker) and there is no send arrow. A badge on the left of the button counts generations currently running.\r\n\r\nFor text-to-speech, the character counter is always visible because text length determines the price. On other surfaces it appears near the right of the bar after roughly three quarters of the model's prompt limit. It turns red over the limit.\r\n\r\n**Collapsing:** the chevron at the top-right of the box collapses it to a single arrow \u2014 nothing else. Click the arrow to bring it back.\r\n\r\nBelow the bar the queue shows pending/active generations with cost and progress.\r\n\r\n### What gets sent (prompt transparency)\r\nUnder the prompt box is a **\"See what gets sent\"** disclosure. Open it and you see the exact text that will be transmitted, produced by the same code that builds the request \u2014 so it can never disagree with what is actually sent. It shows:\r\n\r\n- **Reference numbering** \u2014 `@sarah` becomes `Sarah (image 1)` so the model knows which attached image is which.\r\n- **Key lines for attachments you did NOT mention** \u2014 one short neutral sentence per unmentioned attachment. Mention every reference in your own words and these never generate.\r\n- **The trailing style clause** when a `#style` is attached.\r\n- **The Seed Audio duration append** \u2014 the `\u2026 N seconds` Slates adds to the end of the prompt, which is also what you are billed for.\r\n- **Grid wrapping** when 2\u00D72 or 3\u00D73 is on.\r\n\r\nThe row stays hidden when the composed prompt is identical to what you typed, so it only appears when there is something to show.\r\n\r\n**Unresolved `#tags` and `@mentions` are named, not silently dropped.** If you type `#noir` and there is no saved style called \"noir\", the tag is removed from the text sent to the model (a raw tag confuses every model) \u2014 but the disclosure turns red and says so by name: *\"#noir matches nothing saved \u2014 removed from what gets sent.\"* Save the style, or reword it, and the warning clears.\r\n\r\n**Nothing is ever added that you cannot read here.** No setting in the app injects prompt text; a setting changes *how* a request is made, never *what* you asked for.\r\n\r\n### Prompting guide (on the web)\r\nPer-model prompting guidance lives at <https://slates.video/docs/prompting>, linked from the bottom of Settings. It covers every model Slates offers \u2014 Video, Image, Audio \u2014 with what that model reads, what it ignores, and its gotchas, all on one page so you can compare them. Markdown copy for pasting into an LLM: <https://slates.video/docs/prompting.md>.\r\n\r\nIt is generated from the same source the Slates CLI, the MCP server and Studio Agent are built on, so the guide and the app cannot disagree. It is documentation rather than a control, which is why it is a page on the web and not a panel in the app: it has room to be read, a URL you can send someone, and it is always current rather than frozen at the version you installed.\r\n\r\n### @Mentions\r\nType `@` \u2192 auto-complete shows project characters and environments. Type `#` \u2192 shows styles. Selecting inserts the reference image(s). At send time a mention is rewritten to a numbered citation (`@sarah` \u2192 `Sarah (image 1)`) so the model can tell your attachments apart \u2014 you can read the result in \"See what gets sent\". Nothing else about your wording is rewritten. Prompting works the same as any other AI tool; no special syntax beyond @mentions.\r\n\r\n### Writing the shot list with Studio Agent\r\nThere is no \"Enhance\" button and no \"Generate prompts\" button. **Studio Agent does this work**, because it reads the same shot list you do \u2014 every scene, every beat in order, with its references, its model and its price \u2014 and because you can steer it:\r\n\r\n> \"Write the beats for my current storyboard.\"\r\n> \"Now redo scene 3 handheld, and match its energy to scene 2.\"\r\n> \"SHOT-A4 runs long \u2014 split it after 'and then'.\"\r\n\r\nEvery field it writes is editable by hand, in place, in the storyboard's Document view. Open Studio Agent with **Ctrl+.**\r\n\r\n### Debug Panel (advanced)\r\nA developer panel showing the exact request body, with the ability to override the composed prompt before sending. **There is no toggle button for it on the prompt bar in any build** \u2014 open it with **Ctrl+Shift+D**. For ordinary use, \"See what gets sent\" above is the supported way to inspect a prompt.\r\n\r\n---\r\n\r\n## PROJECTS\r\n\r\n### Structure\r\nEach project = folder on your disk. Subdirectories: images/, videos/, audio/, references/, exports/. Location configurable in Settings \u2192 Projects Directory.\r\n\r\n### Assets\r\nEvery generated or imported file is an asset (image, video, audio). Metadata tracked: prompt, model, settings, cost, dimensions, timestamps. Videos track source image via source_asset_id -- you can see all videos generated from any image.\r\n\r\n### Supported File Formats\r\n- **Images:** PNG, JPEG, WEBP, GIF. Note: HEIC/HEIF (iPhone photos) NOT supported -- convert to JPEG/PNG first.\r\n- **Video:** MP4, MOV, WEBM, AVI, MKV.\r\n- **Audio:** MP3, WAV, OGG, M4A, AAC.\r\n- **Clipboard paste:** Any image format the OS clipboard provides (PNG, JPEG, WEBP, GIF). Pasting works both in the gallery and directly into the prompt box; either way the image becomes a real gallery asset in the folder you're working in (tagged \"Imported\") AND, when pasted into the prompt box, attaches as a reference. Anything generated from it links back to it as a source.\r\n- **Drag and drop:** Any file the browser recognizes as image/* or video/*.\r\n\r\n### Operations\r\nCreate/rename/delete projects. Import external files via drag-and-drop or file picker. Paste images from clipboard. Extract still frames from videos. Relocate project to different disk/folder (all paths auto-update). Cleanup orphaned assets.\r\n\r\n### Moving and copying assets between projects\r\nAssets (images, clips) can be sent to another project three ways: the selection band on the Images/Videos tabs, the right-click menu on any card, or by dragging cards and dropping on a project in the drop palette.\r\n\r\n- **Move** relocates the media files on disk into the destination project's folder. The asset leaves whatever gallery folder it was in and is issued a fresh badge code in the destination.\r\n- **Copy** duplicates it \u2014 new files, new thumbnails, new badge code \u2014 and changes nothing in the source project.\r\n\r\n**Why a move can be refused:** an image another project still builds with (a character/environment/style identity image, or a storyboard frame) cannot leave, because the entity left behind would point at a file it no longer owns. When that happens the dialog lists what's blocking and offers the fix: bring the whole character/environment/style across with all of its images, or copy instead. A storyboard frame is only ever offered a copy \u2014 moving its image out would empty the shot.\r\n\r\nRight-clicking a card that is part of a multi-selection acts on the whole selection (\"Move 5 to Project\u2026\"). Right-clicking a card outside the selection acts on that card alone.\r\n\r\n---\r\n\r\n## SHOTS \u2014 THE STORYBOARD IS THE SHOT LIST\r\n\r\nA **Shot** is the prompt bar, saved: the prompt, every reference with the job it carries, the model, every setting \u2014 and now the beat itself: who speaks, what they say, how it is said, what happens, the prop, the framing and the camera. It lives in the **storyboard**, which is the one place Shots are listed. It is never required: the prompt bar works exactly as it always has for anyone who never touches one.\r\n\r\n### Why it exists\r\nA generation's full recipe was already stored, but only once you had paid for it. A Shot can be written **before anything is generated**, so a whole piece can be planned, read, timed, priced and corrected while it is still free. That is the point of the thing: look at the entire ad or short film \u2014 every cheap asset lined up in the actual flow \u2014 before the videos exist.\r\n\r\n### What a Shot holds\r\nRaw prompt (@mentions intact); the model; aspect ratio / duration / resolution / negative prompt / sound and the rest of the bar's settings; every attachment with its ROLE (plain reference, subject, style, reference video, reference audio, first frame, last frame); and the script layer \u2014 `speaker`, `line`, `delivery`, `action`, `prop`, `shotSize`, `camera`, and a `continues` flag for one sentence running across two cuts. Characters, environments and styles are stored as the ENTITY, not a copied picture, so updating a character updates every Shot that names it.\r\n\r\n**The script fields are for reading and counting. Only the prompt is sent to a model.** Dialogue you want a model to perform still goes in the prompt, verbatim, with its delivery \u2014 writing it in `line` makes it readable and lets Slates check whether it fits the cut, not spoken.\r\n\r\n### Every Shot has an address\r\n`SHOT-A1`, `SHOT-A2` \u2026 per project, never reused \u2014 the same idea as the `IMG-A12` badge on a gallery card. Say it to Claude and you are both pointing at the same row. It is for **this session**, not for retrieval later: there is no shot search and no shot library, because a Shot is workspace state \u2014 alive while you build the piece, worthless once it ships.\r\n\r\n### Making one\r\n- **Save as Shot** on any generated image, clip or track's right-click menu restores that generation and keeps it \u2014 and that generation becomes the Shot's first take, so the row opens showing the result you kept it for. **Reuse Prompt** on those same menus does the restore WITHOUT saving anything.\r\n- Claude can write Shots directly (`slates_create_shot`), including for shots whose image does not exist yet \u2014 and can re-chop them with `slates_split_shot` / `slates_merge_shots`.\r\n- **It files itself.** A saved Shot lands in the scene you have open, else the last scene of the storyboard you were most recently working in; if the project has no storyboard, one appears named after the project. Nothing you save is ever somewhere you have to go and find.\r\n\r\n### There is no save button\r\nSelecting a Shot row **binds** the prompt bar to it. Edits write straight back to that row; `Clear` unbinds and returns the bar to free composing. There is no undo, and none is needed: the generations underneath a row are the permanent record of what actually fired, and the row itself is the working copy.\r\n\r\n### Two ways to look at it\r\nThe storyboard has one toggle and two jobs.\r\n\r\n- **Board** \u2014 arrange. One picture per Shot, dragged into the order you want. Drag one and the whole beat moves with it: the line, the references, the model, the settings, the takes. There is nothing else to drag, so there is never a question of what followed what.\r\n- **Document** \u2014 write. One continuous page: the script, the references beside the words that cite them, the prompts underneath. This is where you read the piece before paying for it.\r\n\r\nInside Document, choose **Script** to edit the words with scene headings and speakers. Delivery notes, model and pricing details, and warnings are hidden. **Shots** shows all layers. Open **Custom** to toggle Scene, Action, Character, Delivery, Dialogue, Shot, References, Prompt, Takes, and Warnings independently. Warnings covers missing models or references, unresolved mentions, model changes, and lines too long for their cut. Hiding a layer changes only what you see; it preserves your text and generation checks. There is a text-size slider and an independent toggle for shot numbers in the margin.\n\nDelivery is an optional performance note, not a required label for every line. TTS sends the authored prompt, or the dialogue when the prompt is empty, verbatim; separate Delivery notes are not added. For speech cues, use the selected TTS model's prompting guide and place supported tags in that spoken text. Script view does not strip inline speech cues from your words.\n\r\n### The header tells you what you are about to make\r\n`5 generations \u00B7 7 cuts \u00B7 54s \u00B7 84 credits`, and beneath it a variety strip like `6/7 wide \u00B7 5 push \u00B7 3 cuts in the loft`.\r\n\r\n**Two counts, because they measure different things.** Rhythm is counted in **cuts**; money is counted in **generations**. A multi-shot generation is several cuts inside one paid call, so mixing them would be wrong. A cut with no model chosen has no duration and shows as `\u2014` rather than `0s` \u2014 a runtime that invented seconds would be a lie about the one number this view exists to give.\r\n\r\n### Splitting and merging \u2014 the chop\r\nPut the caret mid-line and press Enter: the row becomes two, the second inheriting the model, settings and references, and marked as continuing the first if the split lands mid-sentence. Select two adjacent rows and press **Merge**: they become one, references combined, durations summed. **The price and the runtime move as you do it** \u2014 which is the whole reason to make the decision here rather than in a document somewhere else.\r\n\r\nSplit is also the move behind a voiceover that keeps talking while the picture hard-cuts to a new world: split at a word boundary and both rows carry one sentence, each with its own visuals.\n\nThe **Dialogue continues from previous shot** toggle is a planning note that the sentence spans a cut. It does not merge shots, join generated audio, or change generation settings. Its pressed state shows whether the note is set; toggling it does not move the script. **Merge these two shots** is the separate control between adjacent shots that actually combines them. In Document \u2192 Custom, **Scene** toggles scene headings and their controls; the dialogue remains visible even if a hidden heading belonged to a collapsed scene.\n\r\n### Variety, counted and never judged\r\nSlates counts what is in front of it \u2014 shot sizes, camera moves, cast, locations, durations, and any of them repeating three or more times in a row \u2014 and shows the counts. **It never changes anything, never suggests anything and never blocks.** `shotSize` and `camera` are free text: write `long-lens CU, other head blurred` if that is the shot. Anything unrecognised counts as \"other\", which is a fine answer.\r\n\r\nIf a spoken line cannot be read in its cut at any plausible pace, the row says so \u2014 and says it only when the line is genuinely impossible, never when it is merely long.\r\n\r\n### The animatic\r\nPress play on any beat and the storyboard plays as a rough cut: each picture held for **its own cut's duration**, with the line underneath. That tells you the rhythm of the finished piece before a single video exists. A multi-shot generation holds one picture across its internal cuts, and says so. A cut with no duration is held for 3 seconds and marked \u2014 the header leaves it out of the runtime for the same reason it is marked here.\r\n\r\n### Things that stay honest rather than being hidden\r\n- **Deleted references.** If an asset or character a Shot points at is gone, the Shot still loads and the row says how many items were left out of what gets sent \u2014 and editing the row does not quietly drop them.\r\n- **A swapped model.** Changing a Shot's model never rewrites your words \u2014 video models genuinely take different prompt grammars, so the row tells you which model the prompt was written for and leaves the sentence alone. The settings line shows what will actually be sent after the swap, and that is the value it prices.\r\n- **Image Shots carry no role badges.** Image generation sends every reference in one undifferentiated list, so an image Shot can remember that a picture is a style reference but cannot tell the model.\r\n- **The thumbnail is never a question.** Slates picks it \u2014 the first frame, else the first reference, else the newest take \u2014 and you can override it from any reference in the gutter.\r\n\r\n### Firing several\r\nSelect Shots and press **Generate all**: one total, the largest single Shot stated separately, one approval. They run **one at a time**. If one of the selected Shots has been deleted, the whole batch is refused and names it \u2014 nothing fires and nothing is billed. If a generation fails mid-run, the rest still fire, the failure is reported per Shot, and **nothing is retried automatically**.\r\n\r\n### Deleting\r\nDeleting a storyboard tells you how many Shots are attached before it does anything, and deletes them with it. It never moves them somewhere else without asking. A Shot that also lives in another storyboard survives.\r\n\r\n---\r\n\r\n## STORYBOARDING\r\n\r\n### Hierarchy\r\nStoryboard \u2192 Scenes \u2192 **Shots**. A scene is an ordered list of Shots, and a Shot is one beat: its picture, its references and their roles, its model and settings, its prompt, its words, and the generations it has produced. See **SHOTS** above \u2014 that section is the storyboard.\r\n\r\n### What happened to frame types\r\nThere used to be a \"frame type\" on each picture \u2014 first / last / ingredient \u2014 plus a separate motion-prompt box. Both were a second, weaker way of saying what a Shot already says: **a reference's role lives on the Shot** (first frame, last frame, subject, style, plain reference), and the motion prompt was just the Shot's prompt under another name. Existing storyboards were converted automatically and nothing was lost. Pick a picture's job on the Shot's reference rail; write the motion in the Shot's prompt.\r\n\r\n### Grid Exploration\r\n2x2 grid: 4 prompt variations for quick iteration. 3x3 grid: 9 variations for deeper exploration. Select individual cells \u2192 extract to full-resolution images. Tip: 2x2 is usually sufficient and produces better quality. 3x3 can occasionally get proportions slightly wrong when upscaling cells because it faithfully reproduces the lower-resolution proportions. Grid exploration runs on Nano Banana 2 only; no other image model offers it.\r\n\r\n### Storyboard \u2192 Video\r\nSelect frames \u2192 generate video for each \u2192 clips auto-insert into timeline in order with source tracking maintained.\r\n\r\n### The animatic\r\nPlay the storyboard as a rough cut. Each beat is held for **its own duration**, with its line underneath, so what you are watching runs at the finished piece's real length. Space = play/pause. Arrow keys = navigate. Escape = exit.\r\n\r\n### Paste a script\r\nPaste a script into the storyboard and it becomes one row per paragraph \u2014 ALL-CAPS cues become speakers, parentheticals become delivery. It is a plain parse, not a model: nothing is invented, nothing is sent anywhere, and prose that is not screenplay-formatted lands as one row per paragraph for you (or Claude) to chop.\r\n\r\n---\r\n\r\n## VIDEO EDITOR (TIMELINE)\r\n\r\n### Tracks\r\nMulti-track: video tracks + audio tracks stacked vertically. Clips independent per track. Add or remove tracks freely -- layer a music bed, a voiceover, and effects on separate audio tracks. Video assets go on video tracks, audio assets on audio tracks. Overlapping video clips resolve top-track-wins.\r\n\r\n### Audio Mixing\r\nEach track has a volume fader, and the timeline has a master output fader for the final mix. Both range from silent to +12 dB of boost, and both apply to preview playback AND the exported MP4 -- what you hear is what you render. Muting a video track silences its embedded audio but still shows the picture. Use the master fader to prevent clipping when stacking loud tracks.\r\n\r\n### Timeline Settings\r\nResolution and frame rate (24/30/60) are auto-managed: the first video clip sets both, and a later higher-resolution clip raises the canvas. All clips are conformed to the timeline frame rate on export. Changing the frame rate after clips are placed retimes them.\r\n\r\n### Clip Properties\r\nSource asset, in/out points (frame-level precision), duration, scale (fit/fill/custom %), position (X/Y offset), opacity (0-100%).\r\n\r\n### Tools\r\n- **Select (V):** Click/drag clips, view/edit properties\r\n- **Razor (C):** Split clip at playhead into two clips\r\n- **Slip (S):** Adjust clip in/out points without moving its position\r\n- **Snap toggle:** Snap to playhead/clip boundaries\r\n\r\n### Markers\r\nColor-coded timeline markers (6+ colors) with optional labels. Use for scene breaks, cue points, notes.\r\n\r\n### Playback & Navigation\r\nSpace = play/pause. Left/Right arrows = frame-by-frame. Up/Down = +-1 second. Page Up/Down = jump by screen width. Home/End = start/end of timeline.\r\n\r\n### Zoom\r\nCtrl+Plus = zoom in (finer precision). Ctrl+Minus = zoom out (see more timeline).\r\n\r\n### Undo/Redo\r\n50-step history. Ctrl+Z = undo. Ctrl+Shift+Z = redo.\r\n\r\n---\r\n\r\n## EXPORT\r\n\r\n### Video Export (FFmpeg)\r\nExport timeline \u2192 MP4 (H.264). Configure: resolution, frame rate, bitrate, output location. All visible tracks rendered, muted tracks excluded, clip in/out points respected. FFmpeg is bundled -- no separate install needed.\r\n\r\n### DaVinci Resolve XML Export\r\nGenerates XML project file containing: clip references (paths to source videos), timeline structure (tracks, clips), clip properties (scale, position, opacity, in/out points), timeline markers.\r\n\r\n**Importing into DaVinci Resolve:** File \u2192 Import \u2192 Timeline. DaVinci reads the XML and reconstructs your timeline with all clips, properties, and markers intact. From there you can color grade and export your final master.\r\n\r\nExports saved to project's exports/ directory with timestamped filenames.\r\n\r\n---\r\n\r\n## CHARACTERS, ENVIRONMENTS & STYLES\r\n\r\n### Characters\r\nCreate character with name + description. Generate character sheet (license required): AI generates a turnaround with multiple angles for consistency. Generate expression sheet: same character with different facial expressions. Use `@character_name` in any prompt to attach reference images.\r\n\r\n**Tips for consistency:** Experiment with character sheet generation using both the existing project style and photorealistic style. Sometimes a single well-chosen image works better than a full sheet -- especially if the character is already in the same style, lighting, and clothing as your project. You can manually assign any image as a character reference instead of generating a sheet.\r\n\r\n**Voice.** A character can carry one voice clip, the same way it carries one identity image \u2014 a shortcut for reusing a voice, never a requirement for speaking in one. **Add voice** / **Change voice** on the card opens the same voice picker the prompt box uses (a menu off the button, not a pop-up; the other cards stay on screen): a clip from the project attaches as it is; a preset or a described voice renders the character speaking a fixed audition line on Inworld TTS-2 and attaches that clip, at the credit cost the picker states first. Right-click the voice card \u2192 **Remove voice** detaches it without deleting the clip. The character's voice then shows under **Clips** in the Voice lane's picker, and mentioning the character (`@name`) in a Seed Audio prompt attaches the clip as a reference, so the scene casts that voice.\r\n\r\n### Environments\r\nCreate environment with name + description. Generate environment grid (license required, 3x3): 9 variations. Extract individual cells to full-resolution images. Use `@environment_name` in prompts.\r\n\r\n**Tip:** Like characters, sometimes a single strong environment image gives better consistency than a grid of 9. Experiment with both approaches.\r\n\r\n### Styles\r\nCreate style with name + description + upload reference image. Use `#style_name` in prompts. Key visual auto-attachment option for consistent look across all frames.\r\n\r\n---\r\n\r\n## SETTINGS\r\n\r\n### Generation\r\nEvery generation runs on Slates Credits \u2014 there are no API keys to configure. The Generate button shows the exact credit cost before each generation, and failed generations refund immediately.\r\n\r\n### Other Settings\r\n- **Projects Directory:** Where project folders live on disk. Changeable anytime.\r\n- **Default Model:** Pre-selected model for new generations. Override per-generation.\r\n- **Default Quality/Resolution:** Pre-selected resolution. Override per-generation.\r\n- **Grid Size:** Default 2x2 or 3x3 for grid exploration.\r\n- **Auto Naming:** Automatically name generated assets.\r\n- **Prompting guide:** A link at the bottom of Settings to <https://slates.video/docs/prompting> \u2014 per-model prompting guidance for every model (see PROMPT SYSTEM above).\r\n\r\n---\r\n\r\n## ACCOUNT & BILLING\r\n\r\n### Login\r\nEmail-only, no password. Enter email \u2192 receive magic link \u2192 click to log in. First login creates account automatically. Session persists across restarts.\r\n\r\n### License\r\nUnlocks: character sheet generation and environment grid generation. Includes 12 months of updates (Slates Pro includes lifetime updates). Major upgrades discounted after.\r\n\r\n### Credits\r\n\r\n<!-- BEGIN:GENERATED credits -->\nCredits are what every generation is paid with. They are pay-as-you-go, they never expire, and the exact cost of a generation is shown on the Generate button before you commit.\n\n- A **Slates Standard** license ($149 one time) starts you with **1,000 credits**.\n- **Slates Pro** ($297 one time, or $97 to upgrade later) starts you with **3,000 credits**.\n\n| Pack | Credits (Standard) | Credits per dollar | Versus the smallest pack |\n|------|--------------------|--------------------|--------------------------|\n| $10 | 250 | 25.0 | standard rate |\n| $25 | 650 | 26.0 | +4% more credits |\n| $50 | 1,375 | 27.5 | +10% more credits |\n| $100 | 3,000 | 30.0 | +20% more credits |\n| $250 | 8,000 | 32.0 | +28% more credits |\n| $500 | 17,000 | 34.0 | +36% more credits |\n| $1,000 | 35,000 | 35.0 | +40% more credits |\n\nPacks up to $500 are open to everyone; the $1,000 pack is offered inside the app to licensed accounts. Slates Pro receives more credits than the Standard column above on every pack, for life.\n<!-- END:GENERATED credits -->\r\n\r\nCredits NEVER expire, there is no monthly reset, and failed generations refund immediately. You can also turn on auto-topup so your balance refills when it runs low.\r\n\r\n### Standard vs Pro\r\n- **Standard:** the app, every AI model, and pay-as-you-go credits that never expire, plus 12 months of updates.\r\n- **Slates Pro:** everything in Standard, plus our lowest credit rate on every pack, forever (buy the smallest pack and pay the largest pack's rate), **4K video generation**, a priority generation queue, early access to every new model on release day, and lifetime updates. The more you top up, the more the better rate adds up.\r\n\r\nEvery AI model is available on both tiers. The only capability gated to Pro is generating 4K video; 4K images are open to everyone, and exporting your timeline at 4K is available on every tier.\r\n\r\n### 30-Day Guarantee\r\nFull refund within 30 days, no questions asked.\r\n\r\n---\r\n\r\n## KEYBOARD SHORTCUTS\r\n\r\n| Key | Action |\r\n|-----|--------|\r\n| Space | Play/pause |\r\n| V | Select tool |\r\n| C | Razor tool |\r\n| S | Slip tool / snap toggle |\r\n| M | Add marker |\r\n| Left/Right | Frame-by-frame |\r\n| Up/Down | Seek +-1 second |\r\n| Ctrl+Z | Undo |\r\n| Ctrl+Shift+Z | Redo |\r\n| Ctrl+Plus/Minus | Zoom timeline |\r\n| Delete/Backspace | Delete selected clip |\r\n| Escape | Close modal/viewer/slideshow |\r\n| Ctrl+Enter | Submit generation |\r\n| Home/End | Jump to timeline start/end |\r\n| Page Up/Down | Jump by screen width |\r\n\r\n---\r\n\r\n## OFFLINE USAGE\r\n\r\nThe app launches and works offline for everything except AI generation and login. Specifically:\r\n\r\n**Works offline:** Opening projects, viewing all assets (images/videos), editing timeline (trim, reorder, split clips), adding markers, slideshow playback, FFmpeg export to MP4, DaVinci XML export.\r\n\r\n**Requires internet:** AI generation (all models), login/signup, credit purchases, credit balance sync, license validation (only checked on first generation attempt per session, then cached), auto-updater.\r\n\r\nIf you lose internet mid-session, you can keep editing and exporting. Generation will fail until connectivity returns.\r\n\r\n---\r\n\r\n## GENERATION RECOVERY\r\n\r\nIf the app closes during a generation: on restart, Slates detects in-flight jobs, polls the AI provider, and downloads completed results automatically. Nothing is lost. Recovering generations show at 5% in the queue until status is confirmed. Works for every model.\r\n\r\n---\r\n\r\n## TROUBLESHOOTING\r\n\r\n**\"Insufficient credits\"** \u2014 Your credit balance is too low for this generation. Buy more credits in the app (packs from $10 to $500) or turn on auto-topup. The exact cost of any generation is shown on the Generate button before you commit.\r\n\r\n**\"Input was rejected by Kling\"** \u2014 Image may not meet quality requirements (character visibility, proportions, content policy). Try a different image or prompt.\r\n\r\n**\"Failed to upload image to FAL CDN\"** \u2014 Network issue during reference image upload. Check internet connection, retry.\r\n\r\n**\"Generation failed\" / \"Proxy generation failed\"** \u2014 Generic error from the AI provider. Usually temporary. Retry. If persistent, try a different model.\r\n\r\n**\"Source asset not found\" / \"Source video asset not found\" / \"Target image asset not found\"** \u2014 The image or video you're trying to use was deleted or moved. Re-import or select a different asset.\r\n\r\n**\"Invalid audio source\"** \u2014 Lip sync: either enter TTS text or upload an audio file. One is required.\r\n\r\n**\"TTS response missing audio URL\"** \u2014 Text-to-speech failed during lip sync. Retry.\r\n\r\n**Generation stuck** \u2014 Restart app. Recovery system polls providers and picks up where it left off.\r\n\r\n**API rate limit** \u2014 Too many requests (limit: 20 generations/minute). Wait 1-2 minutes, retry.\r\n\r\n**Project files missing** \u2014 Project folder was moved/deleted outside the app. Use project relocation in Settings to re-point to the correct folder.\r\n\r\n**License shows \"revoked\"** \u2014 Contact support. Character sheets and environment grids unavailable until resolved.\r\n\r\n**Session expired** \u2014 Magic link session timed out. Log in again via Settings.\r\n\r\n**iPhone photos won't import** \u2014 iPhones save photos as HEIC/HEIF format, which Slates doesn't support. Convert to JPEG or PNG first (most photo apps and online converters can do this).\r\n\r\n---\r\n\r\n## PRIVACY & DATA\r\n\r\n- Generated files stay on YOUR machine. Slates servers never store your videos/images.\r\n- No prompts logged server-side.\r\n- File uploads go directly to the AI provider via pre-signed URLs. Slates servers never buffer your media.\r\n- Server stores only: email, license status, credit balance, transaction history, session tokens.\r\n- Stripe handles all payment data. Slates never sees your card number.\r\n\r\n---\r\n\r\n## SYSTEM REQUIREMENTS\r\n\r\n- Windows 10/11 or macOS 12+\r\n- Internet connection required for AI generation (not for editing/exporting)\r\n- Disk space for project files (AI videos are typically 5-50MB each)\r\n- FFmpeg bundled with app (no separate install needed)\r\n- No GPU required (all AI processing happens in the cloud)\r\n\r\n---\r\n\r\n## COMMON TASKS (STEP-BY-STEP)\r\n\r\n### Generate an Image\r\n1. Open the floating prompt box (visible on every page).\r\n2. Enter your prompt describing the image.\r\n3. Select an image model (Nano Banana 2 recommended).\r\n4. Choose aspect ratio and resolution.\r\n5. Press Ctrl+Enter or click Generate.\r\n\r\n### Generate Video From an Image\r\n1. In the prompt box, attach a start image.\r\n2. Write a prompt describing the desired motion/action.\r\n3. Select a video model (Kling V3.0 Omni in ingredients mode recommended).\r\n4. Choose duration (Kling bills per second from 3s up, so shorter is always cheaper), aspect ratio, and resolution.\r\n5. Click Generate.\r\n\r\n### Use a Character Reference for Consistency\r\n1. Create a character in your project (name + description).\r\n2. Either generate a character sheet OR manually assign a single image as the character reference.\r\n3. In the prompt box, type `@` and select your character from auto-complete.\r\n4. The reference image is attached automatically. Generate normally.\r\n\r\n### Export to DaVinci Resolve for Color Grading\r\n1. In the video editor, finalize your timeline (clips, markers, timing).\r\n2. Click Export \u2192 DaVinci Resolve XML.\r\n3. Choose output location. File saves to exports/ directory.\r\n4. In DaVinci Resolve: File \u2192 Import \u2192 Timeline. Select the XML file.\r\n5. Your timeline loads with all clips, properties, and markers intact. Grade and export.\r\n\r\n### Extract a Still Frame From a Video\r\n1. Hover over any video clip in the gallery.\r\n2. Camera icon = extract **current frame**. Dropdown arrow next to it = **First frame** or **Last frame**.\r\n3. Extracted image saves to your project gallery. Use as start/end image for I2V, character reference, or storyboard frame.\r\n\r\nKey workflow: extract a clip's last frame \u2192 use it as the start image for the next generation \u2192 seamless visual continuity between scenes.\r\n\r\n### Buy More Credits\r\n1. Open Settings \u2192 Credits, or the credit badge in the top nav.\r\n2. Pick a pack (packs run from $10 to $500, plus a $1,000 pack offered in the app to licensed accounts; bigger packs give more credits per dollar).\r\n3. Pay via Stripe. Credits are added to your balance instantly and never expire.\r\n4. Optional: turn on auto-topup so your balance refills automatically when it runs low.\r\n\r\n---\r\n\r\n## COMMON QUESTIONS\r\n\r\n**Q: Which model should I use for most videos?**\r\nA: Seedance 2.0 is the default and the one to reach for when physics, scale, effects or hero shots matter. For everyday shots built from a start image, Kling V3.0 Omni in ingredients mode is the best balance of cost and quality, which is why the step-by-step guides above use it.\r\n\r\n**Q: What's the best image model?**\r\nA: GPT Image 2.5 is the strongest image model in the app, and the best available anywhere right now. It holds a long instruction more faithfully than anything else here, and it is the only model that renders words in the picture reliably. The two seats cost the same, so the only trade is time: **Flare** is the fast one and is best for drafts and exploring, **Sunburst** is the higher-quality one and is what finals, hero frames and reference-heavy edits should end up on. The quality knob has five settings spanning about 36\u00D7 from cheapest to dearest \u2014 `max` is four times `high`, though the smaller steps are uneven: **`medium`** is for drafts and where iteration belongs, **`high` is the default** and where most finished work should sit, and **`max` is the best output you can get** \u2014 finished frames, client deliverables, anything with exact text. `xhigh` sits just under `max` for about half the price and is worth trying before you jump to the top. Going past `high` should be a deliberate choice rather than a habit. Go to **4K only once you already know your references and your prompt are solid** \u2014 it is the most expensive setting on the model and the worst place to discover the composition was wrong. Prove the shot at 3K first, then re-run the settled prompt at 4K. If you used GPT Image 2 before, note the quality names all shifted by one: its `medium` is now `high`, its `high` is now `max`. Nano Banana 2 is the picker's DEFAULT rather than the best: it is fast, takes the most reference images, and is the only model with grid exploration. Nano Banana Pro is the hero-frame step up when composition, cinematic lighting and skin have to be perfect. Prices for every tier are in the model table above.\r\n\r\n**Q: Do I need to set up API keys?**\r\nA: No. There are no API keys in Slates \u2014 every generation runs on Slates Credits, which come with your license and never expire.\r\n\r\n**Q: How much does a generation cost?**\r\nA: It depends on the model, resolution, and length. The exact credit cost is always shown on the Generate button before you commit, so there are no surprises.\r\n\r\n**Q: Can I use Slates offline?**\r\nA: Yes for viewing projects, editing timeline, and exporting. No for AI generation -- that requires internet.\r\n\r\n**Q: Do credits expire?**\r\nA: No. Credits never expire.\r\n\r\n**Q: What happens if I close the app during a generation?**\r\nA: Nothing is lost. On restart, Slates detects in-flight jobs and downloads completed results automatically.\r\n\r\n---\r\n\r\n## FEATURES NOT IN SLATES\r\n\r\nThe following are NOT available. Do not suggest them:\r\n\r\n- Bring-your-own API keys (BYOK) \u2014 every generation runs on Slates Credits; there is no key-entry option\r\n- Local/on-device GPU inference (all AI runs in the cloud)\r\n- Built-in music generation (use external tools like Suno, import audio)\r\n- A voice picker for Seed Audio \u2014 you describe the voice you want in words instead. The preset voice shelf belongs to the Voice lane (Inworld TTS-2), and Kling Lip-Sync keeps its own small fixed list: six English/UK voices plus a storyteller, with a speed control\r\n- Automatic video editing from a script\r\n- A prompt \"Enhance\" button \u2014 ask Studio Agent to rewrite a prompt instead\r\n- A settings/gear panel on the prompt box \u2014 every parameter is a dropdown on the bar\r\n- Cloud project storage (all files are local)\r\n- Real-time collaboration / multi-user editing\r\n- Mobile app (desktop only: Windows and macOS)\r\n- HEIC/HEIF image import (convert to JPEG/PNG first)\r\n- Storyboard JSON export (import only)\r\n\r\n---\r\n\r\n## VERSION\r\n\r\n<!-- BEGIN:GENERATED version -->\nSlates Reference Version: 1.5.6\nLast Updated: 2026-09-09\n\nThis document is generated. Its source of truth is `slate/docs/slates-llm-manual.md`; its model tables and credit costs are derived from the Slates model registry and pricing tables at build time, so they cannot be typed by hand.\n\nIf the user asks about a feature not documented here, it may have been added after this version. The current copy is always at <https://slates.video/slates-reference.md>.\n<!-- END:GENERATED version -->\r\n\r\nIf this document didn't answer your question, email hello@slates.video so we can help and improve the app.\r\n\r\n</slates_reference>\r\n";
|
|
1
|
+
export declare const APP_MANUAL = "# Slates \u2014 Complete Reference for AI Assistants\r\n\r\n<!-- The heading above must stay first: shipped app builds reject this file if it\r\n does not begin with one. See slate/CLAUDE.md; check:llm-docs enforces it. -->\r\n\r\n<system_role>\r\nYou are a support assistant for the Slates desktop application. Answer user questions using ONLY the information in the <slates_reference> below. Be concise and direct. Use numbered steps for procedures. Use bullet points for explanations when helpful.\r\n</system_role>\r\n\r\n<rules>\r\n- If the answer cannot be found in the <slates_reference>, say: \"That isn't covered in the Slates reference.\" Do not guess or invent features.\r\n- If the user asks how to do something, give step-by-step instructions from the workflows and features described here.\r\n- If the user reports an error, check the TROUBLESHOOTING section first.\r\n- Refer to the FEATURES NOT IN SLATES section before answering questions about capabilities that might not exist.\r\n- Quote the exact error message when referencing troubleshooting entries.\r\n</rules>\r\n\r\n<slates_reference>\r\n\r\n<!-- BEGIN:GENERATED header -->\n# SLATES v1.5.6 \u2014 Complete Reference\n\n> **Freshness.** Generated from the Slates source of truth for app version **1.5.6**, last changed **2026-09-09**. The canonical copy of this file is <https://slates.video/slates-reference.md>. If a model, price or feature the user mentions is missing below, this copy is out of date: re-fetch that URL before answering, and say so.\n<!-- END:GENERATED header -->\r\n\r\n**What is Slates?** Desktop app (Windows 10/11, macOS 12+) for AI image and video creation. One-time purchase, no subscription. Every license includes 1,000 free credits, and Slates Pro starts with 3,000 credits. Every generation runs on Slates Credits \u2014 there are no API keys to set up, and credits never expire.\r\n\r\n---\r\n\r\n## INTENDED WORKFLOW\r\n\r\nThe designed start-to-finish flow:\r\n\r\n1. **Create project** \u2014 New project with name/description. Creates folder on your disk.\r\n2. **Build visual assets** \u2014 Generate images, create characters (with character sheets for consistency), environments (with environment grids), and styles. This is your visual library.\r\n3. **Create storyboard** \u2014 Add scenes. Every picture you drop in becomes a **Shot**: one beat of the piece, holding its references, its model, its settings and (when you want them) its words.\r\n4. **Write the piece** \u2014 Switch the storyboard to **Document** and write. What is said, what happens, the framing, the prompt each beat will send. Nothing is required and nothing is asked for; a visuals-only piece is finished as it stands.\r\n5. **Read it before you pay for it** \u2014 The header states how many generations, how many cuts, how long it runs and what it will cost. Press play for a rough cut at the real timing. Re-chop with split and merge and watch the price move.\r\n6. **Generate** \u2014 Select the beats you want and fire them in one approved batch. Each result lands under the row that made it.\r\n7. **Organize** \u2014 Switch to **Board** to re-order. Drag a beat and the whole thing moves with it.\r\n8. **Export to timeline** \u2014 Send clips to the built-in multi-track video editor.\r\n9. **Edit** \u2014 Trim, reorder, add markers, adjust timing.\r\n10. **Final export** \u2014 Export to MP4 directly, or export DaVinci Resolve XML for professional color grading.\r\n\r\n---\r\n\r\n## MODEL REFERENCE TABLE\r\n\r\n**Generating 4K video is a Slates Pro feature** \u2014 every tier generates video up to 1080p, and 4K images are open to everyone. Exporting your finished timeline at 4K is available on every tier.\r\n\r\nEvery model runs on Slates Credits. **The exact credit cost appears on the Generate button before anything fires.** The tables below are generated from the app's own model registry and rate tables, so they describe exactly what the model picker offers in this version: aspect ratios, resolutions, durations, reference-image limits, and the credit price of each. Bigger credit packs lower your per-credit cost, and Slates Pro gets the best pack rate on every purchase.\r\n\r\n<!-- BEGIN:GENERATED model-tables -->\n### Image Models\n\n| Model | Aspect Ratios | Resolutions | Max Refs | Credits per image |\n|-------|--------------|-------------|----------|-------------------|\n| **GPT Image 2.5 Flare** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 \u00B7 3K 3 \u00B7 4K 5 (default quality; at max: 2K 8 \u00B7 3K 11 \u00B7 4K 20) |\n| **GPT Image 2.5 Sunburst** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 \u00B7 3K 3 \u00B7 4K 5 (default quality; at max: 2K 8 \u00B7 3K 11 \u00B7 4K 20) |\n| **Nano Banana 2** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 4 \u00B7 2K 6 \u00B7 4K 8 |\n| **NB2 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K | 4 | 1K 2 |\n| **Nano Banana Pro** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 8 \u00B7 2K 8 \u00B7 4K 15 |\n| **FLUX.2 Max** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 4 | 1K 4 \u00B7 2K 5 \u00B7 4K 8 |\n| **Seedream 5 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 2K / 3K / 4K | 10 | 2K 2 \u00B7 3K 2 \u00B7 4K 2 |\n\n### Video Models\n\n| Model | Duration | Aspect Ratios | Resolutions | Max Refs | Audio | Credits per second |\n|-------|----------|--------------|-------------|----------|-------|--------------------|\n| **Seedance 2.0** | 4-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p / 4K | 9 | Included | 480p 3.5 \u00B7 720p 7.5 \u00B7 1080p 18.5 \u00B7 4K 39 |\n| **Seedance 2.5** | 4-30s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 30 | Included | 480p 5.1 \u00B7 720p 11.6 \u00B7 1080p 20.5 |\n| **Seedance 2.5 Edit** | Follows the source clip (4-30s) | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 0 | Included | 480p 6.2 \u00B7 720p 13.9 \u00B7 1080p 24.6 |\n| **Kling V3.0 Standard** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (6.3 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (8.4 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Omni** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (5.6 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Omni Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (7 with audio) \u00B7 4K 21 |\n| **Kling O3 Edit** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 6.3 |\n| **Kling O3 Edit Pro** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 8.4 |\n| **MiniMax H3 Max** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 1080p | 9 | Included | 480p 2.5 \u00B7 768p 4 \u00B7 1080p 8 |\n| **MiniMax H3** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 2K / 4K | 9 | Included | 480p 2.5 \u00B7 768p 3 \u00B7 2K 6.5 \u00B7 4K 8 |\n| **Gemini Omni Flash** | 3-10s | 16:9, 9:16 | 720p | 7 | Included | 720p 6.4 |\n| **Omni Flash Edit** | Follows the source clip (3-10s) | 16:9, 9:16 | 720p | 0 | Included | 720p 6.4 |\n| **LTX-2.5** | 6/8/10/12/14/16/18/20s at 720p/1080p; 6/8/10s at 1440p/4K | 16:9, 9:16 | 720p / 1080p / 1440p / 4K | 0 | Included | 720p 4.5 \u00B7 1080p 6.5 \u00B7 1440p 9.5 \u00B7 4K 15 |\n| **LTX-2.5 Pro** | 6, 8, 10s | 16:9, 9:16 | 720p / 1080p | 0 | Included | 720p 6 \u00B7 1080p 8.5 |\n| **Veo 3.1 Fast** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 5 (7.5 with audio) \u00B7 1080p 5 (7.5 with audio) \u00B7 4K 15 (17.5 with audio) |\n| **Veo 3.1 Standard** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 10 (20 with audio) \u00B7 1080p 10 (20 with audio) \u00B7 4K 20 (30 with audio) |\n\n**Seedance 2.0 \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 27.4 credits per second at 1080p.\n**Seedance 2.5 \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 16.3 credits per second at 720p.\n**Seedance 2.5 Edit \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 19.8 credits per second at 720p.\n\n### Audio Models\n\n| Model | Length | Credits |\n|-------|--------|---------|\n| **Seed Audio 1.0** | 3-120s | 1 at 3s \u00B7 3 at 15s \u00B7 19 at 120s |\n| **Inworld TTS-2** | up to 2,000 characters of text per take | 1 at 250 characters \u00B7 2 at 2,000 characters (billed per 250) |\n| **Sound Effects** | 1-22s | 1 at 1s \u00B7 1 at 4s \u00B7 3 at 22s |\n\n### Tools (Lip Sync, Motion Transfer)\n\nThese are real Kling endpoints that take a clip or a still as their subject, not models you prompt from scratch. Both bill in 5-second blocks.\n\n| Tool | Input | Billed in | Credits per block |\n|------|-------|-----------|-------------------|\n| **Kling Lip Sync** | Video source | 5s block | 4 |\n| **Kling Lip Sync (Avatar v2 Standard)** | Still-image source | 5s block | 14 |\n| **Kling Lip Sync (Avatar v2 Pro)** | Still-image source | 5s block | 29 |\n| **Kling Motion Control Standard** | Still image + reference video | 5s block | 32 |\n| **Kling Motion Control Pro** | Still image + reference video | 5s block | 42 |\n<!-- END:GENERATED model-tables -->\r\n\r\n---\r\n\r\n## WHICH MODEL TO USE\r\n\r\n### Images\r\n\r\n**Nano Banana 2 is the default image model.** It is the best all-round image model in the app: the most reference images of any image model, every aspect ratio, and output up to 4K. Brief it like a creative director rather than with tag soup. It is also the only model that supports the 2x2 / 3x3 grid exploration wrapper.\r\n\r\n- **NB2 Lite** is the fast, cheap draft seat in the same Nano Banana family. Roughly half the price of NB2 full and noticeably faster, 1K output only. Iterate here, finish on NB2.\r\n- **Nano Banana Pro** is the hero-frame and typography tier. Reach for it when spatial composition, cinematic lighting and skin, or fine in-image type have to be perfect. NB2 gets you most of the way there, so this is a deliberate step up, never a default.\r\n- **GPT Image 2.5** is the strongest image model in the app: it follows a long instruction more faithfully than anything else here, and it is the one to pick when the picture simply has to be right. It is also the sharp-text model, which is what makes it the choice for character sheets, shot grids, ordered panels and anything with words in the picture. It comes in two seats that cost exactly the same, and the difference is speed against quality. **Flare** is the fast one: OpenAI describes its quality as comparable to the older GPT Image 2, at roughly half the wait. **Sunburst** is OpenAI's most capable image model, better than GPT Image 2, and deliberately slower. Use Flare while you are still exploring, then re-run the shot you like on Sunburst for the final \u2014 and reach for Sunburst directly when several reference images all have to survive into one frame, or when an edit must change one region and leave identity, geometry and lighting untouched.\n Its quality knob has **five** settings \u2014 `low`, `medium`, `high`, `xhigh`, `max` \u2014 spanning about 36\u00D7 from cheapest to dearest, which makes it the biggest cost lever on the model. The steps are uneven rather than a constant multiplier: `max` is four times `high`, but `xhigh` is only about 1.8 times it. **`high` is the default and the everyday setting.** `medium` is for drafts and is cheap enough to iterate on freely. `max` is the ceiling, for finished frames and exact character-level text; `xhigh` sits just under it for about half the price and is worth trying first. Go past `high` deliberately, not by habit.\n 4K is worth it only once your references and prompt are already settled: prove the shot at 3K, then re-run the finished prompt at 4K. Iterating at 4K is the most common way to waste credits on this model. Its 4K tier is API-only, so even a paid ChatGPT account cannot render it. Note that its resolution tiers are **not** a price ladder: the pixel classes are token-priced by OpenAI, so the cheapest seat is not the smallest one. Read the prices in the table above rather than assuming.\n **If you have used GPT Image 2 before, the quality names all shifted by one.** What it called `medium` is now called `high`, and what it called `high` is now `max` \u2014 the same pictures at the same prices, renamed. A remembered setting will quietly buy you a cheaper tier than it used to.\r\n **It is the only image model that can give you a transparent background.** Set **Background** to *Transparent* on the prompt bar and you get a real alpha channel \u2014 a cut-out for a logo, sticker or overlay \u2014 rather than a painted-in backdrop. *Auto* is the default and lets the model decide from your prompt; *Opaque* forces a filled background. It costs nothing either way. Slates always saves PNG, which is what carries the transparency, so there is nothing else to set.\r\n **Square and 4:3 frames cost more than 16:9 on this model, and only on this model.** OpenAI charges by image tokens rather than by pixels, and a square frame uses about 1.8 times the tokens of a 16:9 frame the same size, and 4:3 or 3:4 about 1.37 times. The credit prices in the table above are the 16:9 numbers; pick 1:1 or 4:3 and the price on the Generate button goes up to match. 9:16 costs the same as 16:9. Every other image model charges the same whatever the shape.\r\n- **FLUX.2 Max** and **Seedream 5 Lite** are the less content-restricted options. Seedream is flat-priced at every resolution it offers, so there is no reason to pick a lower one. Both auto-route to their edit endpoint when you attach reference images.\r\n\r\n### Video\r\n\r\n**Seedance 2.0 is the default video model.** Reach for it the moment physics, effects, destruction or scale matter, and for hero shots. It takes many reference images, generates native audio at no extra cost, and is the only Seedance seat that reaches 4K (generating 4K video needs Slates Pro; timeline export at 4K does not). It is also the cheaper of the two seats at every resolution they share. A face in a reference image routes it to a different provider, which is why **Seedance 2.0 \u00B7 Face** is its own row in the model picker at its own price.\r\n\r\n- **Seedance 2.5 is a second seat, not an upgrade.** It buys much longer single takes and far more reference images, plus better prompt adherence. What it gives up is 4K, and it costs more than 2.0 at every resolution the two share \u2014 so 2.0 stays the model for 4K, and for the same resolution at a lower price. Because 2.5 runs longer, a long clip on 2.5 can cost more than a shorter, higher-resolution one on 2.0 \u2014 read the Generate button, not the resolution. **Seedance 2.5 Edit** is its clip-editing row: attach a clip, describe the change, and the output length follows the source.\r\n- **Kling** is the cost-effective workhorse and the most flexible family: strong start-frame adherence for identity, layout and text, acting, dialogue, multi-shot (up to 6 cuts), and the widest range of clip lengths. **Kling V3.0 Omni** adds multi-character dialogue in English, Chinese, Japanese, Korean and Spanish. Standard and Pro are the same model at two fidelity and price tiers. **Kling O3 Edit** takes an existing clip and changes what you describe, with subject and style reference images, while the original audio is preserved verbatim. Kling is also the only engine behind the Lip Sync and Motion Control tools.\r\n- **MiniMax H3** is the seat to pick when the SOUND is part of what you are writing. Every other video model treats audio as a switch; H3 takes it as three separate instructions in one prompt \u2014 the lines and action sounds tied to a moment, the ambience running underneath, and a score only the audience hears \u2014 and generates all of it with the picture in a single pass. It is also the only model where you say how much of a reference should survive, including moving one subject's characteristic onto a different subject. It runs 5-15 seconds at 480p, 768p, 2K or 4K, and takes up to nine reference images plus reference video and audio. Two things to watch: **the first five reference images are free and every one after that costs extra**, so attach what the shot needs rather than the maximum; and 2K and 4K are upscales of a 768p render rather than larger generations \u2014 in our own testing the 2K pass showed more artifacting than the 768p original it was built from, at more than twice the price. Generate and judge at 768p; step up only when a delivery spec demands the pixels.\r\n- **MiniMax H3 Max** is the same model post-trained by fal for SPEED, and it is the more expensive seat, not the cheaper one. It is dramatically faster: on the same 5-second 768p prompt it finished in about 5 seconds against about 57 seconds for H3 \u2014 roughly 12x (measured 2026-08-27). It stops at 768p and costs more per second than H3 at the resolution they share. It still animates a start frame and an end frame, so image-to-video works normally; what it does not have is the reference set \u2014 the extra identity, style and environment images plus reference video and audio that base H3 reads. Pick it when a fast turnaround on a text-to-video or start-frame shot is worth paying for; pick H3 for resolution, references, or the same tier at a lower price.\r\n- **LTX-2.5** is the VOLUME seat \u2014 the cheapest native 1080p second in the catalogue, with synchronised audio included free at every resolution, so it is the model to reach for when the job is many takes rather than one hero shot. Two things are unique to it. It makes the LONGEST clips of anything here, up to 20 seconds, and it is the only model that reaches 1440p. It is also the only one with native MULTISHOT: a single generation can carry two to four connected shots that hold the character, lighting and voice across the cuts, which everywhere else means generating separate clips and watching identity drift between them. Its constraints are unusually sharp, though. Durations are EVEN NUMBERS ONLY starting at six \u2014 6, 8, 10, 12, 14, 16, 18, 20, with no 5-second or 7-second clip \u2014 and above 1080p that ceiling drops to 10 seconds. Aspect ratios are 16:9 and 9:16 only. And it takes FRAMES, not references: a start frame and an optional end frame that generates a transition between them, but no identity, style or environment reference images at all, so cross-shot character consistency belongs on MiniMax H3 or Kling. Because sound is generated in the same pass, write the audio into the prompt and anchor every cue to something visible \u2014 anything unanchored gets invented.\r\n- **LTX-2.5 Pro** is the fidelity seat of that pair, and it is NOT simply a better LTX. It renders the picture with more compute on busy frames, but on a narrower envelope than the base row: 720p and 1080p only (no 1440p, no 4K) and 6, 8 or 10 seconds only, for about a third more per second. Reaching for it because the name says Pro costs more AND takes away the reach. Pick it when a specific shot needs the extra fidelity and fits inside 1080p and ten seconds; pick base LTX for length, resolution and volume.\r\n- **Gemini Omni Flash** is the cheap 720p seat with native synced audio included in one pass. **Omni Flash Edit** is the prompt-only clip editor: no reference images, one short instruction plus \"Keep everything else the same.\" Long descriptive prompts destroy it.\r\n- **Veo 3.1** is niche and is never a default. Pick it only when you specifically want Google's audio pass. It has the fewest aspect ratios and reference slots of any video model, fixed durations, and the highest per-clip cost.\r\n\r\nBoth edit models take an existing clip as their canvas, so their output length follows the source clip rather than a duration you choose.\r\n\r\n### Audio\r\n\r\nAudio is a third media type alongside images and video \u2014 generated as its own asset, shown in the gallery's **Audio** tab, and dragged onto an audio track in the timeline. This is separate from the audio some VIDEO models generate *inside* a clip (see AUDIO IN GENERATION below): use a video model when the sound must be locked to what is on screen, and these when you need audio you can move, trim, re-use, or layer.\r\n\r\n**Seed Audio 1.0 is the default.** A room with dialogue *and* clatter *and* ambience is one generation, not three layered ones, and because it is cheap you can run five takes and keep the best. It makes a whole audio SCENE from one plain sentence. **It has no length setting of its own** \u2014 Slates writes your chosen duration into the prompt, and that is what you are charged. You describe the voice in words; there is no voice list to pick from. Set **Languages** to Mixed if one scene needs more than one language (it costs the same).\r\n\r\n**Sound Effects** makes one effect, or a seamless loop. It is the only surface with an exact duration, so an effect can land on a specific frame. Describe the physical cause (\"heavy oak door slams shut in a stone hallway\"), not the label (\"door sound\"). **Loop** makes it seamless for beds; **Wording** controls how literally your description is followed. Seed Audio is actually the better tool for *long* ambience beds, so the two are not redundant in the direction you would expect.\r\n\r\nKling's `SFX:` / `Ambient noise:` prompt syntax belongs to video prompts and makes Seed Audio results *worse* \u2014 write plain sentences there instead.\r\n\r\n**Inworld TTS-2 is the voice seat** \u2014 type the words, pick a voice, press Generate. In the prompt box's Audio lane pick **Voice** in the model picker; the prompt is the exact text that gets spoken (nothing is added or rewritten \u2014 open \"See what gets sent\" to confirm), and the **Voice** control on the bar opens the voice picker: **Presets** (ready-made voices with gender, accent and age filters \u2014 every one plays the same audition line, so you compare voices rather than scripts), **Clips** (any character's voice, or any audio clip in the project, cloned for the take), or **Describe** (a voice in words). The character counter beside the bar is the bill: the generated audio table above gives the text cap and billing buckets, and the Generate button shows the price. Direction goes in square brackets (`[whispering] \u2026`) \u2014 anything in parentheses is read aloud. Cloning a real person's voice needs their permission. Right-click any audio clip in the Audio tab \u2192 **Use as voice** lands you on the Voice lane with that clip as the voice. Studio Agent and the MCP/CLI do the same through `slates_generate_audio` (a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId`, or a `voiceDescription`).\r\n\r\n**There is no music generation.** For a song, use an external tool and import the audio (see PROJECTS \u2192 Supported File Formats). For spoken lines inside a scene, let Seed Audio perform them, put them in the video prompt on a model with native audio (Seedance, Kling Omni, Omni Flash, Veo), or use Kling Lip-Sync's text-to-speech against a shot.\r\n\r\n### Tools\r\n\r\n**Tools** is not a model family. It is two real Kling endpoints that take a clip or a still as their subject: **Kling Lip Sync** and **Kling Motion Control**. Both bill in 5-second blocks and are described under GENERATION MODES below.\r\n\r\n**How pricing works:** every generation is priced in Slates Credits, and the exact cost is shown on the Generate button before you commit. Bigger credit packs give more credits per dollar; Slates Pro locks in the best pack rate on every purchase, forever.\r\n\r\n---\r\n\r\n## GENERATION MODES\r\n\r\n### Create Image\r\nPrompt \u2192 select image model \u2192 set aspect ratio + resolution \u2192 generate. Batch grids available for quick iteration.\r\n\r\n### Text-to-Video\r\nPrompt \u2192 select video model \u2192 set duration + aspect ratio + resolution \u2192 generate. Output: MP4.\r\n\r\n### Image-to-Video (I2V)\r\nAttach start image + prompt \u2192 select model \u2192 generate video from that image. Optional: attach end image (Veo) for guided transitions.\r\n\r\n### Ingredients / References\r\nUse @character_name, @environment_name, or #style_name in prompt to attach reference images for visual consistency. Kling: up to 4 total references. Veo: up to 3. Nano Banana 2: up to 14. The @mentions auto-complete from your project's characters/environments/styles.\r\n\r\n### Lip Sync\r\n**Kling only.** Pick **Kling Lip Sync** under the Tools family in the model picker. Source: video or still image.\r\n\r\n- **Audio source** \u2014 Text to speech (type the line; six English/UK voices plus a storyteller, with a speed control) OR Upload audio (bring your own recording, max 5MB).\r\n- **Avatar tier** \u2014 only appears for a still-image source: Avatar v2 Standard (the value tier) or Avatar v2 Pro (higher fidelity, higher rate).\r\n- 5s output blocks.\r\n\r\nWorks very well with human-like characters. Less reliable with animals or non-human characters.\r\n\r\n### Motion Transfer\r\n**Kling only.** Pick **Kling Motion Control** under the Tools family. Target: still image (your character). Source: reference video (the motion).\r\n\r\n- **Engine** \u2014 Kling MC Standard (value tier) or Kling MC Pro (higher fidelity, higher rate).\r\n- **Orientation** \u2014 *Match video* copies skeleton and depth from the clip (best for dancing, walking, full-body action; driving clips up to 30s). *Match image* keeps your character's pose and angle and uses the video only as motion hints (best for close-ups; up to 10s).\r\n- 5s output.\r\n\r\n> **Note:** these two tools used to offer a second \"Seedance 2.0\" engine. It was not a separate engine \u2014 picking it made Slates write a sentence into your prompt that you never saw, which is no longer allowed anywhere in the app (see \"What gets sent\" below). Both tools are now Kling endpoints only.\r\n\r\n### Edit Image\r\nRight-click any image asset \u2192 open viewer \u2192 switch to Edit Mode. Enter an edit prompt describing the changes you want. Select edit model: Nano Banana 2 (supports up to 14 reference images), FLUX.2 Max, or Seedream 5 Lite. Choose resolution and aspect ratio. The result saves as a new asset with the original preserved. Useful for refining generated images without starting from scratch.\r\n\r\n### Edit Video (Kling O3 Edit / Omni Flash Edit)\r\nRight-click any video clip (gallery or timeline) \u2192 \"Edit with AI\". The clip attaches to the prompt box as the source; describe the CHANGE, not the whole scene (\"replace the man with @marcus\", \"make it a rainy night, keep everything else\"). Two engines in the model picker:\r\n- **Kling O3 Edit (default):** attach subject images (role: Subject) to swap someone in, or style images (role: Style) for a look \u2014 max 4 combined refs. Clips 3-15s. Original audio preserved.\r\n- **Omni Flash Edit (cheapest):** prompt only \u2014 no reference images; keep instructions simple and add \"Keep everything else the same.\" Clips 3-10s, 720p output.\r\n\r\nOutput length follows the source clip; the credit cost (clip seconds, rounded up, at the per-second rate) shows on the Generate button. The edited clip saves as a NEW asset linked to the original \u2014 chain edits freely. Trim longer clips on the timeline first.\r\n\r\n### Multi-Shot (Kling V3.0/Omni)\r\nEnable multi-shot toggle \u2192 multiple scene prompts in one generation, each with different framing. 6-axis camera controls per shot. Results can be hit-or-miss, but worth trying for quick multi-cut sequences. For more reliable results, most users prefer generating multiple short 5s clips separately using Kling V3.0 Omni in ingredients mode and assembling them on the timeline.\r\n\r\n---\r\n\r\n### Generate Audio\r\n\r\nSwitch the prompt box's lane pill from Image/Video to **Audio**, pick a surface, and generate. The prompt box offers three: **Seed Audio 1.0**, **Voice** (Inworld TTS-2) and **Sound Effects**. The result lands in the gallery's Audio tab as its own asset with a waveform and an inline player, and can be dragged onto an audio track in the timeline.\r\n\r\n- **Scene (Seed Audio 1.0)** \u2014 one plain sentence describing the moment. Set **Length**; Slates writes it into the prompt for you and that is exactly what you're billed for (open \"See what gets sent\" under the prompt box to read the appended text). **Say the crowd/room size out loud** \u2014 \"applause\" returns a full auditorium when you meant three people at an open mic. Ask for a few seconds more than the clip needs so the edit has fade handles. Describe the voice you want in the sentence itself (\"a weary dock foreman in his fifties, gravel in his voice\") \u2014 there is no voice picker.\r\n- **Voice (Inworld TTS-2)** \u2014 the prompt is the words to be spoken, verbatim. Pick the voice with the **Voice** control on the bar (presets you can play first, any clip in the project, a character's voice, or a description); the character counter is the bill. See MODELS \u2192 Inworld TTS-2 for direction tags and the cloning rules.\r\n- **Sound Effect** \u2014 describe the physical cause and set the length to roughly the event (\u22481s for an impact, 2\u20134s for a whoosh, 8\u201322s + **Loop** for a bed). **Wording** sets how literally the description is followed: Interpretive, Balanced (default), or Literal.\n\n#### Use your own voice recording\n\nIn the bottom prompt box, choose **Audio**, then **Inworld TTS-2** in the model picker. Open **Voice \u2192 Clips \u2192 Import voice clip** and select your recording. The import adds an audio asset to this project without generating anything. Click its play button to audition it, then click the recording's name to choose it. Type the words you want spoken in the prompt box and press **Generate**, which shows the price. The new take appears in **Gallery \u2192 Audio**.\n\nUse a clean recording of one speaker whose voice you have permission to use. Slates clones the recording for each take; there is no separate training wizard or persistent vendor voice to manage. **Presets** lets you audition ready-made voices; **Describe** lets you write a voice description and choose **Use this description**. Choosing in the prompt box sets up the next take; only Generate spends credits.\n\nTo attach your recording or a generated take to a character, open **Gallery \u2192 Characters**, then **Add voice** (or **Change voice**) on that character's card. Choose **Clips** and click the clip's name. Attaching an existing clip is free. The card displays its waveform and player. That character's voice is also listed under **Voice \u2192 Clips \u2192 Characters** in the prompt box. Selecting a preset or description from a character card generates and attaches a take; read the cost shown in that picker before choosing.\n\nFor **Inworld TTS-2**, choose the character under **Voice \u2192 Clips \u2192 Characters**. Typing `@name` does not choose a TTS voice: the prompt is spoken verbatim. Automatic voice attachment from a character mention belongs to **Seed Audio** only.\n\r\n---\r\n\r\n## AUDIO IN GENERATION\r\n\r\nThis section is about audio generated **inside a video clip**. For audio as its own asset, see Generate Audio above.\r\n\r\n**Veo 3.1 native audio:** Generates audio WITH video. Prompt syntax: `\"Hello!\"` for dialogue, `SFX: [sound]` for effects, `Ambient noise: [description]` for ambience. Max 10s dialogue. Add `(no subtitles)` to suppress text overlays.\r\n\r\n**Kling V3.0 Omni dialogue:** Multi-character dialogue with distinct voices. Languages: EN, ZH, JA, KO, ES. `Background music: [description]` for music. Max 10s dialogue.\r\n\r\n**Kling V3.0 sound co-generation:** Synchronized sound effects generated with video.\r\n\r\n\u26A0\uFE0F **This prompt syntax is video-only.** `SFX:`, `Ambient noise:` and `Background music:` are Kling/Veo conventions \u2014 the audio models above have no parser for them and will treat them as words in the scene.\r\n\r\n---\r\n\r\n## PROMPT SYSTEM\r\n\r\n### Unified Create Surface (roles + model-on-button)\r\nThe old mode tabs (text-to-video / frames-to-video / ingredients / create-image) are ONE \"Create\" surface. An **Image | Video | Audio pill** on the prompt bar switches your output lane \u2014 it remembers and restores the last model you used in each lane (pick Seedance once and the Video lane stays Seedance until you change it). Only the active lane shows its name; the other two are icons. Every attachment in the reference tray carries a tappable ROLE badge \u2014 Reference / First frame / Last frame / Subject / Style \u2014 you say what each attachment is; nothing is inferred. Your model choice sticks across generations and workflow actions (\"use as first frame\" keeps your chosen video model). Attaching a video via \"Edit with AI\" flips the surface into Edit Video mode. Lip Sync and Motion Transfer live under the **Tools** family in the model picker.\r\n\r\n### Floating Prompt Box \u2014 the bar holds everything\r\nPersistent across all pages. **There is no settings panel and no gear button.** Everything sits on one bottom bar, left to right:\r\n\r\n1. **Media toggle** \u2014 Image / Video / Audio.\r\n2. **Model picker** \u2014 a searchable menu plus a detached submenu. The main menu lists model families with a vendor glyph tile each; picking one opens that family's models beside it, every row carrying capability chips (resolution, clip length, audio, references) and its per-unit rate. Type to search across every model. The submenu is anchored to the row you opened it from, so it never travels. The trigger on the bar shows the model name and nothing else \u2014 no chevron, no resolution appended. Everything listed is a real model or endpoint.\r\n3. **Parameter controls** \u2014 one per setting the chosen model actually has (resolution, aspect, duration, length, quality, count, grid, face-in-reference, audio, loop, and so on). The trigger shows the current value; the explanation lives *inside* the menu as a subtitle under each option, along with what that option costs. A setting with only one possible value still shows, muted and non-interactive, so the row never changes shape.\r\n4. **Sliders for ranges.** A setting with a long list of steps (video duration, audio length) opens a ruler instead of a many-row menu. The handle moves between the values the model actually declares, so it cannot land on one the model will not accept, and the price for the selected value is shown on the ruler.\r\n5. **`More \u25BE`** \u2014 if the model has more parameters than fit the current window width, the extras are collected into a generated `More` dropdown automatically. Widen the window (or close the Studio Agent panel) and they move back onto the bar.\r\n6. **Generate** \u2014 reads `Generate \u00B7 <cost>`. **Cost only** \u2014 the model name is not repeated on the button (it's in the model picker) and there is no send arrow. A badge on the left of the button counts generations currently running.\r\n\r\nFor text-to-speech, the character counter is always visible because text length determines the price. On other surfaces it appears near the right of the bar after roughly three quarters of the model's prompt limit. It turns red over the limit.\r\n\r\n**Collapsing:** the chevron at the top-right of the box collapses it to a single arrow \u2014 nothing else. Click the arrow to bring it back.\r\n\r\nBelow the bar the queue shows pending/active generations with cost and progress.\r\n\r\n### What gets sent (prompt transparency)\r\nUnder the prompt box is a **\"See what gets sent\"** disclosure. Open it and you see the exact text that will be transmitted, produced by the same code that builds the request \u2014 so it can never disagree with what is actually sent. It shows:\r\n\r\n- **Reference numbering** \u2014 `@sarah` becomes `Sarah (image 1)` so the model knows which attached image is which.\r\n- **Key lines for attachments you did NOT mention** \u2014 one short neutral sentence per unmentioned attachment. Mention every reference in your own words and these never generate.\r\n- **The trailing style clause** when a `#style` is attached.\r\n- **The Seed Audio duration append** \u2014 the `\u2026 N seconds` Slates adds to the end of the prompt, which is also what you are billed for.\r\n- **Grid wrapping** when 2\u00D72 or 3\u00D73 is on.\r\n\r\nThe row stays hidden when the composed prompt is identical to what you typed, so it only appears when there is something to show.\r\n\r\n**Unresolved `#tags` and `@mentions` are named, not silently dropped.** If you type `#noir` and there is no saved style called \"noir\", the tag is removed from the text sent to the model (a raw tag confuses every model) \u2014 but the disclosure turns red and says so by name: *\"#noir matches nothing saved \u2014 removed from what gets sent.\"* Save the style, or reword it, and the warning clears.\r\n\r\n**Nothing is ever added that you cannot read here.** No setting in the app injects prompt text; a setting changes *how* a request is made, never *what* you asked for.\r\n\r\n### Prompting guide (on the web)\r\nPer-model prompting guidance lives at <https://slates.video/docs/prompting>, linked from the bottom of Settings. It covers every model Slates offers \u2014 Video, Image, Audio \u2014 with what that model reads, what it ignores, and its gotchas, all on one page so you can compare them. Markdown copy for pasting into an LLM: <https://slates.video/docs/prompting.md>.\r\n\r\nIt is generated from the same source the Slates CLI, the MCP server and Studio Agent are built on, so the guide and the app cannot disagree. It is documentation rather than a control, which is why it is a page on the web and not a panel in the app: it has room to be read, a URL you can send someone, and it is always current rather than frozen at the version you installed.\r\n\r\n### @Mentions\r\nType `@` \u2192 auto-complete shows project characters and environments. Type `#` \u2192 shows styles. Selecting inserts the reference image(s). At send time a mention is rewritten to a numbered citation (`@sarah` \u2192 `Sarah (image 1)`) so the model can tell your attachments apart \u2014 you can read the result in \"See what gets sent\". Nothing else about your wording is rewritten. Prompting works the same as any other AI tool; no special syntax beyond @mentions.\r\n\r\n### Writing the shot list with Studio Agent\r\nThere is no \"Enhance\" button and no \"Generate prompts\" button. **Studio Agent does this work**, because it reads the same shot list you do \u2014 every scene, every beat in order, with its references, its model and its price \u2014 and because you can steer it:\r\n\r\n> \"Write the beats for my current storyboard.\"\r\n> \"Now redo scene 3 handheld, and match its energy to scene 2.\"\r\n> \"SHOT-A4 runs long \u2014 split it after 'and then'.\"\r\n\r\nEvery field it writes is editable by hand, in place, in the storyboard's Document view. Open Studio Agent with **Ctrl+.**\r\n\r\n### Debug Panel (advanced)\r\nA developer panel showing the exact request body, with the ability to override the composed prompt before sending. **There is no toggle button for it on the prompt bar in any build** \u2014 open it with **Ctrl+Shift+D**. For ordinary use, \"See what gets sent\" above is the supported way to inspect a prompt.\r\n\r\n---\r\n\r\n## PROJECTS\r\n\r\n### Structure\r\nEach project = folder on your disk. Subdirectories: images/, videos/, audio/, references/, exports/. Location configurable in Settings \u2192 Projects Directory.\r\n\r\n### Assets\r\nEvery generated or imported file is an asset (image, video, audio). Metadata tracked: prompt, model, settings, cost, dimensions, timestamps. Videos track source image via source_asset_id -- you can see all videos generated from any image.\r\n\r\n### Supported File Formats\r\n- **Images:** PNG, JPEG, WEBP, GIF. Note: HEIC/HEIF (iPhone photos) NOT supported -- convert to JPEG/PNG first.\r\n- **Video:** MP4, MOV, WEBM, AVI, MKV.\r\n- **Audio:** MP3, WAV, OGG, M4A, AAC.\r\n- **Clipboard paste:** Any image format the OS clipboard provides (PNG, JPEG, WEBP, GIF). Pasting works both in the gallery and directly into the prompt box; either way the image becomes a real gallery asset in the folder you're working in (tagged \"Imported\") AND, when pasted into the prompt box, attaches as a reference. Anything generated from it links back to it as a source.\r\n- **Drag and drop:** Any file the browser recognizes as image/* or video/*.\r\n\r\n### Operations\r\nCreate/rename/delete projects. Import external files via drag-and-drop or file picker. Paste images from clipboard. Extract still frames from videos. Relocate project to different disk/folder (all paths auto-update). Cleanup orphaned assets.\r\n\r\n### Moving and copying assets between projects\r\nAssets (images, clips) can be sent to another project three ways: the selection band on the Images/Videos tabs, the right-click menu on any card, or by dragging cards and dropping on a project in the drop palette.\r\n\r\n- **Move** relocates the media files on disk into the destination project's folder. The asset leaves whatever gallery folder it was in and is issued a fresh badge code in the destination.\r\n- **Copy** duplicates it \u2014 new files, new thumbnails, new badge code \u2014 and changes nothing in the source project.\r\n\r\n**Why a move can be refused:** an image another project still builds with (a character/environment/style identity image, or a storyboard frame) cannot leave, because the entity left behind would point at a file it no longer owns. When that happens the dialog lists what's blocking and offers the fix: bring the whole character/environment/style across with all of its images, or copy instead. A storyboard frame is only ever offered a copy \u2014 moving its image out would empty the shot.\r\n\r\nRight-clicking a card that is part of a multi-selection acts on the whole selection (\"Move 5 to Project\u2026\"). Right-clicking a card outside the selection acts on that card alone.\r\n\r\n---\r\n\r\n## SHOTS \u2014 THE STORYBOARD IS THE SHOT LIST\r\n\r\nA **Shot** is the prompt bar, saved: the prompt, every reference with the job it carries, the model, every setting \u2014 and now the beat itself: who speaks, what they say, how it is said, what happens, the prop, the framing and the camera. It lives in the **storyboard**, which is the one place Shots are listed. It is never required: the prompt bar works exactly as it always has for anyone who never touches one.\r\n\r\n### Why it exists\r\nA generation's full recipe was already stored, but only once you had paid for it. A Shot can be written **before anything is generated**, so a whole piece can be planned, read, timed, priced and corrected while it is still free. That is the point of the thing: look at the entire ad or short film \u2014 every cheap asset lined up in the actual flow \u2014 before the videos exist.\r\n\r\n### What a Shot holds\r\nRaw prompt (@mentions intact); the model; aspect ratio / duration / resolution / negative prompt / sound and the rest of the bar's settings; every attachment with its ROLE (plain reference, subject, style, reference video, reference audio, first frame, last frame); and the script layer \u2014 `speaker`, `line`, `delivery`, `action`, `prop`, `shotSize`, `camera`, and a `continues` flag for one sentence running across two cuts. Characters, environments and styles are stored as the ENTITY, not a copied picture, so updating a character updates every Shot that names it.\r\n\r\n**The script fields are for reading and counting. Only the prompt is sent to a model.** Dialogue you want a model to perform still goes in the prompt, verbatim, with its delivery \u2014 writing it in `line` makes it readable and lets Slates check whether it fits the cut, not spoken.\r\n\r\n### Every Shot has an address\r\n`SHOT-A1`, `SHOT-A2` \u2026 per project, never reused \u2014 the same idea as the `IMG-A12` badge on a gallery card. Say it to Claude and you are both pointing at the same row. It is for **this session**, not for retrieval later: there is no shot search and no shot library, because a Shot is workspace state \u2014 alive while you build the piece, worthless once it ships.\r\n\r\n### Making one\r\n- **Save as Shot** on any generated image, clip or track's right-click menu restores that generation and keeps it \u2014 and that generation becomes the Shot's first take, so the row opens showing the result you kept it for. **Reuse Prompt** on those same menus does the restore WITHOUT saving anything.\r\n- Claude can write Shots directly (`slates_create_shot`), including for shots whose image does not exist yet \u2014 and can re-chop them with `slates_split_shot` / `slates_merge_shots`.\r\n- **It files itself.** A saved Shot lands in the scene you have open, else the last scene of the storyboard you were most recently working in; if the project has no storyboard, one appears named after the project. Nothing you save is ever somewhere you have to go and find.\r\n\r\n### There is no save button\r\nSelecting a Shot row **binds** the prompt bar to it. Edits write straight back to that row; `Clear` unbinds and returns the bar to free composing. There is no undo, and none is needed: the generations underneath a row are the permanent record of what actually fired, and the row itself is the working copy.\r\n\r\n### Two ways to look at it\r\nThe storyboard has one toggle and two jobs.\r\n\r\n- **Board** \u2014 arrange. One picture per Shot, dragged into the order you want. Drag one and the whole beat moves with it: the line, the references, the model, the settings, the takes. There is nothing else to drag, so there is never a question of what followed what.\r\n- **Document** \u2014 write. One continuous page: the script, the references beside the words that cite them, the prompts underneath. This is where you read the piece before paying for it.\r\n\r\nInside Document, choose **Script** to edit the words with scene headings and speakers. Delivery notes, model and pricing details, and warnings are hidden. **Shots** shows all layers. Open **Custom** to toggle Scene, Action, Character, Delivery, Dialogue, Shot, References, Prompt, Takes, and Warnings independently. Warnings covers missing models or references, unresolved mentions, model changes, and lines too long for their cut. Hiding a layer changes only what you see; it preserves your text and generation checks. There is a text-size slider and an independent toggle for shot numbers in the margin.\n\nDelivery is an optional performance note, not a required label for every line. TTS sends the authored prompt, or the dialogue when the prompt is empty, verbatim; separate Delivery notes are not added. For speech cues, use the selected TTS model's prompting guide and place supported tags in that spoken text. Script view does not strip inline speech cues from your words.\n\r\n### The header tells you what you are about to make\r\n`5 generations \u00B7 7 cuts \u00B7 54s \u00B7 84 credits`, and beneath it a variety strip like `6/7 wide \u00B7 5 push \u00B7 3 cuts in the loft`.\r\n\r\n**Two counts, because they measure different things.** Rhythm is counted in **cuts**; money is counted in **generations**. A multi-shot generation is several cuts inside one paid call, so mixing them would be wrong. A cut with no model chosen has no duration and shows as `\u2014` rather than `0s` \u2014 a runtime that invented seconds would be a lie about the one number this view exists to give.\r\n\r\n### Splitting and merging \u2014 the chop\r\nPut the caret mid-line and press Enter: the row becomes two, the second inheriting the model, settings and references, and marked as continuing the first if the split lands mid-sentence. Select two adjacent rows and press **Merge**: they become one, references combined, durations summed. **The price and the runtime move as you do it** \u2014 which is the whole reason to make the decision here rather than in a document somewhere else.\r\n\r\nSplit is also the move behind a voiceover that keeps talking while the picture hard-cuts to a new world: split at a word boundary and both rows carry one sentence, each with its own visuals.\n\nThe **Dialogue continues from previous shot** toggle is a planning note that the sentence spans a cut. It does not merge shots, join generated audio, or change generation settings. Its pressed state shows whether the note is set; toggling it does not move the script. **Merge these two shots** is the separate control between adjacent shots that actually combines them. In Document \u2192 Custom, **Scene** toggles scene headings and their controls; the dialogue remains visible even if a hidden heading belonged to a collapsed scene.\n\r\n### Variety, counted and never judged\r\nSlates counts what is in front of it \u2014 shot sizes, camera moves, cast, locations, durations, and any of them repeating three or more times in a row \u2014 and shows the counts. **It never changes anything, never suggests anything and never blocks.** `shotSize` and `camera` are free text: write `long-lens CU, other head blurred` if that is the shot. Anything unrecognised counts as \"other\", which is a fine answer.\r\n\r\nIf a spoken line cannot be read in its cut at any plausible pace, the row says so \u2014 and says it only when the line is genuinely impossible, never when it is merely long.\r\n\r\n### The animatic\r\nPress play on any beat and the storyboard plays as a rough cut: each picture held for **its own cut's duration**, with the line underneath. That tells you the rhythm of the finished piece before a single video exists. A multi-shot generation holds one picture across its internal cuts, and says so. A cut with no duration is held for 3 seconds and marked \u2014 the header leaves it out of the runtime for the same reason it is marked here.\r\n\r\n### Things that stay honest rather than being hidden\r\n- **Deleted references.** If an asset or character a Shot points at is gone, the Shot still loads and the row says how many items were left out of what gets sent \u2014 and editing the row does not quietly drop them.\r\n- **A swapped model.** Changing a Shot's model never rewrites your words \u2014 video models genuinely take different prompt grammars, so the row tells you which model the prompt was written for and leaves the sentence alone. The settings line shows what will actually be sent after the swap, and that is the value it prices.\r\n- **Image Shots carry no role badges.** Image generation sends every reference in one undifferentiated list, so an image Shot can remember that a picture is a style reference but cannot tell the model.\r\n- **The thumbnail is never a question.** Slates picks it \u2014 the first frame, else the first reference, else the newest take \u2014 and you can override it from any reference in the gutter.\r\n\r\n### Firing several\r\nSelect Shots and press **Generate all**: one total, the largest single Shot stated separately, one approval. They run **one at a time**. If one of the selected Shots has been deleted, the whole batch is refused and names it \u2014 nothing fires and nothing is billed. If a generation fails mid-run, the rest still fire, the failure is reported per Shot, and **nothing is retried automatically**.\r\n\r\n### Deleting\r\nDeleting a storyboard tells you how many Shots are attached before it does anything, and deletes them with it. It never moves them somewhere else without asking. A Shot that also lives in another storyboard survives.\r\n\r\n---\r\n\r\n## STORYBOARDING\r\n\r\n### Hierarchy\r\nStoryboard \u2192 Scenes \u2192 **Shots**. A scene is an ordered list of Shots, and a Shot is one beat: its picture, its references and their roles, its model and settings, its prompt, its words, and the generations it has produced. See **SHOTS** above \u2014 that section is the storyboard.\r\n\r\n### What happened to frame types\r\nThere used to be a \"frame type\" on each picture \u2014 first / last / ingredient \u2014 plus a separate motion-prompt box. Both were a second, weaker way of saying what a Shot already says: **a reference's role lives on the Shot** (first frame, last frame, subject, style, plain reference), and the motion prompt was just the Shot's prompt under another name. Existing storyboards were converted automatically and nothing was lost. Pick a picture's job on the Shot's reference rail; write the motion in the Shot's prompt.\r\n\r\n### Grid Exploration\r\n2x2 grid: 4 prompt variations for quick iteration. 3x3 grid: 9 variations for deeper exploration. Select individual cells \u2192 extract to full-resolution images. Tip: 2x2 is usually sufficient and produces better quality. 3x3 can occasionally get proportions slightly wrong when upscaling cells because it faithfully reproduces the lower-resolution proportions. Grid exploration runs on Nano Banana 2 only; no other image model offers it.\r\n\r\n### Storyboard \u2192 Video\r\nSelect frames \u2192 generate video for each \u2192 clips auto-insert into timeline in order with source tracking maintained.\r\n\r\n### The animatic\r\nPlay the storyboard as a rough cut. Each beat is held for **its own duration**, with its line underneath, so what you are watching runs at the finished piece's real length. Space = play/pause. Arrow keys = navigate. Escape = exit.\r\n\r\n### Paste a script\r\nPaste a script into the storyboard and it becomes one row per paragraph \u2014 ALL-CAPS cues become speakers, parentheticals become delivery. It is a plain parse, not a model: nothing is invented, nothing is sent anywhere, and prose that is not screenplay-formatted lands as one row per paragraph for you (or Claude) to chop.\r\n\r\n---\r\n\r\n## VIDEO EDITOR (TIMELINE)\r\n\r\n### Tracks\r\nMulti-track: video tracks + audio tracks stacked vertically. Clips independent per track. Add or remove tracks freely -- layer a music bed, a voiceover, and effects on separate audio tracks. Video assets go on video tracks, audio assets on audio tracks. Overlapping video clips resolve top-track-wins.\r\n\r\n### Audio Mixing\r\nEach track has a volume fader, and the timeline has a master output fader for the final mix. Both range from silent to +12 dB of boost, and both apply to preview playback AND the exported MP4 -- what you hear is what you render. Muting a video track silences its embedded audio but still shows the picture. Use the master fader to prevent clipping when stacking loud tracks.\r\n\r\n### Timeline Settings\r\nResolution and frame rate (24/30/60) are auto-managed: the first video clip sets both, and a later higher-resolution clip raises the canvas. All clips are conformed to the timeline frame rate on export. Changing the frame rate after clips are placed retimes them.\r\n\r\n### Clip Properties\r\nSource asset, in/out points (frame-level precision), duration, scale (fit/fill/custom %), position (X/Y offset), opacity (0-100%).\r\n\r\n### Tools\r\n- **Select (V):** Click/drag clips, view/edit properties\r\n- **Razor (C):** Split clip at playhead into two clips\r\n- **Slip (S):** Adjust clip in/out points without moving its position\r\n- **Snap toggle:** Snap to playhead/clip boundaries\r\n\r\n### Markers\r\nColor-coded timeline markers (6+ colors) with optional labels. Use for scene breaks, cue points, notes.\r\n\r\n### Playback & Navigation\r\nSpace = play/pause. Left/Right arrows = frame-by-frame. Up/Down = +-1 second. Page Up/Down = jump by screen width. Home/End = start/end of timeline.\r\n\r\n### Zoom\r\nCtrl+Plus = zoom in (finer precision). Ctrl+Minus = zoom out (see more timeline).\r\n\r\n### Undo/Redo\r\n50-step history. Ctrl+Z = undo. Ctrl+Shift+Z = redo.\r\n\r\n---\r\n\r\n## EXPORT\r\n\r\n### Video Export (FFmpeg)\r\nExport timeline \u2192 MP4 (H.264). Configure: resolution, frame rate, bitrate, output location. All visible tracks rendered, muted tracks excluded, clip in/out points respected. FFmpeg is bundled -- no separate install needed.\r\n\r\n### DaVinci Resolve XML Export\r\nGenerates XML project file containing: clip references (paths to source videos), timeline structure (tracks, clips), clip properties (scale, position, opacity, in/out points), timeline markers.\r\n\r\n**Importing into DaVinci Resolve:** File \u2192 Import \u2192 Timeline. DaVinci reads the XML and reconstructs your timeline with all clips, properties, and markers intact. From there you can color grade and export your final master.\r\n\r\nExports saved to project's exports/ directory with timestamped filenames.\r\n\r\n---\r\n\r\n## CHARACTERS, ENVIRONMENTS & STYLES\r\n\r\n### Characters\r\nCreate character with name + description. Generate character sheet (license required): AI generates a turnaround with multiple angles for consistency. Generate expression sheet: same character with different facial expressions. Use `@character_name` in any prompt to attach reference images.\r\n\r\n**Tips for consistency:** Experiment with character sheet generation using both the existing project style and photorealistic style. Sometimes a single well-chosen image works better than a full sheet -- especially if the character is already in the same style, lighting, and clothing as your project. You can manually assign any image as a character reference instead of generating a sheet.\r\n\r\n**Voice.** A character can carry one voice clip, the same way it carries one identity image \u2014 a shortcut for reusing a voice, never a requirement for speaking in one. **Add voice** / **Change voice** on the card opens the same voice picker the prompt box uses (a menu off the button, not a pop-up; the other cards stay on screen): a clip from the project attaches as it is; a preset or a described voice renders the character speaking a fixed audition line on Inworld TTS-2 and attaches that clip, at the credit cost the picker states first. Right-click the voice card \u2192 **Remove voice** detaches it without deleting the clip. The character's voice then shows under **Clips** in the Voice lane's picker, and mentioning the character (`@name`) in a Seed Audio prompt attaches the clip as a reference, so the scene casts that voice.\r\n\r\n### Environments\r\nCreate environment with name + description. Generate environment grid (license required, 3x3): 9 variations. Extract individual cells to full-resolution images. Use `@environment_name` in prompts.\r\n\r\n**Tip:** Like characters, sometimes a single strong environment image gives better consistency than a grid of 9. Experiment with both approaches.\r\n\r\n### Styles\r\nCreate style with name + description + upload reference image. Use `#style_name` in prompts. Key visual auto-attachment option for consistent look across all frames.\r\n\r\n---\r\n\r\n## SETTINGS\r\n\r\n### Generation\r\nEvery generation runs on Slates Credits \u2014 there are no API keys to configure. The Generate button shows the exact credit cost before each generation, and failed generations refund immediately.\r\n\r\n### Other Settings\r\n- **Projects Directory:** Where project folders live on disk. Changeable anytime.\r\n- **Default Model:** Pre-selected model for new generations. Override per-generation.\r\n- **Default Quality/Resolution:** Pre-selected resolution. Override per-generation.\r\n- **Grid Size:** Default 2x2 or 3x3 for grid exploration.\r\n- **Auto Naming:** Automatically name generated assets.\r\n- **Prompting guide:** A link at the bottom of Settings to <https://slates.video/docs/prompting> \u2014 per-model prompting guidance for every model (see PROMPT SYSTEM above).\r\n\r\n---\r\n\r\n## ACCOUNT & BILLING\r\n\r\n### Login\r\nEmail-only, no password. Enter email \u2192 receive magic link \u2192 click to log in. First login creates account automatically. Session persists across restarts.\r\n\r\n### License\r\nUnlocks: character sheet generation and environment grid generation. Includes 12 months of updates (Slates Pro includes lifetime updates). Major upgrades discounted after.\r\n\r\n### Credits\r\n\r\n<!-- BEGIN:GENERATED credits -->\nCredits are what every generation is paid with. They are pay-as-you-go, they never expire, and the exact cost of a generation is shown on the Generate button before you commit.\n\n- A **Slates Standard** license ($149 one time) starts you with **1,000 credits**.\n- **Slates Pro** ($297 one time, or $97 to upgrade later) starts you with **3,000 credits**.\n\n| Pack | Credits (Standard) | Credits per dollar | Versus the smallest pack |\n|------|--------------------|--------------------|--------------------------|\n| $10 | 250 | 25.0 | standard rate |\n| $25 | 650 | 26.0 | +4% more credits |\n| $50 | 1,375 | 27.5 | +10% more credits |\n| $100 | 3,000 | 30.0 | +20% more credits |\n| $250 | 8,000 | 32.0 | +28% more credits |\n| $500 | 17,000 | 34.0 | +36% more credits |\n| $1,000 | 35,000 | 35.0 | +40% more credits |\n\nPacks up to $500 are open to everyone; the $1,000 pack is offered inside the app to licensed accounts. Slates Pro receives more credits than the Standard column above on every pack, for life.\n<!-- END:GENERATED credits -->\r\n\r\nCredits NEVER expire, there is no monthly reset, and failed generations refund immediately. You can also turn on auto-topup so your balance refills when it runs low.\r\n\r\n### Standard vs Pro\r\n- **Standard:** the app, every AI model, and pay-as-you-go credits that never expire, plus 12 months of updates.\r\n- **Slates Pro:** everything in Standard, plus our lowest credit rate on every pack, forever (buy the smallest pack and pay the largest pack's rate), **4K video generation**, a priority generation queue, early access to every new model on release day, and lifetime updates. The more you top up, the more the better rate adds up.\r\n\r\nEvery AI model is available on both tiers. The only capability gated to Pro is generating 4K video; 4K images are open to everyone, and exporting your timeline at 4K is available on every tier.\r\n\r\n### 30-Day Guarantee\r\nFull refund within 30 days, no questions asked.\r\n\r\n---\r\n\r\n## KEYBOARD SHORTCUTS\r\n\r\n| Key | Action |\r\n|-----|--------|\r\n| Space | Play/pause |\r\n| V | Select tool |\r\n| C | Razor tool |\r\n| S | Slip tool / snap toggle |\r\n| M | Add marker |\r\n| Left/Right | Frame-by-frame |\r\n| Up/Down | Seek +-1 second |\r\n| Ctrl+Z | Undo |\r\n| Ctrl+Shift+Z | Redo |\r\n| Ctrl+Plus/Minus | Zoom timeline |\r\n| Delete/Backspace | Delete selected clip |\r\n| Escape | Close modal/viewer/slideshow |\r\n| Ctrl+Enter | Submit generation |\r\n| Home/End | Jump to timeline start/end |\r\n| Page Up/Down | Jump by screen width |\r\n\r\n---\r\n\r\n## OFFLINE USAGE\r\n\r\nThe app launches and works offline for everything except AI generation and login. Specifically:\r\n\r\n**Works offline:** Opening projects, viewing all assets (images/videos), editing timeline (trim, reorder, split clips), adding markers, slideshow playback, FFmpeg export to MP4, DaVinci XML export.\r\n\r\n**Requires internet:** AI generation (all models), login/signup, credit purchases, credit balance sync, license validation (only checked on first generation attempt per session, then cached), auto-updater.\r\n\r\nIf you lose internet mid-session, you can keep editing and exporting. Generation will fail until connectivity returns.\r\n\r\n---\r\n\r\n## GENERATION RECOVERY\r\n\r\nIf the app closes during a generation: on restart, Slates detects in-flight jobs, polls the AI provider, and downloads completed results automatically. Nothing is lost. Recovering generations show at 5% in the queue until status is confirmed. Works for every model.\r\n\r\n---\r\n\r\n## TROUBLESHOOTING\r\n\r\n**\"Insufficient credits\"** \u2014 Your credit balance is too low for this generation. Buy more credits in the app (packs from $10 to $500) or turn on auto-topup. The exact cost of any generation is shown on the Generate button before you commit.\r\n\r\n**\"Input was rejected by Kling\"** \u2014 Image may not meet quality requirements (character visibility, proportions, content policy). Try a different image or prompt.\r\n\r\n**\"Failed to upload image to FAL CDN\"** \u2014 Network issue during reference image upload. Check internet connection, retry.\r\n\r\n**\"Generation failed\" / \"Proxy generation failed\"** \u2014 Generic error from the AI provider. Usually temporary. Retry. If persistent, try a different model.\r\n\r\n**\"Source asset not found\" / \"Source video asset not found\" / \"Target image asset not found\"** \u2014 The image or video you're trying to use was deleted or moved. Re-import or select a different asset.\r\n\r\n**\"Invalid audio source\"** \u2014 Lip sync: either enter TTS text or upload an audio file. One is required.\r\n\r\n**\"TTS response missing audio URL\"** \u2014 Text-to-speech failed during lip sync. Retry.\r\n\r\n**Generation stuck** \u2014 Restart app. Recovery system polls providers and picks up where it left off.\r\n\r\n**API rate limit** \u2014 Too many requests (limit: 20 generations/minute). Wait 1-2 minutes, retry.\r\n\r\n**Project files missing** \u2014 Project folder was moved/deleted outside the app. Use project relocation in Settings to re-point to the correct folder.\r\n\r\n**License shows \"revoked\"** \u2014 Contact support. Character sheets and environment grids unavailable until resolved.\r\n\r\n**Session expired** \u2014 Magic link session timed out. Log in again via Settings.\r\n\r\n**iPhone photos won't import** \u2014 iPhones save photos as HEIC/HEIF format, which Slates doesn't support. Convert to JPEG or PNG first (most photo apps and online converters can do this).\r\n\r\n---\r\n\r\n## PRIVACY & DATA\r\n\r\n- Generated files stay on YOUR machine. Slates servers never store your videos/images.\r\n- No prompts logged server-side.\r\n- File uploads go directly to the AI provider via pre-signed URLs. Slates servers never buffer your media.\r\n- Server stores only: email, license status, credit balance, transaction history, session tokens.\r\n- Stripe handles all payment data. Slates never sees your card number.\r\n\r\n---\r\n\r\n## SYSTEM REQUIREMENTS\r\n\r\n- Windows 10/11 or macOS 12+\r\n- Internet connection required for AI generation (not for editing/exporting)\r\n- Disk space for project files (AI videos are typically 5-50MB each)\r\n- FFmpeg bundled with app (no separate install needed)\r\n- No GPU required (all AI processing happens in the cloud)\r\n\r\n---\r\n\r\n## COMMON TASKS (STEP-BY-STEP)\r\n\r\n### Generate an Image\r\n1. Open the floating prompt box (visible on every page).\r\n2. Enter your prompt describing the image.\r\n3. Select an image model (Nano Banana 2 recommended).\r\n4. Choose aspect ratio and resolution.\r\n5. Press Ctrl+Enter or click Generate.\r\n\r\n### Generate Video From an Image\r\n1. In the prompt box, attach a start image.\r\n2. Write a prompt describing the desired motion/action.\r\n3. Select a video model (Kling V3.0 Omni in ingredients mode recommended).\r\n4. Choose duration (Kling bills per second from 3s up, so shorter is always cheaper), aspect ratio, and resolution.\r\n5. Click Generate.\r\n\r\n### Use a Character Reference for Consistency\r\n1. Create a character in your project (name + description).\r\n2. Either generate a character sheet OR manually assign a single image as the character reference.\r\n3. In the prompt box, type `@` and select your character from auto-complete.\r\n4. The reference image is attached automatically. Generate normally.\r\n\r\n### Export to DaVinci Resolve for Color Grading\r\n1. In the video editor, finalize your timeline (clips, markers, timing).\r\n2. Click Export \u2192 DaVinci Resolve XML.\r\n3. Choose output location. File saves to exports/ directory.\r\n4. In DaVinci Resolve: File \u2192 Import \u2192 Timeline. Select the XML file.\r\n5. Your timeline loads with all clips, properties, and markers intact. Grade and export.\r\n\r\n### Extract a Still Frame From a Video\r\n1. Hover over any video clip in the gallery.\r\n2. Camera icon = extract **current frame**. Dropdown arrow next to it = **First frame** or **Last frame**.\r\n3. Extracted image saves to your project gallery. Use as start/end image for I2V, character reference, or storyboard frame.\r\n\r\nKey workflow: extract a clip's last frame \u2192 use it as the start image for the next generation \u2192 seamless visual continuity between scenes.\r\n\r\n### Buy More Credits\r\n1. Open Settings \u2192 Credits, or the credit badge in the top nav.\r\n2. Pick a pack (packs run from $10 to $500, plus a $1,000 pack offered in the app to licensed accounts; bigger packs give more credits per dollar).\r\n3. Pay via Stripe. Credits are added to your balance instantly and never expire.\r\n4. Optional: turn on auto-topup so your balance refills automatically when it runs low.\r\n\r\n---\r\n\r\n## COMMON QUESTIONS\r\n\r\n**Q: Which model should I use for most videos?**\r\nA: Seedance 2.0 is the default and the one to reach for when physics, scale, effects or hero shots matter. For everyday shots built from a start image, Kling V3.0 Omni in ingredients mode is the best balance of cost and quality, which is why the step-by-step guides above use it.\r\n\r\n**Q: What's the best image model?**\r\nA: GPT Image 2.5 is the strongest image model in the app, and the best available anywhere right now. It holds a long instruction more faithfully than anything else here, and it is the only model that renders words in the picture reliably. The two seats cost the same, so the only trade is time: **Flare** is the fast one and is best for drafts and exploring, **Sunburst** is the higher-quality one and is what finals, hero frames and reference-heavy edits should end up on. The quality knob has five settings spanning about 36\u00D7 from cheapest to dearest \u2014 `max` is four times `high`, though the smaller steps are uneven: **`medium`** is for drafts and where iteration belongs, **`high` is the default** and where most finished work should sit, and **`max` is the best output you can get** \u2014 finished frames, client deliverables, anything with exact text. `xhigh` sits just under `max` for about half the price and is worth trying before you jump to the top. Going past `high` should be a deliberate choice rather than a habit. Go to **4K only once you already know your references and your prompt are solid** \u2014 it is the most expensive setting on the model and the worst place to discover the composition was wrong. Prove the shot at 3K first, then re-run the settled prompt at 4K. If you used GPT Image 2 before, note the quality names all shifted by one: its `medium` is now `high`, its `high` is now `max`. Nano Banana 2 is the picker's DEFAULT rather than the best: it is fast, takes the most reference images, and is the only model with grid exploration. Nano Banana Pro is the hero-frame step up when composition, cinematic lighting and skin have to be perfect. Prices for every tier are in the model table above.\r\n\r\n**Q: Do I need to set up API keys?**\r\nA: No. There are no API keys in Slates \u2014 every generation runs on Slates Credits, which come with your license and never expire.\r\n\r\n**Q: How much does a generation cost?**\r\nA: It depends on the model, resolution, and length. The exact credit cost is always shown on the Generate button before you commit, so there are no surprises.\r\n\r\n**Q: Can I use Slates offline?**\r\nA: Yes for viewing projects, editing timeline, and exporting. No for AI generation -- that requires internet.\r\n\r\n**Q: Do credits expire?**\r\nA: No. Credits never expire.\r\n\r\n**Q: What happens if I close the app during a generation?**\r\nA: Nothing is lost. On restart, Slates detects in-flight jobs and downloads completed results automatically.\r\n\r\n---\r\n\r\n## FEATURES NOT IN SLATES\r\n\r\nThe following are NOT available. Do not suggest them:\r\n\r\n- Bring-your-own API keys (BYOK) \u2014 every generation runs on Slates Credits; there is no key-entry option\r\n- Local/on-device GPU inference (all AI runs in the cloud)\r\n- Built-in music generation (use external tools like Suno, import audio)\r\n- A voice picker for Seed Audio \u2014 you describe the voice you want in words instead. The preset voice shelf belongs to the Voice lane (Inworld TTS-2), and Kling Lip-Sync keeps its own small fixed list: six English/UK voices plus a storyteller, with a speed control\r\n- Automatic video editing from a script\r\n- A prompt \"Enhance\" button \u2014 ask Studio Agent to rewrite a prompt instead\r\n- A settings/gear panel on the prompt box \u2014 every parameter is a dropdown on the bar\r\n- Cloud project storage (all files are local)\r\n- Real-time collaboration / multi-user editing\r\n- Mobile app (desktop only: Windows and macOS)\r\n- HEIC/HEIF image import (convert to JPEG/PNG first)\r\n- Storyboard JSON export (import only)\r\n\r\n---\r\n\r\n## VERSION\r\n\r\n<!-- BEGIN:GENERATED version -->\nSlates Reference Version: 1.5.6\nLast Updated: 2026-09-09\n\nThis document is generated. Its source of truth is `slate/docs/slates-llm-manual.md`; its model tables and credit costs are derived from the Slates model registry and pricing tables at build time, so they cannot be typed by hand.\n\nIf the user asks about a feature not documented here, it may have been added after this version. The current copy is always at <https://slates.video/slates-reference.md>.\n<!-- END:GENERATED version -->\r\n\r\nIf this document didn't answer your question, email hello@slates.video so we can help and improve the app.\r\n\r\n</slates_reference>\r\n";
|
|
2
2
|
//# sourceMappingURL=content.d.ts.map
|
package/dist/manual/content.js
CHANGED
|
@@ -1,3 +1,3 @@
|
|
|
1
1
|
// GENERATED by slates-web/scripts/build-llm-docs.mjs. Edit slate/docs/slates-llm-manual.md.
|
|
2
|
-
export const APP_MANUAL = "# Slates — Complete Reference for AI Assistants\r\n\r\n<!-- The heading above must stay first: shipped app builds reject this file if it\r\n does not begin with one. See slate/CLAUDE.md; check:llm-docs enforces it. -->\r\n\r\n<system_role>\r\nYou are a support assistant for the Slates desktop application. Answer user questions using ONLY the information in the <slates_reference> below. Be concise and direct. Use numbered steps for procedures. Use bullet points for explanations when helpful.\r\n</system_role>\r\n\r\n<rules>\r\n- If the answer cannot be found in the <slates_reference>, say: \"That isn't covered in the Slates reference.\" Do not guess or invent features.\r\n- If the user asks how to do something, give step-by-step instructions from the workflows and features described here.\r\n- If the user reports an error, check the TROUBLESHOOTING section first.\r\n- Refer to the FEATURES NOT IN SLATES section before answering questions about capabilities that might not exist.\r\n- Quote the exact error message when referencing troubleshooting entries.\r\n</rules>\r\n\r\n<slates_reference>\r\n\r\n<!-- BEGIN:GENERATED header -->\n# SLATES v1.5.6 — Complete Reference\n\n> **Freshness.** Generated from the Slates source of truth for app version **1.5.6**, last changed **2026-09-09**. The canonical copy of this file is <https://slates.video/slates-reference.md>. If a model, price or feature the user mentions is missing below, this copy is out of date: re-fetch that URL before answering, and say so.\n<!-- END:GENERATED header -->\r\n\r\n**What is Slates?** Desktop app (Windows 10/11, macOS 12+) for AI image and video creation. One-time purchase, no subscription. Every license includes 1,000 free credits, and Slates Pro starts with 3,000 credits. Every generation runs on Slates Credits — there are no API keys to set up, and credits never expire.\r\n\r\n---\r\n\r\n## INTENDED WORKFLOW\r\n\r\nThe designed start-to-finish flow:\r\n\r\n1. **Create project** — New project with name/description. Creates folder on your disk.\r\n2. **Build visual assets** — Generate images, create characters (with character sheets for consistency), environments (with environment grids), and styles. This is your visual library.\r\n3. **Create storyboard** — Add scenes. Every picture you drop in becomes a **Shot**: one beat of the piece, holding its references, its model, its settings and (when you want them) its words.\r\n4. **Write the piece** — Switch the storyboard to **Document** and write. What is said, what happens, the framing, the prompt each beat will send. Nothing is required and nothing is asked for; a visuals-only piece is finished as it stands.\r\n5. **Read it before you pay for it** — The header states how many generations, how many cuts, how long it runs and what it will cost. Press play for a rough cut at the real timing. Re-chop with split and merge and watch the price move.\r\n6. **Generate** — Select the beats you want and fire them in one approved batch. Each result lands under the row that made it.\r\n7. **Organize** — Switch to **Board** to re-order. Drag a beat and the whole thing moves with it.\r\n8. **Export to timeline** — Send clips to the built-in multi-track video editor.\r\n9. **Edit** — Trim, reorder, add markers, adjust timing.\r\n10. **Final export** — Export to MP4 directly, or export DaVinci Resolve XML for professional color grading.\r\n\r\n---\r\n\r\n## MODEL REFERENCE TABLE\r\n\r\n**Generating 4K video is a Slates Pro feature** — every tier generates video up to 1080p, and 4K images are open to everyone. Exporting your finished timeline at 4K is available on every tier.\r\n\r\nEvery model runs on Slates Credits. **The exact credit cost appears on the Generate button before anything fires.** The tables below are generated from the app's own model registry and rate tables, so they describe exactly what the model picker offers in this version: aspect ratios, resolutions, durations, reference-image limits, and the credit price of each. Bigger credit packs lower your per-credit cost, and Slates Pro gets the best pack rate on every purchase.\r\n\r\n<!-- BEGIN:GENERATED model-tables -->\n### Image Models\n\n| Model | Aspect Ratios | Resolutions | Max Refs | Credits per image |\n|-------|--------------|-------------|----------|-------------------|\n| **GPT Image 2.5 Flare** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 · 3K 3 · 4K 5 (default quality; at max: 2K 8 · 3K 11 · 4K 20) |\n| **GPT Image 2.5 Sunburst** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 · 3K 3 · 4K 5 (default quality; at max: 2K 8 · 3K 11 · 4K 20) |\n| **Nano Banana 2** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 4 · 2K 6 · 4K 8 |\n| **NB2 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K | 4 | 1K 2 |\n| **Nano Banana Pro** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 8 · 2K 8 · 4K 15 |\n| **FLUX.2 Max** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 4 | 1K 4 · 2K 5 · 4K 8 |\n| **Seedream 5 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 2K / 3K / 4K | 10 | 2K 2 · 3K 2 · 4K 2 |\n\n### Video Models\n\n| Model | Duration | Aspect Ratios | Resolutions | Max Refs | Audio | Credits per second |\n|-------|----------|--------------|-------------|----------|-------|--------------------|\n| **Seedance 2.0** | 4-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p / 4K | 9 | Included | 480p 3.5 · 720p 7.5 · 1080p 18.5 · 4K 39 |\n| **Seedance 2.5** | 4-30s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 30 | Included | 480p 5.1 · 720p 11.6 · 1080p 20.5 |\n| **Seedance 2.5 Edit** | Follows the source clip (4-30s) | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 0 | Included | 480p 6.2 · 720p 13.9 · 1080p 24.6 |\n| **Kling V3.0 Standard** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (6.3 with audio) · 4K 21 |\n| **Kling V3.0 Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (8.4 with audio) · 4K 21 |\n| **Kling V3.0 Omni** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (5.6 with audio) · 4K 21 |\n| **Kling V3.0 Omni Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (7 with audio) · 4K 21 |\n| **Kling O3 Edit** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 6.3 |\n| **Kling O3 Edit Pro** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 8.4 |\n| **MiniMax H3** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 2K / 4K | 9 | Included | 480p 2.5 · 768p 3 · 2K 6.5 · 4K 8 |\n| **MiniMax H3 Max** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 1080p | 9 | Included | 480p 2.5 · 768p 4 · 1080p 8 |\n| **Gemini Omni Flash** | 3-10s | 16:9, 9:16 | 720p | 7 | Included | 720p 6.4 |\n| **Omni Flash Edit** | Follows the source clip (3-10s) | 16:9, 9:16 | 720p | 0 | Included | 720p 6.4 |\n| **LTX-2.5** | 6/8/10/12/14/16/18/20s at 720p/1080p; 6/8/10s at 1440p/4K | 16:9, 9:16 | 720p / 1080p / 1440p / 4K | 0 | Included | 720p 4.5 · 1080p 6.5 · 1440p 9.5 · 4K 15 |\n| **LTX-2.5 Pro** | 6, 8, 10s | 16:9, 9:16 | 720p / 1080p | 0 | Included | 720p 6 · 1080p 8.5 |\n| **Veo 3.1 Fast** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 5 (7.5 with audio) · 1080p 5 (7.5 with audio) · 4K 15 (17.5 with audio) |\n| **Veo 3.1 Standard** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 10 (20 with audio) · 1080p 10 (20 with audio) · 4K 20 (30 with audio) |\n\n**Seedance 2.0 · Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 27.4 credits per second at 1080p.\n**Seedance 2.5 · Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 16.3 credits per second at 720p.\n**Seedance 2.5 Edit · Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 19.8 credits per second at 720p.\n\n### Audio Models\n\n| Model | Length | Credits |\n|-------|--------|---------|\n| **Seed Audio 1.0** | 3-120s | 1 at 3s · 3 at 15s · 19 at 120s |\n| **Inworld TTS-2** | up to 2,000 characters of text per take | 1 at 250 characters · 2 at 2,000 characters (billed per 250) |\n| **Sound Effects** | 1-22s | 1 at 1s · 1 at 4s · 3 at 22s |\n\n### Tools (Lip Sync, Motion Transfer)\n\nThese are real Kling endpoints that take a clip or a still as their subject, not models you prompt from scratch. Both bill in 5-second blocks.\n\n| Tool | Input | Billed in | Credits per block |\n|------|-------|-----------|-------------------|\n| **Kling Lip Sync** | Video source | 5s block | 4 |\n| **Kling Lip Sync (Avatar v2 Standard)** | Still-image source | 5s block | 14 |\n| **Kling Lip Sync (Avatar v2 Pro)** | Still-image source | 5s block | 29 |\n| **Kling Motion Control Standard** | Still image + reference video | 5s block | 32 |\n| **Kling Motion Control Pro** | Still image + reference video | 5s block | 42 |\n<!-- END:GENERATED model-tables -->\r\n\r\n---\r\n\r\n## WHICH MODEL TO USE\r\n\r\n### Images\r\n\r\n**Nano Banana 2 is the default image model.** It is the best all-round image model in the app: the most reference images of any image model, every aspect ratio, and output up to 4K. Brief it like a creative director rather than with tag soup. It is also the only model that supports the 2x2 / 3x3 grid exploration wrapper.\r\n\r\n- **NB2 Lite** is the fast, cheap draft seat in the same Nano Banana family. Roughly half the price of NB2 full and noticeably faster, 1K output only. Iterate here, finish on NB2.\r\n- **Nano Banana Pro** is the hero-frame and typography tier. Reach for it when spatial composition, cinematic lighting and skin, or fine in-image type have to be perfect. NB2 gets you most of the way there, so this is a deliberate step up, never a default.\r\n- **GPT Image 2.5** is the strongest image model in the app: it follows a long instruction more faithfully than anything else here, and it is the one to pick when the picture simply has to be right. It is also the sharp-text model, which is what makes it the choice for character sheets, shot grids, ordered panels and anything with words in the picture. It comes in two seats that cost exactly the same, and the difference is speed against quality. **Flare** is the fast one: OpenAI describes its quality as comparable to the older GPT Image 2, at roughly half the wait. **Sunburst** is OpenAI's most capable image model, better than GPT Image 2, and deliberately slower. Use Flare while you are still exploring, then re-run the shot you like on Sunburst for the final — and reach for Sunburst directly when several reference images all have to survive into one frame, or when an edit must change one region and leave identity, geometry and lighting untouched.\n Its quality knob has **five** settings — `low`, `medium`, `high`, `xhigh`, `max` — spanning about 36× from cheapest to dearest, which makes it the biggest cost lever on the model. The steps are uneven rather than a constant multiplier: `max` is four times `high`, but `xhigh` is only about 1.8 times it. **`high` is the default and the everyday setting.** `medium` is for drafts and is cheap enough to iterate on freely. `max` is the ceiling, for finished frames and exact character-level text; `xhigh` sits just under it for about half the price and is worth trying first. Go past `high` deliberately, not by habit.\n 4K is worth it only once your references and prompt are already settled: prove the shot at 3K, then re-run the finished prompt at 4K. Iterating at 4K is the most common way to waste credits on this model. Its 4K tier is API-only, so even a paid ChatGPT account cannot render it. Note that its resolution tiers are **not** a price ladder: the pixel classes are token-priced by OpenAI, so the cheapest seat is not the smallest one. Read the prices in the table above rather than assuming.\n **If you have used GPT Image 2 before, the quality names all shifted by one.** What it called `medium` is now called `high`, and what it called `high` is now `max` — the same pictures at the same prices, renamed. A remembered setting will quietly buy you a cheaper tier than it used to.\r\n **It is the only image model that can give you a transparent background.** Set **Background** to *Transparent* on the prompt bar and you get a real alpha channel — a cut-out for a logo, sticker or overlay — rather than a painted-in backdrop. *Auto* is the default and lets the model decide from your prompt; *Opaque* forces a filled background. It costs nothing either way. Slates always saves PNG, which is what carries the transparency, so there is nothing else to set.\r\n **Square and 4:3 frames cost more than 16:9 on this model, and only on this model.** OpenAI charges by image tokens rather than by pixels, and a square frame uses about 1.8 times the tokens of a 16:9 frame the same size, and 4:3 or 3:4 about 1.37 times. The credit prices in the table above are the 16:9 numbers; pick 1:1 or 4:3 and the price on the Generate button goes up to match. 9:16 costs the same as 16:9. Every other image model charges the same whatever the shape.\r\n- **FLUX.2 Max** and **Seedream 5 Lite** are the less content-restricted options. Seedream is flat-priced at every resolution it offers, so there is no reason to pick a lower one. Both auto-route to their edit endpoint when you attach reference images.\r\n\r\n### Video\r\n\r\n**Seedance 2.0 is the default video model.** Reach for it the moment physics, effects, destruction or scale matter, and for hero shots. It takes many reference images, generates native audio at no extra cost, and is the only Seedance seat that reaches 4K (generating 4K video needs Slates Pro; timeline export at 4K does not). It is also the cheaper of the two seats at every resolution they share. A face in a reference image routes it to a different provider, which is why **Seedance 2.0 · Face** is its own row in the model picker at its own price.\r\n\r\n- **Seedance 2.5 is a second seat, not an upgrade.** It buys much longer single takes and far more reference images, plus better prompt adherence. What it gives up is 4K, and it costs more than 2.0 at every resolution the two share — so 2.0 stays the model for 4K, and for the same resolution at a lower price. Because 2.5 runs longer, a long clip on 2.5 can cost more than a shorter, higher-resolution one on 2.0 — read the Generate button, not the resolution. **Seedance 2.5 Edit** is its clip-editing row: attach a clip, describe the change, and the output length follows the source.\r\n- **Kling** is the cost-effective workhorse and the most flexible family: strong start-frame adherence for identity, layout and text, acting, dialogue, multi-shot (up to 6 cuts), and the widest range of clip lengths. **Kling V3.0 Omni** adds multi-character dialogue in English, Chinese, Japanese, Korean and Spanish. Standard and Pro are the same model at two fidelity and price tiers. **Kling O3 Edit** takes an existing clip and changes what you describe, with subject and style reference images, while the original audio is preserved verbatim. Kling is also the only engine behind the Lip Sync and Motion Control tools.\r\n- **MiniMax H3** is the seat to pick when the SOUND is part of what you are writing. Every other video model treats audio as a switch; H3 takes it as three separate instructions in one prompt — the lines and action sounds tied to a moment, the ambience running underneath, and a score only the audience hears — and generates all of it with the picture in a single pass. It is also the only model where you say how much of a reference should survive, including moving one subject's characteristic onto a different subject. It runs 5-15 seconds at 480p, 768p, 2K or 4K, and takes up to nine reference images plus reference video and audio. Two things to watch: **the first five reference images are free and every one after that costs extra**, so attach what the shot needs rather than the maximum; and 2K and 4K are upscales of a 768p render rather than larger generations — in our own testing the 2K pass showed more artifacting than the 768p original it was built from, at more than twice the price. Generate and judge at 768p; step up only when a delivery spec demands the pixels.\r\n- **MiniMax H3 Max** is the same model post-trained by fal for SPEED, and it is the more expensive seat, not the cheaper one. It is dramatically faster: on the same 5-second 768p prompt it finished in about 5 seconds against about 57 seconds for H3 — roughly 12x (measured 2026-08-27). It stops at 768p and costs more per second than H3 at the resolution they share. It still animates a start frame and an end frame, so image-to-video works normally; what it does not have is the reference set — the extra identity, style and environment images plus reference video and audio that base H3 reads. Pick it when a fast turnaround on a text-to-video or start-frame shot is worth paying for; pick H3 for resolution, references, or the same tier at a lower price.\r\n- **LTX-2.5** is the VOLUME seat — the cheapest native 1080p second in the catalogue, with synchronised audio included free at every resolution, so it is the model to reach for when the job is many takes rather than one hero shot. Two things are unique to it. It makes the LONGEST clips of anything here, up to 20 seconds, and it is the only model that reaches 1440p. It is also the only one with native MULTISHOT: a single generation can carry two to four connected shots that hold the character, lighting and voice across the cuts, which everywhere else means generating separate clips and watching identity drift between them. Its constraints are unusually sharp, though. Durations are EVEN NUMBERS ONLY starting at six — 6, 8, 10, 12, 14, 16, 18, 20, with no 5-second or 7-second clip — and above 1080p that ceiling drops to 10 seconds. Aspect ratios are 16:9 and 9:16 only. And it takes FRAMES, not references: a start frame and an optional end frame that generates a transition between them, but no identity, style or environment reference images at all, so cross-shot character consistency belongs on MiniMax H3 or Kling. Because sound is generated in the same pass, write the audio into the prompt and anchor every cue to something visible — anything unanchored gets invented.\r\n- **LTX-2.5 Pro** is the fidelity seat of that pair, and it is NOT simply a better LTX. It renders the picture with more compute on busy frames, but on a narrower envelope than the base row: 720p and 1080p only (no 1440p, no 4K) and 6, 8 or 10 seconds only, for about a third more per second. Reaching for it because the name says Pro costs more AND takes away the reach. Pick it when a specific shot needs the extra fidelity and fits inside 1080p and ten seconds; pick base LTX for length, resolution and volume.\r\n- **Gemini Omni Flash** is the cheap 720p seat with native synced audio included in one pass. **Omni Flash Edit** is the prompt-only clip editor: no reference images, one short instruction plus \"Keep everything else the same.\" Long descriptive prompts destroy it.\r\n- **Veo 3.1** is niche and is never a default. Pick it only when you specifically want Google's audio pass. It has the fewest aspect ratios and reference slots of any video model, fixed durations, and the highest per-clip cost.\r\n\r\nBoth edit models take an existing clip as their canvas, so their output length follows the source clip rather than a duration you choose.\r\n\r\n### Audio\r\n\r\nAudio is a third media type alongside images and video — generated as its own asset, shown in the gallery's **Audio** tab, and dragged onto an audio track in the timeline. This is separate from the audio some VIDEO models generate *inside* a clip (see AUDIO IN GENERATION below): use a video model when the sound must be locked to what is on screen, and these when you need audio you can move, trim, re-use, or layer.\r\n\r\n**Seed Audio 1.0 is the default.** A room with dialogue *and* clatter *and* ambience is one generation, not three layered ones, and because it is cheap you can run five takes and keep the best. It makes a whole audio SCENE from one plain sentence. **It has no length setting of its own** — Slates writes your chosen duration into the prompt, and that is what you are charged. You describe the voice in words; there is no voice list to pick from. Set **Languages** to Mixed if one scene needs more than one language (it costs the same).\r\n\r\n**Sound Effects** makes one effect, or a seamless loop. It is the only surface with an exact duration, so an effect can land on a specific frame. Describe the physical cause (\"heavy oak door slams shut in a stone hallway\"), not the label (\"door sound\"). **Loop** makes it seamless for beds; **Wording** controls how literally your description is followed. Seed Audio is actually the better tool for *long* ambience beds, so the two are not redundant in the direction you would expect.\r\n\r\nKling's `SFX:` / `Ambient noise:` prompt syntax belongs to video prompts and makes Seed Audio results *worse* — write plain sentences there instead.\r\n\r\n**Inworld TTS-2 is the voice seat** — type the words, pick a voice, press Generate. In the prompt box's Audio lane pick **Voice** in the model picker; the prompt is the exact text that gets spoken (nothing is added or rewritten — open \"See what gets sent\" to confirm), and the **Voice** control on the bar opens the voice picker: **Presets** (ready-made voices with gender, accent and age filters — every one plays the same audition line, so you compare voices rather than scripts), **Clips** (any character's voice, or any audio clip in the project, cloned for the take), or **Describe** (a voice in words). The character counter beside the bar is the bill: the generated audio table above gives the text cap and billing buckets, and the Generate button shows the price. Direction goes in square brackets (`[whispering] …`) — anything in parentheses is read aloud. Cloning a real person's voice needs their permission. Right-click any audio clip in the Audio tab → **Use as voice** lands you on the Voice lane with that clip as the voice. Studio Agent and the MCP/CLI do the same through `slates_generate_audio` (a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId`, or a `voiceDescription`).\r\n\r\n**There is no music generation.** For a song, use an external tool and import the audio (see PROJECTS → Supported File Formats). For spoken lines inside a scene, let Seed Audio perform them, put them in the video prompt on a model with native audio (Seedance, Kling Omni, Omni Flash, Veo), or use Kling Lip-Sync's text-to-speech against a shot.\r\n\r\n### Tools\r\n\r\n**Tools** is not a model family. It is two real Kling endpoints that take a clip or a still as their subject: **Kling Lip Sync** and **Kling Motion Control**. Both bill in 5-second blocks and are described under GENERATION MODES below.\r\n\r\n**How pricing works:** every generation is priced in Slates Credits, and the exact cost is shown on the Generate button before you commit. Bigger credit packs give more credits per dollar; Slates Pro locks in the best pack rate on every purchase, forever.\r\n\r\n---\r\n\r\n## GENERATION MODES\r\n\r\n### Create Image\r\nPrompt → select image model → set aspect ratio + resolution → generate. Batch grids available for quick iteration.\r\n\r\n### Text-to-Video\r\nPrompt → select video model → set duration + aspect ratio + resolution → generate. Output: MP4.\r\n\r\n### Image-to-Video (I2V)\r\nAttach start image + prompt → select model → generate video from that image. Optional: attach end image (Veo) for guided transitions.\r\n\r\n### Ingredients / References\r\nUse @character_name, @environment_name, or #style_name in prompt to attach reference images for visual consistency. Kling: up to 4 total references. Veo: up to 3. Nano Banana 2: up to 14. The @mentions auto-complete from your project's characters/environments/styles.\r\n\r\n### Lip Sync\r\n**Kling only.** Pick **Kling Lip Sync** under the Tools family in the model picker. Source: video or still image.\r\n\r\n- **Audio source** — Text to speech (type the line; six English/UK voices plus a storyteller, with a speed control) OR Upload audio (bring your own recording, max 5MB).\r\n- **Avatar tier** — only appears for a still-image source: Avatar v2 Standard (the value tier) or Avatar v2 Pro (higher fidelity, higher rate).\r\n- 5s output blocks.\r\n\r\nWorks very well with human-like characters. Less reliable with animals or non-human characters.\r\n\r\n### Motion Transfer\r\n**Kling only.** Pick **Kling Motion Control** under the Tools family. Target: still image (your character). Source: reference video (the motion).\r\n\r\n- **Engine** — Kling MC Standard (value tier) or Kling MC Pro (higher fidelity, higher rate).\r\n- **Orientation** — *Match video* copies skeleton and depth from the clip (best for dancing, walking, full-body action; driving clips up to 30s). *Match image* keeps your character's pose and angle and uses the video only as motion hints (best for close-ups; up to 10s).\r\n- 5s output.\r\n\r\n> **Note:** these two tools used to offer a second \"Seedance 2.0\" engine. It was not a separate engine — picking it made Slates write a sentence into your prompt that you never saw, which is no longer allowed anywhere in the app (see \"What gets sent\" below). Both tools are now Kling endpoints only.\r\n\r\n### Edit Image\r\nRight-click any image asset → open viewer → switch to Edit Mode. Enter an edit prompt describing the changes you want. Select edit model: Nano Banana 2 (supports up to 14 reference images), FLUX.2 Max, or Seedream 5 Lite. Choose resolution and aspect ratio. The result saves as a new asset with the original preserved. Useful for refining generated images without starting from scratch.\r\n\r\n### Edit Video (Kling O3 Edit / Omni Flash Edit)\r\nRight-click any video clip (gallery or timeline) → \"Edit with AI\". The clip attaches to the prompt box as the source; describe the CHANGE, not the whole scene (\"replace the man with @marcus\", \"make it a rainy night, keep everything else\"). Two engines in the model picker:\r\n- **Kling O3 Edit (default):** attach subject images (role: Subject) to swap someone in, or style images (role: Style) for a look — max 4 combined refs. Clips 3-15s. Original audio preserved.\r\n- **Omni Flash Edit (cheapest):** prompt only — no reference images; keep instructions simple and add \"Keep everything else the same.\" Clips 3-10s, 720p output.\r\n\r\nOutput length follows the source clip; the credit cost (clip seconds, rounded up, at the per-second rate) shows on the Generate button. The edited clip saves as a NEW asset linked to the original — chain edits freely. Trim longer clips on the timeline first.\r\n\r\n### Multi-Shot (Kling V3.0/Omni)\r\nEnable multi-shot toggle → multiple scene prompts in one generation, each with different framing. 6-axis camera controls per shot. Results can be hit-or-miss, but worth trying for quick multi-cut sequences. For more reliable results, most users prefer generating multiple short 5s clips separately using Kling V3.0 Omni in ingredients mode and assembling them on the timeline.\r\n\r\n---\r\n\r\n### Generate Audio\r\n\r\nSwitch the prompt box's lane pill from Image/Video to **Audio**, pick a surface, and generate. The prompt box offers three: **Seed Audio 1.0**, **Voice** (Inworld TTS-2) and **Sound Effects**. The result lands in the gallery's Audio tab as its own asset with a waveform and an inline player, and can be dragged onto an audio track in the timeline.\r\n\r\n- **Scene (Seed Audio 1.0)** — one plain sentence describing the moment. Set **Length**; Slates writes it into the prompt for you and that is exactly what you're billed for (open \"See what gets sent\" under the prompt box to read the appended text). **Say the crowd/room size out loud** — \"applause\" returns a full auditorium when you meant three people at an open mic. Ask for a few seconds more than the clip needs so the edit has fade handles. Describe the voice you want in the sentence itself (\"a weary dock foreman in his fifties, gravel in his voice\") — there is no voice picker.\r\n- **Voice (Inworld TTS-2)** — the prompt is the words to be spoken, verbatim. Pick the voice with the **Voice** control on the bar (presets you can play first, any clip in the project, a character's voice, or a description); the character counter is the bill. See MODELS → Inworld TTS-2 for direction tags and the cloning rules.\r\n- **Sound Effect** — describe the physical cause and set the length to roughly the event (≈1s for an impact, 2–4s for a whoosh, 8–22s + **Loop** for a bed). **Wording** sets how literally the description is followed: Interpretive, Balanced (default), or Literal.\n\n#### Use your own voice recording\n\nIn the bottom prompt box, choose **Audio**, then **Inworld TTS-2** in the model picker. Open **Voice → Clips → Import voice clip** and select your recording. The import adds an audio asset to this project without generating anything. Click its play button to audition it, then click the recording's name to choose it. Type the words you want spoken in the prompt box and press **Generate**, which shows the price. The new take appears in **Gallery → Audio**.\n\nUse a clean recording of one speaker whose voice you have permission to use. Slates clones the recording for each take; there is no separate training wizard or persistent vendor voice to manage. **Presets** lets you audition ready-made voices; **Describe** lets you write a voice description and choose **Use this description**. Choosing in the prompt box sets up the next take; only Generate spends credits.\n\nTo attach your recording or a generated take to a character, open **Gallery → Characters**, then **Add voice** (or **Change voice**) on that character's card. Choose **Clips** and click the clip's name. Attaching an existing clip is free. The card displays its waveform and player. That character's voice is also listed under **Voice → Clips → Characters** in the prompt box. Selecting a preset or description from a character card generates and attaches a take; read the cost shown in that picker before choosing.\n\nFor **Inworld TTS-2**, choose the character under **Voice → Clips → Characters**. Typing `@name` does not choose a TTS voice: the prompt is spoken verbatim. Automatic voice attachment from a character mention belongs to **Seed Audio** only.\n\r\n---\r\n\r\n## AUDIO IN GENERATION\r\n\r\nThis section is about audio generated **inside a video clip**. For audio as its own asset, see Generate Audio above.\r\n\r\n**Veo 3.1 native audio:** Generates audio WITH video. Prompt syntax: `\"Hello!\"` for dialogue, `SFX: [sound]` for effects, `Ambient noise: [description]` for ambience. Max 10s dialogue. Add `(no subtitles)` to suppress text overlays.\r\n\r\n**Kling V3.0 Omni dialogue:** Multi-character dialogue with distinct voices. Languages: EN, ZH, JA, KO, ES. `Background music: [description]` for music. Max 10s dialogue.\r\n\r\n**Kling V3.0 sound co-generation:** Synchronized sound effects generated with video.\r\n\r\n⚠️ **This prompt syntax is video-only.** `SFX:`, `Ambient noise:` and `Background music:` are Kling/Veo conventions — the audio models above have no parser for them and will treat them as words in the scene.\r\n\r\n---\r\n\r\n## PROMPT SYSTEM\r\n\r\n### Unified Create Surface (roles + model-on-button)\r\nThe old mode tabs (text-to-video / frames-to-video / ingredients / create-image) are ONE \"Create\" surface. An **Image | Video | Audio pill** on the prompt bar switches your output lane — it remembers and restores the last model you used in each lane (pick Seedance once and the Video lane stays Seedance until you change it). Only the active lane shows its name; the other two are icons. Every attachment in the reference tray carries a tappable ROLE badge — Reference / First frame / Last frame / Subject / Style — you say what each attachment is; nothing is inferred. Your model choice sticks across generations and workflow actions (\"use as first frame\" keeps your chosen video model). Attaching a video via \"Edit with AI\" flips the surface into Edit Video mode. Lip Sync and Motion Transfer live under the **Tools** family in the model picker.\r\n\r\n### Floating Prompt Box — the bar holds everything\r\nPersistent across all pages. **There is no settings panel and no gear button.** Everything sits on one bottom bar, left to right:\r\n\r\n1. **Media toggle** — Image / Video / Audio.\r\n2. **Model picker** — a searchable menu plus a detached submenu. The main menu lists model families with a vendor glyph tile each; picking one opens that family's models beside it, every row carrying capability chips (resolution, clip length, audio, references) and its per-unit rate. Type to search across every model. The submenu is anchored to the row you opened it from, so it never travels. The trigger on the bar shows the model name and nothing else — no chevron, no resolution appended. Everything listed is a real model or endpoint.\r\n3. **Parameter controls** — one per setting the chosen model actually has (resolution, aspect, duration, length, quality, count, grid, face-in-reference, audio, loop, and so on). The trigger shows the current value; the explanation lives *inside* the menu as a subtitle under each option, along with what that option costs. A setting with only one possible value still shows, muted and non-interactive, so the row never changes shape.\r\n4. **Sliders for ranges.** A setting with a long list of steps (video duration, audio length) opens a ruler instead of a many-row menu. The handle moves between the values the model actually declares, so it cannot land on one the model will not accept, and the price for the selected value is shown on the ruler.\r\n5. **`More ▾`** — if the model has more parameters than fit the current window width, the extras are collected into a generated `More` dropdown automatically. Widen the window (or close the Studio Agent panel) and they move back onto the bar.\r\n6. **Generate** — reads `Generate · <cost>`. **Cost only** — the model name is not repeated on the button (it's in the model picker) and there is no send arrow. A badge on the left of the button counts generations currently running.\r\n\r\nFor text-to-speech, the character counter is always visible because text length determines the price. On other surfaces it appears near the right of the bar after roughly three quarters of the model's prompt limit. It turns red over the limit.\r\n\r\n**Collapsing:** the chevron at the top-right of the box collapses it to a single arrow — nothing else. Click the arrow to bring it back.\r\n\r\nBelow the bar the queue shows pending/active generations with cost and progress.\r\n\r\n### What gets sent (prompt transparency)\r\nUnder the prompt box is a **\"See what gets sent\"** disclosure. Open it and you see the exact text that will be transmitted, produced by the same code that builds the request — so it can never disagree with what is actually sent. It shows:\r\n\r\n- **Reference numbering** — `@sarah` becomes `Sarah (image 1)` so the model knows which attached image is which.\r\n- **Key lines for attachments you did NOT mention** — one short neutral sentence per unmentioned attachment. Mention every reference in your own words and these never generate.\r\n- **The trailing style clause** when a `#style` is attached.\r\n- **The Seed Audio duration append** — the `… N seconds` Slates adds to the end of the prompt, which is also what you are billed for.\r\n- **Grid wrapping** when 2×2 or 3×3 is on.\r\n\r\nThe row stays hidden when the composed prompt is identical to what you typed, so it only appears when there is something to show.\r\n\r\n**Unresolved `#tags` and `@mentions` are named, not silently dropped.** If you type `#noir` and there is no saved style called \"noir\", the tag is removed from the text sent to the model (a raw tag confuses every model) — but the disclosure turns red and says so by name: *\"#noir matches nothing saved — removed from what gets sent.\"* Save the style, or reword it, and the warning clears.\r\n\r\n**Nothing is ever added that you cannot read here.** No setting in the app injects prompt text; a setting changes *how* a request is made, never *what* you asked for.\r\n\r\n### Prompting guide (on the web)\r\nPer-model prompting guidance lives at <https://slates.video/docs/prompting>, linked from the bottom of Settings. It covers every model Slates offers — Video, Image, Audio — with what that model reads, what it ignores, and its gotchas, all on one page so you can compare them. Markdown copy for pasting into an LLM: <https://slates.video/docs/prompting.md>.\r\n\r\nIt is generated from the same source the Slates CLI, the MCP server and Studio Agent are built on, so the guide and the app cannot disagree. It is documentation rather than a control, which is why it is a page on the web and not a panel in the app: it has room to be read, a URL you can send someone, and it is always current rather than frozen at the version you installed.\r\n\r\n### @Mentions\r\nType `@` → auto-complete shows project characters and environments. Type `#` → shows styles. Selecting inserts the reference image(s). At send time a mention is rewritten to a numbered citation (`@sarah` → `Sarah (image 1)`) so the model can tell your attachments apart — you can read the result in \"See what gets sent\". Nothing else about your wording is rewritten. Prompting works the same as any other AI tool; no special syntax beyond @mentions.\r\n\r\n### Writing the shot list with Studio Agent\r\nThere is no \"Enhance\" button and no \"Generate prompts\" button. **Studio Agent does this work**, because it reads the same shot list you do — every scene, every beat in order, with its references, its model and its price — and because you can steer it:\r\n\r\n> \"Write the beats for my current storyboard.\"\r\n> \"Now redo scene 3 handheld, and match its energy to scene 2.\"\r\n> \"SHOT-A4 runs long — split it after 'and then'.\"\r\n\r\nEvery field it writes is editable by hand, in place, in the storyboard's Document view. Open Studio Agent with **Ctrl+.**\r\n\r\n### Debug Panel (advanced)\r\nA developer panel showing the exact request body, with the ability to override the composed prompt before sending. **There is no toggle button for it on the prompt bar in any build** — open it with **Ctrl+Shift+D**. For ordinary use, \"See what gets sent\" above is the supported way to inspect a prompt.\r\n\r\n---\r\n\r\n## PROJECTS\r\n\r\n### Structure\r\nEach project = folder on your disk. Subdirectories: images/, videos/, audio/, references/, exports/. Location configurable in Settings → Projects Directory.\r\n\r\n### Assets\r\nEvery generated or imported file is an asset (image, video, audio). Metadata tracked: prompt, model, settings, cost, dimensions, timestamps. Videos track source image via source_asset_id -- you can see all videos generated from any image.\r\n\r\n### Supported File Formats\r\n- **Images:** PNG, JPEG, WEBP, GIF. Note: HEIC/HEIF (iPhone photos) NOT supported -- convert to JPEG/PNG first.\r\n- **Video:** MP4, MOV, WEBM, AVI, MKV.\r\n- **Audio:** MP3, WAV, OGG, M4A, AAC.\r\n- **Clipboard paste:** Any image format the OS clipboard provides (PNG, JPEG, WEBP, GIF). Pasting works both in the gallery and directly into the prompt box; either way the image becomes a real gallery asset in the folder you're working in (tagged \"Imported\") AND, when pasted into the prompt box, attaches as a reference. Anything generated from it links back to it as a source.\r\n- **Drag and drop:** Any file the browser recognizes as image/* or video/*.\r\n\r\n### Operations\r\nCreate/rename/delete projects. Import external files via drag-and-drop or file picker. Paste images from clipboard. Extract still frames from videos. Relocate project to different disk/folder (all paths auto-update). Cleanup orphaned assets.\r\n\r\n### Moving and copying assets between projects\r\nAssets (images, clips) can be sent to another project three ways: the selection band on the Images/Videos tabs, the right-click menu on any card, or by dragging cards and dropping on a project in the drop palette.\r\n\r\n- **Move** relocates the media files on disk into the destination project's folder. The asset leaves whatever gallery folder it was in and is issued a fresh badge code in the destination.\r\n- **Copy** duplicates it — new files, new thumbnails, new badge code — and changes nothing in the source project.\r\n\r\n**Why a move can be refused:** an image another project still builds with (a character/environment/style identity image, or a storyboard frame) cannot leave, because the entity left behind would point at a file it no longer owns. When that happens the dialog lists what's blocking and offers the fix: bring the whole character/environment/style across with all of its images, or copy instead. A storyboard frame is only ever offered a copy — moving its image out would empty the shot.\r\n\r\nRight-clicking a card that is part of a multi-selection acts on the whole selection (\"Move 5 to Project…\"). Right-clicking a card outside the selection acts on that card alone.\r\n\r\n---\r\n\r\n## SHOTS — THE STORYBOARD IS THE SHOT LIST\r\n\r\nA **Shot** is the prompt bar, saved: the prompt, every reference with the job it carries, the model, every setting — and now the beat itself: who speaks, what they say, how it is said, what happens, the prop, the framing and the camera. It lives in the **storyboard**, which is the one place Shots are listed. It is never required: the prompt bar works exactly as it always has for anyone who never touches one.\r\n\r\n### Why it exists\r\nA generation's full recipe was already stored, but only once you had paid for it. A Shot can be written **before anything is generated**, so a whole piece can be planned, read, timed, priced and corrected while it is still free. That is the point of the thing: look at the entire ad or short film — every cheap asset lined up in the actual flow — before the videos exist.\r\n\r\n### What a Shot holds\r\nRaw prompt (@mentions intact); the model; aspect ratio / duration / resolution / negative prompt / sound and the rest of the bar's settings; every attachment with its ROLE (plain reference, subject, style, reference video, reference audio, first frame, last frame); and the script layer — `speaker`, `line`, `delivery`, `action`, `prop`, `shotSize`, `camera`, and a `continues` flag for one sentence running across two cuts. Characters, environments and styles are stored as the ENTITY, not a copied picture, so updating a character updates every Shot that names it.\r\n\r\n**The script fields are for reading and counting. Only the prompt is sent to a model.** Dialogue you want a model to perform still goes in the prompt, verbatim, with its delivery — writing it in `line` makes it readable and lets Slates check whether it fits the cut, not spoken.\r\n\r\n### Every Shot has an address\r\n`SHOT-A1`, `SHOT-A2` … per project, never reused — the same idea as the `IMG-A12` badge on a gallery card. Say it to Claude and you are both pointing at the same row. It is for **this session**, not for retrieval later: there is no shot search and no shot library, because a Shot is workspace state — alive while you build the piece, worthless once it ships.\r\n\r\n### Making one\r\n- **Save as Shot** on any generated image, clip or track's right-click menu restores that generation and keeps it — and that generation becomes the Shot's first take, so the row opens showing the result you kept it for. **Reuse Prompt** on those same menus does the restore WITHOUT saving anything.\r\n- Claude can write Shots directly (`slates_create_shot`), including for shots whose image does not exist yet — and can re-chop them with `slates_split_shot` / `slates_merge_shots`.\r\n- **It files itself.** A saved Shot lands in the scene you have open, else the last scene of the storyboard you were most recently working in; if the project has no storyboard, one appears named after the project. Nothing you save is ever somewhere you have to go and find.\r\n\r\n### There is no save button\r\nSelecting a Shot row **binds** the prompt bar to it. Edits write straight back to that row; `Clear` unbinds and returns the bar to free composing. There is no undo, and none is needed: the generations underneath a row are the permanent record of what actually fired, and the row itself is the working copy.\r\n\r\n### Two ways to look at it\r\nThe storyboard has one toggle and two jobs.\r\n\r\n- **Board** — arrange. One picture per Shot, dragged into the order you want. Drag one and the whole beat moves with it: the line, the references, the model, the settings, the takes. There is nothing else to drag, so there is never a question of what followed what.\r\n- **Document** — write. One continuous page: the script, the references beside the words that cite them, the prompts underneath. This is where you read the piece before paying for it.\r\n\r\nInside Document, choose **Script** to edit the words with scene headings and speakers. Delivery notes, model and pricing details, and warnings are hidden. **Shots** shows all layers. Open **Custom** to toggle Scene, Action, Character, Delivery, Dialogue, Shot, References, Prompt, Takes, and Warnings independently. Warnings covers missing models or references, unresolved mentions, model changes, and lines too long for their cut. Hiding a layer changes only what you see; it preserves your text and generation checks. There is a text-size slider and an independent toggle for shot numbers in the margin.\n\nDelivery is an optional performance note, not a required label for every line. TTS sends the authored prompt, or the dialogue when the prompt is empty, verbatim; separate Delivery notes are not added. For speech cues, use the selected TTS model's prompting guide and place supported tags in that spoken text. Script view does not strip inline speech cues from your words.\n\r\n### The header tells you what you are about to make\r\n`5 generations · 7 cuts · 54s · 84 credits`, and beneath it a variety strip like `6/7 wide · 5 push · 3 cuts in the loft`.\r\n\r\n**Two counts, because they measure different things.** Rhythm is counted in **cuts**; money is counted in **generations**. A multi-shot generation is several cuts inside one paid call, so mixing them would be wrong. A cut with no model chosen has no duration and shows as `—` rather than `0s` — a runtime that invented seconds would be a lie about the one number this view exists to give.\r\n\r\n### Splitting and merging — the chop\r\nPut the caret mid-line and press Enter: the row becomes two, the second inheriting the model, settings and references, and marked as continuing the first if the split lands mid-sentence. Select two adjacent rows and press **Merge**: they become one, references combined, durations summed. **The price and the runtime move as you do it** — which is the whole reason to make the decision here rather than in a document somewhere else.\r\n\r\nSplit is also the move behind a voiceover that keeps talking while the picture hard-cuts to a new world: split at a word boundary and both rows carry one sentence, each with its own visuals.\n\nThe **Dialogue continues from previous shot** toggle is a planning note that the sentence spans a cut. It does not merge shots, join generated audio, or change generation settings. Its pressed state shows whether the note is set; toggling it does not move the script. **Merge these two shots** is the separate control between adjacent shots that actually combines them. In Document → Custom, **Scene** toggles scene headings and their controls; the dialogue remains visible even if a hidden heading belonged to a collapsed scene.\n\r\n### Variety, counted and never judged\r\nSlates counts what is in front of it — shot sizes, camera moves, cast, locations, durations, and any of them repeating three or more times in a row — and shows the counts. **It never changes anything, never suggests anything and never blocks.** `shotSize` and `camera` are free text: write `long-lens CU, other head blurred` if that is the shot. Anything unrecognised counts as \"other\", which is a fine answer.\r\n\r\nIf a spoken line cannot be read in its cut at any plausible pace, the row says so — and says it only when the line is genuinely impossible, never when it is merely long.\r\n\r\n### The animatic\r\nPress play on any beat and the storyboard plays as a rough cut: each picture held for **its own cut's duration**, with the line underneath. That tells you the rhythm of the finished piece before a single video exists. A multi-shot generation holds one picture across its internal cuts, and says so. A cut with no duration is held for 3 seconds and marked — the header leaves it out of the runtime for the same reason it is marked here.\r\n\r\n### Things that stay honest rather than being hidden\r\n- **Deleted references.** If an asset or character a Shot points at is gone, the Shot still loads and the row says how many items were left out of what gets sent — and editing the row does not quietly drop them.\r\n- **A swapped model.** Changing a Shot's model never rewrites your words — video models genuinely take different prompt grammars, so the row tells you which model the prompt was written for and leaves the sentence alone. The settings line shows what will actually be sent after the swap, and that is the value it prices.\r\n- **Image Shots carry no role badges.** Image generation sends every reference in one undifferentiated list, so an image Shot can remember that a picture is a style reference but cannot tell the model.\r\n- **The thumbnail is never a question.** Slates picks it — the first frame, else the first reference, else the newest take — and you can override it from any reference in the gutter.\r\n\r\n### Firing several\r\nSelect Shots and press **Generate all**: one total, the largest single Shot stated separately, one approval. They run **one at a time**. If one of the selected Shots has been deleted, the whole batch is refused and names it — nothing fires and nothing is billed. If a generation fails mid-run, the rest still fire, the failure is reported per Shot, and **nothing is retried automatically**.\r\n\r\n### Deleting\r\nDeleting a storyboard tells you how many Shots are attached before it does anything, and deletes them with it. It never moves them somewhere else without asking. A Shot that also lives in another storyboard survives.\r\n\r\n---\r\n\r\n## STORYBOARDING\r\n\r\n### Hierarchy\r\nStoryboard → Scenes → **Shots**. A scene is an ordered list of Shots, and a Shot is one beat: its picture, its references and their roles, its model and settings, its prompt, its words, and the generations it has produced. See **SHOTS** above — that section is the storyboard.\r\n\r\n### What happened to frame types\r\nThere used to be a \"frame type\" on each picture — first / last / ingredient — plus a separate motion-prompt box. Both were a second, weaker way of saying what a Shot already says: **a reference's role lives on the Shot** (first frame, last frame, subject, style, plain reference), and the motion prompt was just the Shot's prompt under another name. Existing storyboards were converted automatically and nothing was lost. Pick a picture's job on the Shot's reference rail; write the motion in the Shot's prompt.\r\n\r\n### Grid Exploration\r\n2x2 grid: 4 prompt variations for quick iteration. 3x3 grid: 9 variations for deeper exploration. Select individual cells → extract to full-resolution images. Tip: 2x2 is usually sufficient and produces better quality. 3x3 can occasionally get proportions slightly wrong when upscaling cells because it faithfully reproduces the lower-resolution proportions. Grid exploration runs on Nano Banana 2 only; no other image model offers it.\r\n\r\n### Storyboard → Video\r\nSelect frames → generate video for each → clips auto-insert into timeline in order with source tracking maintained.\r\n\r\n### The animatic\r\nPlay the storyboard as a rough cut. Each beat is held for **its own duration**, with its line underneath, so what you are watching runs at the finished piece's real length. Space = play/pause. Arrow keys = navigate. Escape = exit.\r\n\r\n### Paste a script\r\nPaste a script into the storyboard and it becomes one row per paragraph — ALL-CAPS cues become speakers, parentheticals become delivery. It is a plain parse, not a model: nothing is invented, nothing is sent anywhere, and prose that is not screenplay-formatted lands as one row per paragraph for you (or Claude) to chop.\r\n\r\n---\r\n\r\n## VIDEO EDITOR (TIMELINE)\r\n\r\n### Tracks\r\nMulti-track: video tracks + audio tracks stacked vertically. Clips independent per track. Add or remove tracks freely -- layer a music bed, a voiceover, and effects on separate audio tracks. Video assets go on video tracks, audio assets on audio tracks. Overlapping video clips resolve top-track-wins.\r\n\r\n### Audio Mixing\r\nEach track has a volume fader, and the timeline has a master output fader for the final mix. Both range from silent to +12 dB of boost, and both apply to preview playback AND the exported MP4 -- what you hear is what you render. Muting a video track silences its embedded audio but still shows the picture. Use the master fader to prevent clipping when stacking loud tracks.\r\n\r\n### Timeline Settings\r\nResolution and frame rate (24/30/60) are auto-managed: the first video clip sets both, and a later higher-resolution clip raises the canvas. All clips are conformed to the timeline frame rate on export. Changing the frame rate after clips are placed retimes them.\r\n\r\n### Clip Properties\r\nSource asset, in/out points (frame-level precision), duration, scale (fit/fill/custom %), position (X/Y offset), opacity (0-100%).\r\n\r\n### Tools\r\n- **Select (V):** Click/drag clips, view/edit properties\r\n- **Razor (C):** Split clip at playhead into two clips\r\n- **Slip (S):** Adjust clip in/out points without moving its position\r\n- **Snap toggle:** Snap to playhead/clip boundaries\r\n\r\n### Markers\r\nColor-coded timeline markers (6+ colors) with optional labels. Use for scene breaks, cue points, notes.\r\n\r\n### Playback & Navigation\r\nSpace = play/pause. Left/Right arrows = frame-by-frame. Up/Down = +-1 second. Page Up/Down = jump by screen width. Home/End = start/end of timeline.\r\n\r\n### Zoom\r\nCtrl+Plus = zoom in (finer precision). Ctrl+Minus = zoom out (see more timeline).\r\n\r\n### Undo/Redo\r\n50-step history. Ctrl+Z = undo. Ctrl+Shift+Z = redo.\r\n\r\n---\r\n\r\n## EXPORT\r\n\r\n### Video Export (FFmpeg)\r\nExport timeline → MP4 (H.264). Configure: resolution, frame rate, bitrate, output location. All visible tracks rendered, muted tracks excluded, clip in/out points respected. FFmpeg is bundled -- no separate install needed.\r\n\r\n### DaVinci Resolve XML Export\r\nGenerates XML project file containing: clip references (paths to source videos), timeline structure (tracks, clips), clip properties (scale, position, opacity, in/out points), timeline markers.\r\n\r\n**Importing into DaVinci Resolve:** File → Import → Timeline. DaVinci reads the XML and reconstructs your timeline with all clips, properties, and markers intact. From there you can color grade and export your final master.\r\n\r\nExports saved to project's exports/ directory with timestamped filenames.\r\n\r\n---\r\n\r\n## CHARACTERS, ENVIRONMENTS & STYLES\r\n\r\n### Characters\r\nCreate character with name + description. Generate character sheet (license required): AI generates a turnaround with multiple angles for consistency. Generate expression sheet: same character with different facial expressions. Use `@character_name` in any prompt to attach reference images.\r\n\r\n**Tips for consistency:** Experiment with character sheet generation using both the existing project style and photorealistic style. Sometimes a single well-chosen image works better than a full sheet -- especially if the character is already in the same style, lighting, and clothing as your project. You can manually assign any image as a character reference instead of generating a sheet.\r\n\r\n**Voice.** A character can carry one voice clip, the same way it carries one identity image — a shortcut for reusing a voice, never a requirement for speaking in one. **Add voice** / **Change voice** on the card opens the same voice picker the prompt box uses (a menu off the button, not a pop-up; the other cards stay on screen): a clip from the project attaches as it is; a preset or a described voice renders the character speaking a fixed audition line on Inworld TTS-2 and attaches that clip, at the credit cost the picker states first. Right-click the voice card → **Remove voice** detaches it without deleting the clip. The character's voice then shows under **Clips** in the Voice lane's picker, and mentioning the character (`@name`) in a Seed Audio prompt attaches the clip as a reference, so the scene casts that voice.\r\n\r\n### Environments\r\nCreate environment with name + description. Generate environment grid (license required, 3x3): 9 variations. Extract individual cells to full-resolution images. Use `@environment_name` in prompts.\r\n\r\n**Tip:** Like characters, sometimes a single strong environment image gives better consistency than a grid of 9. Experiment with both approaches.\r\n\r\n### Styles\r\nCreate style with name + description + upload reference image. Use `#style_name` in prompts. Key visual auto-attachment option for consistent look across all frames.\r\n\r\n---\r\n\r\n## SETTINGS\r\n\r\n### Generation\r\nEvery generation runs on Slates Credits — there are no API keys to configure. The Generate button shows the exact credit cost before each generation, and failed generations refund immediately.\r\n\r\n### Other Settings\r\n- **Projects Directory:** Where project folders live on disk. Changeable anytime.\r\n- **Default Model:** Pre-selected model for new generations. Override per-generation.\r\n- **Default Quality/Resolution:** Pre-selected resolution. Override per-generation.\r\n- **Grid Size:** Default 2x2 or 3x3 for grid exploration.\r\n- **Auto Naming:** Automatically name generated assets.\r\n- **Prompting guide:** A link at the bottom of Settings to <https://slates.video/docs/prompting> — per-model prompting guidance for every model (see PROMPT SYSTEM above).\r\n\r\n---\r\n\r\n## ACCOUNT & BILLING\r\n\r\n### Login\r\nEmail-only, no password. Enter email → receive magic link → click to log in. First login creates account automatically. Session persists across restarts.\r\n\r\n### License\r\nUnlocks: character sheet generation and environment grid generation. Includes 12 months of updates (Slates Pro includes lifetime updates). Major upgrades discounted after.\r\n\r\n### Credits\r\n\r\n<!-- BEGIN:GENERATED credits -->\nCredits are what every generation is paid with. They are pay-as-you-go, they never expire, and the exact cost of a generation is shown on the Generate button before you commit.\n\n- A **Slates Standard** license ($149 one time) starts you with **1,000 credits**.\n- **Slates Pro** ($297 one time, or $97 to upgrade later) starts you with **3,000 credits**.\n\n| Pack | Credits (Standard) | Credits per dollar | Versus the smallest pack |\n|------|--------------------|--------------------|--------------------------|\n| $10 | 250 | 25.0 | standard rate |\n| $25 | 650 | 26.0 | +4% more credits |\n| $50 | 1,375 | 27.5 | +10% more credits |\n| $100 | 3,000 | 30.0 | +20% more credits |\n| $250 | 8,000 | 32.0 | +28% more credits |\n| $500 | 17,000 | 34.0 | +36% more credits |\n| $1,000 | 35,000 | 35.0 | +40% more credits |\n\nPacks up to $500 are open to everyone; the $1,000 pack is offered inside the app to licensed accounts. Slates Pro receives more credits than the Standard column above on every pack, for life.\n<!-- END:GENERATED credits -->\r\n\r\nCredits NEVER expire, there is no monthly reset, and failed generations refund immediately. You can also turn on auto-topup so your balance refills when it runs low.\r\n\r\n### Standard vs Pro\r\n- **Standard:** the app, every AI model, and pay-as-you-go credits that never expire, plus 12 months of updates.\r\n- **Slates Pro:** everything in Standard, plus our lowest credit rate on every pack, forever (buy the smallest pack and pay the largest pack's rate), **4K video generation**, a priority generation queue, early access to every new model on release day, and lifetime updates. The more you top up, the more the better rate adds up.\r\n\r\nEvery AI model is available on both tiers. The only capability gated to Pro is generating 4K video; 4K images are open to everyone, and exporting your timeline at 4K is available on every tier.\r\n\r\n### 30-Day Guarantee\r\nFull refund within 30 days, no questions asked.\r\n\r\n---\r\n\r\n## KEYBOARD SHORTCUTS\r\n\r\n| Key | Action |\r\n|-----|--------|\r\n| Space | Play/pause |\r\n| V | Select tool |\r\n| C | Razor tool |\r\n| S | Slip tool / snap toggle |\r\n| M | Add marker |\r\n| Left/Right | Frame-by-frame |\r\n| Up/Down | Seek +-1 second |\r\n| Ctrl+Z | Undo |\r\n| Ctrl+Shift+Z | Redo |\r\n| Ctrl+Plus/Minus | Zoom timeline |\r\n| Delete/Backspace | Delete selected clip |\r\n| Escape | Close modal/viewer/slideshow |\r\n| Ctrl+Enter | Submit generation |\r\n| Home/End | Jump to timeline start/end |\r\n| Page Up/Down | Jump by screen width |\r\n\r\n---\r\n\r\n## OFFLINE USAGE\r\n\r\nThe app launches and works offline for everything except AI generation and login. Specifically:\r\n\r\n**Works offline:** Opening projects, viewing all assets (images/videos), editing timeline (trim, reorder, split clips), adding markers, slideshow playback, FFmpeg export to MP4, DaVinci XML export.\r\n\r\n**Requires internet:** AI generation (all models), login/signup, credit purchases, credit balance sync, license validation (only checked on first generation attempt per session, then cached), auto-updater.\r\n\r\nIf you lose internet mid-session, you can keep editing and exporting. Generation will fail until connectivity returns.\r\n\r\n---\r\n\r\n## GENERATION RECOVERY\r\n\r\nIf the app closes during a generation: on restart, Slates detects in-flight jobs, polls the AI provider, and downloads completed results automatically. Nothing is lost. Recovering generations show at 5% in the queue until status is confirmed. Works for every model.\r\n\r\n---\r\n\r\n## TROUBLESHOOTING\r\n\r\n**\"Insufficient credits\"** — Your credit balance is too low for this generation. Buy more credits in the app (packs from $10 to $500) or turn on auto-topup. The exact cost of any generation is shown on the Generate button before you commit.\r\n\r\n**\"Input was rejected by Kling\"** — Image may not meet quality requirements (character visibility, proportions, content policy). Try a different image or prompt.\r\n\r\n**\"Failed to upload image to FAL CDN\"** — Network issue during reference image upload. Check internet connection, retry.\r\n\r\n**\"Generation failed\" / \"Proxy generation failed\"** — Generic error from the AI provider. Usually temporary. Retry. If persistent, try a different model.\r\n\r\n**\"Source asset not found\" / \"Source video asset not found\" / \"Target image asset not found\"** — The image or video you're trying to use was deleted or moved. Re-import or select a different asset.\r\n\r\n**\"Invalid audio source\"** — Lip sync: either enter TTS text or upload an audio file. One is required.\r\n\r\n**\"TTS response missing audio URL\"** — Text-to-speech failed during lip sync. Retry.\r\n\r\n**Generation stuck** — Restart app. Recovery system polls providers and picks up where it left off.\r\n\r\n**API rate limit** — Too many requests (limit: 20 generations/minute). Wait 1-2 minutes, retry.\r\n\r\n**Project files missing** — Project folder was moved/deleted outside the app. Use project relocation in Settings to re-point to the correct folder.\r\n\r\n**License shows \"revoked\"** — Contact support. Character sheets and environment grids unavailable until resolved.\r\n\r\n**Session expired** — Magic link session timed out. Log in again via Settings.\r\n\r\n**iPhone photos won't import** — iPhones save photos as HEIC/HEIF format, which Slates doesn't support. Convert to JPEG or PNG first (most photo apps and online converters can do this).\r\n\r\n---\r\n\r\n## PRIVACY & DATA\r\n\r\n- Generated files stay on YOUR machine. Slates servers never store your videos/images.\r\n- No prompts logged server-side.\r\n- File uploads go directly to the AI provider via pre-signed URLs. Slates servers never buffer your media.\r\n- Server stores only: email, license status, credit balance, transaction history, session tokens.\r\n- Stripe handles all payment data. Slates never sees your card number.\r\n\r\n---\r\n\r\n## SYSTEM REQUIREMENTS\r\n\r\n- Windows 10/11 or macOS 12+\r\n- Internet connection required for AI generation (not for editing/exporting)\r\n- Disk space for project files (AI videos are typically 5-50MB each)\r\n- FFmpeg bundled with app (no separate install needed)\r\n- No GPU required (all AI processing happens in the cloud)\r\n\r\n---\r\n\r\n## COMMON TASKS (STEP-BY-STEP)\r\n\r\n### Generate an Image\r\n1. Open the floating prompt box (visible on every page).\r\n2. Enter your prompt describing the image.\r\n3. Select an image model (Nano Banana 2 recommended).\r\n4. Choose aspect ratio and resolution.\r\n5. Press Ctrl+Enter or click Generate.\r\n\r\n### Generate Video From an Image\r\n1. In the prompt box, attach a start image.\r\n2. Write a prompt describing the desired motion/action.\r\n3. Select a video model (Kling V3.0 Omni in ingredients mode recommended).\r\n4. Choose duration (Kling bills per second from 3s up, so shorter is always cheaper), aspect ratio, and resolution.\r\n5. Click Generate.\r\n\r\n### Use a Character Reference for Consistency\r\n1. Create a character in your project (name + description).\r\n2. Either generate a character sheet OR manually assign a single image as the character reference.\r\n3. In the prompt box, type `@` and select your character from auto-complete.\r\n4. The reference image is attached automatically. Generate normally.\r\n\r\n### Export to DaVinci Resolve for Color Grading\r\n1. In the video editor, finalize your timeline (clips, markers, timing).\r\n2. Click Export → DaVinci Resolve XML.\r\n3. Choose output location. File saves to exports/ directory.\r\n4. In DaVinci Resolve: File → Import → Timeline. Select the XML file.\r\n5. Your timeline loads with all clips, properties, and markers intact. Grade and export.\r\n\r\n### Extract a Still Frame From a Video\r\n1. Hover over any video clip in the gallery.\r\n2. Camera icon = extract **current frame**. Dropdown arrow next to it = **First frame** or **Last frame**.\r\n3. Extracted image saves to your project gallery. Use as start/end image for I2V, character reference, or storyboard frame.\r\n\r\nKey workflow: extract a clip's last frame → use it as the start image for the next generation → seamless visual continuity between scenes.\r\n\r\n### Buy More Credits\r\n1. Open Settings → Credits, or the credit badge in the top nav.\r\n2. Pick a pack (packs run from $10 to $500, plus a $1,000 pack offered in the app to licensed accounts; bigger packs give more credits per dollar).\r\n3. Pay via Stripe. Credits are added to your balance instantly and never expire.\r\n4. Optional: turn on auto-topup so your balance refills automatically when it runs low.\r\n\r\n---\r\n\r\n## COMMON QUESTIONS\r\n\r\n**Q: Which model should I use for most videos?**\r\nA: Seedance 2.0 is the default and the one to reach for when physics, scale, effects or hero shots matter. For everyday shots built from a start image, Kling V3.0 Omni in ingredients mode is the best balance of cost and quality, which is why the step-by-step guides above use it.\r\n\r\n**Q: What's the best image model?**\r\nA: GPT Image 2.5 is the strongest image model in the app, and the best available anywhere right now. It holds a long instruction more faithfully than anything else here, and it is the only model that renders words in the picture reliably. The two seats cost the same, so the only trade is time: **Flare** is the fast one and is best for drafts and exploring, **Sunburst** is the higher-quality one and is what finals, hero frames and reference-heavy edits should end up on. The quality knob has five settings spanning about 36× from cheapest to dearest — `max` is four times `high`, though the smaller steps are uneven: **`medium`** is for drafts and where iteration belongs, **`high` is the default** and where most finished work should sit, and **`max` is the best output you can get** — finished frames, client deliverables, anything with exact text. `xhigh` sits just under `max` for about half the price and is worth trying before you jump to the top. Going past `high` should be a deliberate choice rather than a habit. Go to **4K only once you already know your references and your prompt are solid** — it is the most expensive setting on the model and the worst place to discover the composition was wrong. Prove the shot at 3K first, then re-run the settled prompt at 4K. If you used GPT Image 2 before, note the quality names all shifted by one: its `medium` is now `high`, its `high` is now `max`. Nano Banana 2 is the picker's DEFAULT rather than the best: it is fast, takes the most reference images, and is the only model with grid exploration. Nano Banana Pro is the hero-frame step up when composition, cinematic lighting and skin have to be perfect. Prices for every tier are in the model table above.\r\n\r\n**Q: Do I need to set up API keys?**\r\nA: No. There are no API keys in Slates — every generation runs on Slates Credits, which come with your license and never expire.\r\n\r\n**Q: How much does a generation cost?**\r\nA: It depends on the model, resolution, and length. The exact credit cost is always shown on the Generate button before you commit, so there are no surprises.\r\n\r\n**Q: Can I use Slates offline?**\r\nA: Yes for viewing projects, editing timeline, and exporting. No for AI generation -- that requires internet.\r\n\r\n**Q: Do credits expire?**\r\nA: No. Credits never expire.\r\n\r\n**Q: What happens if I close the app during a generation?**\r\nA: Nothing is lost. On restart, Slates detects in-flight jobs and downloads completed results automatically.\r\n\r\n---\r\n\r\n## FEATURES NOT IN SLATES\r\n\r\nThe following are NOT available. Do not suggest them:\r\n\r\n- Bring-your-own API keys (BYOK) — every generation runs on Slates Credits; there is no key-entry option\r\n- Local/on-device GPU inference (all AI runs in the cloud)\r\n- Built-in music generation (use external tools like Suno, import audio)\r\n- A voice picker for Seed Audio — you describe the voice you want in words instead. The preset voice shelf belongs to the Voice lane (Inworld TTS-2), and Kling Lip-Sync keeps its own small fixed list: six English/UK voices plus a storyteller, with a speed control\r\n- Automatic video editing from a script\r\n- A prompt \"Enhance\" button — ask Studio Agent to rewrite a prompt instead\r\n- A settings/gear panel on the prompt box — every parameter is a dropdown on the bar\r\n- Cloud project storage (all files are local)\r\n- Real-time collaboration / multi-user editing\r\n- Mobile app (desktop only: Windows and macOS)\r\n- HEIC/HEIF image import (convert to JPEG/PNG first)\r\n- Storyboard JSON export (import only)\r\n\r\n---\r\n\r\n## VERSION\r\n\r\n<!-- BEGIN:GENERATED version -->\nSlates Reference Version: 1.5.6\nLast Updated: 2026-09-09\n\nThis document is generated. Its source of truth is `slate/docs/slates-llm-manual.md`; its model tables and credit costs are derived from the Slates model registry and pricing tables at build time, so they cannot be typed by hand.\n\nIf the user asks about a feature not documented here, it may have been added after this version. The current copy is always at <https://slates.video/slates-reference.md>.\n<!-- END:GENERATED version -->\r\n\r\nIf this document didn't answer your question, email hello@slates.video so we can help and improve the app.\r\n\r\n</slates_reference>\r\n";
|
|
2
|
+
export const APP_MANUAL = "# Slates — Complete Reference for AI Assistants\r\n\r\n<!-- The heading above must stay first: shipped app builds reject this file if it\r\n does not begin with one. See slate/CLAUDE.md; check:llm-docs enforces it. -->\r\n\r\n<system_role>\r\nYou are a support assistant for the Slates desktop application. Answer user questions using ONLY the information in the <slates_reference> below. Be concise and direct. Use numbered steps for procedures. Use bullet points for explanations when helpful.\r\n</system_role>\r\n\r\n<rules>\r\n- If the answer cannot be found in the <slates_reference>, say: \"That isn't covered in the Slates reference.\" Do not guess or invent features.\r\n- If the user asks how to do something, give step-by-step instructions from the workflows and features described here.\r\n- If the user reports an error, check the TROUBLESHOOTING section first.\r\n- Refer to the FEATURES NOT IN SLATES section before answering questions about capabilities that might not exist.\r\n- Quote the exact error message when referencing troubleshooting entries.\r\n</rules>\r\n\r\n<slates_reference>\r\n\r\n<!-- BEGIN:GENERATED header -->\n# SLATES v1.5.6 — Complete Reference\n\n> **Freshness.** Generated from the Slates source of truth for app version **1.5.6**, last changed **2026-09-09**. The canonical copy of this file is <https://slates.video/slates-reference.md>. If a model, price or feature the user mentions is missing below, this copy is out of date: re-fetch that URL before answering, and say so.\n<!-- END:GENERATED header -->\r\n\r\n**What is Slates?** Desktop app (Windows 10/11, macOS 12+) for AI image and video creation. One-time purchase, no subscription. Every license includes 1,000 free credits, and Slates Pro starts with 3,000 credits. Every generation runs on Slates Credits — there are no API keys to set up, and credits never expire.\r\n\r\n---\r\n\r\n## INTENDED WORKFLOW\r\n\r\nThe designed start-to-finish flow:\r\n\r\n1. **Create project** — New project with name/description. Creates folder on your disk.\r\n2. **Build visual assets** — Generate images, create characters (with character sheets for consistency), environments (with environment grids), and styles. This is your visual library.\r\n3. **Create storyboard** — Add scenes. Every picture you drop in becomes a **Shot**: one beat of the piece, holding its references, its model, its settings and (when you want them) its words.\r\n4. **Write the piece** — Switch the storyboard to **Document** and write. What is said, what happens, the framing, the prompt each beat will send. Nothing is required and nothing is asked for; a visuals-only piece is finished as it stands.\r\n5. **Read it before you pay for it** — The header states how many generations, how many cuts, how long it runs and what it will cost. Press play for a rough cut at the real timing. Re-chop with split and merge and watch the price move.\r\n6. **Generate** — Select the beats you want and fire them in one approved batch. Each result lands under the row that made it.\r\n7. **Organize** — Switch to **Board** to re-order. Drag a beat and the whole thing moves with it.\r\n8. **Export to timeline** — Send clips to the built-in multi-track video editor.\r\n9. **Edit** — Trim, reorder, add markers, adjust timing.\r\n10. **Final export** — Export to MP4 directly, or export DaVinci Resolve XML for professional color grading.\r\n\r\n---\r\n\r\n## MODEL REFERENCE TABLE\r\n\r\n**Generating 4K video is a Slates Pro feature** — every tier generates video up to 1080p, and 4K images are open to everyone. Exporting your finished timeline at 4K is available on every tier.\r\n\r\nEvery model runs on Slates Credits. **The exact credit cost appears on the Generate button before anything fires.** The tables below are generated from the app's own model registry and rate tables, so they describe exactly what the model picker offers in this version: aspect ratios, resolutions, durations, reference-image limits, and the credit price of each. Bigger credit packs lower your per-credit cost, and Slates Pro gets the best pack rate on every purchase.\r\n\r\n<!-- BEGIN:GENERATED model-tables -->\n### Image Models\n\n| Model | Aspect Ratios | Resolutions | Max Refs | Credits per image |\n|-------|--------------|-------------|----------|-------------------|\n| **GPT Image 2.5 Flare** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 · 3K 3 · 4K 5 (default quality; at max: 2K 8 · 3K 11 · 4K 20) |\n| **GPT Image 2.5 Sunburst** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 · 3K 3 · 4K 5 (default quality; at max: 2K 8 · 3K 11 · 4K 20) |\n| **Nano Banana 2** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 4 · 2K 6 · 4K 8 |\n| **NB2 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K | 4 | 1K 2 |\n| **Nano Banana Pro** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 8 · 2K 8 · 4K 15 |\n| **FLUX.2 Max** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 4 | 1K 4 · 2K 5 · 4K 8 |\n| **Seedream 5 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 2K / 3K / 4K | 10 | 2K 2 · 3K 2 · 4K 2 |\n\n### Video Models\n\n| Model | Duration | Aspect Ratios | Resolutions | Max Refs | Audio | Credits per second |\n|-------|----------|--------------|-------------|----------|-------|--------------------|\n| **Seedance 2.0** | 4-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p / 4K | 9 | Included | 480p 3.5 · 720p 7.5 · 1080p 18.5 · 4K 39 |\n| **Seedance 2.5** | 4-30s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 30 | Included | 480p 5.1 · 720p 11.6 · 1080p 20.5 |\n| **Seedance 2.5 Edit** | Follows the source clip (4-30s) | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 0 | Included | 480p 6.2 · 720p 13.9 · 1080p 24.6 |\n| **Kling V3.0 Standard** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (6.3 with audio) · 4K 21 |\n| **Kling V3.0 Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (8.4 with audio) · 4K 21 |\n| **Kling V3.0 Omni** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (5.6 with audio) · 4K 21 |\n| **Kling V3.0 Omni Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (7 with audio) · 4K 21 |\n| **Kling O3 Edit** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 6.3 |\n| **Kling O3 Edit Pro** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 8.4 |\n| **MiniMax H3 Max** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 1080p | 9 | Included | 480p 2.5 · 768p 4 · 1080p 8 |\n| **MiniMax H3** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 2K / 4K | 9 | Included | 480p 2.5 · 768p 3 · 2K 6.5 · 4K 8 |\n| **Gemini Omni Flash** | 3-10s | 16:9, 9:16 | 720p | 7 | Included | 720p 6.4 |\n| **Omni Flash Edit** | Follows the source clip (3-10s) | 16:9, 9:16 | 720p | 0 | Included | 720p 6.4 |\n| **LTX-2.5** | 6/8/10/12/14/16/18/20s at 720p/1080p; 6/8/10s at 1440p/4K | 16:9, 9:16 | 720p / 1080p / 1440p / 4K | 0 | Included | 720p 4.5 · 1080p 6.5 · 1440p 9.5 · 4K 15 |\n| **LTX-2.5 Pro** | 6, 8, 10s | 16:9, 9:16 | 720p / 1080p | 0 | Included | 720p 6 · 1080p 8.5 |\n| **Veo 3.1 Fast** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 5 (7.5 with audio) · 1080p 5 (7.5 with audio) · 4K 15 (17.5 with audio) |\n| **Veo 3.1 Standard** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 10 (20 with audio) · 1080p 10 (20 with audio) · 4K 20 (30 with audio) |\n\n**Seedance 2.0 · Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 27.4 credits per second at 1080p.\n**Seedance 2.5 · Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 16.3 credits per second at 720p.\n**Seedance 2.5 Edit · Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 19.8 credits per second at 720p.\n\n### Audio Models\n\n| Model | Length | Credits |\n|-------|--------|---------|\n| **Seed Audio 1.0** | 3-120s | 1 at 3s · 3 at 15s · 19 at 120s |\n| **Inworld TTS-2** | up to 2,000 characters of text per take | 1 at 250 characters · 2 at 2,000 characters (billed per 250) |\n| **Sound Effects** | 1-22s | 1 at 1s · 1 at 4s · 3 at 22s |\n\n### Tools (Lip Sync, Motion Transfer)\n\nThese are real Kling endpoints that take a clip or a still as their subject, not models you prompt from scratch. Both bill in 5-second blocks.\n\n| Tool | Input | Billed in | Credits per block |\n|------|-------|-----------|-------------------|\n| **Kling Lip Sync** | Video source | 5s block | 4 |\n| **Kling Lip Sync (Avatar v2 Standard)** | Still-image source | 5s block | 14 |\n| **Kling Lip Sync (Avatar v2 Pro)** | Still-image source | 5s block | 29 |\n| **Kling Motion Control Standard** | Still image + reference video | 5s block | 32 |\n| **Kling Motion Control Pro** | Still image + reference video | 5s block | 42 |\n<!-- END:GENERATED model-tables -->\r\n\r\n---\r\n\r\n## WHICH MODEL TO USE\r\n\r\n### Images\r\n\r\n**Nano Banana 2 is the default image model.** It is the best all-round image model in the app: the most reference images of any image model, every aspect ratio, and output up to 4K. Brief it like a creative director rather than with tag soup. It is also the only model that supports the 2x2 / 3x3 grid exploration wrapper.\r\n\r\n- **NB2 Lite** is the fast, cheap draft seat in the same Nano Banana family. Roughly half the price of NB2 full and noticeably faster, 1K output only. Iterate here, finish on NB2.\r\n- **Nano Banana Pro** is the hero-frame and typography tier. Reach for it when spatial composition, cinematic lighting and skin, or fine in-image type have to be perfect. NB2 gets you most of the way there, so this is a deliberate step up, never a default.\r\n- **GPT Image 2.5** is the strongest image model in the app: it follows a long instruction more faithfully than anything else here, and it is the one to pick when the picture simply has to be right. It is also the sharp-text model, which is what makes it the choice for character sheets, shot grids, ordered panels and anything with words in the picture. It comes in two seats that cost exactly the same, and the difference is speed against quality. **Flare** is the fast one: OpenAI describes its quality as comparable to the older GPT Image 2, at roughly half the wait. **Sunburst** is OpenAI's most capable image model, better than GPT Image 2, and deliberately slower. Use Flare while you are still exploring, then re-run the shot you like on Sunburst for the final — and reach for Sunburst directly when several reference images all have to survive into one frame, or when an edit must change one region and leave identity, geometry and lighting untouched.\n Its quality knob has **five** settings — `low`, `medium`, `high`, `xhigh`, `max` — spanning about 36× from cheapest to dearest, which makes it the biggest cost lever on the model. The steps are uneven rather than a constant multiplier: `max` is four times `high`, but `xhigh` is only about 1.8 times it. **`high` is the default and the everyday setting.** `medium` is for drafts and is cheap enough to iterate on freely. `max` is the ceiling, for finished frames and exact character-level text; `xhigh` sits just under it for about half the price and is worth trying first. Go past `high` deliberately, not by habit.\n 4K is worth it only once your references and prompt are already settled: prove the shot at 3K, then re-run the finished prompt at 4K. Iterating at 4K is the most common way to waste credits on this model. Its 4K tier is API-only, so even a paid ChatGPT account cannot render it. Note that its resolution tiers are **not** a price ladder: the pixel classes are token-priced by OpenAI, so the cheapest seat is not the smallest one. Read the prices in the table above rather than assuming.\n **If you have used GPT Image 2 before, the quality names all shifted by one.** What it called `medium` is now called `high`, and what it called `high` is now `max` — the same pictures at the same prices, renamed. A remembered setting will quietly buy you a cheaper tier than it used to.\r\n **It is the only image model that can give you a transparent background.** Set **Background** to *Transparent* on the prompt bar and you get a real alpha channel — a cut-out for a logo, sticker or overlay — rather than a painted-in backdrop. *Auto* is the default and lets the model decide from your prompt; *Opaque* forces a filled background. It costs nothing either way. Slates always saves PNG, which is what carries the transparency, so there is nothing else to set.\r\n **Square and 4:3 frames cost more than 16:9 on this model, and only on this model.** OpenAI charges by image tokens rather than by pixels, and a square frame uses about 1.8 times the tokens of a 16:9 frame the same size, and 4:3 or 3:4 about 1.37 times. The credit prices in the table above are the 16:9 numbers; pick 1:1 or 4:3 and the price on the Generate button goes up to match. 9:16 costs the same as 16:9. Every other image model charges the same whatever the shape.\r\n- **FLUX.2 Max** and **Seedream 5 Lite** are the less content-restricted options. Seedream is flat-priced at every resolution it offers, so there is no reason to pick a lower one. Both auto-route to their edit endpoint when you attach reference images.\r\n\r\n### Video\r\n\r\n**Seedance 2.0 is the default video model.** Reach for it the moment physics, effects, destruction or scale matter, and for hero shots. It takes many reference images, generates native audio at no extra cost, and is the only Seedance seat that reaches 4K (generating 4K video needs Slates Pro; timeline export at 4K does not). It is also the cheaper of the two seats at every resolution they share. A face in a reference image routes it to a different provider, which is why **Seedance 2.0 · Face** is its own row in the model picker at its own price.\r\n\r\n- **Seedance 2.5 is a second seat, not an upgrade.** It buys much longer single takes and far more reference images, plus better prompt adherence. What it gives up is 4K, and it costs more than 2.0 at every resolution the two share — so 2.0 stays the model for 4K, and for the same resolution at a lower price. Because 2.5 runs longer, a long clip on 2.5 can cost more than a shorter, higher-resolution one on 2.0 — read the Generate button, not the resolution. **Seedance 2.5 Edit** is its clip-editing row: attach a clip, describe the change, and the output length follows the source.\r\n- **Kling** is the cost-effective workhorse and the most flexible family: strong start-frame adherence for identity, layout and text, acting, dialogue, multi-shot (up to 6 cuts), and the widest range of clip lengths. **Kling V3.0 Omni** adds multi-character dialogue in English, Chinese, Japanese, Korean and Spanish. Standard and Pro are the same model at two fidelity and price tiers. **Kling O3 Edit** takes an existing clip and changes what you describe, with subject and style reference images, while the original audio is preserved verbatim. Kling is also the only engine behind the Lip Sync and Motion Control tools.\r\n- **MiniMax H3** is the seat to pick when the SOUND is part of what you are writing. Every other video model treats audio as a switch; H3 takes it as three separate instructions in one prompt — the lines and action sounds tied to a moment, the ambience running underneath, and a score only the audience hears — and generates all of it with the picture in a single pass. It is also the only model where you say how much of a reference should survive, including moving one subject's characteristic onto a different subject. It runs 5-15 seconds at 480p, 768p, 2K or 4K, and takes up to nine reference images plus reference video and audio. Two things to watch: **the first five reference images are free and every one after that costs extra**, so attach what the shot needs rather than the maximum; and 2K and 4K are upscales of a 768p render rather than larger generations — in our own testing the 2K pass showed more artifacting than the 768p original it was built from, at more than twice the price. Generate and judge at 768p; step up only when a delivery spec demands the pixels.\r\n- **MiniMax H3 Max** is the same model post-trained by fal for SPEED, and it is the more expensive seat, not the cheaper one. It is dramatically faster: on the same 5-second 768p prompt it finished in about 5 seconds against about 57 seconds for H3 — roughly 12x (measured 2026-08-27). It stops at 768p and costs more per second than H3 at the resolution they share. It still animates a start frame and an end frame, so image-to-video works normally; what it does not have is the reference set — the extra identity, style and environment images plus reference video and audio that base H3 reads. Pick it when a fast turnaround on a text-to-video or start-frame shot is worth paying for; pick H3 for resolution, references, or the same tier at a lower price.\r\n- **LTX-2.5** is the VOLUME seat — the cheapest native 1080p second in the catalogue, with synchronised audio included free at every resolution, so it is the model to reach for when the job is many takes rather than one hero shot. Two things are unique to it. It makes the LONGEST clips of anything here, up to 20 seconds, and it is the only model that reaches 1440p. It is also the only one with native MULTISHOT: a single generation can carry two to four connected shots that hold the character, lighting and voice across the cuts, which everywhere else means generating separate clips and watching identity drift between them. Its constraints are unusually sharp, though. Durations are EVEN NUMBERS ONLY starting at six — 6, 8, 10, 12, 14, 16, 18, 20, with no 5-second or 7-second clip — and above 1080p that ceiling drops to 10 seconds. Aspect ratios are 16:9 and 9:16 only. And it takes FRAMES, not references: a start frame and an optional end frame that generates a transition between them, but no identity, style or environment reference images at all, so cross-shot character consistency belongs on MiniMax H3 or Kling. Because sound is generated in the same pass, write the audio into the prompt and anchor every cue to something visible — anything unanchored gets invented.\r\n- **LTX-2.5 Pro** is the fidelity seat of that pair, and it is NOT simply a better LTX. It renders the picture with more compute on busy frames, but on a narrower envelope than the base row: 720p and 1080p only (no 1440p, no 4K) and 6, 8 or 10 seconds only, for about a third more per second. Reaching for it because the name says Pro costs more AND takes away the reach. Pick it when a specific shot needs the extra fidelity and fits inside 1080p and ten seconds; pick base LTX for length, resolution and volume.\r\n- **Gemini Omni Flash** is the cheap 720p seat with native synced audio included in one pass. **Omni Flash Edit** is the prompt-only clip editor: no reference images, one short instruction plus \"Keep everything else the same.\" Long descriptive prompts destroy it.\r\n- **Veo 3.1** is niche and is never a default. Pick it only when you specifically want Google's audio pass. It has the fewest aspect ratios and reference slots of any video model, fixed durations, and the highest per-clip cost.\r\n\r\nBoth edit models take an existing clip as their canvas, so their output length follows the source clip rather than a duration you choose.\r\n\r\n### Audio\r\n\r\nAudio is a third media type alongside images and video — generated as its own asset, shown in the gallery's **Audio** tab, and dragged onto an audio track in the timeline. This is separate from the audio some VIDEO models generate *inside* a clip (see AUDIO IN GENERATION below): use a video model when the sound must be locked to what is on screen, and these when you need audio you can move, trim, re-use, or layer.\r\n\r\n**Seed Audio 1.0 is the default.** A room with dialogue *and* clatter *and* ambience is one generation, not three layered ones, and because it is cheap you can run five takes and keep the best. It makes a whole audio SCENE from one plain sentence. **It has no length setting of its own** — Slates writes your chosen duration into the prompt, and that is what you are charged. You describe the voice in words; there is no voice list to pick from. Set **Languages** to Mixed if one scene needs more than one language (it costs the same).\r\n\r\n**Sound Effects** makes one effect, or a seamless loop. It is the only surface with an exact duration, so an effect can land on a specific frame. Describe the physical cause (\"heavy oak door slams shut in a stone hallway\"), not the label (\"door sound\"). **Loop** makes it seamless for beds; **Wording** controls how literally your description is followed. Seed Audio is actually the better tool for *long* ambience beds, so the two are not redundant in the direction you would expect.\r\n\r\nKling's `SFX:` / `Ambient noise:` prompt syntax belongs to video prompts and makes Seed Audio results *worse* — write plain sentences there instead.\r\n\r\n**Inworld TTS-2 is the voice seat** — type the words, pick a voice, press Generate. In the prompt box's Audio lane pick **Voice** in the model picker; the prompt is the exact text that gets spoken (nothing is added or rewritten — open \"See what gets sent\" to confirm), and the **Voice** control on the bar opens the voice picker: **Presets** (ready-made voices with gender, accent and age filters — every one plays the same audition line, so you compare voices rather than scripts), **Clips** (any character's voice, or any audio clip in the project, cloned for the take), or **Describe** (a voice in words). The character counter beside the bar is the bill: the generated audio table above gives the text cap and billing buckets, and the Generate button shows the price. Direction goes in square brackets (`[whispering] …`) — anything in parentheses is read aloud. Cloning a real person's voice needs their permission. Right-click any audio clip in the Audio tab → **Use as voice** lands you on the Voice lane with that clip as the voice. Studio Agent and the MCP/CLI do the same through `slates_generate_audio` (a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId`, or a `voiceDescription`).\r\n\r\n**There is no music generation.** For a song, use an external tool and import the audio (see PROJECTS → Supported File Formats). For spoken lines inside a scene, let Seed Audio perform them, put them in the video prompt on a model with native audio (Seedance, Kling Omni, Omni Flash, Veo), or use Kling Lip-Sync's text-to-speech against a shot.\r\n\r\n### Tools\r\n\r\n**Tools** is not a model family. It is two real Kling endpoints that take a clip or a still as their subject: **Kling Lip Sync** and **Kling Motion Control**. Both bill in 5-second blocks and are described under GENERATION MODES below.\r\n\r\n**How pricing works:** every generation is priced in Slates Credits, and the exact cost is shown on the Generate button before you commit. Bigger credit packs give more credits per dollar; Slates Pro locks in the best pack rate on every purchase, forever.\r\n\r\n---\r\n\r\n## GENERATION MODES\r\n\r\n### Create Image\r\nPrompt → select image model → set aspect ratio + resolution → generate. Batch grids available for quick iteration.\r\n\r\n### Text-to-Video\r\nPrompt → select video model → set duration + aspect ratio + resolution → generate. Output: MP4.\r\n\r\n### Image-to-Video (I2V)\r\nAttach start image + prompt → select model → generate video from that image. Optional: attach end image (Veo) for guided transitions.\r\n\r\n### Ingredients / References\r\nUse @character_name, @environment_name, or #style_name in prompt to attach reference images for visual consistency. Kling: up to 4 total references. Veo: up to 3. Nano Banana 2: up to 14. The @mentions auto-complete from your project's characters/environments/styles.\r\n\r\n### Lip Sync\r\n**Kling only.** Pick **Kling Lip Sync** under the Tools family in the model picker. Source: video or still image.\r\n\r\n- **Audio source** — Text to speech (type the line; six English/UK voices plus a storyteller, with a speed control) OR Upload audio (bring your own recording, max 5MB).\r\n- **Avatar tier** — only appears for a still-image source: Avatar v2 Standard (the value tier) or Avatar v2 Pro (higher fidelity, higher rate).\r\n- 5s output blocks.\r\n\r\nWorks very well with human-like characters. Less reliable with animals or non-human characters.\r\n\r\n### Motion Transfer\r\n**Kling only.** Pick **Kling Motion Control** under the Tools family. Target: still image (your character). Source: reference video (the motion).\r\n\r\n- **Engine** — Kling MC Standard (value tier) or Kling MC Pro (higher fidelity, higher rate).\r\n- **Orientation** — *Match video* copies skeleton and depth from the clip (best for dancing, walking, full-body action; driving clips up to 30s). *Match image* keeps your character's pose and angle and uses the video only as motion hints (best for close-ups; up to 10s).\r\n- 5s output.\r\n\r\n> **Note:** these two tools used to offer a second \"Seedance 2.0\" engine. It was not a separate engine — picking it made Slates write a sentence into your prompt that you never saw, which is no longer allowed anywhere in the app (see \"What gets sent\" below). Both tools are now Kling endpoints only.\r\n\r\n### Edit Image\r\nRight-click any image asset → open viewer → switch to Edit Mode. Enter an edit prompt describing the changes you want. Select edit model: Nano Banana 2 (supports up to 14 reference images), FLUX.2 Max, or Seedream 5 Lite. Choose resolution and aspect ratio. The result saves as a new asset with the original preserved. Useful for refining generated images without starting from scratch.\r\n\r\n### Edit Video (Kling O3 Edit / Omni Flash Edit)\r\nRight-click any video clip (gallery or timeline) → \"Edit with AI\". The clip attaches to the prompt box as the source; describe the CHANGE, not the whole scene (\"replace the man with @marcus\", \"make it a rainy night, keep everything else\"). Two engines in the model picker:\r\n- **Kling O3 Edit (default):** attach subject images (role: Subject) to swap someone in, or style images (role: Style) for a look — max 4 combined refs. Clips 3-15s. Original audio preserved.\r\n- **Omni Flash Edit (cheapest):** prompt only — no reference images; keep instructions simple and add \"Keep everything else the same.\" Clips 3-10s, 720p output.\r\n\r\nOutput length follows the source clip; the credit cost (clip seconds, rounded up, at the per-second rate) shows on the Generate button. The edited clip saves as a NEW asset linked to the original — chain edits freely. Trim longer clips on the timeline first.\r\n\r\n### Multi-Shot (Kling V3.0/Omni)\r\nEnable multi-shot toggle → multiple scene prompts in one generation, each with different framing. 6-axis camera controls per shot. Results can be hit-or-miss, but worth trying for quick multi-cut sequences. For more reliable results, most users prefer generating multiple short 5s clips separately using Kling V3.0 Omni in ingredients mode and assembling them on the timeline.\r\n\r\n---\r\n\r\n### Generate Audio\r\n\r\nSwitch the prompt box's lane pill from Image/Video to **Audio**, pick a surface, and generate. The prompt box offers three: **Seed Audio 1.0**, **Voice** (Inworld TTS-2) and **Sound Effects**. The result lands in the gallery's Audio tab as its own asset with a waveform and an inline player, and can be dragged onto an audio track in the timeline.\r\n\r\n- **Scene (Seed Audio 1.0)** — one plain sentence describing the moment. Set **Length**; Slates writes it into the prompt for you and that is exactly what you're billed for (open \"See what gets sent\" under the prompt box to read the appended text). **Say the crowd/room size out loud** — \"applause\" returns a full auditorium when you meant three people at an open mic. Ask for a few seconds more than the clip needs so the edit has fade handles. Describe the voice you want in the sentence itself (\"a weary dock foreman in his fifties, gravel in his voice\") — there is no voice picker.\r\n- **Voice (Inworld TTS-2)** — the prompt is the words to be spoken, verbatim. Pick the voice with the **Voice** control on the bar (presets you can play first, any clip in the project, a character's voice, or a description); the character counter is the bill. See MODELS → Inworld TTS-2 for direction tags and the cloning rules.\r\n- **Sound Effect** — describe the physical cause and set the length to roughly the event (≈1s for an impact, 2–4s for a whoosh, 8–22s + **Loop** for a bed). **Wording** sets how literally the description is followed: Interpretive, Balanced (default), or Literal.\n\n#### Use your own voice recording\n\nIn the bottom prompt box, choose **Audio**, then **Inworld TTS-2** in the model picker. Open **Voice → Clips → Import voice clip** and select your recording. The import adds an audio asset to this project without generating anything. Click its play button to audition it, then click the recording's name to choose it. Type the words you want spoken in the prompt box and press **Generate**, which shows the price. The new take appears in **Gallery → Audio**.\n\nUse a clean recording of one speaker whose voice you have permission to use. Slates clones the recording for each take; there is no separate training wizard or persistent vendor voice to manage. **Presets** lets you audition ready-made voices; **Describe** lets you write a voice description and choose **Use this description**. Choosing in the prompt box sets up the next take; only Generate spends credits.\n\nTo attach your recording or a generated take to a character, open **Gallery → Characters**, then **Add voice** (or **Change voice**) on that character's card. Choose **Clips** and click the clip's name. Attaching an existing clip is free. The card displays its waveform and player. That character's voice is also listed under **Voice → Clips → Characters** in the prompt box. Selecting a preset or description from a character card generates and attaches a take; read the cost shown in that picker before choosing.\n\nFor **Inworld TTS-2**, choose the character under **Voice → Clips → Characters**. Typing `@name` does not choose a TTS voice: the prompt is spoken verbatim. Automatic voice attachment from a character mention belongs to **Seed Audio** only.\n\r\n---\r\n\r\n## AUDIO IN GENERATION\r\n\r\nThis section is about audio generated **inside a video clip**. For audio as its own asset, see Generate Audio above.\r\n\r\n**Veo 3.1 native audio:** Generates audio WITH video. Prompt syntax: `\"Hello!\"` for dialogue, `SFX: [sound]` for effects, `Ambient noise: [description]` for ambience. Max 10s dialogue. Add `(no subtitles)` to suppress text overlays.\r\n\r\n**Kling V3.0 Omni dialogue:** Multi-character dialogue with distinct voices. Languages: EN, ZH, JA, KO, ES. `Background music: [description]` for music. Max 10s dialogue.\r\n\r\n**Kling V3.0 sound co-generation:** Synchronized sound effects generated with video.\r\n\r\n⚠️ **This prompt syntax is video-only.** `SFX:`, `Ambient noise:` and `Background music:` are Kling/Veo conventions — the audio models above have no parser for them and will treat them as words in the scene.\r\n\r\n---\r\n\r\n## PROMPT SYSTEM\r\n\r\n### Unified Create Surface (roles + model-on-button)\r\nThe old mode tabs (text-to-video / frames-to-video / ingredients / create-image) are ONE \"Create\" surface. An **Image | Video | Audio pill** on the prompt bar switches your output lane — it remembers and restores the last model you used in each lane (pick Seedance once and the Video lane stays Seedance until you change it). Only the active lane shows its name; the other two are icons. Every attachment in the reference tray carries a tappable ROLE badge — Reference / First frame / Last frame / Subject / Style — you say what each attachment is; nothing is inferred. Your model choice sticks across generations and workflow actions (\"use as first frame\" keeps your chosen video model). Attaching a video via \"Edit with AI\" flips the surface into Edit Video mode. Lip Sync and Motion Transfer live under the **Tools** family in the model picker.\r\n\r\n### Floating Prompt Box — the bar holds everything\r\nPersistent across all pages. **There is no settings panel and no gear button.** Everything sits on one bottom bar, left to right:\r\n\r\n1. **Media toggle** — Image / Video / Audio.\r\n2. **Model picker** — a searchable menu plus a detached submenu. The main menu lists model families with a vendor glyph tile each; picking one opens that family's models beside it, every row carrying capability chips (resolution, clip length, audio, references) and its per-unit rate. Type to search across every model. The submenu is anchored to the row you opened it from, so it never travels. The trigger on the bar shows the model name and nothing else — no chevron, no resolution appended. Everything listed is a real model or endpoint.\r\n3. **Parameter controls** — one per setting the chosen model actually has (resolution, aspect, duration, length, quality, count, grid, face-in-reference, audio, loop, and so on). The trigger shows the current value; the explanation lives *inside* the menu as a subtitle under each option, along with what that option costs. A setting with only one possible value still shows, muted and non-interactive, so the row never changes shape.\r\n4. **Sliders for ranges.** A setting with a long list of steps (video duration, audio length) opens a ruler instead of a many-row menu. The handle moves between the values the model actually declares, so it cannot land on one the model will not accept, and the price for the selected value is shown on the ruler.\r\n5. **`More ▾`** — if the model has more parameters than fit the current window width, the extras are collected into a generated `More` dropdown automatically. Widen the window (or close the Studio Agent panel) and they move back onto the bar.\r\n6. **Generate** — reads `Generate · <cost>`. **Cost only** — the model name is not repeated on the button (it's in the model picker) and there is no send arrow. A badge on the left of the button counts generations currently running.\r\n\r\nFor text-to-speech, the character counter is always visible because text length determines the price. On other surfaces it appears near the right of the bar after roughly three quarters of the model's prompt limit. It turns red over the limit.\r\n\r\n**Collapsing:** the chevron at the top-right of the box collapses it to a single arrow — nothing else. Click the arrow to bring it back.\r\n\r\nBelow the bar the queue shows pending/active generations with cost and progress.\r\n\r\n### What gets sent (prompt transparency)\r\nUnder the prompt box is a **\"See what gets sent\"** disclosure. Open it and you see the exact text that will be transmitted, produced by the same code that builds the request — so it can never disagree with what is actually sent. It shows:\r\n\r\n- **Reference numbering** — `@sarah` becomes `Sarah (image 1)` so the model knows which attached image is which.\r\n- **Key lines for attachments you did NOT mention** — one short neutral sentence per unmentioned attachment. Mention every reference in your own words and these never generate.\r\n- **The trailing style clause** when a `#style` is attached.\r\n- **The Seed Audio duration append** — the `… N seconds` Slates adds to the end of the prompt, which is also what you are billed for.\r\n- **Grid wrapping** when 2×2 or 3×3 is on.\r\n\r\nThe row stays hidden when the composed prompt is identical to what you typed, so it only appears when there is something to show.\r\n\r\n**Unresolved `#tags` and `@mentions` are named, not silently dropped.** If you type `#noir` and there is no saved style called \"noir\", the tag is removed from the text sent to the model (a raw tag confuses every model) — but the disclosure turns red and says so by name: *\"#noir matches nothing saved — removed from what gets sent.\"* Save the style, or reword it, and the warning clears.\r\n\r\n**Nothing is ever added that you cannot read here.** No setting in the app injects prompt text; a setting changes *how* a request is made, never *what* you asked for.\r\n\r\n### Prompting guide (on the web)\r\nPer-model prompting guidance lives at <https://slates.video/docs/prompting>, linked from the bottom of Settings. It covers every model Slates offers — Video, Image, Audio — with what that model reads, what it ignores, and its gotchas, all on one page so you can compare them. Markdown copy for pasting into an LLM: <https://slates.video/docs/prompting.md>.\r\n\r\nIt is generated from the same source the Slates CLI, the MCP server and Studio Agent are built on, so the guide and the app cannot disagree. It is documentation rather than a control, which is why it is a page on the web and not a panel in the app: it has room to be read, a URL you can send someone, and it is always current rather than frozen at the version you installed.\r\n\r\n### @Mentions\r\nType `@` → auto-complete shows project characters and environments. Type `#` → shows styles. Selecting inserts the reference image(s). At send time a mention is rewritten to a numbered citation (`@sarah` → `Sarah (image 1)`) so the model can tell your attachments apart — you can read the result in \"See what gets sent\". Nothing else about your wording is rewritten. Prompting works the same as any other AI tool; no special syntax beyond @mentions.\r\n\r\n### Writing the shot list with Studio Agent\r\nThere is no \"Enhance\" button and no \"Generate prompts\" button. **Studio Agent does this work**, because it reads the same shot list you do — every scene, every beat in order, with its references, its model and its price — and because you can steer it:\r\n\r\n> \"Write the beats for my current storyboard.\"\r\n> \"Now redo scene 3 handheld, and match its energy to scene 2.\"\r\n> \"SHOT-A4 runs long — split it after 'and then'.\"\r\n\r\nEvery field it writes is editable by hand, in place, in the storyboard's Document view. Open Studio Agent with **Ctrl+.**\r\n\r\n### Debug Panel (advanced)\r\nA developer panel showing the exact request body, with the ability to override the composed prompt before sending. **There is no toggle button for it on the prompt bar in any build** — open it with **Ctrl+Shift+D**. For ordinary use, \"See what gets sent\" above is the supported way to inspect a prompt.\r\n\r\n---\r\n\r\n## PROJECTS\r\n\r\n### Structure\r\nEach project = folder on your disk. Subdirectories: images/, videos/, audio/, references/, exports/. Location configurable in Settings → Projects Directory.\r\n\r\n### Assets\r\nEvery generated or imported file is an asset (image, video, audio). Metadata tracked: prompt, model, settings, cost, dimensions, timestamps. Videos track source image via source_asset_id -- you can see all videos generated from any image.\r\n\r\n### Supported File Formats\r\n- **Images:** PNG, JPEG, WEBP, GIF. Note: HEIC/HEIF (iPhone photos) NOT supported -- convert to JPEG/PNG first.\r\n- **Video:** MP4, MOV, WEBM, AVI, MKV.\r\n- **Audio:** MP3, WAV, OGG, M4A, AAC.\r\n- **Clipboard paste:** Any image format the OS clipboard provides (PNG, JPEG, WEBP, GIF). Pasting works both in the gallery and directly into the prompt box; either way the image becomes a real gallery asset in the folder you're working in (tagged \"Imported\") AND, when pasted into the prompt box, attaches as a reference. Anything generated from it links back to it as a source.\r\n- **Drag and drop:** Any file the browser recognizes as image/* or video/*.\r\n\r\n### Operations\r\nCreate/rename/delete projects. Import external files via drag-and-drop or file picker. Paste images from clipboard. Extract still frames from videos. Relocate project to different disk/folder (all paths auto-update). Cleanup orphaned assets.\r\n\r\n### Moving and copying assets between projects\r\nAssets (images, clips) can be sent to another project three ways: the selection band on the Images/Videos tabs, the right-click menu on any card, or by dragging cards and dropping on a project in the drop palette.\r\n\r\n- **Move** relocates the media files on disk into the destination project's folder. The asset leaves whatever gallery folder it was in and is issued a fresh badge code in the destination.\r\n- **Copy** duplicates it — new files, new thumbnails, new badge code — and changes nothing in the source project.\r\n\r\n**Why a move can be refused:** an image another project still builds with (a character/environment/style identity image, or a storyboard frame) cannot leave, because the entity left behind would point at a file it no longer owns. When that happens the dialog lists what's blocking and offers the fix: bring the whole character/environment/style across with all of its images, or copy instead. A storyboard frame is only ever offered a copy — moving its image out would empty the shot.\r\n\r\nRight-clicking a card that is part of a multi-selection acts on the whole selection (\"Move 5 to Project…\"). Right-clicking a card outside the selection acts on that card alone.\r\n\r\n---\r\n\r\n## SHOTS — THE STORYBOARD IS THE SHOT LIST\r\n\r\nA **Shot** is the prompt bar, saved: the prompt, every reference with the job it carries, the model, every setting — and now the beat itself: who speaks, what they say, how it is said, what happens, the prop, the framing and the camera. It lives in the **storyboard**, which is the one place Shots are listed. It is never required: the prompt bar works exactly as it always has for anyone who never touches one.\r\n\r\n### Why it exists\r\nA generation's full recipe was already stored, but only once you had paid for it. A Shot can be written **before anything is generated**, so a whole piece can be planned, read, timed, priced and corrected while it is still free. That is the point of the thing: look at the entire ad or short film — every cheap asset lined up in the actual flow — before the videos exist.\r\n\r\n### What a Shot holds\r\nRaw prompt (@mentions intact); the model; aspect ratio / duration / resolution / negative prompt / sound and the rest of the bar's settings; every attachment with its ROLE (plain reference, subject, style, reference video, reference audio, first frame, last frame); and the script layer — `speaker`, `line`, `delivery`, `action`, `prop`, `shotSize`, `camera`, and a `continues` flag for one sentence running across two cuts. Characters, environments and styles are stored as the ENTITY, not a copied picture, so updating a character updates every Shot that names it.\r\n\r\n**The script fields are for reading and counting. Only the prompt is sent to a model.** Dialogue you want a model to perform still goes in the prompt, verbatim, with its delivery — writing it in `line` makes it readable and lets Slates check whether it fits the cut, not spoken.\r\n\r\n### Every Shot has an address\r\n`SHOT-A1`, `SHOT-A2` … per project, never reused — the same idea as the `IMG-A12` badge on a gallery card. Say it to Claude and you are both pointing at the same row. It is for **this session**, not for retrieval later: there is no shot search and no shot library, because a Shot is workspace state — alive while you build the piece, worthless once it ships.\r\n\r\n### Making one\r\n- **Save as Shot** on any generated image, clip or track's right-click menu restores that generation and keeps it — and that generation becomes the Shot's first take, so the row opens showing the result you kept it for. **Reuse Prompt** on those same menus does the restore WITHOUT saving anything.\r\n- Claude can write Shots directly (`slates_create_shot`), including for shots whose image does not exist yet — and can re-chop them with `slates_split_shot` / `slates_merge_shots`.\r\n- **It files itself.** A saved Shot lands in the scene you have open, else the last scene of the storyboard you were most recently working in; if the project has no storyboard, one appears named after the project. Nothing you save is ever somewhere you have to go and find.\r\n\r\n### There is no save button\r\nSelecting a Shot row **binds** the prompt bar to it. Edits write straight back to that row; `Clear` unbinds and returns the bar to free composing. There is no undo, and none is needed: the generations underneath a row are the permanent record of what actually fired, and the row itself is the working copy.\r\n\r\n### Two ways to look at it\r\nThe storyboard has one toggle and two jobs.\r\n\r\n- **Board** — arrange. One picture per Shot, dragged into the order you want. Drag one and the whole beat moves with it: the line, the references, the model, the settings, the takes. There is nothing else to drag, so there is never a question of what followed what.\r\n- **Document** — write. One continuous page: the script, the references beside the words that cite them, the prompts underneath. This is where you read the piece before paying for it.\r\n\r\nInside Document, choose **Script** to edit the words with scene headings and speakers. Delivery notes, model and pricing details, and warnings are hidden. **Shots** shows all layers. Open **Custom** to toggle Scene, Action, Character, Delivery, Dialogue, Shot, References, Prompt, Takes, and Warnings independently. Warnings covers missing models or references, unresolved mentions, model changes, and lines too long for their cut. Hiding a layer changes only what you see; it preserves your text and generation checks. There is a text-size slider and an independent toggle for shot numbers in the margin.\n\nDelivery is an optional performance note, not a required label for every line. TTS sends the authored prompt, or the dialogue when the prompt is empty, verbatim; separate Delivery notes are not added. For speech cues, use the selected TTS model's prompting guide and place supported tags in that spoken text. Script view does not strip inline speech cues from your words.\n\r\n### The header tells you what you are about to make\r\n`5 generations · 7 cuts · 54s · 84 credits`, and beneath it a variety strip like `6/7 wide · 5 push · 3 cuts in the loft`.\r\n\r\n**Two counts, because they measure different things.** Rhythm is counted in **cuts**; money is counted in **generations**. A multi-shot generation is several cuts inside one paid call, so mixing them would be wrong. A cut with no model chosen has no duration and shows as `—` rather than `0s` — a runtime that invented seconds would be a lie about the one number this view exists to give.\r\n\r\n### Splitting and merging — the chop\r\nPut the caret mid-line and press Enter: the row becomes two, the second inheriting the model, settings and references, and marked as continuing the first if the split lands mid-sentence. Select two adjacent rows and press **Merge**: they become one, references combined, durations summed. **The price and the runtime move as you do it** — which is the whole reason to make the decision here rather than in a document somewhere else.\r\n\r\nSplit is also the move behind a voiceover that keeps talking while the picture hard-cuts to a new world: split at a word boundary and both rows carry one sentence, each with its own visuals.\n\nThe **Dialogue continues from previous shot** toggle is a planning note that the sentence spans a cut. It does not merge shots, join generated audio, or change generation settings. Its pressed state shows whether the note is set; toggling it does not move the script. **Merge these two shots** is the separate control between adjacent shots that actually combines them. In Document → Custom, **Scene** toggles scene headings and their controls; the dialogue remains visible even if a hidden heading belonged to a collapsed scene.\n\r\n### Variety, counted and never judged\r\nSlates counts what is in front of it — shot sizes, camera moves, cast, locations, durations, and any of them repeating three or more times in a row — and shows the counts. **It never changes anything, never suggests anything and never blocks.** `shotSize` and `camera` are free text: write `long-lens CU, other head blurred` if that is the shot. Anything unrecognised counts as \"other\", which is a fine answer.\r\n\r\nIf a spoken line cannot be read in its cut at any plausible pace, the row says so — and says it only when the line is genuinely impossible, never when it is merely long.\r\n\r\n### The animatic\r\nPress play on any beat and the storyboard plays as a rough cut: each picture held for **its own cut's duration**, with the line underneath. That tells you the rhythm of the finished piece before a single video exists. A multi-shot generation holds one picture across its internal cuts, and says so. A cut with no duration is held for 3 seconds and marked — the header leaves it out of the runtime for the same reason it is marked here.\r\n\r\n### Things that stay honest rather than being hidden\r\n- **Deleted references.** If an asset or character a Shot points at is gone, the Shot still loads and the row says how many items were left out of what gets sent — and editing the row does not quietly drop them.\r\n- **A swapped model.** Changing a Shot's model never rewrites your words — video models genuinely take different prompt grammars, so the row tells you which model the prompt was written for and leaves the sentence alone. The settings line shows what will actually be sent after the swap, and that is the value it prices.\r\n- **Image Shots carry no role badges.** Image generation sends every reference in one undifferentiated list, so an image Shot can remember that a picture is a style reference but cannot tell the model.\r\n- **The thumbnail is never a question.** Slates picks it — the first frame, else the first reference, else the newest take — and you can override it from any reference in the gutter.\r\n\r\n### Firing several\r\nSelect Shots and press **Generate all**: one total, the largest single Shot stated separately, one approval. They run **one at a time**. If one of the selected Shots has been deleted, the whole batch is refused and names it — nothing fires and nothing is billed. If a generation fails mid-run, the rest still fire, the failure is reported per Shot, and **nothing is retried automatically**.\r\n\r\n### Deleting\r\nDeleting a storyboard tells you how many Shots are attached before it does anything, and deletes them with it. It never moves them somewhere else without asking. A Shot that also lives in another storyboard survives.\r\n\r\n---\r\n\r\n## STORYBOARDING\r\n\r\n### Hierarchy\r\nStoryboard → Scenes → **Shots**. A scene is an ordered list of Shots, and a Shot is one beat: its picture, its references and their roles, its model and settings, its prompt, its words, and the generations it has produced. See **SHOTS** above — that section is the storyboard.\r\n\r\n### What happened to frame types\r\nThere used to be a \"frame type\" on each picture — first / last / ingredient — plus a separate motion-prompt box. Both were a second, weaker way of saying what a Shot already says: **a reference's role lives on the Shot** (first frame, last frame, subject, style, plain reference), and the motion prompt was just the Shot's prompt under another name. Existing storyboards were converted automatically and nothing was lost. Pick a picture's job on the Shot's reference rail; write the motion in the Shot's prompt.\r\n\r\n### Grid Exploration\r\n2x2 grid: 4 prompt variations for quick iteration. 3x3 grid: 9 variations for deeper exploration. Select individual cells → extract to full-resolution images. Tip: 2x2 is usually sufficient and produces better quality. 3x3 can occasionally get proportions slightly wrong when upscaling cells because it faithfully reproduces the lower-resolution proportions. Grid exploration runs on Nano Banana 2 only; no other image model offers it.\r\n\r\n### Storyboard → Video\r\nSelect frames → generate video for each → clips auto-insert into timeline in order with source tracking maintained.\r\n\r\n### The animatic\r\nPlay the storyboard as a rough cut. Each beat is held for **its own duration**, with its line underneath, so what you are watching runs at the finished piece's real length. Space = play/pause. Arrow keys = navigate. Escape = exit.\r\n\r\n### Paste a script\r\nPaste a script into the storyboard and it becomes one row per paragraph — ALL-CAPS cues become speakers, parentheticals become delivery. It is a plain parse, not a model: nothing is invented, nothing is sent anywhere, and prose that is not screenplay-formatted lands as one row per paragraph for you (or Claude) to chop.\r\n\r\n---\r\n\r\n## VIDEO EDITOR (TIMELINE)\r\n\r\n### Tracks\r\nMulti-track: video tracks + audio tracks stacked vertically. Clips independent per track. Add or remove tracks freely -- layer a music bed, a voiceover, and effects on separate audio tracks. Video assets go on video tracks, audio assets on audio tracks. Overlapping video clips resolve top-track-wins.\r\n\r\n### Audio Mixing\r\nEach track has a volume fader, and the timeline has a master output fader for the final mix. Both range from silent to +12 dB of boost, and both apply to preview playback AND the exported MP4 -- what you hear is what you render. Muting a video track silences its embedded audio but still shows the picture. Use the master fader to prevent clipping when stacking loud tracks.\r\n\r\n### Timeline Settings\r\nResolution and frame rate (24/30/60) are auto-managed: the first video clip sets both, and a later higher-resolution clip raises the canvas. All clips are conformed to the timeline frame rate on export. Changing the frame rate after clips are placed retimes them.\r\n\r\n### Clip Properties\r\nSource asset, in/out points (frame-level precision), duration, scale (fit/fill/custom %), position (X/Y offset), opacity (0-100%).\r\n\r\n### Tools\r\n- **Select (V):** Click/drag clips, view/edit properties\r\n- **Razor (C):** Split clip at playhead into two clips\r\n- **Slip (S):** Adjust clip in/out points without moving its position\r\n- **Snap toggle:** Snap to playhead/clip boundaries\r\n\r\n### Markers\r\nColor-coded timeline markers (6+ colors) with optional labels. Use for scene breaks, cue points, notes.\r\n\r\n### Playback & Navigation\r\nSpace = play/pause. Left/Right arrows = frame-by-frame. Up/Down = +-1 second. Page Up/Down = jump by screen width. Home/End = start/end of timeline.\r\n\r\n### Zoom\r\nCtrl+Plus = zoom in (finer precision). Ctrl+Minus = zoom out (see more timeline).\r\n\r\n### Undo/Redo\r\n50-step history. Ctrl+Z = undo. Ctrl+Shift+Z = redo.\r\n\r\n---\r\n\r\n## EXPORT\r\n\r\n### Video Export (FFmpeg)\r\nExport timeline → MP4 (H.264). Configure: resolution, frame rate, bitrate, output location. All visible tracks rendered, muted tracks excluded, clip in/out points respected. FFmpeg is bundled -- no separate install needed.\r\n\r\n### DaVinci Resolve XML Export\r\nGenerates XML project file containing: clip references (paths to source videos), timeline structure (tracks, clips), clip properties (scale, position, opacity, in/out points), timeline markers.\r\n\r\n**Importing into DaVinci Resolve:** File → Import → Timeline. DaVinci reads the XML and reconstructs your timeline with all clips, properties, and markers intact. From there you can color grade and export your final master.\r\n\r\nExports saved to project's exports/ directory with timestamped filenames.\r\n\r\n---\r\n\r\n## CHARACTERS, ENVIRONMENTS & STYLES\r\n\r\n### Characters\r\nCreate character with name + description. Generate character sheet (license required): AI generates a turnaround with multiple angles for consistency. Generate expression sheet: same character with different facial expressions. Use `@character_name` in any prompt to attach reference images.\r\n\r\n**Tips for consistency:** Experiment with character sheet generation using both the existing project style and photorealistic style. Sometimes a single well-chosen image works better than a full sheet -- especially if the character is already in the same style, lighting, and clothing as your project. You can manually assign any image as a character reference instead of generating a sheet.\r\n\r\n**Voice.** A character can carry one voice clip, the same way it carries one identity image — a shortcut for reusing a voice, never a requirement for speaking in one. **Add voice** / **Change voice** on the card opens the same voice picker the prompt box uses (a menu off the button, not a pop-up; the other cards stay on screen): a clip from the project attaches as it is; a preset or a described voice renders the character speaking a fixed audition line on Inworld TTS-2 and attaches that clip, at the credit cost the picker states first. Right-click the voice card → **Remove voice** detaches it without deleting the clip. The character's voice then shows under **Clips** in the Voice lane's picker, and mentioning the character (`@name`) in a Seed Audio prompt attaches the clip as a reference, so the scene casts that voice.\r\n\r\n### Environments\r\nCreate environment with name + description. Generate environment grid (license required, 3x3): 9 variations. Extract individual cells to full-resolution images. Use `@environment_name` in prompts.\r\n\r\n**Tip:** Like characters, sometimes a single strong environment image gives better consistency than a grid of 9. Experiment with both approaches.\r\n\r\n### Styles\r\nCreate style with name + description + upload reference image. Use `#style_name` in prompts. Key visual auto-attachment option for consistent look across all frames.\r\n\r\n---\r\n\r\n## SETTINGS\r\n\r\n### Generation\r\nEvery generation runs on Slates Credits — there are no API keys to configure. The Generate button shows the exact credit cost before each generation, and failed generations refund immediately.\r\n\r\n### Other Settings\r\n- **Projects Directory:** Where project folders live on disk. Changeable anytime.\r\n- **Default Model:** Pre-selected model for new generations. Override per-generation.\r\n- **Default Quality/Resolution:** Pre-selected resolution. Override per-generation.\r\n- **Grid Size:** Default 2x2 or 3x3 for grid exploration.\r\n- **Auto Naming:** Automatically name generated assets.\r\n- **Prompting guide:** A link at the bottom of Settings to <https://slates.video/docs/prompting> — per-model prompting guidance for every model (see PROMPT SYSTEM above).\r\n\r\n---\r\n\r\n## ACCOUNT & BILLING\r\n\r\n### Login\r\nEmail-only, no password. Enter email → receive magic link → click to log in. First login creates account automatically. Session persists across restarts.\r\n\r\n### License\r\nUnlocks: character sheet generation and environment grid generation. Includes 12 months of updates (Slates Pro includes lifetime updates). Major upgrades discounted after.\r\n\r\n### Credits\r\n\r\n<!-- BEGIN:GENERATED credits -->\nCredits are what every generation is paid with. They are pay-as-you-go, they never expire, and the exact cost of a generation is shown on the Generate button before you commit.\n\n- A **Slates Standard** license ($149 one time) starts you with **1,000 credits**.\n- **Slates Pro** ($297 one time, or $97 to upgrade later) starts you with **3,000 credits**.\n\n| Pack | Credits (Standard) | Credits per dollar | Versus the smallest pack |\n|------|--------------------|--------------------|--------------------------|\n| $10 | 250 | 25.0 | standard rate |\n| $25 | 650 | 26.0 | +4% more credits |\n| $50 | 1,375 | 27.5 | +10% more credits |\n| $100 | 3,000 | 30.0 | +20% more credits |\n| $250 | 8,000 | 32.0 | +28% more credits |\n| $500 | 17,000 | 34.0 | +36% more credits |\n| $1,000 | 35,000 | 35.0 | +40% more credits |\n\nPacks up to $500 are open to everyone; the $1,000 pack is offered inside the app to licensed accounts. Slates Pro receives more credits than the Standard column above on every pack, for life.\n<!-- END:GENERATED credits -->\r\n\r\nCredits NEVER expire, there is no monthly reset, and failed generations refund immediately. You can also turn on auto-topup so your balance refills when it runs low.\r\n\r\n### Standard vs Pro\r\n- **Standard:** the app, every AI model, and pay-as-you-go credits that never expire, plus 12 months of updates.\r\n- **Slates Pro:** everything in Standard, plus our lowest credit rate on every pack, forever (buy the smallest pack and pay the largest pack's rate), **4K video generation**, a priority generation queue, early access to every new model on release day, and lifetime updates. The more you top up, the more the better rate adds up.\r\n\r\nEvery AI model is available on both tiers. The only capability gated to Pro is generating 4K video; 4K images are open to everyone, and exporting your timeline at 4K is available on every tier.\r\n\r\n### 30-Day Guarantee\r\nFull refund within 30 days, no questions asked.\r\n\r\n---\r\n\r\n## KEYBOARD SHORTCUTS\r\n\r\n| Key | Action |\r\n|-----|--------|\r\n| Space | Play/pause |\r\n| V | Select tool |\r\n| C | Razor tool |\r\n| S | Slip tool / snap toggle |\r\n| M | Add marker |\r\n| Left/Right | Frame-by-frame |\r\n| Up/Down | Seek +-1 second |\r\n| Ctrl+Z | Undo |\r\n| Ctrl+Shift+Z | Redo |\r\n| Ctrl+Plus/Minus | Zoom timeline |\r\n| Delete/Backspace | Delete selected clip |\r\n| Escape | Close modal/viewer/slideshow |\r\n| Ctrl+Enter | Submit generation |\r\n| Home/End | Jump to timeline start/end |\r\n| Page Up/Down | Jump by screen width |\r\n\r\n---\r\n\r\n## OFFLINE USAGE\r\n\r\nThe app launches and works offline for everything except AI generation and login. Specifically:\r\n\r\n**Works offline:** Opening projects, viewing all assets (images/videos), editing timeline (trim, reorder, split clips), adding markers, slideshow playback, FFmpeg export to MP4, DaVinci XML export.\r\n\r\n**Requires internet:** AI generation (all models), login/signup, credit purchases, credit balance sync, license validation (only checked on first generation attempt per session, then cached), auto-updater.\r\n\r\nIf you lose internet mid-session, you can keep editing and exporting. Generation will fail until connectivity returns.\r\n\r\n---\r\n\r\n## GENERATION RECOVERY\r\n\r\nIf the app closes during a generation: on restart, Slates detects in-flight jobs, polls the AI provider, and downloads completed results automatically. Nothing is lost. Recovering generations show at 5% in the queue until status is confirmed. Works for every model.\r\n\r\n---\r\n\r\n## TROUBLESHOOTING\r\n\r\n**\"Insufficient credits\"** — Your credit balance is too low for this generation. Buy more credits in the app (packs from $10 to $500) or turn on auto-topup. The exact cost of any generation is shown on the Generate button before you commit.\r\n\r\n**\"Input was rejected by Kling\"** — Image may not meet quality requirements (character visibility, proportions, content policy). Try a different image or prompt.\r\n\r\n**\"Failed to upload image to FAL CDN\"** — Network issue during reference image upload. Check internet connection, retry.\r\n\r\n**\"Generation failed\" / \"Proxy generation failed\"** — Generic error from the AI provider. Usually temporary. Retry. If persistent, try a different model.\r\n\r\n**\"Source asset not found\" / \"Source video asset not found\" / \"Target image asset not found\"** — The image or video you're trying to use was deleted or moved. Re-import or select a different asset.\r\n\r\n**\"Invalid audio source\"** — Lip sync: either enter TTS text or upload an audio file. One is required.\r\n\r\n**\"TTS response missing audio URL\"** — Text-to-speech failed during lip sync. Retry.\r\n\r\n**Generation stuck** — Restart app. Recovery system polls providers and picks up where it left off.\r\n\r\n**API rate limit** — Too many requests (limit: 20 generations/minute). Wait 1-2 minutes, retry.\r\n\r\n**Project files missing** — Project folder was moved/deleted outside the app. Use project relocation in Settings to re-point to the correct folder.\r\n\r\n**License shows \"revoked\"** — Contact support. Character sheets and environment grids unavailable until resolved.\r\n\r\n**Session expired** — Magic link session timed out. Log in again via Settings.\r\n\r\n**iPhone photos won't import** — iPhones save photos as HEIC/HEIF format, which Slates doesn't support. Convert to JPEG or PNG first (most photo apps and online converters can do this).\r\n\r\n---\r\n\r\n## PRIVACY & DATA\r\n\r\n- Generated files stay on YOUR machine. Slates servers never store your videos/images.\r\n- No prompts logged server-side.\r\n- File uploads go directly to the AI provider via pre-signed URLs. Slates servers never buffer your media.\r\n- Server stores only: email, license status, credit balance, transaction history, session tokens.\r\n- Stripe handles all payment data. Slates never sees your card number.\r\n\r\n---\r\n\r\n## SYSTEM REQUIREMENTS\r\n\r\n- Windows 10/11 or macOS 12+\r\n- Internet connection required for AI generation (not for editing/exporting)\r\n- Disk space for project files (AI videos are typically 5-50MB each)\r\n- FFmpeg bundled with app (no separate install needed)\r\n- No GPU required (all AI processing happens in the cloud)\r\n\r\n---\r\n\r\n## COMMON TASKS (STEP-BY-STEP)\r\n\r\n### Generate an Image\r\n1. Open the floating prompt box (visible on every page).\r\n2. Enter your prompt describing the image.\r\n3. Select an image model (Nano Banana 2 recommended).\r\n4. Choose aspect ratio and resolution.\r\n5. Press Ctrl+Enter or click Generate.\r\n\r\n### Generate Video From an Image\r\n1. In the prompt box, attach a start image.\r\n2. Write a prompt describing the desired motion/action.\r\n3. Select a video model (Kling V3.0 Omni in ingredients mode recommended).\r\n4. Choose duration (Kling bills per second from 3s up, so shorter is always cheaper), aspect ratio, and resolution.\r\n5. Click Generate.\r\n\r\n### Use a Character Reference for Consistency\r\n1. Create a character in your project (name + description).\r\n2. Either generate a character sheet OR manually assign a single image as the character reference.\r\n3. In the prompt box, type `@` and select your character from auto-complete.\r\n4. The reference image is attached automatically. Generate normally.\r\n\r\n### Export to DaVinci Resolve for Color Grading\r\n1. In the video editor, finalize your timeline (clips, markers, timing).\r\n2. Click Export → DaVinci Resolve XML.\r\n3. Choose output location. File saves to exports/ directory.\r\n4. In DaVinci Resolve: File → Import → Timeline. Select the XML file.\r\n5. Your timeline loads with all clips, properties, and markers intact. Grade and export.\r\n\r\n### Extract a Still Frame From a Video\r\n1. Hover over any video clip in the gallery.\r\n2. Camera icon = extract **current frame**. Dropdown arrow next to it = **First frame** or **Last frame**.\r\n3. Extracted image saves to your project gallery. Use as start/end image for I2V, character reference, or storyboard frame.\r\n\r\nKey workflow: extract a clip's last frame → use it as the start image for the next generation → seamless visual continuity between scenes.\r\n\r\n### Buy More Credits\r\n1. Open Settings → Credits, or the credit badge in the top nav.\r\n2. Pick a pack (packs run from $10 to $500, plus a $1,000 pack offered in the app to licensed accounts; bigger packs give more credits per dollar).\r\n3. Pay via Stripe. Credits are added to your balance instantly and never expire.\r\n4. Optional: turn on auto-topup so your balance refills automatically when it runs low.\r\n\r\n---\r\n\r\n## COMMON QUESTIONS\r\n\r\n**Q: Which model should I use for most videos?**\r\nA: Seedance 2.0 is the default and the one to reach for when physics, scale, effects or hero shots matter. For everyday shots built from a start image, Kling V3.0 Omni in ingredients mode is the best balance of cost and quality, which is why the step-by-step guides above use it.\r\n\r\n**Q: What's the best image model?**\r\nA: GPT Image 2.5 is the strongest image model in the app, and the best available anywhere right now. It holds a long instruction more faithfully than anything else here, and it is the only model that renders words in the picture reliably. The two seats cost the same, so the only trade is time: **Flare** is the fast one and is best for drafts and exploring, **Sunburst** is the higher-quality one and is what finals, hero frames and reference-heavy edits should end up on. The quality knob has five settings spanning about 36× from cheapest to dearest — `max` is four times `high`, though the smaller steps are uneven: **`medium`** is for drafts and where iteration belongs, **`high` is the default** and where most finished work should sit, and **`max` is the best output you can get** — finished frames, client deliverables, anything with exact text. `xhigh` sits just under `max` for about half the price and is worth trying before you jump to the top. Going past `high` should be a deliberate choice rather than a habit. Go to **4K only once you already know your references and your prompt are solid** — it is the most expensive setting on the model and the worst place to discover the composition was wrong. Prove the shot at 3K first, then re-run the settled prompt at 4K. If you used GPT Image 2 before, note the quality names all shifted by one: its `medium` is now `high`, its `high` is now `max`. Nano Banana 2 is the picker's DEFAULT rather than the best: it is fast, takes the most reference images, and is the only model with grid exploration. Nano Banana Pro is the hero-frame step up when composition, cinematic lighting and skin have to be perfect. Prices for every tier are in the model table above.\r\n\r\n**Q: Do I need to set up API keys?**\r\nA: No. There are no API keys in Slates — every generation runs on Slates Credits, which come with your license and never expire.\r\n\r\n**Q: How much does a generation cost?**\r\nA: It depends on the model, resolution, and length. The exact credit cost is always shown on the Generate button before you commit, so there are no surprises.\r\n\r\n**Q: Can I use Slates offline?**\r\nA: Yes for viewing projects, editing timeline, and exporting. No for AI generation -- that requires internet.\r\n\r\n**Q: Do credits expire?**\r\nA: No. Credits never expire.\r\n\r\n**Q: What happens if I close the app during a generation?**\r\nA: Nothing is lost. On restart, Slates detects in-flight jobs and downloads completed results automatically.\r\n\r\n---\r\n\r\n## FEATURES NOT IN SLATES\r\n\r\nThe following are NOT available. Do not suggest them:\r\n\r\n- Bring-your-own API keys (BYOK) — every generation runs on Slates Credits; there is no key-entry option\r\n- Local/on-device GPU inference (all AI runs in the cloud)\r\n- Built-in music generation (use external tools like Suno, import audio)\r\n- A voice picker for Seed Audio — you describe the voice you want in words instead. The preset voice shelf belongs to the Voice lane (Inworld TTS-2), and Kling Lip-Sync keeps its own small fixed list: six English/UK voices plus a storyteller, with a speed control\r\n- Automatic video editing from a script\r\n- A prompt \"Enhance\" button — ask Studio Agent to rewrite a prompt instead\r\n- A settings/gear panel on the prompt box — every parameter is a dropdown on the bar\r\n- Cloud project storage (all files are local)\r\n- Real-time collaboration / multi-user editing\r\n- Mobile app (desktop only: Windows and macOS)\r\n- HEIC/HEIF image import (convert to JPEG/PNG first)\r\n- Storyboard JSON export (import only)\r\n\r\n---\r\n\r\n## VERSION\r\n\r\n<!-- BEGIN:GENERATED version -->\nSlates Reference Version: 1.5.6\nLast Updated: 2026-09-09\n\nThis document is generated. Its source of truth is `slate/docs/slates-llm-manual.md`; its model tables and credit costs are derived from the Slates model registry and pricing tables at build time, so they cannot be typed by hand.\n\nIf the user asks about a feature not documented here, it may have been added after this version. The current copy is always at <https://slates.video/slates-reference.md>.\n<!-- END:GENERATED version -->\r\n\r\nIf this document didn't answer your question, email hello@slates.video so we can help and improve the app.\r\n\r\n</slates_reference>\r\n";
|
|
3
3
|
//# sourceMappingURL=content.js.map
|
package/dist/skills/content.js
CHANGED
|
@@ -20,7 +20,7 @@ export const SKILLS = {
|
|
|
20
20
|
"slates-prompting-kling-v3": "---\nname: slates-prompting-kling-v3\ndescription: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_generate_video with kling-v3.0-std, kling-v3.0-pro, or kling-v3.0-omni. Kling has dialogue + SFX + ambient native syntax (Omni adds multi-character dialogue and language codes). Multi-shot rules differ from Seedance/Veo — don't cross syntaxes.\n---\n\n# Kling V3.0 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Kling V3.0.** The general default. Define the core subjects clearly at the START and keep those descriptions identical across shots. Up to 15s, up to 6 cuts, and the strongest image-to-video identity hold in the catalogue.\n\n**The five levers**\n1. **Dialogue in quotes** — `Character says, \"exact words here\"`. On Omni, direct the voice with `Gender + Age + Voice quality + Speech rate + Emotional tone + Language`: `[Character A: Detective, mid-40s, raspy, slow cadence, weary]: \"I've seen this before.\"`\n2. **Unique speaker labels, no pronouns after the introduction.** `he`, `the agent`, any synonym causes voice drift.\n3. **Sound has real syntax** — `SFX: heavy boots on wet pavement, distant siren wailing`, `Ambient noise: city traffic`, `Background music: low cello`. Always physical-cause specific; `SFX: footsteps` is not enough.\n4. **Motion adverbs modulate energy directly** — `slowly`, `rapidly`, `gently`, `explosively`. One primary camera move per shot, never stacked.\n5. **On image-to-video, do NOT re-describe the image.** It is an anchor; prompt how the scene EVOLVES from it — movement, camera, environmental change.\n\n**Examples**\n- `A detective in a wet grey overcoat stands under a stairwell light. He steps forward slowly as the light flickers. [Character A: Detective, mid-40s, raspy voice, slow cadence, weary]: \"I've seen this before.\" SFX: heavy boots on wet concrete, distant siren wailing. Ambient noise: rain on metal.`\n- `Camera tracks right alongside a cyclist crossing a bridge at dusk. She rises out of the saddle rapidly as the grade steepens. Ambient noise: wind, tyres on wet asphalt, distant traffic.`\n\n**Hard constraint:** `Immediately` (Omni only) removes the natural conversational beat between speakers — use it when timing matters and leave it out when it does not. Kling has a real `negativePrompt` field, unlike Seedance; start from the standard block and layer scene-specific suppressions.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- `SFX: footsteps` and any label-only effect — physical-cause specificity or nothing\n- a pronoun or synonym for a speaker after the first introduction (`he`, `the agent`) — it causes voice drift; repeat the full label\n- `single continuous take` — Seedance's phrase, and it fights Kling's multi-shot\n<!-- @banned:end -->\n\nKuaishou's video model. Three tiers: `kling-v3.0-std` (general use, no audio), `kling-v3.0-pro` (higher visual quality, no audio), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).\n\nUp to 15s. Multi-shot supported (up to 6 cuts in 15s total). Strong on image-to-video — preserves identity, layout, and text from the input image well.\n\n## Subject definition rule (verbatim, fal blog)\n\n> \"Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots.\"\n\n## Dialogue syntax\n\n```\nCharacter says, \"exact words here\"\n```\n\nUse quotation marks for precise speech. Languages (Omni only): EN, ZH, JA, KO, ES.\n\n## Voice direction formula (Omni)\n\n```\nGender + Age Range + Voice Quality + Speech Rate + Emotional Tone + Language\n```\n\nExample:\n```\n[Character A: Detective, mid-40s, raspy voice, slow cadence, weary]: \"I've seen this before.\"\n```\n\nTone phrases that fire:\n- `speaking in a hushed, trembling whisper`\n- `shouting with commanding authority`\n- `clear, fearful voice`\n- `with a trembling voice, \"I'm scared\"`\n\n## The `Immediately` keyword (Omni only)\n\nWithout `Immediately`, Kling adds a natural conversational beat between speakers. With it, dialogue is back-to-back. Use when timing matters.\n\n```\n[Alice]: \"Get down!\" Immediately, [Bob]: \"Where?\"\n```\n\n## Speaker label discipline\n\nUnique labels per character. **No pronouns or synonyms after first introduction** — they cause voice drift.\n\n✅ `[Character A: Black-suited Agent]` ... `[Character A: Black-suited Agent]: \"Stop.\"`\n❌ `[Agent]... then he says...`\n\n## Multi-character dialogue (Omni)\n\n```\nAlice says in English, \"Hello!\" Then Bob replies in Spanish, \"¡Hola!\"\n```\n\n## Sound effects, ambient noise, music\n\n```\nSFX: thunder cracks, footsteps approaching\nAmbient noise: city traffic, birds chirping, ocean waves\nBackground music: tense orchestral strings, low cello\n```\n\nSFX accepts physical-cause specificity:\n- ✅ `SFX: heavy boots on wet pavement, distant siren wailing`\n- ❌ `SFX: footsteps`\n\n## Image-to-video guidance\n\n**Verbatim (fal blog):**\n> \"Treat the input image as an anchor. Kling 3.0 excels at preserving the identity, layout, and text details. Focus prompts on how the scene evolves *from* the image: subtle movements, camera motion, or environmental changes.\"\n\n**Don't re-describe what's already in the image.** Focus on motion, changes, evolution.\n\n## Multi-shot — what makes them hit\n\n**Hard cap: total duration ≤ 15s across all shots. Max 6 cuts.**\n\nHit conditions:\n- Shot labels are explicit: `Shot 1:`, `Shot 2:`\n- One primary action per shot\n- Subject described identically in each shot block\n- Camera move per shot is **one verb**, not a chain\n- Per-shot blocks: 30-60 words\n\nMiss conditions:\n- Compressing narrative into one paragraph\n- Pronoun-only references after the first shot\n- Mixing camera moves within a shot (\"pan then orbit then push in\")\n- Extreme wide → extreme close in adjacent shots without reference images\n\n## Element references (Omni)\n\nUpload 2-4 multi-angle reference photos per character/object. Tag inline:\n\n```\n@element1 is the protagonist (refs: front, side, back angles).\n@element2 is the antagonist.\n```\n\n## Reference discipline (character / environment refs)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Kling specifically\n\n- **Kling's consistency lever is \"lock the subject with a fixed label reused verbatim.\"** That is Kling's phrasing for rules 2 and 3, and it is stricter than the others: **pronoun and synonym drift breaks it**, so the exact same label must appear on every single mention — not \"he\", not \"the detective\" after you named him. Reusing the label verbatim is the whole game. Slates composes this for you from `@mentions`.\n- **Element references are the transport for rule 1** — 2-4 multi-angle photos per character/object, tagged `@element1` / `@element2` (see Element references above). The cap is 4 combined refs on the edit path.\n\n## Negative prompting — has a real field\n\nKling exposes `negative_prompt` on the fal endpoint (different from Seedance which has none). Default block to start from:\n\n```\nblurry, low quality, watermark, text overlay, distorted hands, extra fingers,\nduplicate limbs, unnatural skin texture, overly saturated colors, lens flare,\nfloating objects, inconsistent shadows, jittery, flickering, morphing face\n```\n\nLayer scene-specific suppressions on top.\n\n## Cinematic tactics\n\n- **Motion adverb precision** modulates motion energy directly: `slowly`, `rapidly`, `gently`, `explosively`\n- **Camera vocabulary that registers as instructions:** profile shot, tracking, following, freezing, panning, \"moving in sync with the subject\"\n- **One primary camera move per shot** — never stack\n\n## Tier choice\n\n- **Standard**: general use, no audio\n- **Pro**: higher visual quality, no audio\n- **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references\n\nPick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.\n\n## Benchmark prompt structure\n\n```\n[Character A: <role>, <voice quality>]: \"<line>.\" Immediately, [Character B: <role>, <voice quality>]: \"<reply>.\"\nAmbient noise: <soundscape>.\nCamera <single move>.\n```\n\nCinematic example (paraphrasing fal blog patterns):\n> \"Shot 1: Wide establishing shot of a neon-lit alleyway in heavy rain, steam rising from grates. Camera slowly tracks forward.\n> Shot 2: Medium shot of a detective in a trench coat ducking under an awning, water dripping from his hat brim. [Detective: weary, raspy]: 'I knew she'd come back.' Ambient noise: distant traffic, rain on metal.\n> Shot 3: Close-up on his eyes, narrowing as headlights flash across his face.\"\n\n<!-- slates-only -->\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with `firstFrameAssetId` or `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside cost + `requires_confirm: true`. Look at them, revise prompt if needed, then re-call with `confirm=true`. Kling Omni multi-character with several ingredient images especially benefits — confirm each character image lands cleanly before spending.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Detective Closeup`. The user sees that code as a gallery badge.\n\n- ✅ \"I'm anchoring on **IMG-A12** as the detective and **IMG-A18** as the alleyway environment — Omni will handle the line delivery in EN.\"\n- ❌ \"I'm using the detective image and the alley one...\" (which alley? Three exist.)\n<!-- /slates-only -->\n\n## Video-to-video EDIT<!-- slates-only --> (`slates_edit_video`)<!-- /slates-only --> — @Video1 / @ElementN / @ImageN\n\nKling O3 edit takes an EXISTING 3-15s clip and changes only what the prompt names — character swap, environment change, style transfer — in one pass, no masking. Original motion, camera, and audio are preserved by default. Its notation is Kling's own, different from the \"image N\" naming used everywhere else:\n\n- **`@Video1`** — the source clip (always; the transport anchors the instruction to it).\n- **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images<!-- slates-only --> (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically)<!-- /slates-only -->.\n- **`@Image1..`** — style/appearance references<!-- slates-only --> (pass as `styleAssetIds`)<!-- /slates-only -->.\n- Max **4 combined** element + image refs per edit.\n\n**Prompt shape — the change, not the whole scene:**\n\n```\nReplace the man in @Video1 with @Element1, keeping his walk cycle, the camera move, and the rain unchanged.\n```\n\n```\nEdit @Video1: turn the daytime street into a neon-lit Tokyo alley at night, wet asphalt reflections. Apply the visual style of @Image1. Keep the subject and camera motion exactly as they are.\n```\n\nRules:\n- Name what CHANGES; explicitly state what stays (\"keep the motion / camera / everything else unchanged\") — the model preserves better when told to.\n- One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).\n- Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.\n- Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.\n- Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.\n\n## Sources\n\n- [fal.ai — Kling 3.0 Prompting Guide](https://blog.fal.ai/kling-3-0-prompting-guide/)\n- [Vidguru — Kling 3.0 Omni Guide](https://www.vidguru.ai/blog/kling-3.0-omni-guide.html)\n- [AcceptPrompt — Kling 3 Prompt Guide](https://www.acceptprompt.com/blog/kling-3-prompt-guide)\n- [DataCamp — Kling 3.0 Tutorial](https://www.datacamp.com/tutorial/kling-3-0)\n",
|
|
21
21
|
"slates-prompting-lip-sync": "---\nname: slates-prompting-lip-sync\ndescription: How to set up lip-sync — Kling-only (dedicated lip-sync and avatar endpoints, 5-second outputs). Read before calling slates_generate_lip_sync. Two flows — video→video re-dub and image→video avatar — with different inputs, pricing, and gotchas. Voice catalog, framing rules, audio file constraints, and which tier to pick. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.\n---\n\n# Lip-sync — setup guide\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Lip-sync (Kling only).** Two different flows with different inputs and different prices; every output is 5 seconds.\n\n**The five levers**\n1. **Pick `sourceType` deliberately** — `video` re-dubs an existing talking head (cheapest); `image` animates a still portrait (avatar-standard, then avatar-pro only on the final selected take).\n2. **The `prompt` on the avatar flows is SCENE CONTEXT, not motion direction.** Ambience, lighting, micro-expression: `Soft rim light`, `warm office`, `cool blue evening light through a window`, `gentle confident smile between sentences`, `focused intent expression`.\n3. **Clean the audio before uploading** — `noise-reduced`, `levelled`. Lip detection is sensitive, and a raw recording is the most common cause of a bad take.\n4. **Iterate on the SOURCE or the AUDIO, never on a refinement prompt** — there is not one. If the output is wrong, change the input.\n5. **Use avatar-standard for first-pass dialogue takes**, and switch to pro only once the line is locked. Facial fidelity is not visible until then.\n\n**Examples**\n- `Soft rim light, warm office, gentle confident smile between sentences.`\n- `Cool blue evening light through a window, focused intent expression.` (Or `.` — an empty prompt is fine when you have nothing to add.)\n\n**Hard constraint:** it is Kling-only and always 5 seconds. For a generated PERFORMANCE instead — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal — that is a normal Seedance video generation with the clip attached as a video reference, not a mode of this tool. A real recording, or a cloned/cast voice rendered on `inworld-tts-2`, for production; this tool's built-in TTS is for scratch.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** — the avatar prompt is scene context and motion verbs are ignored:\n- `turns her head`, `raises an eyebrow`, `hand gestures`, `nods`, `walks`\n- `reader_en_m-v1` — listed in fal's docs, returns \"Voice id not found\" in production\n<!-- @banned:end -->\n\n**This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; every entry is a real endpoint and every output is 5 seconds.\n\n| Flow | Source | Model | Cost | Use case |\n|------|--------|-------|-----------|----------|\n| Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |\n| Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |\n| Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |\n\nPick `sourceType` deliberately — it decides the pricing tier and the underlying endpoint.\n\n## Want Seedance instead? That is a video generation, not a mode here\n\nSeedance can generate the performance rather than bolting a mouth onto finished pixels — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal, and a video source keeps its own voice. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the clip (or portrait) attached as a video/ingredient reference and the dialogue written into the prompt yourself.\n\nThat is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a \"video 1\" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.\n\n- Driving clips must be 2–15s; output duration is whatever you set (4–15s).\n- Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming.\n- Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.\n\nEverything below is about the Kling tool.\n\n## Choosing video vs avatar\n\nUse **video** (re-dub) when:\n- A talking-head clip already exists (Slates-generated, recorded, or imported)\n- The mouth/face is already moving and only the audio needs to change\n- ~4 credits is hard to beat for short dialogue replacement\n\nUse **avatar** when:\n- Only a still portrait exists\n- The character needs to come alive from a single image\n- Identity + face fidelity matter (avatar-pro for hero shots, standard for everything else)\n\n## Source asset constraints\n\n### Video flow (`sourceType: 'video'`)\n- Format: mp4 or mov\n- Duration: 2–10s (lip-sync output is always 5s — long videos get trimmed)\n- Resolution: 720p or 1080p (480p will be rejected)\n- Max file size: 100MB\n- Face must be visible and roughly facing camera. Profile shots fail.\n- Existing audio is replaced.\n\n### Avatar flow (`sourceType: 'image'`)\n- Min 512×512, PNG/JPG/WebP\n- **Face occupies 60–70% of frame.** This is the single biggest avatar quality lever.\n- Eyes open, mouth neutral, looking near-camera. Side profile = bad output.\n- Single subject, clean background. Group photos confuse the face anchor.\n\n## Audio source\n\nTwo ways to drive the lips:\n\n### TTS (`audioMethod: 'tts'`)\n- Pass `ttsText` (the words spoken)\n- Optional: `ttsVoice` (default `oversea_male1`), `ttsLanguage` (default EN), `ttsSpeed` (default 1.0)\n- **Hard cap: 120 characters of text.** Longer = silently truncated.\n- Languages: EN, ZH, JA, KO, ES\n\n### Upload (`audioMethod: 'upload'`)\n- Pass `audioFilePath` — absolute path to an audio file on the user's machine\n- Format: mp3, wav, m4a, ogg, aac\n- Max 5MB\n- Duration: 2–60s (output is 5s — longer audio gets trimmed)\n- Single clean voice. Music underneath, multiple speakers, or noisy mics produce garbage lips.\n\nPrefer upload for production-quality voice. TTS for fast iteration / placeholder dialogue.\n\n## Voice catalog (TTS)\n\nReliable English voices (verified working on the fal endpoint as of 2026):\n\n| Voice ID | Description |\n|----------|-------------|\n| `oversea_male1` | Male, English — default, stable |\n| `commercial_lady_en_f-v1` | Female commercial English |\n| `uk_boy1` | Young man, UK accent |\n| `uk_man2` | Man, UK accent |\n| `uk_oldman3` | Older man, UK accent |\n| `calm_story1` | Storyteller / narrator |\n\nAvoid `reader_en_m-v1` — listed in fal.ai docs but returns \"Voice id not found\" in production.\n\nFull 48-voice list (ZH, JA, KO included): https://fal.ai/models/fal-ai/kling-video/lipsync/text-to-video/api\n\n## Speech-rate notes\n\n`ttsSpeed` range 0.5–2.0:\n- 0.8–1.0: natural conversational\n- 1.1–1.3: punchy ad delivery\n- 1.4+: rushed, clips consonants\n- 0.6–0.7: slow, weighty (good for dramatic lines)\n\nDefault 1.0 unless the line specifically calls for slower or faster cadence.\n\n## Avatar prompt usage\n\nThe `prompt` parameter on avatar-v2 (standard + pro) is **scene context**, not motion direction. The mouth animation comes from the audio — the prompt sets ambiance, lighting, micro-expression.\n\nGood:\n- `Soft rim light, warm office, gentle confident smile between sentences.`\n- `Cool blue evening light through a window, focused intent expression.`\n\nBad (the model ignores motion verbs):\n- ❌ `She turns her head, raises an eyebrow, then speaks.`\n- ❌ `Hand gestures while talking.`\n\nDefault `\".\"` is fine if you have nothing useful to add.\n\n## Tier selection — avatar standard vs pro\n\n**Use standard** when:\n- Drafts, A/B testing voices, internal review reels\n- Wide / medium shots where face isn't the focal point\n- Cost matters more than micro-expression fidelity\n\n**Use pro** when:\n- Final ads where the avatar's face fills the screen\n- The character is named / branded — identity drift kills the take\n- You're already paying tens of credits for the surrounding video pipeline\n\nDon't default to pro. The ~15-credit delta per take adds up across iteration.\n\n## Common failure modes\n\n| Symptom | Likely cause | Fix |\n|---------|--------------|-----|\n| Lip movement looks \"rubber\" / disconnected | Source face <60% of frame | Re-crop the still tighter |\n| Voice doesn't match character age/gender | Default voice id used | Pick from voice catalog |\n| Output truncated mid-word | TTS text >120 chars | Shorten or chain two takes |\n| Garbled mouth on uploaded audio | Background music / multi-voice | Use clean dialogue-only audio |\n| \"Voice id not found\" 422 | Hit `reader_en_m-v1` | Switch to `oversea_male1` |\n| Avatar eyes drift / cross | Source had closed/angled eyes | Pick a frame with neutral open eyes |\n| Generation completes but lips don't move | Profile shot / face >70° off-axis | Use a near-frontal portrait |\n\n## Cost discipline\n\n- Video re-dub at ~4 credits is the cheapest dialogue iteration in the entire Slates stack — use it for voice A/B testing\n- Avatar standard at ~14 credits is fine for medium use\n- Avatar pro at ~29 credits trips the confirm gate — explicit user OK required every time\n- All 5s. There is no shorter option.\n\n## Workflow patterns\n\n**Voice A/B test (cheap):**\n1. Generate one base talking-head video clip with Veo or Seedance (~40 credits)\n2. Run `slates_generate_lip_sync` with `sourceType: 'video'` against 3–5 different `ttsVoice` values\n3. Total cost: ~40 + (5 × ~4) ≈ 60 credits to compare voices\n\n**Brand avatar from a single portrait:**\n1. Generate or upload the hero portrait (face fills frame, eyes open, neutral mouth)\n2. Avatar standard for first-pass dialogue takes\n3. Avatar pro only on the final selected take\n\n**Avoid:**\n- Avatar pro on first iteration (waste — facial fidelity isn't visible until you've locked the line)\n- TTS for final ads (production should use real voice or cloned voice — the upload flow)\n- Uploading raw recordings — clean noise + level the file first, lip detection is sensitive\n\n## Confirm gate: cost + codes, no inline preview\n\nLip-sync is mechanical — the model re-syncs the chosen source to the chosen audio. The confirm response carries the source asset's code so you can announce it in chat.\n\n- ✅ \"Lip-syncing **IMG-A12 — Founder Portrait** to the new line. ~29 credits on avatar-pro. Confirm?\"\n- ❌ \"Using the founder image...\" (which? Three exist.)\n\nDon't second-guess the source. If the output is wrong, iterate on source choice or audio, not on a refinement prompt (there isn't one).\n\n## Sources\n\n- [fal.ai — Kling LipSync API](https://fal.ai/models/fal-ai/kling-video/lipsync/text-to-video/api)\n- [fal.ai — AI Avatar v2 Standard](https://fal.ai/models/fal-ai/kling-video/ai-avatar/v2/standard/api)\n- [fal.ai — AI Avatar v2 Pro](https://fal.ai/models/fal-ai/kling-video/ai-avatar/v2/pro/api)\n",
|
|
22
22
|
"slates-prompting-ltx-2-5": "---\nname: slates-prompting-ltx-2-5\ndescription: How to prompt LTX-2.5 and LTX-2.5 Pro. Read before calling slates_generate_video with model ltx-2-5 or ltx-2-5-pro. LTX scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE — Lightricks ranks the prompt sound, camera, character detail, shot type and scene, then scene dressing, all in one flowing paragraph. It is also the catalogue's native MULTISHOT seat: one generation carries two to four connected shots holding character, light and voice across the cuts. Base ltx-2-5 is the distilled build — 720p/1080p/1440p/4K, clips of 6 to 20 seconds in EVEN steps, and the cheapest native 1080p second in Slates; ltx-2-5-pro is the full diffusion build and is NOT a superset, reaching only 1080p and 10 seconds for about a third more money. Three hazards live here: durations are even numbers only from six (there is no 5s or 7s clip), the model has NO reference endpoint at all so identity references are unavailable, and any sound not anchored to something in frame gets invented for you.\n---\n\n# LTX-2.5 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — LTX-2.5.** It scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE. Lightricks' own priority order: sound, camera, character detail, shot type and scene, then scene dressing. One flowing paragraph, not labelled sections. When a prompt sprawls, cut from the bottom.\n\n**The five levers**\n1. **Lead with sound, and anchor every sound to something in frame.** The test is \"visible, or at least locatable\" — `the rope creaks against the cleat`, `the hull knocking hollow against the fenders`, `rain on the awning`. A distant whistle is fine IF you have named the post it comes from.\n2. **Camera second**, because framing decides the visual weight of the shot — `low camera at the gunwale`, `slow drift right`, `static medium behind the counter`.\n3. **Character detail as physical ACTION**, not as adjectives about a person — `she braces a boot on the rail and hauls`, `his hands counting notes`.\n4. **It is the native MULTISHOT seat** — one generation carries two to four connected shots holding character, light and voice across the cuts. Write the cuts.\n5. **Quote dialogue and name the language and accent** — `in English with a slight German accent` — `\"We should not have come back,\" in English with a slight German accent.`\n\n**Examples**\n- `The rope creaks against the cleat as she leans back, gulls calling somewhere off the port bow, the hull knocking hollow against the fenders. Low camera at the gunwale, slow drift right. She braces a boot on the rail and hauls, twice, then stops.`\n- `A till drawer bangs shut, a fan ticks against its cage, rain on the awning outside. Static medium behind the counter, then cut to a close-up of his hands counting notes, then cut wide as he looks up at the door.`\n\n**Hard constraint:** durations are EVEN numbers from six — there is no 5s or 7s clip. There is NO reference endpoint at all, so identity references are unavailable; use MiniMax H3 or Kling when a character must hold across shots. Any sound not anchored to something in frame gets invented for you. And never write mood adjectives as sound: \"tense atmosphere\", \"a sense of dread\" and \"ominous ambience\" produce nothing usable — the fix is one more moving object in frame with a sound attached to it.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** — mood adjectives standing in for sound produce nothing usable:\n- `tense atmosphere`, `a sense of dread`, `ominous ambience`, `eerie silence`\n- an unanchored sound: name the thing in frame it comes from, or cut it\n<!-- @banned:end -->\n\nLTX-2.5 generates picture and sound **in a single pass**, with a Gemma-4 12B text encoder reading\none flowing paragraph. That single fact drives everything below: the prompt is not a shot\ndescription with audio bolted on, it is **a scene where the sound is load-bearing** — and\nLightricks' own priority order puts sound first, ahead of the camera.\n\nTwo seats, and the naming is a trap:\n\n| | `ltx-2-5` (base) | `ltx-2-5-pro` |\n|---|---|---|\n| Build | Distilled, 8-step | Full diffusion (\"Diffusion Fidelity Rendering\") |\n| Resolutions | 720p / 1080p / **1440p** / 4K | 720p / 1080p |\n| Durations | 6–20s, even steps | 6 / 8 / 10s |\n| Price | $0.09–$0.30 per second | $0.12–$0.17 per second |\n| Reach for it when | iterating, long takes, 4K delivery, batch volume | one dense final render inside 1080p and 10s |\n\n**Pro is not \"base plus more.\"** It buys picture quality on a *narrower* envelope — it cannot make\na 1440p frame and it cannot make a 12-second clip. Reaching for it out of habit costs a third more\n*and* takes away the reach.\n\n---\n\n## 1. The six parts, in priority order, in one paragraph\n\nLightricks ranks the elements of an LTX prompt like this. When a prompt sprawls, **cut from the\nbottom.**\n\n1. **Sound** — highest priority; the model scores the picture as it draws it.\n2. **Camera** — framing decides visual weight and the feel of the shot.\n3. **Character detail** — expressed as physical action.\n4. **Shot type and scene** — the action itself.\n5. **Scene dressing** — the first thing to trim.\n\nWrite it as **one flowing paragraph**, not a list of labelled sections. LTX is not Seedance (eight\nengineering slots) and not H3 (three separate audio layers) — it wants continuous prose.\n\n---\n\n## 2. Sound: anchor it or it gets invented\n\n**Write the audio line last, then go back and check every cue has a source you could point at.**\nAnything unanchored, the model invents for you.\n\nThe test is **\"visible, or at least locatable.\"** A distant whistle is fine *if* you have named the\nmarshal's post it comes from. A \"distant whistle\" with nothing to attach to is a coin flip.\n\n> the rope creaks against the cleat as she leans back, gulls calling somewhere off the port bow,\n> the hull knocking hollow against the fenders\n\n**Never write mood adjectives as sound.** \"Tense atmosphere\", \"a sense of dread\" and \"ominous\nambience\" produce nothing usable. If a scene feels thin, the fix is **one more moving object in\nframe with a sound attached to it** — never another adjective.\n\n### Dialogue\n\nQuote it, and name the language and accent:\n\n> \"We should not have come back,\" in English with a slight German accent.\n\nTwo rules that decide whether the lip sync lands:\n\n- **Give the character a beat of stillness before they speak.** The sync needs something to lock\n against; a character already mid-motion when the line starts drifts.\n- **Describe the beat structure** — when they look, how long they wait, when they speak, where they\n look afterwards.\n\nSlates pins the frame rate at 25fps, which is also what Lightricks recommends for dialogue: at 50fps\nthe performance \"pulls toward a video look.\"\n\n---\n\n## 3. Character emotion is physical\n\nThe model renders actions. It does not render adjectives.\n\n| Instead of | Write |\n|---|---|\n| she looks anxious | her jaw sets, she turns the ring on her finger twice |\n| he seems exhausted | he blinks slowly and lets his shoulder take the doorframe |\n| a tense standoff | neither moves; his thumb finds the strap and stays there |\n\n---\n\n## 4. Multishot — the thing this model is uniquely for\n\n**One LTX generation can carry several connected shots**, holding character, environment, lighting,\nvoice and style across every cut. Nothing else in the catalogue does this natively; everywhere else\nyou generate separate clips and stitch them, and identity drifts between them.\n\n**Working range is two to four shots.** Three is the comfortable stopping point.\n\nAt **every** transition you must supply four things:\n\n1. **Name the edit in the prose** — \"hard cut\", \"dissolve\", \"match cut\".\n2. **Re-establish the shot completely** — scale, angle, lens and light all reset at a cut. A cut is\n not a continuation.\n3. **Re-identify recurring characters by their original descriptor.** \"The woman in the bronze\n gown\", never \"she\". Pronouns lose the character across a cut — this is the single most common\n multishot failure.\n4. **State what the sound does at the cut.** Silence is not assumed; if the room tone should drop\n out, say so.\n\nA shape that works:\n\n> Wide establishing shot of the workshop, dust in the window light, a lathe turning somewhere off\n> frame — hard cut — macro close-up of the brass fitting as it seats, the turning noise gone,\n> replaced by a single dry click — match cut — medium shot of the woman in the bronze gown stepping\n> back, the room tone returning underneath her.\n\n---\n\n## 5. Camera: write it, don't enumerate it\n\nfal exposes a `camera_motion` enum (dolly in/out/left/right, jib up/down, static, focus shift).\n**Slates does not surface it, deliberately** — and prose is the better instrument anyway:\n\n- **A written move can be tied to a specific moment.** \"A slow push-in that settles as she reaches\n the door, then holds\" is not expressible as an enum value.\n- **For multishot it would be actively wrong** — one enum value would impose a single camera\n behaviour on three shots that each want their own.\n\nSo name the lens, the framing, the move, and **the moment the move resolves**.\n\n---\n\n## 6. The hard constraints\n\n### Durations are even numbers only, starting at six\n\n**6, 8, 10, 12, 14, 16, 18, 20.** There is no 5-second LTX clip and no odd duration of any length.\nAsking for 7s is not a rounding matter — that generation does not exist.\n\n**And the long end is 1080p-and-below only.** At 1440p and 4K the ceiling drops to **6, 8 or 10**.\n\nfal's own default is `auto`, which lets the model pick the length from the described action.\n**Slates always sends an explicit length instead**, so what you choose is what you are billed for.\nChoose the length the beat needs.\n\n### Aspect ratios: 16:9 and 9:16, and nothing else\n\nThe narrowest set in the catalogue alongside Veo. Square, 4:5 and 21:9 are not available on this\nmodel at any resolution.\n\n### Frames, not references\n\nLTX takes a **start frame** and an **optional end frame** (which generates a transition between the\ntwo). It has **no reference-to-video endpoint at all** — no identity references, no style\nreferences, no environment references, no reference video, no reference audio.\n\n**For character consistency across separate shots, use MiniMax H3 or Kling.** Within a single LTX\ngeneration, use multishot instead — that is precisely the gap it fills.\n\nIn image-to-video, **do not cut away from the opening frame too early.** You have paid for that\nframe; let it play before the first move.\n\n### Do not ask for text on screen\n\nNeither the spelling nor its stability from frame to frame can be relied on. Signage, labels,\ncaptions and lower-thirds belong in post.\n\n---\n\n## 7. Audio is free here, and that changes the routing\n\nNative synchronised audio is **included at every resolution on both seats**, with no surcharge and\nno toggle that costs money — unlike Kling, where sound is a paid dimension. A 6-second 1080p LTX\nclip **with sound** is 39 credits.\n\nCombined with 1080p at $0.13/s — the cheapest native 1080p second in Slates — this makes LTX **the\ncoverage seat**: the one to reach for when the job is many takes rather than one hero shot, when a\nsequence needs its own sound, or when the credit budget is the binding constraint.\n\nRoute away from it when you need identity references (H3, Kling), a ratio other than 16:9 or 9:16\n(Seedance, Kling), or authored multi-layer audio direction (H3).\n",
|
|
23
|
-
"slates-prompting-minimax-h3": "---\nname: slates-prompting-minimax-h3\ndescription: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09), so the seats differ on ladder and price, not on what they accept. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; 4 free then +1 on Max), and audio written into the wrong section is dropped or duplicated.\n---\n\n# MiniMax H3 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — MiniMax H3.** The only seat where audio is AUTHORED rather than toggled: dialogue, scene sound and score are three separate sections of the prompt, generated in one pass, and putting a sound in the wrong section drops or doubles it.\n\n**The five levers**\n1. **Write the three audio layers separately** — `Scene sound:` for what is in the room, `Score:` for what only the audience hears, and the dialogue quoted inline. Section decides attribution.\n2. **Quote dialogue and name the language** — `says in English`, `speaks in Spanish`. Eleven languages are stably supported; the language is part of the instruction, not an afterthought.\n3. **Declare the reference RELATIONSHIP**, which no other seat has: `kept whole`, `partly kept`, `transferred`, or `a loose echo`. An undeclared reference is a guess.\n4. **Give a beat of stillness before a line** — `sits still for a beat, then looks up`. The sync needs something to lock against; a character already mid-motion when the line starts drifts.\n5. **Describe the beat structure** — `waits`, `then speaks`, `under the last three seconds`. H3 is a timeline, so write one.\n\n**Examples**\n- `A woman sits still at a kitchen table for a beat, then looks up. She says in English, \"You said Tuesday.\" Scene sound: a fridge hum, a spoon set down on formica. Score: none.`\n- `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, \"No es el alternador.\" Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`\n\n**Hard constraint:** the two seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p rather than 4K, takes the same 9+3+3 references, and costs MORE at the tier they share — it is a speed pick, never the cheap one. H3's top two resolution tiers are UPSCALES of the native render: judge at native. Reference images past the free allowance are a paid dimension of the cost key (5 free on base, 4 on Max) — declare the count when quoting.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- a sound written into the wrong audio section — it is dropped, doubled, or attributed to the wrong layer\n- `background music` as a bare instruction: the score is its own authored layer, audience-only, and it is named as such\n- an undeclared reference relationship — say kept whole, partly kept, transferred, or a loose echo\n<!-- @banned:end -->\n\nH3 is an **omni transformer**: it generates picture and sound in the same pass, at 24fps with\n32kHz stereo, 5–15 seconds, in 11 stably-supported languages (Arabic, Chinese, English, French,\nGerman, Italian, Japanese, Korean, Portuguese, Russian, Spanish). That single fact drives\neverything below — the prompt is not a shot description with sound bolted on, it is a **timeline\nwith three audio layers you author separately**.\n\n**Two seats, one grammar.** Everything in this file applies to both. They differ only in what the\nendpoint accepts:\n\n| | `minimax-h3` | `minimax-h3-max` |\n|---|---|---|\n| Resolution | 480p / 768p / **2K / 4K** | 480p / 768p / **1080p** |\n| References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) |\n| Frames | start and/or end | start and/or end |\n| Price at 768p | **$0.060/s** | $0.080/s |\n| Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) |\n\n**Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it\ntops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth\npaying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.\n\n**The speed is measured, not claimed** (2026-08-27, same prompt and params on both rows): a 5-second\n768p text-to-video finished in **4.8 seconds** on Max against **57 seconds** on base H3 — roughly\n**12x**, queue to finished file. fal advertises \"under 3 seconds\"; the literal claim did not hold at\n4.8s wall-clock, but the order of magnitude did. For iteration loops and client-present work that gap\nis the entire reason the seat exists.\n\n---\n\n## The one thing that makes H3 different: audio is a THREE-LAYER instruction\n\nEvery other video seat treats sound as on or off. H3 splits it, and the split is enforced by where\nyou write each thing. Get the section wrong and the sound is dropped, doubled, or attributed to the\nwrong source.\n\n| Layer | What belongs in it | Where it goes |\n|---|---|---|\n| **Synchronised events** | dialogue, singing, and any sound tied to a specific shot or action | the **body** of the prompt, on the beat it lands |\n| **Scene sound** | ambience and physical sounds that run across the whole clip — room tone, rain, traffic, a ventilation hum | the **soundscape** section |\n| **Score** | music the characters cannot hear; audience-only | the **music** section |\n\n**Three rules, all from MiniMax's own guide:**\n\n1. **Dialogue and singing NEVER go in the soundscape section.** They are synchronised events; they\n belong in the body, at the moment they happen.\n2. **Diegetic music — music the characters can hear** (a radio in the scene, a busker) — also\n belongs in the **body**, not in the score section. The score section is audience-only.\n3. **Write the score in instrumental terms, not mood words.** Name the instruments, the tempo, and\n how it develops. *\"A restrained solo-piano score at a slow tempo, sustained low cello underneath,\n no swell\"* — not *\"emotional music\"*.\n\nUse **N/A** for a section only when silence or absence is genuinely what the shot wants. An empty\nscore section is a real choice; a vague one is a wasted layer.\n\n### The shape, in the one prompt field\n\nSlates sends one prompt string, so write the three layers as labelled paragraphs in this order:\n\n```\n[Shot 1] Live-action, cinematic. A medium-wide shot frames a baker opening the shutters of a\nsmall street bakery before sunrise. The camera pushes in with small amplitude at slow speed as\nthe middle-aged baker with a calm, slightly raspy voice places a fresh loaf on the counter and\nsays: \"First batch of the morning.\" [Shot 2] At 00:05.000, the camera cuts to a close-up of\nsteam rising from the sliced bread while his final words carry over from the previous shot.\n\nSoundscape: wooden shutters scrape open over a quiet street, trays clink softly inside, a\ndoorbell rings once, then light footsteps and the crisp sound of bread being sliced.\n\nScore: a soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes,\ngentle fade at the end.\n```\n\n**Body target: 350–500 words** for a reference-carrying shot. Dialogue-heavy content prioritises\nfitting the complete spoken timeline over hitting a word count.\n\n🚨 **Slates disables the provider's prompt expander.** H3's API can rewrite your prompt before\ngeneration; Slates turns that off, because a model rewriting the user's words invisibly is banned\noutright (prompt transparency: what the composer shows is what the model gets). The practical\nconsequence is on you: **nothing will pad a thin prompt.** Write the whole body.\n\n---\n\n## Shots and timing\n\nThe first shot carries **no timestamp**. Every later shot opens with the bracket and a cut time\nthat increases and stays inside the clip length:\n\n```\n[Shot 2] At 00:03.500, the camera cuts to ...\n```\n\nTransition verbs the model knows: **cuts to · transitions to · changes to · switches to**.\n\n**Dialogue that continues across a cut** needs the continuity said out loud — *\"his final words\ncarry over from the previous shot\"* — or the line restarts. **Speech that ends abruptly** should be\ndescribed as cut off rather than trailed off.\n\n---\n\n## Camera — write the move into the sentence\n\nThe model has a named motion vocabulary:\n\n> Zoom In / Zoom Out · Push In / Pull Out · Pan Left / Pan Right · Truck Left / Truck Right ·\n> Tilt Up / Tilt Down · Pedestal Up / Pedestal Down · Arc Shot · Tracking Shot · Static Shot ·\n> Shake Slightly / Shake Strongly · POV · Roll Clockwise / Roll Counterclockwise\n\nModify with **amplitude** (`with small amplitude` / `with large amplitude`) and **speed**\n(`at slow speed` / `at fast speed`).\n\n🚨 **Integrate the motion into the sentence — never stack labels.** MiniMax's own example:\n*\"The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.\"*\nNot *\"Push In. Small amplitude. Slow.\"*\n\n---\n\n## Speakers and dialogue\n\nGive each speaking character a stable identity in the prose and keep it: describe the voice once\n(*\"a young woman with a quiet, breathy voice\"*), then refer back to the same description at every\nline. Identification, delivery and action sit **outside** the quoted line; the line itself is only\nthe words.\n\n```\nThe young woman with a quiet, breathy voice says: \"I get off at the next station.\"\n```\n\n**Voiceover** needs two things — the phrase *\"says in an off-screen voiceover\"* **and** an explicit\nstatement that the lips stay closed. Without the second half the model animates a mouth.\n\n```\nThe man says in an off-screen voiceover: \"I still remember that road.\" — his lips remain\ncompletely closed.\n```\n\n**On-screen text** — signs, banners, labels, subtitles, neon — goes in double quotes with the\noriginal wording preserved exactly: *A red neon sign reading \"Open Late\" glows above the doorway.*\n\n---\n\n## References — H3's real differentiator is the declared RELATIONSHIP\n\n*(BOTH rows. `minimax-h3-max` gained the reference set on 2026-09-09; its free allowance is\nFOUR images rather than the base row's five.)*\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### Cite references by number — Slates already does it for you\n\nH3 on fal takes references as **typed slots** and expects the prompt to name them by modality and\norder: **`image 1`, `image 2`, `video 1`, `audio 1`**. That is exactly what the Slates composer\nemits from your `@mentions` and `#tags` (`Marcus (image 1) in the workshop (image 2)`), in the\nexact order it sends them.\n\n🚨 **Do NOT hand-write angle-bracket reference tags.** MiniMax's own model-card grammar uses\n`<Subject N>` / `<Picture N>` / `<Video N>` / `<Audio N>` labels; the fal endpoints Slates calls do\nnot — they build the binding from the typed slots and ask for plain numbered prose. Typing the tags\nyourself puts literal angle brackets in the prompt the model reads.\n\n### State how much of each reference survives\n\nThis is the lever no other model in the catalogue gives you. Say, in plain words, what each\nreference is FOR and how much of it should carry through:\n\n| Intent | Say something like |\n|---|---|\n| Keep it whole | *\"Keep the woman in image 1 exactly as she appears — hair, cardigan, necklace.\"* |\n| Keep part of it | *\"Use the café in image 2 for the brick wall and the sofa; the lighting is late evening, not daylight.\"* |\n| **Move a trait onto someone else** | *\"Give the man in image 3 the weathered leather texture of the jacket in image 4.\"* |\n| Loose echo | *\"Match the general palette and grain of image 5; nothing else from it.\"* |\n\nThe third row is the one with no equivalent anywhere else in Slates: **transferring a characteristic\nonto a different subject** is a first-class thing H3 understands. Reach for H3 when that is the job.\n\n**Audio references** bind a voice or a texture without copying the words. Say which speaker an\naudio reference is for (*\"the woman in image 1 speaks in the voice timbre of audio 1\"*), and when\nyou are referencing only the timbre, **do not carry the reference clip's original dialogue into your\nprompt** — write the new line. When you genuinely want the same words re-performed, quote them\nexactly and say so.\n\n**An audio reference cannot travel alone** — H3 refuses a reference set that is audio only. Pair it\nwith at least one image or video reference.\n\n### 💸 Reference images past the free allowance are billed — and the two rows differ\n\nOn `minimax-h3` the first **5** are free and each additional image adds **4 credits**. On\n`minimax-h3-max` the first **4** are free and each additional image adds **1 credit** — fal prices\nMax's references by token rather than per image, and Slates normalises every Max reference to\n1024x1024 so that per-image number is exact. Both rows take **9** images, at every resolution and\nevery length. Four extra images on a 10s\n768p clip add 16 credits to a 30-credit generation: **more than half again**, for references that\noften make the output worse rather than better (see the 2–4 rule above).\n\nAttach the references the shot needs, not the ceiling. Call\n`slates_estimate_generation_cost` with `referenceImages` set to the real count before a\nreference-heavy job — a quote that omits it under-reports the bill.\n\n---\n\n## Frames\n\n`minimax-h3` and `minimax-h3-max` both take a **start frame**, an **end frame**, or both. With an\nend frame, land it explicitly: describe the final pose, spacing and composition as the thing the\nshot **settles into** at the end, rather than hoping the model finds it.\n\n> *\"…she rotates the handle into the final angle and settles into the pose, spacing and composition\n> of image 2 at the end of the shot.\"*\n\n**Frames and references are mutually exclusive** on both rows — they are different endpoints, and\nthe reference endpoint has no frame slots at all. Slates refuses the combination rather than\ndropping one side.\n\n---\n\n## Cost discipline\n\n| Combination | Credits |\n|---|---:|\n| `minimax-h3` · 768p · 5s | 15 |\n| `minimax-h3` · 768p · 10s | 30 |\n| `minimax-h3` · 2K · 10s | 65 |\n| `minimax-h3` · 4K · 10s | 80 |\n| `minimax-h3-max` · 768p · 10s | 40 |\n| `minimax-h3-max` · 1080p · 10s | 80 |\n| `minimax-h3` — every reference image past the **fifth** | **+4** |\n| `minimax-h3-max` — every reference image past the **fourth** | **+1** |\n\n**768p is the default for a reason.** It is the tier the model natively generates.\n\n🚨 **2K and 4K are UPSCALES of a 768p render, not larger generations.** fal's own schema says so:\n*\"480P and 768P are native generation modes; 2K and 4K upscale a 768P base result.\"* The upscaler\n(H3-Regenerate-2K) is a separate stage bolted onto a finished take — it can enlarge detail but it\ncannot add information.\n\n**In our own test (2026-08-27, same prompt, same seed) the 2K pass came back with MORE artifacting\nthan the 768p original it was built from**, while costing 33 credits for a 5-second take against 15,\nand taking nearly twice as long to return. One shot, so treat it as a warning rather than a law —\nbut the mechanism explains it, and the burden of proof is on 2K.\n\n**So: generate at 768p and judge it at 768p.** Reach for 2K or 4K only when a delivery spec demands\nthe pixels, and expect to be paying for size rather than quality — a post-production upscale from a\nclean 768p master is very often the better result. **4K video is Pro-only** (the server returns\n`PRO_REQUIRED` for a base account); 2K is open to every tier.\n\n---\n\n## Quick checklist\n\n- Body written as a timeline, first shot untimestamped, later shots on `[Shot N] At MM:SS.mmm`.\n- Camera motion written **into** a sentence with amplitude and speed.\n- Dialogue and diegetic music in the body; ambience in the soundscape section; audience-only score\n in the score section, described by instrument and tempo.\n- Voiceover carries both the off-screen phrase and the closed-lips statement.\n- References cited as `image 1` / `video 1` / `audio 1`, each with a stated job and a stated degree\n of retention. No angle-bracket tags.\n- Reference count is deliberate — you are paying 4 credits for each one past the fifth.\n- Frames **or** references, never both.\n- The prompt is the prompt: no expander will fill it out for you.\n",
|
|
23
|
+
"slates-prompting-minimax-h3": "---\nname: slates-prompting-minimax-h3\ndescription: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09), so the seats differ on ladder and price, not on what they accept. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; 4 free then +1 on Max), and audio written into the wrong section is dropped or duplicated.\n---\n\n# MiniMax H3 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — MiniMax H3.** The only seat where audio is AUTHORED rather than toggled: dialogue, scene sound and score are three separate sections of the prompt, generated in one pass, and putting a sound in the wrong section drops or doubles it.\n\n**The five levers**\n1. **Write the three audio layers separately** — `Scene sound:` for what is in the room, `Score:` for what only the audience hears, and the dialogue quoted inline. Section decides attribution.\n2. **Quote dialogue and name the language** — `says in English`, `speaks in Spanish`. Eleven languages are stably supported; the language is part of the instruction, not an afterthought.\n3. **Declare the reference RELATIONSHIP**, which no other seat has: `kept whole`, `partly kept`, `transferred`, or `a loose echo`. An undeclared reference is a guess.\n4. **Give a beat of stillness before a line** — `sits still for a beat, then looks up`. The sync needs something to lock against; a character already mid-motion when the line starts drifts.\n5. **Describe the beat structure** — `waits`, `then speaks`, `under the last three seconds`. H3 is a timeline, so write one.\n\n**Examples**\n- `A woman sits still at a kitchen table for a beat, then looks up. She says in English, \"You said Tuesday.\" Scene sound: a fridge hum, a spoon set down on formica. Score: none.`\n- `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, \"No es el alternador.\" Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`\n\n**Hard constraint:** the two seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p rather than 4K, takes the same 9+3+3 references, and costs MORE at the tier they share — it is a speed pick, never the cheap one. H3's top two resolution tiers are UPSCALES of the native render: judge at native. Reference images past the free allowance are a paid dimension of the cost key (5 free on base, 4 on Max) — declare the count when quoting.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- a sound written into the wrong audio section — it is dropped, doubled, or attributed to the wrong layer\n- `background music` as a bare instruction: the score is its own authored layer, audience-only, and it is named as such\n- an undeclared reference relationship — say kept whole, partly kept, transferred, or a loose echo\n<!-- @banned:end -->\n\nH3 is an **omni transformer**: it generates picture and sound in the same pass, at 24fps with\n32kHz stereo, 5–15 seconds, in 11 stably-supported languages (Arabic, Chinese, English, French,\nGerman, Italian, Japanese, Korean, Portuguese, Russian, Spanish). That single fact drives\neverything below — the prompt is not a shot description with sound bolted on, it is a **timeline\nwith three audio layers you author separately**.\n\n**Two seats, one grammar.** Everything in this file applies to both. They differ only in what the\nendpoint accepts:\n\n| | `minimax-h3` | `minimax-h3-max` |\n|---|---|---|\n| Resolution | 480p / 768p / **2K / 4K** | 480p / 768p / **1080p** |\n| References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) |\n| Frames | start and/or end | start and/or end |\n| Price at 768p | **$0.060/s** | $0.080/s |\n| Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) |\n\n**Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it\ntops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth\npaying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.\n\n**The speed is measured, not claimed** (2026-08-27, same prompt and params on both rows): a 5-second\n768p text-to-video finished in **4.8 seconds** on Max against **57 seconds** on base H3 — roughly\n**12x**, queue to finished file. fal advertises \"under 3 seconds\"; the literal claim did not hold at\n4.8s wall-clock, but the order of magnitude did. For iteration loops and client-present work that gap\nis the entire reason the seat exists.\n\n🚨 **Max's known weakness: colour banding in low light (Eric, 2026-09-09).** Certain shots —\nespecially dark or low-key ones — come back with low-bitrate-looking banding across gradients (skies,\nwalls, shadow falloff). It is the one place the seat visibly gives something up. If a shot is dark\nand gradient-heavy, either light it up in the prompt or route to base H3 at 768p; do not fix it by\nreaching for 2K, which adds its own artifacting on top.\n\n---\n\n## The one thing that makes H3 different: audio is a THREE-LAYER instruction\n\nEvery other video seat treats sound as on or off. H3 splits it, and the split is enforced by where\nyou write each thing. Get the section wrong and the sound is dropped, doubled, or attributed to the\nwrong source.\n\n| Layer | What belongs in it | Where it goes |\n|---|---|---|\n| **Synchronised events** | dialogue, singing, and any sound tied to a specific shot or action | the **body** of the prompt, on the beat it lands |\n| **Scene sound** | ambience and physical sounds that run across the whole clip — room tone, rain, traffic, a ventilation hum | the **soundscape** section |\n| **Score** | music the characters cannot hear; audience-only | the **music** section |\n\n**Three rules, all from MiniMax's own guide:**\n\n1. **Dialogue and singing NEVER go in the soundscape section.** They are synchronised events; they\n belong in the body, at the moment they happen.\n2. **Diegetic music — music the characters can hear** (a radio in the scene, a busker) — also\n belongs in the **body**, not in the score section. The score section is audience-only.\n3. **Write the score in instrumental terms, not mood words.** Name the instruments, the tempo, and\n how it develops. *\"A restrained solo-piano score at a slow tempo, sustained low cello underneath,\n no swell\"* — not *\"emotional music\"*.\n\nUse **N/A** for a section only when silence or absence is genuinely what the shot wants. An empty\nscore section is a real choice; a vague one is a wasted layer.\n\n### The shape, in the one prompt field\n\nSlates sends one prompt string, so write the three layers as labelled paragraphs in this order:\n\n```\n[Shot 1] Live-action, cinematic. A medium-wide shot frames a baker opening the shutters of a\nsmall street bakery before sunrise. The camera pushes in with small amplitude at slow speed as\nthe middle-aged baker with a calm, slightly raspy voice places a fresh loaf on the counter and\nsays: \"First batch of the morning.\" [Shot 2] At 00:05.000, the camera cuts to a close-up of\nsteam rising from the sliced bread while his final words carry over from the previous shot.\n\nSoundscape: wooden shutters scrape open over a quiet street, trays clink softly inside, a\ndoorbell rings once, then light footsteps and the crisp sound of bread being sliced.\n\nScore: a soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes,\ngentle fade at the end.\n```\n\n**Body target: 350–500 words** for a reference-carrying shot. Dialogue-heavy content prioritises\nfitting the complete spoken timeline over hitting a word count.\n\n🚨 **Slates disables the provider's prompt expander.** H3's API can rewrite your prompt before\ngeneration; Slates turns that off, because a model rewriting the user's words invisibly is banned\noutright (prompt transparency: what the composer shows is what the model gets). The practical\nconsequence is on you: **nothing will pad a thin prompt.** Write the whole body.\n\n---\n\n## Shots and timing\n\nThe first shot carries **no timestamp**. Every later shot opens with the bracket and a cut time\nthat increases and stays inside the clip length:\n\n```\n[Shot 2] At 00:03.500, the camera cuts to ...\n```\n\nTransition verbs the model knows: **cuts to · transitions to · changes to · switches to**.\n\n**Dialogue that continues across a cut** needs the continuity said out loud — *\"his final words\ncarry over from the previous shot\"* — or the line restarts. **Speech that ends abruptly** should be\ndescribed as cut off rather than trailed off.\n\n---\n\n## Camera — write the move into the sentence\n\nThe model has a named motion vocabulary:\n\n> Zoom In / Zoom Out · Push In / Pull Out · Pan Left / Pan Right · Truck Left / Truck Right ·\n> Tilt Up / Tilt Down · Pedestal Up / Pedestal Down · Arc Shot · Tracking Shot · Static Shot ·\n> Shake Slightly / Shake Strongly · POV · Roll Clockwise / Roll Counterclockwise\n\nModify with **amplitude** (`with small amplitude` / `with large amplitude`) and **speed**\n(`at slow speed` / `at fast speed`).\n\n🚨 **Integrate the motion into the sentence — never stack labels.** MiniMax's own example:\n*\"The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.\"*\nNot *\"Push In. Small amplitude. Slow.\"*\n\n---\n\n## Speakers and dialogue\n\nGive each speaking character a stable identity in the prose and keep it: describe the voice once\n(*\"a young woman with a quiet, breathy voice\"*), then refer back to the same description at every\nline. Identification, delivery and action sit **outside** the quoted line; the line itself is only\nthe words.\n\n```\nThe young woman with a quiet, breathy voice says: \"I get off at the next station.\"\n```\n\n**Voiceover** needs two things — the phrase *\"says in an off-screen voiceover\"* **and** an explicit\nstatement that the lips stay closed. Without the second half the model animates a mouth.\n\n```\nThe man says in an off-screen voiceover: \"I still remember that road.\" — his lips remain\ncompletely closed.\n```\n\n**On-screen text** — signs, banners, labels, subtitles, neon — goes in double quotes with the\noriginal wording preserved exactly: *A red neon sign reading \"Open Late\" glows above the doorway.*\n\n---\n\n## References — H3's real differentiator is the declared RELATIONSHIP\n\n*(BOTH rows. `minimax-h3-max` gained the reference set on 2026-09-09; its free allowance is\nFOUR images rather than the base row's five.)*\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### Cite references by number — Slates already does it for you\n\nH3 on fal takes references as **typed slots** and expects the prompt to name them by modality and\norder: **`image 1`, `image 2`, `video 1`, `audio 1`**. That is exactly what the Slates composer\nemits from your `@mentions` and `#tags` (`Marcus (image 1) in the workshop (image 2)`), in the\nexact order it sends them.\n\n🚨 **Do NOT hand-write angle-bracket reference tags.** MiniMax's own model-card grammar uses\n`<Subject N>` / `<Picture N>` / `<Video N>` / `<Audio N>` labels; the fal endpoints Slates calls do\nnot — they build the binding from the typed slots and ask for plain numbered prose. Typing the tags\nyourself puts literal angle brackets in the prompt the model reads.\n\n### State how much of each reference survives\n\nThis is the lever no other model in the catalogue gives you. Say, in plain words, what each\nreference is FOR and how much of it should carry through:\n\n| Intent | Say something like |\n|---|---|\n| Keep it whole | *\"Keep the woman in image 1 exactly as she appears — hair, cardigan, necklace.\"* |\n| Keep part of it | *\"Use the café in image 2 for the brick wall and the sofa; the lighting is late evening, not daylight.\"* |\n| **Move a trait onto someone else** | *\"Give the man in image 3 the weathered leather texture of the jacket in image 4.\"* |\n| Loose echo | *\"Match the general palette and grain of image 5; nothing else from it.\"* |\n\nThe third row is the one with no equivalent anywhere else in Slates: **transferring a characteristic\nonto a different subject** is a first-class thing H3 understands. Reach for H3 when that is the job.\n\n**Audio references** bind a voice or a texture without copying the words. Say which speaker an\naudio reference is for (*\"the woman in image 1 speaks in the voice timbre of audio 1\"*), and when\nyou are referencing only the timbre, **do not carry the reference clip's original dialogue into your\nprompt** — write the new line. When you genuinely want the same words re-performed, quote them\nexactly and say so.\n\n**An audio reference cannot travel alone** — H3 refuses a reference set that is audio only. Pair it\nwith at least one image or video reference.\n\n### 💸 Reference images past the free allowance are billed — and the two rows differ\n\nOn `minimax-h3` the first **5** are free and each additional image adds **4 credits**. On\n`minimax-h3-max` the first **4** are free and each additional image adds **1 credit** — fal prices\nMax's references by token rather than per image, and Slates normalises every Max reference to\n1024x1024 so that per-image number is exact. Both rows take **9** images, at every resolution and\nevery length. Four extra images on a 10s\n768p clip add 16 credits to a 30-credit generation: **more than half again**, for references that\noften make the output worse rather than better (see the 2–4 rule above).\n\nAttach the references the shot needs, not the ceiling. Call\n`slates_estimate_generation_cost` with `referenceImages` set to the real count before a\nreference-heavy job — a quote that omits it under-reports the bill.\n\n---\n\n## Frames\n\n`minimax-h3` and `minimax-h3-max` both take a **start frame**, an **end frame**, or both. With an\nend frame, land it explicitly: describe the final pose, spacing and composition as the thing the\nshot **settles into** at the end, rather than hoping the model finds it.\n\n> *\"…she rotates the handle into the final angle and settles into the pose, spacing and composition\n> of image 2 at the end of the shot.\"*\n\n**Frames and references are mutually exclusive** on both rows — they are different endpoints, and\nthe reference endpoint has no frame slots at all. Slates refuses the combination rather than\ndropping one side.\n\n---\n\n## Cost discipline\n\n| Combination | Credits |\n|---|---:|\n| `minimax-h3` · 768p · 5s | 15 |\n| `minimax-h3` · 768p · 10s | 30 |\n| `minimax-h3` · 2K · 10s | 65 |\n| `minimax-h3` · 4K · 10s | 80 |\n| `minimax-h3-max` · 768p · 10s | 40 |\n| `minimax-h3-max` · 1080p · 10s | 80 |\n| `minimax-h3` — every reference image past the **fifth** | **+4** |\n| `minimax-h3-max` — every reference image past the **fourth** | **+1** |\n\n**768p is the default for a reason.** It is the tier the model natively generates.\n\n🚨 **2K and 4K are UPSCALES of a 768p render, not larger generations.** fal's own schema says so:\n*\"480P and 768P are native generation modes; 2K and 4K upscale a 768P base result.\"* The upscaler\n(H3-Regenerate-2K) is a separate stage bolted onto a finished take — it can enlarge detail but it\ncannot add information.\n\n**In our own test (2026-08-27, same prompt, same seed) the 2K pass came back with MORE artifacting\nthan the 768p original it was built from**, while costing 33 credits for a 5-second take against 15,\nand taking nearly twice as long to return.\n\n🚨 **Confirmed independently (Eric, 2026-09-09): 2K and 4K carry visible AI noise artifacting and\n\"just look bad\".** That is now TWO separate observations, months apart, pointing the same way — it is\nno longer a single-shot warning. The tiers stay available because a delivery spec sometimes demands\nthe pixels, but **do not route to 2K/4K for quality**: you are paying more, waiting longer, and\nadding artifacts to a 768p render. Upscale in post from a clean 768p master instead.\n\n**So: generate at 768p and judge it at 768p.** Reach for 2K or 4K only when a delivery spec demands\nthe pixels, and expect to be paying for size rather than quality — a post-production upscale from a\nclean 768p master is very often the better result. **4K video is Pro-only** (the server returns\n`PRO_REQUIRED` for a base account); 2K is open to every tier.\n\n---\n\n## Quick checklist\n\n- Body written as a timeline, first shot untimestamped, later shots on `[Shot N] At MM:SS.mmm`.\n- Camera motion written **into** a sentence with amplitude and speed.\n- Dialogue and diegetic music in the body; ambience in the soundscape section; audience-only score\n in the score section, described by instrument and tempo.\n- Voiceover carries both the off-screen phrase and the closed-lips statement.\n- References cited as `image 1` / `video 1` / `audio 1`, each with a stated job and a stated degree\n of retention. No angle-bracket tags.\n- Reference count is deliberate — you are paying 4 credits for each one past the fifth.\n- Frames **or** references, never both.\n- The prompt is the prompt: no expander will fill it out for you.\n",
|
|
24
24
|
"slates-prompting-motion-transfer": "---\nname: slates-prompting-motion-transfer\ndescription: How to set up motion transfer — Kling Motion Control only (std and pro tiers, 5-second outputs). Read before calling slates_generate_motion_transfer. Reference image (character) + driving video (motion source) → new video of the character performing the motion. Asset selection rules, character_orientation, tiers, and prompt usage. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.\n---\n\n# Motion transfer — setup guide\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Motion transfer (Kling Motion Control only).** A target IMAGE (your character) plus a source VIDEO (the motion) produces your character performing that motion. Always 5 seconds.\n\n**The five levers**\n1. **The target image must show body proportions clearly** and the character must occupy more than about 5% of the frame. A tiny figure in a wide shot has nothing to drive.\n2. **Single character in the target.** A group image breaks the identity anchor.\n3. **Choose `characterOrientation` on purpose** — `video` takes the source clip's framing, `image` preserves the portrait's. It is the most-missed choice here.\n\n4. **The prompt is atmosphere only** — `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.` Motion verbs are ignored; the motion is already in the driving video.\n5. **Pick the best 5 seconds of the source up front**, and write only atmosphere: `soft afternoon sunlight`, `vintage warm color grade`, `clean studio backdrop`. The output is 5s regardless, so a long driving clip just wastes the choice.\n\n**Examples**\n- `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.`\n- `Clean studio backdrop, sharp focus on the character.` (Or leave it empty.)\n\n**Hard constraint:** cartoon driving videos fail, and a cropped or partial target character drifts. std is fine while the motion-and-framing combination is still moving; switch to pro once it is locked. For a REGENERATED shot instead — physical contact, cloth and hair, camera motion, native audio — that is a normal Seedance video generation with the driving clip as a video reference, not a mode of this tool.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** — the motion is already in the driving video, so motion verbs are ignored:\n- `spins faster`, `jumps higher`, `add more energy`, `moves quicker`\n<!-- @banned:end -->\n\nTake a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion. **This tool is Kling-only** — it wraps Kling Motion Control and nothing else.\n\n| Tier | Cost | Use case |\n|------|-----------|----------|\n| Kling std (`kling-mc-std-5s`) | ~32 credits / 5s | General motion transfer, budget lane |\n| Kling pro (`kling-mc-pro-5s`) | ~42 credits / 5s | Cleaner anatomy, better identity preservation |\n\nBoth tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)\n\n## Want Seedance instead? That is a video generation, not a mode here\n\nKling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion, and the output carries native audio. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the driving clip attached as a video reference and the character image as an ingredient, then write the prompt yourself:\n\n```\nThe character from image 1 performs the exact motion, choreography, and camera\nmovement from video 1. Preserve the character's identity, appearance, and outfit.\n```\n\nThat is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.\n\n- **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).\n- **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.\n- **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).\n- `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.\n\nEverything below is about the Kling tool.\n\n## Inputs\n\n- `sourceVideoAssetId` — driving video. **Must be a realistic human** with clear proportions. Anime/cartoon/CG driving videos fail.\n- `targetImageAssetId` — character to be animated. Can be any style (cartoon, anime, realistic, painted).\n- Both must already exist as assets in the project. Use `slates_list_assets` to find them or upload first.\n\n## Source video constraints\n\n- Realistic human (not animated, not CG)\n- Entire body OR upper body visible — head must not be obstructed\n- Subject occupies a clear share of the frame\n- Single primary subject. Multi-person driving videos confuse the motion anchor.\n- Clean motion — choppy / cut-edited driving videos produce jittery output\n\nGood driving video sources:\n- Reference dance footage with one subject\n- Walking / gesture / posing clips\n- Talking-head footage when paired with character_orientation: 'video'\n\nBad driving video sources:\n- Music videos with multi-shot edits\n- Anime / animation clips\n- Heavily stylized footage with smoke / particles obscuring the body\n- Footage where the subject's head leaves frame mid-clip\n\n## Target image constraints\n\n- Character body proportions clearly visible\n- Character occupies >5% of image area (not a tiny figure in a wide shot)\n- Single character. Group images break the identity anchor.\n- Any artistic style works — cartoon, anime, painted, realistic, 3D render\n\nAvoid:\n- Extreme close-up of just the face (no body to drive)\n- Character partially cropped at the waist when the driving video is full-body\n- Multiple characters\n\n## character_orientation — the most-missed choice\n\nThis single parameter changes the output dramatically. Pick deliberately.\n\n| Value | Output framing | Max source duration | Best for |\n|-------|----------------|---------------------|----------|\n| `video` | Matches driving video framing | Up to 30s source | Complex full-body motion (dance, action, athletics) |\n| `image` | Matches target image framing | Up to 10s source | Camera moves, simpler motion, preserving original composition |\n\n**Default `video`** when the driving video has the look you want (most cases).\n\nSwitch to `image` when the target image's composition is the brand asset and the motion is secondary (e.g., a hero shot of a character that needs subtle gesture, not a full performance).\n\n## Tier choice — std vs pro\n\n**std (~32 credits)** for:\n- Drafts, motion exploration, blocking\n- Group scenes where the character isn't a hero shot\n- When the budget is tight and the motion is the focus\n\n**pro (~42 credits)** for:\n- Final hero takes\n- Branded characters where identity drift = unacceptable\n- Anatomically complex motion (limbs crossing, fast direction changes)\n- Anime / cartoon target images — pro handles non-realistic styles better\n\nDon't default to pro. The ~10-credit delta compounds fast across iteration.\n\n## Prompt usage (optional)\n\nThe `prompt` field is **scene/style refinement**, not motion direction. The motion comes from the driving video — the prompt sets ambiance, lighting, additional detail.\n\nGood:\n- `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.`\n- `Clean studio backdrop, sharp focus on the character.`\n\nBad (model ignores motion verbs — they're already in the driving video):\n- ❌ `She spins faster and jumps higher.`\n- ❌ `Add more energy to the dance.`\n\nLeave it empty if you don't have a specific atmospheric note.\n\n## Common failure modes\n\n| Symptom | Likely cause | Fix |\n|---------|--------------|-----|\n| Limbs distort / extra fingers | std tier, complex motion | Switch to pro |\n| Character identity drifts | Target image cropped too tight | Use a fuller-body target |\n| Output looks \"stuck\" / minimal motion | Driving video subject too small in frame | Pick a driving video where the subject fills more of the frame |\n| Cartoon target turns realistic | std tier on stylized art | Switch to pro — handles non-realistic styles better |\n| Garbled output entirely | Anime / CG driving video | Use realistic human driving footage |\n| Wrong framing on output | character_orientation set wrong | Try the other value |\n| Background bleeds through character | Target image had complex background | Use a target with cleaner background separation |\n\n## Workflow patterns\n\n**Reference dance to brand character:**\n1. Generate or upload the brand character as a still image (clean background, full body, single subject)\n2. Find driving footage — a clean reference video of the dance you want\n3. Upload both as project assets\n4. Run motion transfer with `motionModel: 'kling-mc-pro'`, `characterOrientation: 'video'`\n5. Total cost: ~42 credits per 5s take\n\n**Subtle motion on a hero portrait:**\n1. Use the locked hero portrait as the target image\n2. Pick a driving video with subtle gesture (head turn, slight posture shift)\n3. `characterOrientation: 'image'` to preserve the portrait's framing\n4. std tier is fine for this case — motion isn't dramatic\n\n**Avoid:**\n- Pro tier on first iteration — waste, switch to it once the motion + framing combo is locked\n- Cartoon driving videos — guaranteed failure\n- Cropped or partial target characters — identity will drift\n- Long driving videos when output is 5s — pick the best 5s of the source upfront\n\n## Cost discipline\n\n- 5 seconds, no shorter option\n- Both tiers trip the confirm gate — every call needs explicit user OK\n- Iteration is expensive: 4 takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.\n- Always run a single std take first to validate the motion + framing combo before committing to pro\n\n## Confirm gate: cost + codes, no inline preview\n\nMotion transfer is mechanical — the model deterministically applies source motion to target image. Both tiers trip the confirm gate; the response includes the asset codes for source and target so you can announce them in chat.\n\n- ✅ \"Transferring motion from **VID-V3** onto **IMG-A12 — Detective Closeup**. ~42 credits, confirm?\"\n- ❌ \"Using the walk video and the detective image...\" (multiple of each in the project.)\n\nDon't second-guess the assets the user picked — the model executes the transfer. If the output is wrong, iterate on motion source or target choice, not on a refinement prompt.\n\n## Sources\n\n- [fal.ai — Kling Motion Control V3 Standard](https://fal.ai/models/fal-ai/kling-video/v3/standard/motion-control)\n- [fal.ai — Kling Motion Control V3 Pro](https://fal.ai/models/fal-ai/kling-video/v3/pro/motion-control)\n",
|
|
25
25
|
"slates-prompting-nano-banana-2": "---\nname: slates-prompting-nano-banana-2\ndescription: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3.1 Flash Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.\n---\n\n# Nano Banana 2 — cinematic & photorealistic prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Nano Banana 2 (Gemini 3.1 Flash Image).** Brief it like a creative director, not a tag list. Structure: `Film still from [director] [genre]. Shot on [camera] with [lens]. [Subject and action]. [3-5 specific visual details]. [Lighting — direction + quality]. [Color palette]. [Film stock]. [1-2 word tone].`\n\n**The five levers**\n1. **Named lens + aperture** beats \"shallow depth of field\" — `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin), `Panavision anamorphic`, `400mm telephoto`.\n2. **Light by direction and quality**, never \"good lighting\" — `hard sidelight from a single window, deep falloff`, `overcast north light`, `practical tungsten spill`.\n3. **A named film stock or sensor** carries a whole palette — `Kodak Portra 400`, `Cinestill 800T`, `ARRI Alexa 65`.\n4. **Composition as a shot** — `low angle`, `aerial view`, `rule of thirds with the subject camera-left`, `foreground occlusion`.\n5. **Positive framing only.** Describe what is there. \"Empty street\", never \"no cars\"; \"unstaged documentary photography\", never \"not anime\".\n\n**Examples**\n- `Film still from a Denis Villeneuve thriller. Shot on ARRI Alexa 65, 85mm f/1.4. A woman in a charcoal wool coat stands at a rain-slick bus stop, breath visible. Hard sodium light from a single overhead lamp, deep falloff into blue night. Kodak Vision3 500T. Isolated.`\n- `Editorial still life on seamless bone paper. 100mm macro, f/8. A cracked ceramic bowl holding three figs. Soft north light from camera-left, one gentle shadow. Muted earth palette. Portra 400 grain. Quiet.`\n\n**Hard constraint:** there is no `negativePrompt` field. Suppress by reframing positively, or inline `without` / `free of`. Knowledge cutoff January 2025 — anything later needs reference images.\n<!-- @card:end -->\n\nNano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the default model behind `slates_generate_image` — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.\n\nKnowledge cutoff: January 2025. Anything after needs explicit reference images.\n\n## Google's 4 official rules (verbatim)\n\n1. **Be specific.** Provide concrete details on subject, lighting, and composition.\n2. **Use positive framing.** Describe what you want, not what you don't want.\n3. **Control the camera.** Use photographic and cinematic terms like \"low angle\" and \"aerial view.\"\n4. **Iterate.** Refine images with follow-up prompts in a conversational manner.\n\n## Official prompt formula\n\n```\n[Subject] + [Action] + [Location/context] + [Composition] + [Style]\n```\n\nFor the cinematic / photoreal use case, expand to:\n\n```\nFilm still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and action]. [3-5 specific visual details]. [LIGHTING — direction + quality]. [COLOR PALETTE]. [FILM STOCK or sensor language]. [1-2 word emotional tone].\n```\n\n## Photorealism positives — what consistently works\n\n> ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern — do not carry it across.**\n> Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.\n> The leak happens in one specific way — you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `slates-prompting-seedance` (Part 3, \"Don't cross-pollinate image-model syntax\").\n\n**Named lenses + apertures** beat generic \"shallow depth of field\":\n- `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`\n- `Panavision anamorphic` for horizontal flares + cinematic width\n- `400mm telephoto` for compression + isolation\n- `24mm` for environmental interiors\n\n**Named cameras / sensors:**\n- `ARRI Alexa 65`, `Hasselblad X2D`, `Canon EOS R5`, `Sony A7III`, `Fujifilm X-T5`\n- \"Specific gear\" beats \"DSLR\"\n\n**Named film stocks** (one per prompt — never mix):\n- `Kodak Portra 400` — natural skin, warm\n- `Fuji Velvia 50` — saturated, landscape\n- `Ilford HP5 Plus` — black and white, gritty grain\n- `CineStill 800T` — tungsten night, halation\n\n**Physics-based lighting** (direction + quality):\n- `Single key light at 45 degrees from upper left`\n- `Late afternoon sun at 15 degrees above horizon`\n- `Color temperature 4500K` beats `slightly warm`\n- `Practicals only — no fill` for Deakins-style realism\n\n**Imperfection vocabulary** (forces away from AI-clean):\n- `visible pores`, `natural skin grain`, `peach fuzz`, `slight hyperpigmentation`\n- `unretouched raw photography`, `ISO noise`, `sweat beading`\n- `crisp catchlights in the eyes`, `skin micro-detail`\n\n**Director references** (use when locking style):\n| Director | Tone | Visual signature |\n|---|---|---|\n| Denis Villeneuve | Cold, vast, existential | Desaturated, overwhelming scale |\n| Roger Deakins | Precise motivated light | Single source, deep shadows, practicals |\n| Emmanuel Lubezki | Natural, spiritual | Available light, golden hour |\n| Bradford Young | Warm darkness | Underexposed, rich shadows, skin tones |\n\n**Genre cues that move the model:**\n- `unstaged documentary photography style`\n- `fashion magazine editorial, shot on medium-format analog film, pronounced grain`\n- `Film still from [Director] [genre]`\n\n## The anti-list — phrases that DEGRADE realism\n\nThese are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is extracted\n by src/prompts/banned-tokens.ts, inlined verbatim into the slates_generate_image\n op description (always in context on both surfaces), and matched against every\n submitted prompt. Editing this list changes what the agent is told AND what it\n is warned about — keep every entry backticked, and keep prose outside the\n backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- `8k`, `4k` (as a quality token)\n- `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone\n- `masterpiece`, `best quality`, `highly detailed`, `ultra-detailed`\n- `trending on ArtStation`, `award-winning`\n- `perfect skin`, `flawless`, `airbrushed`, `smooth skin`\n- `cinematic` standing alone — always specify *which cinema* (director, lens, era, stock)\n- `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue\n<!-- @banned:end -->\n\n## Negative prompting — there is no field\n\nNano Banana 2 has **no `negativePrompt` parameter**. Three patterns to suppress unwanted content:\n\n1. **Positive reframing (preferred):** \"empty street\" not \"no cars\". \"Unstaged documentary photography\" not \"not anime.\"\n2. **Inline `without` / `free of`:** \"without any people, vehicles, or man-made structures\", \"free of text overlays, logos, or watermarks.\"\n3. **Constraint clauses for anatomy/quality:** \"accurate anatomy with five fingers per hand, symmetrical features, natural proportions\"; \"sharp, well-exposed, free of blur or JPEG artifacts.\"\n\nDefault to #1. Reach for #2 only when positive framing can't suppress the unwanted element.\n\n## Reference images\n\n- **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.\n- **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as \"image N\" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **\"assign a distinct name to each character/object\"**. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render the scene's expression\") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n\n### Reference rules (the verified ones)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Nano Banana 2 specifically\n\n- **NB2's own consistency lever is \"assign a distinct name to each character/object.\"** That is Google's phrasing for rule 3 — cite each canonical identity inline by name.\n- **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.\n- **Character consistency is officially \"not 100% perfect\"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.\n- **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.\n\n## Common failure modes + fixes\n\n**Hands:** Append `accurate anatomy with five fingers per hand, symmetrical features, natural proportions, relaxed open palm`. Avoid heavy jewelry, props intersecting fingers, motion blur in references.\n\n**Text in images:** Quote-wrap target text. Specify font (`Century Gothic, 12pt`). Long phrases work; small text degrades. Two-step works best — generate text concepts conversationally first, then ask for the image.\n\n**Left/right confusion:** Default is **viewer's perspective**, not subject's. Append `left and right are from the character's perspective, NOT the camera's` when scene-blocking matters.\n\n**Surreal / absurd prompts trip uncanny valley:** The model drags toward realism. If you want surrealism, lean hard into stylization keywords (`painted`, `illustrated`, `stop-motion`).\n\n**Soft faces / dead eyes:** Add `crisp catchlights in the eyes`, `skin micro-detail`, `peach fuzz visible`. Don't stack quality enhancers — single clean prompt beats multiple re-interpretations.\n\n**Post-cutoff content (anything after Jan 2025):** Use reference images. The model has no knowledge of recent franchises, products, events.\n\n## Resolution tactics\n\n- Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — check current numbers<!-- slates-only --> by calling `slates_estimate_generation_cost`<!-- /slates-only -->. Pick the cheapest resolution that serves the use case.\n- **At 2K and above, the model allocates more tokens to surface detail** — explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.\n- 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.\n- 2K generations vary 20-60s+. Don't time-budget tightly.\n\n## Boring vs cinema — examples\n\n❌ **Boring:** \"Wide shot of a man on a dock looking at the forest.\"\n\n✅ **Cinema:** \"Direct overhead drone shot on weathered dock surface. Single figure standing center frame, climbing up from frame bottom. Boot prints leading away from him toward shore. Pale winter light. Anamorphic lens flare from low sun. Desaturated blue and slate grey palette. Kodak Portra 400 grain. The path already walked by someone else. Map of threat.\"\n\n❌ **Boring:** \"Close up of a woman looking scared.\"\n\n✅ **Cinema:** \"Extreme close on subject's mouth and nose, 135mm f/2.8, shallow depth of field. Breath pluming out, catching cold light from upper-left key. Lips slightly parted, peach fuzz visible. The breath holds. CineStill 800T halation around catchlights. Waiting.\"\n\n## The 3-strike rule\n\nIf three iterations on the same prompt haven't produced what the user wants, stop. Hand back to the user with what you tried and what isn't working. The slot machine doesn't converge — the prompt structure is wrong, not the seed.\n\n## Family variants — Lite and Pro\n\nEverything in this skill applies to the whole Nano Banana family; two variants trade speed/ceiling around NB2 full:\n\n- **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.\n- **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.\n\n<!-- slates-only -->\nRouting between them (and vs GPT Image 2.5 / FLUX / Seedream): `slates-model-selection`.\n<!-- /slates-only -->\n",
|
|
26
26
|
"slates-prompting-omni-flash": "---\nname: slates-prompting-omni-flash\ndescription: How to prompt Gemini Omni Flash (Google, via fal). Read before calling slates_generate_video with omni-flash or slates_edit_video with omni-flash-edit. Cheap 720p tier with native synced audio included — 3-10s, 16:9/9:16 only; t2v, single-start-frame i2v, or reference-to-video with up to 7 reference images. The edit variant is the EDIT-FIDELITY WINNER for footage-synced VFX (receipt 2026-07-09) — but ONLY with short prompts: one change + \"Keep everything else the same.\" Long descriptive prompts destroy fidelity.\n---\n\n# Gemini Omni Flash — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Gemini Omni Flash.** Two different jobs with OPPOSITE prompt rules, and getting them the wrong way round is the whole failure mode.\n\n**The five levers**\n1. **Editing: short prompt, ONE change, nothing else.** Google's own doc says so and a 2026-07-09 receipt confirms it — a long \"keep every frame identical\" preamble produced WORSE drift than two sentences.\n2. **Editing: always end with `Keep everything else the same.`** — the one documented preservation lever.\n\n3. **Editing: describe the EFFECT, never a real object as a metaphor.** \"Candle-like flame\" rendered a literal candle in the subject's hand.\n4. **Editing: no conditional timing cues.** \"…when he calls it, as he walks…\" hard-fails with `invalid_request`. Collapse to one continuous action; the model syncs to the footage's own motion.\n5. **Generation: the opposite — describe fully.** Subject, action, setting, `camera tracking alongside`, `overcast flat light`, tone. Audio is prompt-driven with no parameters: dialogue in quotes, sound in plain language — `rain patters on the tin roof`, `spray from tyres`, `a horn somewhere behind`.\n\n**Examples**\n- Edit: `Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.`\n- Generate: `A courier in a yellow shell jacket weaves between stalled cars on a wet arterial road, camera tracking alongside at shoulder height. Overcast flat light, spray from tyres. Rain patters on car roofs, a horn somewhere behind.`\n\n**Hard constraint:** it is a CHEAP DRAFT seat for generation and the EDIT-fidelity winner for footage-synced VFX — never a hero generation shot. Expect a possible jitter or doubled speech beat in the last half second of an edit: trim the tail rather than burning a re-roll.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use in an EDIT prompt** (each one has a receipt above):\n- a long preservation preamble — it produces WORSE drift than `Keep everything else the same.`\n- a real object as a metaphor: `candle-like`, `flame-like`, `laser-like`\n- a conditional timing cue: `when he`, `as she`, `once they` — these hard-fail, they do not merely drift\n- harm-to-person framing: `ignite`, `catch fire`, `on fire` applied to a person trips the safety filter\n<!-- @banned:end -->\n\nGoogle's fast video generation + editing model (\"Nano Banana Pro for video\" in creator slang — a nickname; it is NOT the NB Pro image model). Carried on fal (`google/gemini-omni-flash*`). 720p only, 24fps, 3–10 second clips, 16:9 or 9:16. **Audio is native and included** — dialogue, SFX, and ambient generate WITH the video at no extra cost.\n\n## Where it routes\n\n- **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.\n- **Cheap drafts and iteration volume** — lowest-cost audio-native video seat (~6.4 cr/s at 720p).\n- **NOT hero GENERATION shots** — Kling 3.0 stays the general gen default, Seedance 2.0 the premium tier; Omni Flash's *generation* quality seat is still unproven.\n\n## Editing (`slates_edit_video`, model `omni-flash-edit`) — THE RULES (receipts, not theory)\n\n1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *\"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes.\"* Live receipt 2026-07-09: a long \"keep every frame/word/movement identical…\" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *\"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.\"*\n2. **Always end with \"Keep everything else the same.\"** — the one documented preservation lever.\n3. **Never name a real-world object as a metaphor.** \"Candle-like flame\" rendered a literal candle in his hand. Describe the effect itself (\"small magical flames on his fingertips\").\n3b. **No conditional timing cues — they HARD-FAIL, not drift.** Receipt 2026-07-09: \"a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…\" → deterministic `invalid_request` (2×, \"could not generate with the given inputs\"); collapsing to one continuous action — \"A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke.\" — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video.\n4. **Safety filter (Google's, strict about harm-to-person):** \"fingertips ignite / catch fire\" → `content_policy_violation`. Frame effects as magical/harmless VFX: \"small magical flames appear on his fingertips\" passed. See slates-content-policy §Gemini for the substitution patterns.\n5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.\n6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.\n7. Source clip 3–10s (trim longer clips first). Output length follows the source; billing per output second, rounded up. Voice editing unsupported — never ask it to change dialogue.\n8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).\n9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.\n\n## Generation (`slates_generate_video`, model `omni-flash`)\n\n- **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.\n- Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.\n- **Name references inline** the standard Slates way (\"Marcus (image 1) walks…\"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.\n- **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language (\"rain patters on the tin roof\"). Negative direction as plain instructions (\"Do not show text\").\n- Duration is an explicit 3–10s integer param; cost scales linearly per second.\n\n## Input conditioning (Slates handles this — know it exists)\n\nPhone footage stores rotation as a metadata flag; models ignore it and edit the raw sideways pixels. Clips must be rotation-normalized (and oversized sources downscaled) before upload — receipt 2026-07-09: a portrait Pixel clip came back sideways until conditioned. If an edit output comes back rotated, the source wasn't normalized.\n\n## Content notes\n\n- Google applies its own safety filters to input images/clips and output. Uploads containing recognizable real people are restricted by Google's policy — though own-footage editing of the uploader passed on our route 2026-07-09. See slates-content-policy.\n- Output carries an invisible SynthID watermark (Google-side, programmatic detection only).\n",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@slatesvideo/shared",
|
|
3
|
-
"version": "0.6.
|
|
3
|
+
"version": "0.6.9",
|
|
4
4
|
"description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -71,6 +71,12 @@ paying for; route to base H3 for anything needing resolution, references, or the
|
|
|
71
71
|
4.8s wall-clock, but the order of magnitude did. For iteration loops and client-present work that gap
|
|
72
72
|
is the entire reason the seat exists.
|
|
73
73
|
|
|
74
|
+
🚨 **Max's known weakness: colour banding in low light (Eric, 2026-09-09).** Certain shots —
|
|
75
|
+
especially dark or low-key ones — come back with low-bitrate-looking banding across gradients (skies,
|
|
76
|
+
walls, shadow falloff). It is the one place the seat visibly gives something up. If a shot is dark
|
|
77
|
+
and gradient-heavy, either light it up in the prompt or route to base H3 at 768p; do not fix it by
|
|
78
|
+
reaching for 2K, which adds its own artifacting on top.
|
|
79
|
+
|
|
74
80
|
---
|
|
75
81
|
|
|
76
82
|
## The one thing that makes H3 different: audio is a THREE-LAYER instruction
|
|
@@ -307,8 +313,13 @@ cannot add information.
|
|
|
307
313
|
|
|
308
314
|
**In our own test (2026-08-27, same prompt, same seed) the 2K pass came back with MORE artifacting
|
|
309
315
|
than the 768p original it was built from**, while costing 33 credits for a 5-second take against 15,
|
|
310
|
-
and taking nearly twice as long to return.
|
|
311
|
-
|
|
316
|
+
and taking nearly twice as long to return.
|
|
317
|
+
|
|
318
|
+
🚨 **Confirmed independently (Eric, 2026-09-09): 2K and 4K carry visible AI noise artifacting and
|
|
319
|
+
"just look bad".** That is now TWO separate observations, months apart, pointing the same way — it is
|
|
320
|
+
no longer a single-shot warning. The tiers stay available because a delivery spec sometimes demands
|
|
321
|
+
the pixels, but **do not route to 2K/4K for quality**: you are paying more, waiting longer, and
|
|
322
|
+
adding artifacts to a 768p render. Upscale in post from a clean 768p master instead.
|
|
312
323
|
|
|
313
324
|
**So: generate at 768p and judge it at 768p.** Reach for 2K or 4K only when a delivery spec demands
|
|
314
325
|
the pixels, and expect to be paying for size rather than quality — a post-production upscale from a
|