@slatesvideo/shared 0.6.7 → 0.6.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/manual/content.d.ts +1 -1
- package/dist/manual/content.js +1 -1
- package/dist/operations/index.js +18 -6
- package/dist/prompts/model-capabilities.js +37 -27
- package/dist/prompts/prompting-tips.js +1 -1
- package/dist/skills/content.js +3 -3
- package/package.json +83 -83
- package/skills/slates-model-selection.md +1 -1
- package/skills/slates-one-prompt-film.md +1 -1
- package/skills/slates-prompting-minimax-h3.md +28 -12
package/dist/manual/content.d.ts
CHANGED
|
@@ -1,2 +1,2 @@
|
|
|
1
|
-
export declare const APP_MANUAL = "# Slates \u2014 Complete Reference for AI Assistants\r\n\r\n<!-- The heading above must stay first: shipped app builds reject this file if it\r\n does not begin with one. See slate/CLAUDE.md; check:llm-docs enforces it. -->\r\n\r\n<system_role>\r\nYou are a support assistant for the Slates desktop application. Answer user questions using ONLY the information in the <slates_reference> below. Be concise and direct. Use numbered steps for procedures. Use bullet points for explanations when helpful.\r\n</system_role>\r\n\r\n<rules>\r\n- If the answer cannot be found in the <slates_reference>, say: \"That isn't covered in the Slates reference.\" Do not guess or invent features.\r\n- If the user asks how to do something, give step-by-step instructions from the workflows and features described here.\r\n- If the user reports an error, check the TROUBLESHOOTING section first.\r\n- Refer to the FEATURES NOT IN SLATES section before answering questions about capabilities that might not exist.\r\n- Quote the exact error message when referencing troubleshooting entries.\r\n</rules>\r\n\r\n<slates_reference>\r\n\r\n<!-- BEGIN:GENERATED header -->\n# SLATES v1.5.6 \u2014 Complete Reference\n\n> **Freshness.** Generated from the Slates source of truth for app version **1.5.6**, last changed **2026-09-09**. The canonical copy of this file is <https://slates.video/slates-reference.md>. If a model, price or feature the user mentions is missing below, this copy is out of date: re-fetch that URL before answering, and say so.\n<!-- END:GENERATED header -->\r\n\r\n**What is Slates?** Desktop app (Windows 10/11, macOS 12+) for AI image and video creation. One-time purchase, no subscription. Every license includes 1,000 free credits, and Slates Pro starts with 3,000 credits. Every generation runs on Slates Credits \u2014 there are no API keys to set up, and credits never expire.\r\n\r\n---\r\n\r\n## INTENDED WORKFLOW\r\n\r\nThe designed start-to-finish flow:\r\n\r\n1. **Create project** \u2014 New project with name/description. Creates folder on your disk.\r\n2. **Build visual assets** \u2014 Generate images, create characters (with character sheets for consistency), environments (with environment grids), and styles. This is your visual library.\r\n3. **Create storyboard** \u2014 Add scenes. Every picture you drop in becomes a **Shot**: one beat of the piece, holding its references, its model, its settings and (when you want them) its words.\r\n4. **Write the piece** \u2014 Switch the storyboard to **Document** and write. What is said, what happens, the framing, the prompt each beat will send. Nothing is required and nothing is asked for; a visuals-only piece is finished as it stands.\r\n5. **Read it before you pay for it** \u2014 The header states how many generations, how many cuts, how long it runs and what it will cost. Press play for a rough cut at the real timing. Re-chop with split and merge and watch the price move.\r\n6. **Generate** \u2014 Select the beats you want and fire them in one approved batch. Each result lands under the row that made it.\r\n7. **Organize** \u2014 Switch to **Board** to re-order. Drag a beat and the whole thing moves with it.\r\n8. **Export to timeline** \u2014 Send clips to the built-in multi-track video editor.\r\n9. **Edit** \u2014 Trim, reorder, add markers, adjust timing.\r\n10. **Final export** \u2014 Export to MP4 directly, or export DaVinci Resolve XML for professional color grading.\r\n\r\n---\r\n\r\n## MODEL REFERENCE TABLE\r\n\r\n**Generating 4K video is a Slates Pro feature** \u2014 every tier generates video up to 1080p, and 4K images are open to everyone. Exporting your finished timeline at 4K is available on every tier.\r\n\r\nEvery model runs on Slates Credits. **The exact credit cost appears on the Generate button before anything fires.** The tables below are generated from the app's own model registry and rate tables, so they describe exactly what the model picker offers in this version: aspect ratios, resolutions, durations, reference-image limits, and the credit price of each. Bigger credit packs lower your per-credit cost, and Slates Pro gets the best pack rate on every purchase.\r\n\r\n<!-- BEGIN:GENERATED model-tables -->\n### Image Models\n\n| Model | Aspect Ratios | Resolutions | Max Refs | Credits per image |\n|-------|--------------|-------------|----------|-------------------|\n| **GPT Image 2.5 Flare** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 \u00B7 3K 3 \u00B7 4K 5 (default quality; at max: 2K 8 \u00B7 3K 11 \u00B7 4K 20) |\n| **GPT Image 2.5 Sunburst** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 \u00B7 3K 3 \u00B7 4K 5 (default quality; at max: 2K 8 \u00B7 3K 11 \u00B7 4K 20) |\n| **Nano Banana 2** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 4 \u00B7 2K 6 \u00B7 4K 8 |\n| **NB2 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K | 4 | 1K 2 |\n| **Nano Banana Pro** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 8 \u00B7 2K 8 \u00B7 4K 15 |\n| **FLUX.2 Max** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 4 | 1K 4 \u00B7 2K 5 \u00B7 4K 8 |\n| **Seedream 5 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 2K / 3K / 4K | 10 | 2K 2 \u00B7 3K 2 \u00B7 4K 2 |\n\n### Video Models\n\n| Model | Duration | Aspect Ratios | Resolutions | Max Refs | Audio | Credits per second |\n|-------|----------|--------------|-------------|----------|-------|--------------------|\n| **Seedance 2.0** | 4-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p / 4K | 9 | Included | 480p 3.5 \u00B7 720p 7.5 \u00B7 1080p 18.5 \u00B7 4K 39 |\n| **Seedance 2.5** | 4-30s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 30 | Included | 480p 5.1 \u00B7 720p 11.6 \u00B7 1080p 20.5 |\n| **Seedance 2.5 Edit** | Follows the source clip (4-30s) | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 0 | Included | 480p 6.2 \u00B7 720p 13.9 \u00B7 1080p 24.6 |\n| **Kling V3.0 Standard** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (6.3 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (8.4 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Omni** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (5.6 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Omni Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (7 with audio) \u00B7 4K 21 |\n| **Kling O3 Edit** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 6.3 |\n| **Kling O3 Edit Pro** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 8.4 |\n| **MiniMax H3** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 2K / 4K | 9 | Included | 480p 2.5 \u00B7 768p 3 \u00B7 2K 6.5 \u00B7 4K 8 |\n| **MiniMax H3 Max** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p | 0 | Included | 480p 2.5 \u00B7 768p 4 |\n| **Gemini Omni Flash** | 3-10s | 16:9, 9:16 | 720p | 7 | Included | 720p 6.4 |\n| **Omni Flash Edit** | Follows the source clip (3-10s) | 16:9, 9:16 | 720p | 0 | Included | 720p 6.4 |\n| **LTX-2.5** | 6/8/10/12/14/16/18/20s at 720p/1080p; 6/8/10s at 1440p/4K | 16:9, 9:16 | 720p / 1080p / 1440p / 4K | 0 | Included | 720p 4.5 \u00B7 1080p 6.5 \u00B7 1440p 9.5 \u00B7 4K 15 |\n| **LTX-2.5 Pro** | 6, 8, 10s | 16:9, 9:16 | 720p / 1080p | 0 | Included | 720p 6 \u00B7 1080p 8.5 |\n| **Veo 3.1 Fast** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 5 (7.5 with audio) \u00B7 1080p 5 (7.5 with audio) \u00B7 4K 15 (17.5 with audio) |\n| **Veo 3.1 Standard** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 10 (20 with audio) \u00B7 1080p 10 (20 with audio) \u00B7 4K 20 (30 with audio) |\n\n**Seedance 2.0 \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 27.4 credits per second at 1080p.\n**Seedance 2.5 \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 16.3 credits per second at 720p.\n**Seedance 2.5 Edit \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 19.8 credits per second at 720p.\n\n### Audio Models\n\n| Model | Length | Credits |\n|-------|--------|---------|\n| **Seed Audio 1.0** | 3-120s | 1 at 3s \u00B7 3 at 15s \u00B7 19 at 120s |\n| **Inworld TTS-2** | up to 2,000 characters of text per take | 1 at 250 characters \u00B7 2 at 2,000 characters (billed per 250) |\n| **Sound Effects** | 1-22s | 1 at 1s \u00B7 1 at 4s \u00B7 3 at 22s |\n\n### Tools (Lip Sync, Motion Transfer)\n\nThese are real Kling endpoints that take a clip or a still as their subject, not models you prompt from scratch. Both bill in 5-second blocks.\n\n| Tool | Input | Billed in | Credits per block |\n|------|-------|-----------|-------------------|\n| **Kling Lip Sync** | Video source | 5s block | 4 |\n| **Kling Lip Sync (Avatar v2 Standard)** | Still-image source | 5s block | 14 |\n| **Kling Lip Sync (Avatar v2 Pro)** | Still-image source | 5s block | 29 |\n| **Kling Motion Control Standard** | Still image + reference video | 5s block | 32 |\n| **Kling Motion Control Pro** | Still image + reference video | 5s block | 42 |\n<!-- END:GENERATED model-tables -->\r\n\r\n---\r\n\r\n## WHICH MODEL TO USE\r\n\r\n### Images\r\n\r\n**Nano Banana 2 is the default image model.** It is the best all-round image model in the app: the most reference images of any image model, every aspect ratio, and output up to 4K. Brief it like a creative director rather than with tag soup. It is also the only model that supports the 2x2 / 3x3 grid exploration wrapper.\r\n\r\n- **NB2 Lite** is the fast, cheap draft seat in the same Nano Banana family. Roughly half the price of NB2 full and noticeably faster, 1K output only. Iterate here, finish on NB2.\r\n- **Nano Banana Pro** is the hero-frame and typography tier. Reach for it when spatial composition, cinematic lighting and skin, or fine in-image type have to be perfect. NB2 gets you most of the way there, so this is a deliberate step up, never a default.\r\n- **GPT Image 2.5** is the strongest image model in the app: it follows a long instruction more faithfully than anything else here, and it is the one to pick when the picture simply has to be right. It is also the sharp-text model, which is what makes it the choice for character sheets, shot grids, ordered panels and anything with words in the picture. It comes in two seats that cost exactly the same, and the difference is speed against quality. **Flare** is the fast one: OpenAI describes its quality as comparable to the older GPT Image 2, at roughly half the wait. **Sunburst** is OpenAI's most capable image model, better than GPT Image 2, and deliberately slower. Use Flare while you are still exploring, then re-run the shot you like on Sunburst for the final \u2014 and reach for Sunburst directly when several reference images all have to survive into one frame, or when an edit must change one region and leave identity, geometry and lighting untouched.\n Its quality knob has **five** settings \u2014 `low`, `medium`, `high`, `xhigh`, `max` \u2014 spanning about 36\u00D7 from cheapest to dearest, which makes it the biggest cost lever on the model. The steps are uneven rather than a constant multiplier: `max` is four times `high`, but `xhigh` is only about 1.8 times it. **`high` is the default and the everyday setting.** `medium` is for drafts and is cheap enough to iterate on freely. `max` is the ceiling, for finished frames and exact character-level text; `xhigh` sits just under it for about half the price and is worth trying first. Go past `high` deliberately, not by habit.\n 4K is worth it only once your references and prompt are already settled: prove the shot at 3K, then re-run the finished prompt at 4K. Iterating at 4K is the most common way to waste credits on this model. Its 4K tier is API-only, so even a paid ChatGPT account cannot render it. Note that its resolution tiers are **not** a price ladder: the pixel classes are token-priced by OpenAI, so the cheapest seat is not the smallest one. Read the prices in the table above rather than assuming.\n **If you have used GPT Image 2 before, the quality names all shifted by one.** What it called `medium` is now called `high`, and what it called `high` is now `max` \u2014 the same pictures at the same prices, renamed. A remembered setting will quietly buy you a cheaper tier than it used to.\r\n **It is the only image model that can give you a transparent background.** Set **Background** to *Transparent* on the prompt bar and you get a real alpha channel \u2014 a cut-out for a logo, sticker or overlay \u2014 rather than a painted-in backdrop. *Auto* is the default and lets the model decide from your prompt; *Opaque* forces a filled background. It costs nothing either way. Slates always saves PNG, which is what carries the transparency, so there is nothing else to set.\r\n **Square and 4:3 frames cost more than 16:9 on this model, and only on this model.** OpenAI charges by image tokens rather than by pixels, and a square frame uses about 1.8 times the tokens of a 16:9 frame the same size, and 4:3 or 3:4 about 1.37 times. The credit prices in the table above are the 16:9 numbers; pick 1:1 or 4:3 and the price on the Generate button goes up to match. 9:16 costs the same as 16:9. Every other image model charges the same whatever the shape.\r\n- **FLUX.2 Max** and **Seedream 5 Lite** are the less content-restricted options. Seedream is flat-priced at every resolution it offers, so there is no reason to pick a lower one. Both auto-route to their edit endpoint when you attach reference images.\r\n\r\n### Video\r\n\r\n**Seedance 2.0 is the default video model.** Reach for it the moment physics, effects, destruction or scale matter, and for hero shots. It takes many reference images, generates native audio at no extra cost, and is the only Seedance seat that reaches 4K (generating 4K video needs Slates Pro; timeline export at 4K does not). It is also the cheaper of the two seats at every resolution they share. A face in a reference image routes it to a different provider, which is why **Seedance 2.0 \u00B7 Face** is its own row in the model picker at its own price.\r\n\r\n- **Seedance 2.5 is a second seat, not an upgrade.** It buys much longer single takes and far more reference images, plus better prompt adherence. What it gives up is 4K, and it costs more than 2.0 at every resolution the two share \u2014 so 2.0 stays the model for 4K, and for the same resolution at a lower price. Because 2.5 runs longer, a long clip on 2.5 can cost more than a shorter, higher-resolution one on 2.0 \u2014 read the Generate button, not the resolution. **Seedance 2.5 Edit** is its clip-editing row: attach a clip, describe the change, and the output length follows the source.\r\n- **Kling** is the cost-effective workhorse and the most flexible family: strong start-frame adherence for identity, layout and text, acting, dialogue, multi-shot (up to 6 cuts), and the widest range of clip lengths. **Kling V3.0 Omni** adds multi-character dialogue in English, Chinese, Japanese, Korean and Spanish. Standard and Pro are the same model at two fidelity and price tiers. **Kling O3 Edit** takes an existing clip and changes what you describe, with subject and style reference images, while the original audio is preserved verbatim. Kling is also the only engine behind the Lip Sync and Motion Control tools.\r\n- **MiniMax H3** is the seat to pick when the SOUND is part of what you are writing. Every other video model treats audio as a switch; H3 takes it as three separate instructions in one prompt \u2014 the lines and action sounds tied to a moment, the ambience running underneath, and a score only the audience hears \u2014 and generates all of it with the picture in a single pass. It is also the only model where you say how much of a reference should survive, including moving one subject's characteristic onto a different subject. It runs 5-15 seconds at 480p, 768p, 2K or 4K, and takes up to nine reference images plus reference video and audio. Two things to watch: **the first five reference images are free and every one after that costs extra**, so attach what the shot needs rather than the maximum; and 2K and 4K are upscales of a 768p render rather than larger generations \u2014 in our own testing the 2K pass showed more artifacting than the 768p original it was built from, at more than twice the price. Generate and judge at 768p; step up only when a delivery spec demands the pixels.\r\n- **MiniMax H3 Max** is the same model post-trained by fal for SPEED, and it is the more expensive seat, not the cheaper one. It is dramatically faster: on the same 5-second 768p prompt it finished in about 5 seconds against about 57 seconds for H3 \u2014 roughly 12x (measured 2026-08-27). It stops at 768p and costs more per second than H3 at the resolution they share. It still animates a start frame and an end frame, so image-to-video works normally; what it does not have is the reference set \u2014 the extra identity, style and environment images plus reference video and audio that base H3 reads. Pick it when a fast turnaround on a text-to-video or start-frame shot is worth paying for; pick H3 for resolution, references, or the same tier at a lower price.\r\n- **LTX-2.5** is the VOLUME seat \u2014 the cheapest native 1080p second in the catalogue, with synchronised audio included free at every resolution, so it is the model to reach for when the job is many takes rather than one hero shot. Two things are unique to it. It makes the LONGEST clips of anything here, up to 20 seconds, and it is the only model that reaches 1440p. It is also the only one with native MULTISHOT: a single generation can carry two to four connected shots that hold the character, lighting and voice across the cuts, which everywhere else means generating separate clips and watching identity drift between them. Its constraints are unusually sharp, though. Durations are EVEN NUMBERS ONLY starting at six \u2014 6, 8, 10, 12, 14, 16, 18, 20, with no 5-second or 7-second clip \u2014 and above 1080p that ceiling drops to 10 seconds. Aspect ratios are 16:9 and 9:16 only. And it takes FRAMES, not references: a start frame and an optional end frame that generates a transition between them, but no identity, style or environment reference images at all, so cross-shot character consistency belongs on MiniMax H3 or Kling. Because sound is generated in the same pass, write the audio into the prompt and anchor every cue to something visible \u2014 anything unanchored gets invented.\r\n- **LTX-2.5 Pro** is the fidelity seat of that pair, and it is NOT simply a better LTX. It renders the picture with more compute on busy frames, but on a narrower envelope than the base row: 720p and 1080p only (no 1440p, no 4K) and 6, 8 or 10 seconds only, for about a third more per second. Reaching for it because the name says Pro costs more AND takes away the reach. Pick it when a specific shot needs the extra fidelity and fits inside 1080p and ten seconds; pick base LTX for length, resolution and volume.\r\n- **Gemini Omni Flash** is the cheap 720p seat with native synced audio included in one pass. **Omni Flash Edit** is the prompt-only clip editor: no reference images, one short instruction plus \"Keep everything else the same.\" Long descriptive prompts destroy it.\r\n- **Veo 3.1** is niche and is never a default. Pick it only when you specifically want Google's audio pass. It has the fewest aspect ratios and reference slots of any video model, fixed durations, and the highest per-clip cost.\r\n\r\nBoth edit models take an existing clip as their canvas, so their output length follows the source clip rather than a duration you choose.\r\n\r\n### Audio\r\n\r\nAudio is a third media type alongside images and video \u2014 generated as its own asset, shown in the gallery's **Audio** tab, and dragged onto an audio track in the timeline. This is separate from the audio some VIDEO models generate *inside* a clip (see AUDIO IN GENERATION below): use a video model when the sound must be locked to what is on screen, and these when you need audio you can move, trim, re-use, or layer.\r\n\r\n**Seed Audio 1.0 is the default.** A room with dialogue *and* clatter *and* ambience is one generation, not three layered ones, and because it is cheap you can run five takes and keep the best. It makes a whole audio SCENE from one plain sentence. **It has no length setting of its own** \u2014 Slates writes your chosen duration into the prompt, and that is what you are charged. You describe the voice in words; there is no voice list to pick from. Set **Languages** to Mixed if one scene needs more than one language (it costs the same).\r\n\r\n**Sound Effects** makes one effect, or a seamless loop. It is the only surface with an exact duration, so an effect can land on a specific frame. Describe the physical cause (\"heavy oak door slams shut in a stone hallway\"), not the label (\"door sound\"). **Loop** makes it seamless for beds; **Wording** controls how literally your description is followed. Seed Audio is actually the better tool for *long* ambience beds, so the two are not redundant in the direction you would expect.\r\n\r\nKling's `SFX:` / `Ambient noise:` prompt syntax belongs to video prompts and makes Seed Audio results *worse* \u2014 write plain sentences there instead.\r\n\r\n**Inworld TTS-2 is the voice seat** \u2014 type the words, pick a voice, press Generate. In the prompt box's Audio lane pick **Voice** in the model picker; the prompt is the exact text that gets spoken (nothing is added or rewritten \u2014 open \"See what gets sent\" to confirm), and the **Voice** control on the bar opens the voice picker: **Presets** (ready-made voices with gender, accent and age filters \u2014 every one plays the same audition line, so you compare voices rather than scripts), **Clips** (any character's voice, or any audio clip in the project, cloned for the take), or **Describe** (a voice in words). The character counter beside the bar is the bill: the generated audio table above gives the text cap and billing buckets, and the Generate button shows the price. Direction goes in square brackets (`[whispering] \u2026`) \u2014 anything in parentheses is read aloud. Cloning a real person's voice needs their permission. Right-click any audio clip in the Audio tab \u2192 **Use as voice** lands you on the Voice lane with that clip as the voice. Studio Agent and the MCP/CLI do the same through `slates_generate_audio` (a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId`, or a `voiceDescription`).\r\n\r\n**There is no music generation.** For a song, use an external tool and import the audio (see PROJECTS \u2192 Supported File Formats). For spoken lines inside a scene, let Seed Audio perform them, put them in the video prompt on a model with native audio (Seedance, Kling Omni, Omni Flash, Veo), or use Kling Lip-Sync's text-to-speech against a shot.\r\n\r\n### Tools\r\n\r\n**Tools** is not a model family. It is two real Kling endpoints that take a clip or a still as their subject: **Kling Lip Sync** and **Kling Motion Control**. Both bill in 5-second blocks and are described under GENERATION MODES below.\r\n\r\n**How pricing works:** every generation is priced in Slates Credits, and the exact cost is shown on the Generate button before you commit. Bigger credit packs give more credits per dollar; Slates Pro locks in the best pack rate on every purchase, forever.\r\n\r\n---\r\n\r\n## GENERATION MODES\r\n\r\n### Create Image\r\nPrompt \u2192 select image model \u2192 set aspect ratio + resolution \u2192 generate. Batch grids available for quick iteration.\r\n\r\n### Text-to-Video\r\nPrompt \u2192 select video model \u2192 set duration + aspect ratio + resolution \u2192 generate. Output: MP4.\r\n\r\n### Image-to-Video (I2V)\r\nAttach start image + prompt \u2192 select model \u2192 generate video from that image. Optional: attach end image (Veo) for guided transitions.\r\n\r\n### Ingredients / References\r\nUse @character_name, @environment_name, or #style_name in prompt to attach reference images for visual consistency. Kling: up to 4 total references. Veo: up to 3. Nano Banana 2: up to 14. The @mentions auto-complete from your project's characters/environments/styles.\r\n\r\n### Lip Sync\r\n**Kling only.** Pick **Kling Lip Sync** under the Tools family in the model picker. Source: video or still image.\r\n\r\n- **Audio source** \u2014 Text to speech (type the line; six English/UK voices plus a storyteller, with a speed control) OR Upload audio (bring your own recording, max 5MB).\r\n- **Avatar tier** \u2014 only appears for a still-image source: Avatar v2 Standard (the value tier) or Avatar v2 Pro (higher fidelity, higher rate).\r\n- 5s output blocks.\r\n\r\nWorks very well with human-like characters. Less reliable with animals or non-human characters.\r\n\r\n### Motion Transfer\r\n**Kling only.** Pick **Kling Motion Control** under the Tools family. Target: still image (your character). Source: reference video (the motion).\r\n\r\n- **Engine** \u2014 Kling MC Standard (value tier) or Kling MC Pro (higher fidelity, higher rate).\r\n- **Orientation** \u2014 *Match video* copies skeleton and depth from the clip (best for dancing, walking, full-body action; driving clips up to 30s). *Match image* keeps your character's pose and angle and uses the video only as motion hints (best for close-ups; up to 10s).\r\n- 5s output.\r\n\r\n> **Note:** these two tools used to offer a second \"Seedance 2.0\" engine. It was not a separate engine \u2014 picking it made Slates write a sentence into your prompt that you never saw, which is no longer allowed anywhere in the app (see \"What gets sent\" below). Both tools are now Kling endpoints only.\r\n\r\n### Edit Image\r\nRight-click any image asset \u2192 open viewer \u2192 switch to Edit Mode. Enter an edit prompt describing the changes you want. Select edit model: Nano Banana 2 (supports up to 14 reference images), FLUX.2 Max, or Seedream 5 Lite. Choose resolution and aspect ratio. The result saves as a new asset with the original preserved. Useful for refining generated images without starting from scratch.\r\n\r\n### Edit Video (Kling O3 Edit / Omni Flash Edit)\r\nRight-click any video clip (gallery or timeline) \u2192 \"Edit with AI\". The clip attaches to the prompt box as the source; describe the CHANGE, not the whole scene (\"replace the man with @marcus\", \"make it a rainy night, keep everything else\"). Two engines in the model picker:\r\n- **Kling O3 Edit (default):** attach subject images (role: Subject) to swap someone in, or style images (role: Style) for a look \u2014 max 4 combined refs. Clips 3-15s. Original audio preserved.\r\n- **Omni Flash Edit (cheapest):** prompt only \u2014 no reference images; keep instructions simple and add \"Keep everything else the same.\" Clips 3-10s, 720p output.\r\n\r\nOutput length follows the source clip; the credit cost (clip seconds, rounded up, at the per-second rate) shows on the Generate button. The edited clip saves as a NEW asset linked to the original \u2014 chain edits freely. Trim longer clips on the timeline first.\r\n\r\n### Multi-Shot (Kling V3.0/Omni)\r\nEnable multi-shot toggle \u2192 multiple scene prompts in one generation, each with different framing. 6-axis camera controls per shot. Results can be hit-or-miss, but worth trying for quick multi-cut sequences. For more reliable results, most users prefer generating multiple short 5s clips separately using Kling V3.0 Omni in ingredients mode and assembling them on the timeline.\r\n\r\n---\r\n\r\n### Generate Audio\r\n\r\nSwitch the prompt box's lane pill from Image/Video to **Audio**, pick a surface, and generate. The prompt box offers three: **Seed Audio 1.0**, **Voice** (Inworld TTS-2) and **Sound Effects**. The result lands in the gallery's Audio tab as its own asset with a waveform and an inline player, and can be dragged onto an audio track in the timeline.\r\n\r\n- **Scene (Seed Audio 1.0)** \u2014 one plain sentence describing the moment. Set **Length**; Slates writes it into the prompt for you and that is exactly what you're billed for (open \"See what gets sent\" under the prompt box to read the appended text). **Say the crowd/room size out loud** \u2014 \"applause\" returns a full auditorium when you meant three people at an open mic. Ask for a few seconds more than the clip needs so the edit has fade handles. Describe the voice you want in the sentence itself (\"a weary dock foreman in his fifties, gravel in his voice\") \u2014 there is no voice picker.\r\n- **Voice (Inworld TTS-2)** \u2014 the prompt is the words to be spoken, verbatim. Pick the voice with the **Voice** control on the bar (presets you can play first, any clip in the project, a character's voice, or a description); the character counter is the bill. See MODELS \u2192 Inworld TTS-2 for direction tags and the cloning rules.\r\n- **Sound Effect** \u2014 describe the physical cause and set the length to roughly the event (\u22481s for an impact, 2\u20134s for a whoosh, 8\u201322s + **Loop** for a bed). **Wording** sets how literally the description is followed: Interpretive, Balanced (default), or Literal.\n\n#### Use your own voice recording\n\nIn the bottom prompt box, choose **Audio**, then **Inworld TTS-2** in the model picker. Open **Voice \u2192 Clips \u2192 Import voice clip** and select your recording. The import adds an audio asset to this project without generating anything. Click its play button to audition it, then click the recording's name to choose it. Type the words you want spoken in the prompt box and press **Generate**, which shows the price. The new take appears in **Gallery \u2192 Audio**.\n\nUse a clean recording of one speaker whose voice you have permission to use. Slates clones the recording for each take; there is no separate training wizard or persistent vendor voice to manage. **Presets** lets you audition ready-made voices; **Describe** lets you write a voice description and choose **Use this description**. Choosing in the prompt box sets up the next take; only Generate spends credits.\n\nTo attach your recording or a generated take to a character, open **Gallery \u2192 Characters**, then **Add voice** (or **Change voice**) on that character's card. Choose **Clips** and click the clip's name. Attaching an existing clip is free. The card displays its waveform and player. That character's voice is also listed under **Voice \u2192 Clips \u2192 Characters** in the prompt box. Selecting a preset or description from a character card generates and attaches a take; read the cost shown in that picker before choosing.\n\nFor **Inworld TTS-2**, choose the character under **Voice \u2192 Clips \u2192 Characters**. Typing `@name` does not choose a TTS voice: the prompt is spoken verbatim. Automatic voice attachment from a character mention belongs to **Seed Audio** only.\n\r\n---\r\n\r\n## AUDIO IN GENERATION\r\n\r\nThis section is about audio generated **inside a video clip**. For audio as its own asset, see Generate Audio above.\r\n\r\n**Veo 3.1 native audio:** Generates audio WITH video. Prompt syntax: `\"Hello!\"` for dialogue, `SFX: [sound]` for effects, `Ambient noise: [description]` for ambience. Max 10s dialogue. Add `(no subtitles)` to suppress text overlays.\r\n\r\n**Kling V3.0 Omni dialogue:** Multi-character dialogue with distinct voices. Languages: EN, ZH, JA, KO, ES. `Background music: [description]` for music. Max 10s dialogue.\r\n\r\n**Kling V3.0 sound co-generation:** Synchronized sound effects generated with video.\r\n\r\n\u26A0\uFE0F **This prompt syntax is video-only.** `SFX:`, `Ambient noise:` and `Background music:` are Kling/Veo conventions \u2014 the audio models above have no parser for them and will treat them as words in the scene.\r\n\r\n---\r\n\r\n## PROMPT SYSTEM\r\n\r\n### Unified Create Surface (roles + model-on-button)\r\nThe old mode tabs (text-to-video / frames-to-video / ingredients / create-image) are ONE \"Create\" surface. An **Image | Video | Audio pill** on the prompt bar switches your output lane \u2014 it remembers and restores the last model you used in each lane (pick Seedance once and the Video lane stays Seedance until you change it). Only the active lane shows its name; the other two are icons. Every attachment in the reference tray carries a tappable ROLE badge \u2014 Reference / First frame / Last frame / Subject / Style \u2014 you say what each attachment is; nothing is inferred. Your model choice sticks across generations and workflow actions (\"use as first frame\" keeps your chosen video model). Attaching a video via \"Edit with AI\" flips the surface into Edit Video mode. Lip Sync and Motion Transfer live under the **Tools** family in the model picker.\r\n\r\n### Floating Prompt Box \u2014 the bar holds everything\r\nPersistent across all pages. **There is no settings panel and no gear button.** Everything sits on one bottom bar, left to right:\r\n\r\n1. **Media toggle** \u2014 Image / Video / Audio.\r\n2. **Model picker** \u2014 a searchable menu plus a detached submenu. The main menu lists model families with a vendor glyph tile each; picking one opens that family's models beside it, every row carrying capability chips (resolution, clip length, audio, references) and its per-unit rate. Type to search across every model. The submenu is anchored to the row you opened it from, so it never travels. The trigger on the bar shows the model name and nothing else \u2014 no chevron, no resolution appended. Everything listed is a real model or endpoint.\r\n3. **Parameter controls** \u2014 one per setting the chosen model actually has (resolution, aspect, duration, length, quality, count, grid, face-in-reference, audio, loop, and so on). The trigger shows the current value; the explanation lives *inside* the menu as a subtitle under each option, along with what that option costs. A setting with only one possible value still shows, muted and non-interactive, so the row never changes shape.\r\n4. **Sliders for ranges.** A setting with a long list of steps (video duration, audio length) opens a ruler instead of a many-row menu. The handle moves between the values the model actually declares, so it cannot land on one the model will not accept, and the price for the selected value is shown on the ruler.\r\n5. **`More \u25BE`** \u2014 if the model has more parameters than fit the current window width, the extras are collected into a generated `More` dropdown automatically. Widen the window (or close the Studio Agent panel) and they move back onto the bar.\r\n6. **Generate** \u2014 reads `Generate \u00B7 <cost>`. **Cost only** \u2014 the model name is not repeated on the button (it's in the model picker) and there is no send arrow. A badge on the left of the button counts generations currently running.\r\n\r\nFor text-to-speech, the character counter is always visible because text length determines the price. On other surfaces it appears near the right of the bar after roughly three quarters of the model's prompt limit. It turns red over the limit.\r\n\r\n**Collapsing:** the chevron at the top-right of the box collapses it to a single arrow \u2014 nothing else. Click the arrow to bring it back.\r\n\r\nBelow the bar the queue shows pending/active generations with cost and progress.\r\n\r\n### What gets sent (prompt transparency)\r\nUnder the prompt box is a **\"See what gets sent\"** disclosure. Open it and you see the exact text that will be transmitted, produced by the same code that builds the request \u2014 so it can never disagree with what is actually sent. It shows:\r\n\r\n- **Reference numbering** \u2014 `@sarah` becomes `Sarah (image 1)` so the model knows which attached image is which.\r\n- **Key lines for attachments you did NOT mention** \u2014 one short neutral sentence per unmentioned attachment. Mention every reference in your own words and these never generate.\r\n- **The trailing style clause** when a `#style` is attached.\r\n- **The Seed Audio duration append** \u2014 the `\u2026 N seconds` Slates adds to the end of the prompt, which is also what you are billed for.\r\n- **Grid wrapping** when 2\u00D72 or 3\u00D73 is on.\r\n\r\nThe row stays hidden when the composed prompt is identical to what you typed, so it only appears when there is something to show.\r\n\r\n**Unresolved `#tags` and `@mentions` are named, not silently dropped.** If you type `#noir` and there is no saved style called \"noir\", the tag is removed from the text sent to the model (a raw tag confuses every model) \u2014 but the disclosure turns red and says so by name: *\"#noir matches nothing saved \u2014 removed from what gets sent.\"* Save the style, or reword it, and the warning clears.\r\n\r\n**Nothing is ever added that you cannot read here.** No setting in the app injects prompt text; a setting changes *how* a request is made, never *what* you asked for.\r\n\r\n### Prompting guide (on the web)\r\nPer-model prompting guidance lives at <https://slates.video/docs/prompting>, linked from the bottom of Settings. It covers every model Slates offers \u2014 Video, Image, Audio \u2014 with what that model reads, what it ignores, and its gotchas, all on one page so you can compare them. Markdown copy for pasting into an LLM: <https://slates.video/docs/prompting.md>.\r\n\r\nIt is generated from the same source the Slates CLI, the MCP server and Studio Agent are built on, so the guide and the app cannot disagree. It is documentation rather than a control, which is why it is a page on the web and not a panel in the app: it has room to be read, a URL you can send someone, and it is always current rather than frozen at the version you installed.\r\n\r\n### @Mentions\r\nType `@` \u2192 auto-complete shows project characters and environments. Type `#` \u2192 shows styles. Selecting inserts the reference image(s). At send time a mention is rewritten to a numbered citation (`@sarah` \u2192 `Sarah (image 1)`) so the model can tell your attachments apart \u2014 you can read the result in \"See what gets sent\". Nothing else about your wording is rewritten. Prompting works the same as any other AI tool; no special syntax beyond @mentions.\r\n\r\n### Writing the shot list with Studio Agent\r\nThere is no \"Enhance\" button and no \"Generate prompts\" button. **Studio Agent does this work**, because it reads the same shot list you do \u2014 every scene, every beat in order, with its references, its model and its price \u2014 and because you can steer it:\r\n\r\n> \"Write the beats for my current storyboard.\"\r\n> \"Now redo scene 3 handheld, and match its energy to scene 2.\"\r\n> \"SHOT-A4 runs long \u2014 split it after 'and then'.\"\r\n\r\nEvery field it writes is editable by hand, in place, in the storyboard's Document view. Open Studio Agent with **Ctrl+.**\r\n\r\n### Debug Panel (advanced)\r\nA developer panel showing the exact request body, with the ability to override the composed prompt before sending. **There is no toggle button for it on the prompt bar in any build** \u2014 open it with **Ctrl+Shift+D**. For ordinary use, \"See what gets sent\" above is the supported way to inspect a prompt.\r\n\r\n---\r\n\r\n## PROJECTS\r\n\r\n### Structure\r\nEach project = folder on your disk. Subdirectories: images/, videos/, audio/, references/, exports/. Location configurable in Settings \u2192 Projects Directory.\r\n\r\n### Assets\r\nEvery generated or imported file is an asset (image, video, audio). Metadata tracked: prompt, model, settings, cost, dimensions, timestamps. Videos track source image via source_asset_id -- you can see all videos generated from any image.\r\n\r\n### Supported File Formats\r\n- **Images:** PNG, JPEG, WEBP, GIF. Note: HEIC/HEIF (iPhone photos) NOT supported -- convert to JPEG/PNG first.\r\n- **Video:** MP4, MOV, WEBM, AVI, MKV.\r\n- **Audio:** MP3, WAV, OGG, M4A, AAC.\r\n- **Clipboard paste:** Any image format the OS clipboard provides (PNG, JPEG, WEBP, GIF). Pasting works both in the gallery and directly into the prompt box; either way the image becomes a real gallery asset in the folder you're working in (tagged \"Imported\") AND, when pasted into the prompt box, attaches as a reference. Anything generated from it links back to it as a source.\r\n- **Drag and drop:** Any file the browser recognizes as image/* or video/*.\r\n\r\n### Operations\r\nCreate/rename/delete projects. Import external files via drag-and-drop or file picker. Paste images from clipboard. Extract still frames from videos. Relocate project to different disk/folder (all paths auto-update). Cleanup orphaned assets.\r\n\r\n### Moving and copying assets between projects\r\nAssets (images, clips) can be sent to another project three ways: the selection band on the Images/Videos tabs, the right-click menu on any card, or by dragging cards and dropping on a project in the drop palette.\r\n\r\n- **Move** relocates the media files on disk into the destination project's folder. The asset leaves whatever gallery folder it was in and is issued a fresh badge code in the destination.\r\n- **Copy** duplicates it \u2014 new files, new thumbnails, new badge code \u2014 and changes nothing in the source project.\r\n\r\n**Why a move can be refused:** an image another project still builds with (a character/environment/style identity image, or a storyboard frame) cannot leave, because the entity left behind would point at a file it no longer owns. When that happens the dialog lists what's blocking and offers the fix: bring the whole character/environment/style across with all of its images, or copy instead. A storyboard frame is only ever offered a copy \u2014 moving its image out would empty the shot.\r\n\r\nRight-clicking a card that is part of a multi-selection acts on the whole selection (\"Move 5 to Project\u2026\"). Right-clicking a card outside the selection acts on that card alone.\r\n\r\n---\r\n\r\n## SHOTS \u2014 THE STORYBOARD IS THE SHOT LIST\r\n\r\nA **Shot** is the prompt bar, saved: the prompt, every reference with the job it carries, the model, every setting \u2014 and now the beat itself: who speaks, what they say, how it is said, what happens, the prop, the framing and the camera. It lives in the **storyboard**, which is the one place Shots are listed. It is never required: the prompt bar works exactly as it always has for anyone who never touches one.\r\n\r\n### Why it exists\r\nA generation's full recipe was already stored, but only once you had paid for it. A Shot can be written **before anything is generated**, so a whole piece can be planned, read, timed, priced and corrected while it is still free. That is the point of the thing: look at the entire ad or short film \u2014 every cheap asset lined up in the actual flow \u2014 before the videos exist.\r\n\r\n### What a Shot holds\r\nRaw prompt (@mentions intact); the model; aspect ratio / duration / resolution / negative prompt / sound and the rest of the bar's settings; every attachment with its ROLE (plain reference, subject, style, reference video, reference audio, first frame, last frame); and the script layer \u2014 `speaker`, `line`, `delivery`, `action`, `prop`, `shotSize`, `camera`, and a `continues` flag for one sentence running across two cuts. Characters, environments and styles are stored as the ENTITY, not a copied picture, so updating a character updates every Shot that names it.\r\n\r\n**The script fields are for reading and counting. Only the prompt is sent to a model.** Dialogue you want a model to perform still goes in the prompt, verbatim, with its delivery \u2014 writing it in `line` makes it readable and lets Slates check whether it fits the cut, not spoken.\r\n\r\n### Every Shot has an address\r\n`SHOT-A1`, `SHOT-A2` \u2026 per project, never reused \u2014 the same idea as the `IMG-A12` badge on a gallery card. Say it to Claude and you are both pointing at the same row. It is for **this session**, not for retrieval later: there is no shot search and no shot library, because a Shot is workspace state \u2014 alive while you build the piece, worthless once it ships.\r\n\r\n### Making one\r\n- **Save as Shot** on any generated image, clip or track's right-click menu restores that generation and keeps it \u2014 and that generation becomes the Shot's first take, so the row opens showing the result you kept it for. **Reuse Prompt** on those same menus does the restore WITHOUT saving anything.\r\n- Claude can write Shots directly (`slates_create_shot`), including for shots whose image does not exist yet \u2014 and can re-chop them with `slates_split_shot` / `slates_merge_shots`.\r\n- **It files itself.** A saved Shot lands in the scene you have open, else the last scene of the storyboard you were most recently working in; if the project has no storyboard, one appears named after the project. Nothing you save is ever somewhere you have to go and find.\r\n\r\n### There is no save button\r\nSelecting a Shot row **binds** the prompt bar to it. Edits write straight back to that row; `Clear` unbinds and returns the bar to free composing. There is no undo, and none is needed: the generations underneath a row are the permanent record of what actually fired, and the row itself is the working copy.\r\n\r\n### Two ways to look at it\r\nThe storyboard has one toggle and two jobs.\r\n\r\n- **Board** \u2014 arrange. One picture per Shot, dragged into the order you want. Drag one and the whole beat moves with it: the line, the references, the model, the settings, the takes. There is nothing else to drag, so there is never a question of what followed what.\r\n- **Document** \u2014 write. One continuous page: the script, the references beside the words that cite them, the prompts underneath. This is where you read the piece before paying for it.\r\n\r\nInside Document, choose **Script** to edit the words with scene headings and speakers. Delivery notes, model and pricing details, and warnings are hidden. **Shots** shows all layers. Open **Custom** to toggle Scene, Action, Character, Delivery, Dialogue, Shot, References, Prompt, Takes, and Warnings independently. Warnings covers missing models or references, unresolved mentions, model changes, and lines too long for their cut. Hiding a layer changes only what you see; it preserves your text and generation checks. There is a text-size slider and an independent toggle for shot numbers in the margin.\n\nDelivery is an optional performance note, not a required label for every line. TTS sends the authored prompt, or the dialogue when the prompt is empty, verbatim; separate Delivery notes are not added. For speech cues, use the selected TTS model's prompting guide and place supported tags in that spoken text. Script view does not strip inline speech cues from your words.\n\r\n### The header tells you what you are about to make\r\n`5 generations \u00B7 7 cuts \u00B7 54s \u00B7 84 credits`, and beneath it a variety strip like `6/7 wide \u00B7 5 push \u00B7 3 cuts in the loft`.\r\n\r\n**Two counts, because they measure different things.** Rhythm is counted in **cuts**; money is counted in **generations**. A multi-shot generation is several cuts inside one paid call, so mixing them would be wrong. A cut with no model chosen has no duration and shows as `\u2014` rather than `0s` \u2014 a runtime that invented seconds would be a lie about the one number this view exists to give.\r\n\r\n### Splitting and merging \u2014 the chop\r\nPut the caret mid-line and press Enter: the row becomes two, the second inheriting the model, settings and references, and marked as continuing the first if the split lands mid-sentence. Select two adjacent rows and press **Merge**: they become one, references combined, durations summed. **The price and the runtime move as you do it** \u2014 which is the whole reason to make the decision here rather than in a document somewhere else.\r\n\r\nSplit is also the move behind a voiceover that keeps talking while the picture hard-cuts to a new world: split at a word boundary and both rows carry one sentence, each with its own visuals.\n\nThe **Dialogue continues from previous shot** toggle is a planning note that the sentence spans a cut. It does not merge shots, join generated audio, or change generation settings. Its pressed state shows whether the note is set; toggling it does not move the script. **Merge these two shots** is the separate control between adjacent shots that actually combines them. In Document \u2192 Custom, **Scene** toggles scene headings and their controls; the dialogue remains visible even if a hidden heading belonged to a collapsed scene.\n\r\n### Variety, counted and never judged\r\nSlates counts what is in front of it \u2014 shot sizes, camera moves, cast, locations, durations, and any of them repeating three or more times in a row \u2014 and shows the counts. **It never changes anything, never suggests anything and never blocks.** `shotSize` and `camera` are free text: write `long-lens CU, other head blurred` if that is the shot. Anything unrecognised counts as \"other\", which is a fine answer.\r\n\r\nIf a spoken line cannot be read in its cut at any plausible pace, the row says so \u2014 and says it only when the line is genuinely impossible, never when it is merely long.\r\n\r\n### The animatic\r\nPress play on any beat and the storyboard plays as a rough cut: each picture held for **its own cut's duration**, with the line underneath. That tells you the rhythm of the finished piece before a single video exists. A multi-shot generation holds one picture across its internal cuts, and says so. A cut with no duration is held for 3 seconds and marked \u2014 the header leaves it out of the runtime for the same reason it is marked here.\r\n\r\n### Things that stay honest rather than being hidden\r\n- **Deleted references.** If an asset or character a Shot points at is gone, the Shot still loads and the row says how many items were left out of what gets sent \u2014 and editing the row does not quietly drop them.\r\n- **A swapped model.** Changing a Shot's model never rewrites your words \u2014 video models genuinely take different prompt grammars, so the row tells you which model the prompt was written for and leaves the sentence alone. The settings line shows what will actually be sent after the swap, and that is the value it prices.\r\n- **Image Shots carry no role badges.** Image generation sends every reference in one undifferentiated list, so an image Shot can remember that a picture is a style reference but cannot tell the model.\r\n- **The thumbnail is never a question.** Slates picks it \u2014 the first frame, else the first reference, else the newest take \u2014 and you can override it from any reference in the gutter.\r\n\r\n### Firing several\r\nSelect Shots and press **Generate all**: one total, the largest single Shot stated separately, one approval. They run **one at a time**. If one of the selected Shots has been deleted, the whole batch is refused and names it \u2014 nothing fires and nothing is billed. If a generation fails mid-run, the rest still fire, the failure is reported per Shot, and **nothing is retried automatically**.\r\n\r\n### Deleting\r\nDeleting a storyboard tells you how many Shots are attached before it does anything, and deletes them with it. It never moves them somewhere else without asking. A Shot that also lives in another storyboard survives.\r\n\r\n---\r\n\r\n## STORYBOARDING\r\n\r\n### Hierarchy\r\nStoryboard \u2192 Scenes \u2192 **Shots**. A scene is an ordered list of Shots, and a Shot is one beat: its picture, its references and their roles, its model and settings, its prompt, its words, and the generations it has produced. See **SHOTS** above \u2014 that section is the storyboard.\r\n\r\n### What happened to frame types\r\nThere used to be a \"frame type\" on each picture \u2014 first / last / ingredient \u2014 plus a separate motion-prompt box. Both were a second, weaker way of saying what a Shot already says: **a reference's role lives on the Shot** (first frame, last frame, subject, style, plain reference), and the motion prompt was just the Shot's prompt under another name. Existing storyboards were converted automatically and nothing was lost. Pick a picture's job on the Shot's reference rail; write the motion in the Shot's prompt.\r\n\r\n### Grid Exploration\r\n2x2 grid: 4 prompt variations for quick iteration. 3x3 grid: 9 variations for deeper exploration. Select individual cells \u2192 extract to full-resolution images. Tip: 2x2 is usually sufficient and produces better quality. 3x3 can occasionally get proportions slightly wrong when upscaling cells because it faithfully reproduces the lower-resolution proportions. Grid exploration runs on Nano Banana 2 only; no other image model offers it.\r\n\r\n### Storyboard \u2192 Video\r\nSelect frames \u2192 generate video for each \u2192 clips auto-insert into timeline in order with source tracking maintained.\r\n\r\n### The animatic\r\nPlay the storyboard as a rough cut. Each beat is held for **its own duration**, with its line underneath, so what you are watching runs at the finished piece's real length. Space = play/pause. Arrow keys = navigate. Escape = exit.\r\n\r\n### Paste a script\r\nPaste a script into the storyboard and it becomes one row per paragraph \u2014 ALL-CAPS cues become speakers, parentheticals become delivery. It is a plain parse, not a model: nothing is invented, nothing is sent anywhere, and prose that is not screenplay-formatted lands as one row per paragraph for you (or Claude) to chop.\r\n\r\n---\r\n\r\n## VIDEO EDITOR (TIMELINE)\r\n\r\n### Tracks\r\nMulti-track: video tracks + audio tracks stacked vertically. Clips independent per track. Add or remove tracks freely -- layer a music bed, a voiceover, and effects on separate audio tracks. Video assets go on video tracks, audio assets on audio tracks. Overlapping video clips resolve top-track-wins.\r\n\r\n### Audio Mixing\r\nEach track has a volume fader, and the timeline has a master output fader for the final mix. Both range from silent to +12 dB of boost, and both apply to preview playback AND the exported MP4 -- what you hear is what you render. Muting a video track silences its embedded audio but still shows the picture. Use the master fader to prevent clipping when stacking loud tracks.\r\n\r\n### Timeline Settings\r\nResolution and frame rate (24/30/60) are auto-managed: the first video clip sets both, and a later higher-resolution clip raises the canvas. All clips are conformed to the timeline frame rate on export. Changing the frame rate after clips are placed retimes them.\r\n\r\n### Clip Properties\r\nSource asset, in/out points (frame-level precision), duration, scale (fit/fill/custom %), position (X/Y offset), opacity (0-100%).\r\n\r\n### Tools\r\n- **Select (V):** Click/drag clips, view/edit properties\r\n- **Razor (C):** Split clip at playhead into two clips\r\n- **Slip (S):** Adjust clip in/out points without moving its position\r\n- **Snap toggle:** Snap to playhead/clip boundaries\r\n\r\n### Markers\r\nColor-coded timeline markers (6+ colors) with optional labels. Use for scene breaks, cue points, notes.\r\n\r\n### Playback & Navigation\r\nSpace = play/pause. Left/Right arrows = frame-by-frame. Up/Down = +-1 second. Page Up/Down = jump by screen width. Home/End = start/end of timeline.\r\n\r\n### Zoom\r\nCtrl+Plus = zoom in (finer precision). Ctrl+Minus = zoom out (see more timeline).\r\n\r\n### Undo/Redo\r\n50-step history. Ctrl+Z = undo. Ctrl+Shift+Z = redo.\r\n\r\n---\r\n\r\n## EXPORT\r\n\r\n### Video Export (FFmpeg)\r\nExport timeline \u2192 MP4 (H.264). Configure: resolution, frame rate, bitrate, output location. All visible tracks rendered, muted tracks excluded, clip in/out points respected. FFmpeg is bundled -- no separate install needed.\r\n\r\n### DaVinci Resolve XML Export\r\nGenerates XML project file containing: clip references (paths to source videos), timeline structure (tracks, clips), clip properties (scale, position, opacity, in/out points), timeline markers.\r\n\r\n**Importing into DaVinci Resolve:** File \u2192 Import \u2192 Timeline. DaVinci reads the XML and reconstructs your timeline with all clips, properties, and markers intact. From there you can color grade and export your final master.\r\n\r\nExports saved to project's exports/ directory with timestamped filenames.\r\n\r\n---\r\n\r\n## CHARACTERS, ENVIRONMENTS & STYLES\r\n\r\n### Characters\r\nCreate character with name + description. Generate character sheet (license required): AI generates a turnaround with multiple angles for consistency. Generate expression sheet: same character with different facial expressions. Use `@character_name` in any prompt to attach reference images.\r\n\r\n**Tips for consistency:** Experiment with character sheet generation using both the existing project style and photorealistic style. Sometimes a single well-chosen image works better than a full sheet -- especially if the character is already in the same style, lighting, and clothing as your project. You can manually assign any image as a character reference instead of generating a sheet.\r\n\r\n**Voice.** A character can carry one voice clip, the same way it carries one identity image \u2014 a shortcut for reusing a voice, never a requirement for speaking in one. **Add voice** / **Change voice** on the card opens the same voice picker the prompt box uses (a menu off the button, not a pop-up; the other cards stay on screen): a clip from the project attaches as it is; a preset or a described voice renders the character speaking a fixed audition line on Inworld TTS-2 and attaches that clip, at the credit cost the picker states first. Right-click the voice card \u2192 **Remove voice** detaches it without deleting the clip. The character's voice then shows under **Clips** in the Voice lane's picker, and mentioning the character (`@name`) in a Seed Audio prompt attaches the clip as a reference, so the scene casts that voice.\r\n\r\n### Environments\r\nCreate environment with name + description. Generate environment grid (license required, 3x3): 9 variations. Extract individual cells to full-resolution images. Use `@environment_name` in prompts.\r\n\r\n**Tip:** Like characters, sometimes a single strong environment image gives better consistency than a grid of 9. Experiment with both approaches.\r\n\r\n### Styles\r\nCreate style with name + description + upload reference image. Use `#style_name` in prompts. Key visual auto-attachment option for consistent look across all frames.\r\n\r\n---\r\n\r\n## SETTINGS\r\n\r\n### Generation\r\nEvery generation runs on Slates Credits \u2014 there are no API keys to configure. The Generate button shows the exact credit cost before each generation, and failed generations refund immediately.\r\n\r\n### Other Settings\r\n- **Projects Directory:** Where project folders live on disk. Changeable anytime.\r\n- **Default Model:** Pre-selected model for new generations. Override per-generation.\r\n- **Default Quality/Resolution:** Pre-selected resolution. Override per-generation.\r\n- **Grid Size:** Default 2x2 or 3x3 for grid exploration.\r\n- **Auto Naming:** Automatically name generated assets.\r\n- **Prompting guide:** A link at the bottom of Settings to <https://slates.video/docs/prompting> \u2014 per-model prompting guidance for every model (see PROMPT SYSTEM above).\r\n\r\n---\r\n\r\n## ACCOUNT & BILLING\r\n\r\n### Login\r\nEmail-only, no password. Enter email \u2192 receive magic link \u2192 click to log in. First login creates account automatically. Session persists across restarts.\r\n\r\n### License\r\nUnlocks: character sheet generation and environment grid generation. Includes 12 months of updates (Slates Pro includes lifetime updates). Major upgrades discounted after.\r\n\r\n### Credits\r\n\r\n<!-- BEGIN:GENERATED credits -->\nCredits are what every generation is paid with. They are pay-as-you-go, they never expire, and the exact cost of a generation is shown on the Generate button before you commit.\n\n- A **Slates Standard** license ($149 one time) starts you with **1,000 credits**.\n- **Slates Pro** ($297 one time, or $97 to upgrade later) starts you with **3,000 credits**.\n\n| Pack | Credits (Standard) | Credits per dollar | Versus the smallest pack |\n|------|--------------------|--------------------|--------------------------|\n| $10 | 250 | 25.0 | standard rate |\n| $25 | 650 | 26.0 | +4% more credits |\n| $50 | 1,375 | 27.5 | +10% more credits |\n| $100 | 3,000 | 30.0 | +20% more credits |\n| $250 | 8,000 | 32.0 | +28% more credits |\n| $500 | 17,000 | 34.0 | +36% more credits |\n| $1,000 | 35,000 | 35.0 | +40% more credits |\n\nPacks up to $500 are open to everyone; the $1,000 pack is offered inside the app to licensed accounts. Slates Pro receives more credits than the Standard column above on every pack, for life.\n<!-- END:GENERATED credits -->\r\n\r\nCredits NEVER expire, there is no monthly reset, and failed generations refund immediately. You can also turn on auto-topup so your balance refills when it runs low.\r\n\r\n### Standard vs Pro\r\n- **Standard:** the app, every AI model, and pay-as-you-go credits that never expire, plus 12 months of updates.\r\n- **Slates Pro:** everything in Standard, plus our lowest credit rate on every pack, forever (buy the smallest pack and pay the largest pack's rate), **4K video generation**, a priority generation queue, early access to every new model on release day, and lifetime updates. The more you top up, the more the better rate adds up.\r\n\r\nEvery AI model is available on both tiers. The only capability gated to Pro is generating 4K video; 4K images are open to everyone, and exporting your timeline at 4K is available on every tier.\r\n\r\n### 30-Day Guarantee\r\nFull refund within 30 days, no questions asked.\r\n\r\n---\r\n\r\n## KEYBOARD SHORTCUTS\r\n\r\n| Key | Action |\r\n|-----|--------|\r\n| Space | Play/pause |\r\n| V | Select tool |\r\n| C | Razor tool |\r\n| S | Slip tool / snap toggle |\r\n| M | Add marker |\r\n| Left/Right | Frame-by-frame |\r\n| Up/Down | Seek +-1 second |\r\n| Ctrl+Z | Undo |\r\n| Ctrl+Shift+Z | Redo |\r\n| Ctrl+Plus/Minus | Zoom timeline |\r\n| Delete/Backspace | Delete selected clip |\r\n| Escape | Close modal/viewer/slideshow |\r\n| Ctrl+Enter | Submit generation |\r\n| Home/End | Jump to timeline start/end |\r\n| Page Up/Down | Jump by screen width |\r\n\r\n---\r\n\r\n## OFFLINE USAGE\r\n\r\nThe app launches and works offline for everything except AI generation and login. Specifically:\r\n\r\n**Works offline:** Opening projects, viewing all assets (images/videos), editing timeline (trim, reorder, split clips), adding markers, slideshow playback, FFmpeg export to MP4, DaVinci XML export.\r\n\r\n**Requires internet:** AI generation (all models), login/signup, credit purchases, credit balance sync, license validation (only checked on first generation attempt per session, then cached), auto-updater.\r\n\r\nIf you lose internet mid-session, you can keep editing and exporting. Generation will fail until connectivity returns.\r\n\r\n---\r\n\r\n## GENERATION RECOVERY\r\n\r\nIf the app closes during a generation: on restart, Slates detects in-flight jobs, polls the AI provider, and downloads completed results automatically. Nothing is lost. Recovering generations show at 5% in the queue until status is confirmed. Works for every model.\r\n\r\n---\r\n\r\n## TROUBLESHOOTING\r\n\r\n**\"Insufficient credits\"** \u2014 Your credit balance is too low for this generation. Buy more credits in the app (packs from $10 to $500) or turn on auto-topup. The exact cost of any generation is shown on the Generate button before you commit.\r\n\r\n**\"Input was rejected by Kling\"** \u2014 Image may not meet quality requirements (character visibility, proportions, content policy). Try a different image or prompt.\r\n\r\n**\"Failed to upload image to FAL CDN\"** \u2014 Network issue during reference image upload. Check internet connection, retry.\r\n\r\n**\"Generation failed\" / \"Proxy generation failed\"** \u2014 Generic error from the AI provider. Usually temporary. Retry. If persistent, try a different model.\r\n\r\n**\"Source asset not found\" / \"Source video asset not found\" / \"Target image asset not found\"** \u2014 The image or video you're trying to use was deleted or moved. Re-import or select a different asset.\r\n\r\n**\"Invalid audio source\"** \u2014 Lip sync: either enter TTS text or upload an audio file. One is required.\r\n\r\n**\"TTS response missing audio URL\"** \u2014 Text-to-speech failed during lip sync. Retry.\r\n\r\n**Generation stuck** \u2014 Restart app. Recovery system polls providers and picks up where it left off.\r\n\r\n**API rate limit** \u2014 Too many requests (limit: 20 generations/minute). Wait 1-2 minutes, retry.\r\n\r\n**Project files missing** \u2014 Project folder was moved/deleted outside the app. Use project relocation in Settings to re-point to the correct folder.\r\n\r\n**License shows \"revoked\"** \u2014 Contact support. Character sheets and environment grids unavailable until resolved.\r\n\r\n**Session expired** \u2014 Magic link session timed out. Log in again via Settings.\r\n\r\n**iPhone photos won't import** \u2014 iPhones save photos as HEIC/HEIF format, which Slates doesn't support. Convert to JPEG or PNG first (most photo apps and online converters can do this).\r\n\r\n---\r\n\r\n## PRIVACY & DATA\r\n\r\n- Generated files stay on YOUR machine. Slates servers never store your videos/images.\r\n- No prompts logged server-side.\r\n- File uploads go directly to the AI provider via pre-signed URLs. Slates servers never buffer your media.\r\n- Server stores only: email, license status, credit balance, transaction history, session tokens.\r\n- Stripe handles all payment data. Slates never sees your card number.\r\n\r\n---\r\n\r\n## SYSTEM REQUIREMENTS\r\n\r\n- Windows 10/11 or macOS 12+\r\n- Internet connection required for AI generation (not for editing/exporting)\r\n- Disk space for project files (AI videos are typically 5-50MB each)\r\n- FFmpeg bundled with app (no separate install needed)\r\n- No GPU required (all AI processing happens in the cloud)\r\n\r\n---\r\n\r\n## COMMON TASKS (STEP-BY-STEP)\r\n\r\n### Generate an Image\r\n1. Open the floating prompt box (visible on every page).\r\n2. Enter your prompt describing the image.\r\n3. Select an image model (Nano Banana 2 recommended).\r\n4. Choose aspect ratio and resolution.\r\n5. Press Ctrl+Enter or click Generate.\r\n\r\n### Generate Video From an Image\r\n1. In the prompt box, attach a start image.\r\n2. Write a prompt describing the desired motion/action.\r\n3. Select a video model (Kling V3.0 Omni in ingredients mode recommended).\r\n4. Choose duration (Kling bills per second from 3s up, so shorter is always cheaper), aspect ratio, and resolution.\r\n5. Click Generate.\r\n\r\n### Use a Character Reference for Consistency\r\n1. Create a character in your project (name + description).\r\n2. Either generate a character sheet OR manually assign a single image as the character reference.\r\n3. In the prompt box, type `@` and select your character from auto-complete.\r\n4. The reference image is attached automatically. Generate normally.\r\n\r\n### Export to DaVinci Resolve for Color Grading\r\n1. In the video editor, finalize your timeline (clips, markers, timing).\r\n2. Click Export \u2192 DaVinci Resolve XML.\r\n3. Choose output location. File saves to exports/ directory.\r\n4. In DaVinci Resolve: File \u2192 Import \u2192 Timeline. Select the XML file.\r\n5. Your timeline loads with all clips, properties, and markers intact. Grade and export.\r\n\r\n### Extract a Still Frame From a Video\r\n1. Hover over any video clip in the gallery.\r\n2. Camera icon = extract **current frame**. Dropdown arrow next to it = **First frame** or **Last frame**.\r\n3. Extracted image saves to your project gallery. Use as start/end image for I2V, character reference, or storyboard frame.\r\n\r\nKey workflow: extract a clip's last frame \u2192 use it as the start image for the next generation \u2192 seamless visual continuity between scenes.\r\n\r\n### Buy More Credits\r\n1. Open Settings \u2192 Credits, or the credit badge in the top nav.\r\n2. Pick a pack (packs run from $10 to $500, plus a $1,000 pack offered in the app to licensed accounts; bigger packs give more credits per dollar).\r\n3. Pay via Stripe. Credits are added to your balance instantly and never expire.\r\n4. Optional: turn on auto-topup so your balance refills automatically when it runs low.\r\n\r\n---\r\n\r\n## COMMON QUESTIONS\r\n\r\n**Q: Which model should I use for most videos?**\r\nA: Seedance 2.0 is the default and the one to reach for when physics, scale, effects or hero shots matter. For everyday shots built from a start image, Kling V3.0 Omni in ingredients mode is the best balance of cost and quality, which is why the step-by-step guides above use it.\r\n\r\n**Q: What's the best image model?**\r\nA: GPT Image 2.5 is the strongest image model in the app, and the best available anywhere right now. It holds a long instruction more faithfully than anything else here, and it is the only model that renders words in the picture reliably. The two seats cost the same, so the only trade is time: **Flare** is the fast one and is best for drafts and exploring, **Sunburst** is the higher-quality one and is what finals, hero frames and reference-heavy edits should end up on. The quality knob has five settings spanning about 36\u00D7 from cheapest to dearest \u2014 `max` is four times `high`, though the smaller steps are uneven: **`medium`** is for drafts and where iteration belongs, **`high` is the default** and where most finished work should sit, and **`max` is the best output you can get** \u2014 finished frames, client deliverables, anything with exact text. `xhigh` sits just under `max` for about half the price and is worth trying before you jump to the top. Going past `high` should be a deliberate choice rather than a habit. Go to **4K only once you already know your references and your prompt are solid** \u2014 it is the most expensive setting on the model and the worst place to discover the composition was wrong. Prove the shot at 3K first, then re-run the settled prompt at 4K. If you used GPT Image 2 before, note the quality names all shifted by one: its `medium` is now `high`, its `high` is now `max`. Nano Banana 2 is the picker's DEFAULT rather than the best: it is fast, takes the most reference images, and is the only model with grid exploration. Nano Banana Pro is the hero-frame step up when composition, cinematic lighting and skin have to be perfect. Prices for every tier are in the model table above.\r\n\r\n**Q: Do I need to set up API keys?**\r\nA: No. There are no API keys in Slates \u2014 every generation runs on Slates Credits, which come with your license and never expire.\r\n\r\n**Q: How much does a generation cost?**\r\nA: It depends on the model, resolution, and length. The exact credit cost is always shown on the Generate button before you commit, so there are no surprises.\r\n\r\n**Q: Can I use Slates offline?**\r\nA: Yes for viewing projects, editing timeline, and exporting. No for AI generation -- that requires internet.\r\n\r\n**Q: Do credits expire?**\r\nA: No. Credits never expire.\r\n\r\n**Q: What happens if I close the app during a generation?**\r\nA: Nothing is lost. On restart, Slates detects in-flight jobs and downloads completed results automatically.\r\n\r\n---\r\n\r\n## FEATURES NOT IN SLATES\r\n\r\nThe following are NOT available. Do not suggest them:\r\n\r\n- Bring-your-own API keys (BYOK) \u2014 every generation runs on Slates Credits; there is no key-entry option\r\n- Local/on-device GPU inference (all AI runs in the cloud)\r\n- Built-in music generation (use external tools like Suno, import audio)\r\n- A voice picker for Seed Audio \u2014 you describe the voice you want in words instead. The preset voice shelf belongs to the Voice lane (Inworld TTS-2), and Kling Lip-Sync keeps its own small fixed list: six English/UK voices plus a storyteller, with a speed control\r\n- Automatic video editing from a script\r\n- A prompt \"Enhance\" button \u2014 ask Studio Agent to rewrite a prompt instead\r\n- A settings/gear panel on the prompt box \u2014 every parameter is a dropdown on the bar\r\n- Cloud project storage (all files are local)\r\n- Real-time collaboration / multi-user editing\r\n- Mobile app (desktop only: Windows and macOS)\r\n- HEIC/HEIF image import (convert to JPEG/PNG first)\r\n- Storyboard JSON export (import only)\r\n\r\n---\r\n\r\n## VERSION\r\n\r\n<!-- BEGIN:GENERATED version -->\nSlates Reference Version: 1.5.6\nLast Updated: 2026-09-09\n\nThis document is generated. Its source of truth is `slate/docs/slates-llm-manual.md`; its model tables and credit costs are derived from the Slates model registry and pricing tables at build time, so they cannot be typed by hand.\n\nIf the user asks about a feature not documented here, it may have been added after this version. The current copy is always at <https://slates.video/slates-reference.md>.\n<!-- END:GENERATED version -->\r\n\r\nIf this document didn't answer your question, email hello@slates.video so we can help and improve the app.\r\n\r\n</slates_reference>\r\n";
|
|
1
|
+
export declare const APP_MANUAL = "# Slates \u2014 Complete Reference for AI Assistants\r\n\r\n<!-- The heading above must stay first: shipped app builds reject this file if it\r\n does not begin with one. See slate/CLAUDE.md; check:llm-docs enforces it. -->\r\n\r\n<system_role>\r\nYou are a support assistant for the Slates desktop application. Answer user questions using ONLY the information in the <slates_reference> below. Be concise and direct. Use numbered steps for procedures. Use bullet points for explanations when helpful.\r\n</system_role>\r\n\r\n<rules>\r\n- If the answer cannot be found in the <slates_reference>, say: \"That isn't covered in the Slates reference.\" Do not guess or invent features.\r\n- If the user asks how to do something, give step-by-step instructions from the workflows and features described here.\r\n- If the user reports an error, check the TROUBLESHOOTING section first.\r\n- Refer to the FEATURES NOT IN SLATES section before answering questions about capabilities that might not exist.\r\n- Quote the exact error message when referencing troubleshooting entries.\r\n</rules>\r\n\r\n<slates_reference>\r\n\r\n<!-- BEGIN:GENERATED header -->\n# SLATES v1.5.6 \u2014 Complete Reference\n\n> **Freshness.** Generated from the Slates source of truth for app version **1.5.6**, last changed **2026-09-09**. The canonical copy of this file is <https://slates.video/slates-reference.md>. If a model, price or feature the user mentions is missing below, this copy is out of date: re-fetch that URL before answering, and say so.\n<!-- END:GENERATED header -->\r\n\r\n**What is Slates?** Desktop app (Windows 10/11, macOS 12+) for AI image and video creation. One-time purchase, no subscription. Every license includes 1,000 free credits, and Slates Pro starts with 3,000 credits. Every generation runs on Slates Credits \u2014 there are no API keys to set up, and credits never expire.\r\n\r\n---\r\n\r\n## INTENDED WORKFLOW\r\n\r\nThe designed start-to-finish flow:\r\n\r\n1. **Create project** \u2014 New project with name/description. Creates folder on your disk.\r\n2. **Build visual assets** \u2014 Generate images, create characters (with character sheets for consistency), environments (with environment grids), and styles. This is your visual library.\r\n3. **Create storyboard** \u2014 Add scenes. Every picture you drop in becomes a **Shot**: one beat of the piece, holding its references, its model, its settings and (when you want them) its words.\r\n4. **Write the piece** \u2014 Switch the storyboard to **Document** and write. What is said, what happens, the framing, the prompt each beat will send. Nothing is required and nothing is asked for; a visuals-only piece is finished as it stands.\r\n5. **Read it before you pay for it** \u2014 The header states how many generations, how many cuts, how long it runs and what it will cost. Press play for a rough cut at the real timing. Re-chop with split and merge and watch the price move.\r\n6. **Generate** \u2014 Select the beats you want and fire them in one approved batch. Each result lands under the row that made it.\r\n7. **Organize** \u2014 Switch to **Board** to re-order. Drag a beat and the whole thing moves with it.\r\n8. **Export to timeline** \u2014 Send clips to the built-in multi-track video editor.\r\n9. **Edit** \u2014 Trim, reorder, add markers, adjust timing.\r\n10. **Final export** \u2014 Export to MP4 directly, or export DaVinci Resolve XML for professional color grading.\r\n\r\n---\r\n\r\n## MODEL REFERENCE TABLE\r\n\r\n**Generating 4K video is a Slates Pro feature** \u2014 every tier generates video up to 1080p, and 4K images are open to everyone. Exporting your finished timeline at 4K is available on every tier.\r\n\r\nEvery model runs on Slates Credits. **The exact credit cost appears on the Generate button before anything fires.** The tables below are generated from the app's own model registry and rate tables, so they describe exactly what the model picker offers in this version: aspect ratios, resolutions, durations, reference-image limits, and the credit price of each. Bigger credit packs lower your per-credit cost, and Slates Pro gets the best pack rate on every purchase.\r\n\r\n<!-- BEGIN:GENERATED model-tables -->\n### Image Models\n\n| Model | Aspect Ratios | Resolutions | Max Refs | Credits per image |\n|-------|--------------|-------------|----------|-------------------|\n| **GPT Image 2.5 Flare** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 \u00B7 3K 3 \u00B7 4K 5 (default quality; at max: 2K 8 \u00B7 3K 11 \u00B7 4K 20) |\n| **GPT Image 2.5 Sunburst** | 1:1, 16:9, 9:16, 4:3, 3:4 | 2K / 3K / 4K | 16 | 2K 2 \u00B7 3K 3 \u00B7 4K 5 (default quality; at max: 2K 8 \u00B7 3K 11 \u00B7 4K 20) |\n| **Nano Banana 2** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 4 \u00B7 2K 6 \u00B7 4K 8 |\n| **NB2 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K | 4 | 1K 2 |\n| **Nano Banana Pro** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 14 | 1K 8 \u00B7 2K 8 \u00B7 4K 15 |\n| **FLUX.2 Max** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 1K / 2K / 4K | 4 | 1K 4 \u00B7 2K 5 \u00B7 4K 8 |\n| **Seedream 5 Lite** | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 21:9 | 2K / 3K / 4K | 10 | 2K 2 \u00B7 3K 2 \u00B7 4K 2 |\n\n### Video Models\n\n| Model | Duration | Aspect Ratios | Resolutions | Max Refs | Audio | Credits per second |\n|-------|----------|--------------|-------------|----------|-------|--------------------|\n| **Seedance 2.0** | 4-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p / 4K | 9 | Included | 480p 3.5 \u00B7 720p 7.5 \u00B7 1080p 18.5 \u00B7 4K 39 |\n| **Seedance 2.5** | 4-30s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 30 | Included | 480p 5.1 \u00B7 720p 11.6 \u00B7 1080p 20.5 |\n| **Seedance 2.5 Edit** | Follows the source clip (4-30s) | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 720p / 1080p | 0 | Included | 480p 6.2 \u00B7 720p 13.9 \u00B7 1080p 24.6 |\n| **Kling V3.0 Standard** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (6.3 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (8.4 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Omni** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 4.2 (5.6 with audio) \u00B7 4K 21 |\n| **Kling V3.0 Omni Pro** | 3-15s | 16:9, 9:16, 1:1 | 1080p / 4K | 4 | Optional, costs more | 1080p 5.6 (7 with audio) \u00B7 4K 21 |\n| **Kling O3 Edit** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 6.3 |\n| **Kling O3 Edit Pro** | Follows the source clip (3-15s) | 16:9, 9:16, 1:1 | 1080p | 4 | Included | 1080p 8.4 |\n| **MiniMax H3 Max** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 1080p | 9 | Included | 480p 2.5 \u00B7 768p 4 \u00B7 1080p 8 |\n| **MiniMax H3** | 5-15s | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 480p / 768p / 2K / 4K | 9 | Included | 480p 2.5 \u00B7 768p 3 \u00B7 2K 6.5 \u00B7 4K 8 |\n| **Gemini Omni Flash** | 3-10s | 16:9, 9:16 | 720p | 7 | Included | 720p 6.4 |\n| **Omni Flash Edit** | Follows the source clip (3-10s) | 16:9, 9:16 | 720p | 0 | Included | 720p 6.4 |\n| **LTX-2.5** | 6/8/10/12/14/16/18/20s at 720p/1080p; 6/8/10s at 1440p/4K | 16:9, 9:16 | 720p / 1080p / 1440p / 4K | 0 | Included | 720p 4.5 \u00B7 1080p 6.5 \u00B7 1440p 9.5 \u00B7 4K 15 |\n| **LTX-2.5 Pro** | 6, 8, 10s | 16:9, 9:16 | 720p / 1080p | 0 | Included | 720p 6 \u00B7 1080p 8.5 |\n| **Veo 3.1 Fast** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 5 (7.5 with audio) \u00B7 1080p 5 (7.5 with audio) \u00B7 4K 15 (17.5 with audio) |\n| **Veo 3.1 Standard** | 4/6/8s at 720p; 8s at 1080p/4K | 16:9, 9:16 | 720p / 1080p / 4K | 3 | Optional, costs more | 720p 10 (20 with audio) \u00B7 1080p 10 (20 with audio) \u00B7 4K 20 (30 with audio) |\n\n**Seedance 2.0 \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 27.4 credits per second at 1080p.\n**Seedance 2.5 \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 16.3 credits per second at 720p.\n**Seedance 2.5 Edit \u00B7 Face** is a separate row in the model picker (a face in a reference image routes to a different provider, which costs more): 19.8 credits per second at 720p.\n\n### Audio Models\n\n| Model | Length | Credits |\n|-------|--------|---------|\n| **Seed Audio 1.0** | 3-120s | 1 at 3s \u00B7 3 at 15s \u00B7 19 at 120s |\n| **Inworld TTS-2** | up to 2,000 characters of text per take | 1 at 250 characters \u00B7 2 at 2,000 characters (billed per 250) |\n| **Sound Effects** | 1-22s | 1 at 1s \u00B7 1 at 4s \u00B7 3 at 22s |\n\n### Tools (Lip Sync, Motion Transfer)\n\nThese are real Kling endpoints that take a clip or a still as their subject, not models you prompt from scratch. Both bill in 5-second blocks.\n\n| Tool | Input | Billed in | Credits per block |\n|------|-------|-----------|-------------------|\n| **Kling Lip Sync** | Video source | 5s block | 4 |\n| **Kling Lip Sync (Avatar v2 Standard)** | Still-image source | 5s block | 14 |\n| **Kling Lip Sync (Avatar v2 Pro)** | Still-image source | 5s block | 29 |\n| **Kling Motion Control Standard** | Still image + reference video | 5s block | 32 |\n| **Kling Motion Control Pro** | Still image + reference video | 5s block | 42 |\n<!-- END:GENERATED model-tables -->\r\n\r\n---\r\n\r\n## WHICH MODEL TO USE\r\n\r\n### Images\r\n\r\n**Nano Banana 2 is the default image model.** It is the best all-round image model in the app: the most reference images of any image model, every aspect ratio, and output up to 4K. Brief it like a creative director rather than with tag soup. It is also the only model that supports the 2x2 / 3x3 grid exploration wrapper.\r\n\r\n- **NB2 Lite** is the fast, cheap draft seat in the same Nano Banana family. Roughly half the price of NB2 full and noticeably faster, 1K output only. Iterate here, finish on NB2.\r\n- **Nano Banana Pro** is the hero-frame and typography tier. Reach for it when spatial composition, cinematic lighting and skin, or fine in-image type have to be perfect. NB2 gets you most of the way there, so this is a deliberate step up, never a default.\r\n- **GPT Image 2.5** is the strongest image model in the app: it follows a long instruction more faithfully than anything else here, and it is the one to pick when the picture simply has to be right. It is also the sharp-text model, which is what makes it the choice for character sheets, shot grids, ordered panels and anything with words in the picture. It comes in two seats that cost exactly the same, and the difference is speed against quality. **Flare** is the fast one: OpenAI describes its quality as comparable to the older GPT Image 2, at roughly half the wait. **Sunburst** is OpenAI's most capable image model, better than GPT Image 2, and deliberately slower. Use Flare while you are still exploring, then re-run the shot you like on Sunburst for the final \u2014 and reach for Sunburst directly when several reference images all have to survive into one frame, or when an edit must change one region and leave identity, geometry and lighting untouched.\n Its quality knob has **five** settings \u2014 `low`, `medium`, `high`, `xhigh`, `max` \u2014 spanning about 36\u00D7 from cheapest to dearest, which makes it the biggest cost lever on the model. The steps are uneven rather than a constant multiplier: `max` is four times `high`, but `xhigh` is only about 1.8 times it. **`high` is the default and the everyday setting.** `medium` is for drafts and is cheap enough to iterate on freely. `max` is the ceiling, for finished frames and exact character-level text; `xhigh` sits just under it for about half the price and is worth trying first. Go past `high` deliberately, not by habit.\n 4K is worth it only once your references and prompt are already settled: prove the shot at 3K, then re-run the finished prompt at 4K. Iterating at 4K is the most common way to waste credits on this model. Its 4K tier is API-only, so even a paid ChatGPT account cannot render it. Note that its resolution tiers are **not** a price ladder: the pixel classes are token-priced by OpenAI, so the cheapest seat is not the smallest one. Read the prices in the table above rather than assuming.\n **If you have used GPT Image 2 before, the quality names all shifted by one.** What it called `medium` is now called `high`, and what it called `high` is now `max` \u2014 the same pictures at the same prices, renamed. A remembered setting will quietly buy you a cheaper tier than it used to.\r\n **It is the only image model that can give you a transparent background.** Set **Background** to *Transparent* on the prompt bar and you get a real alpha channel \u2014 a cut-out for a logo, sticker or overlay \u2014 rather than a painted-in backdrop. *Auto* is the default and lets the model decide from your prompt; *Opaque* forces a filled background. It costs nothing either way. Slates always saves PNG, which is what carries the transparency, so there is nothing else to set.\r\n **Square and 4:3 frames cost more than 16:9 on this model, and only on this model.** OpenAI charges by image tokens rather than by pixels, and a square frame uses about 1.8 times the tokens of a 16:9 frame the same size, and 4:3 or 3:4 about 1.37 times. The credit prices in the table above are the 16:9 numbers; pick 1:1 or 4:3 and the price on the Generate button goes up to match. 9:16 costs the same as 16:9. Every other image model charges the same whatever the shape.\r\n- **FLUX.2 Max** and **Seedream 5 Lite** are the less content-restricted options. Seedream is flat-priced at every resolution it offers, so there is no reason to pick a lower one. Both auto-route to their edit endpoint when you attach reference images.\r\n\r\n### Video\r\n\r\n**Seedance 2.0 is the default video model.** Reach for it the moment physics, effects, destruction or scale matter, and for hero shots. It takes many reference images, generates native audio at no extra cost, and is the only Seedance seat that reaches 4K (generating 4K video needs Slates Pro; timeline export at 4K does not). It is also the cheaper of the two seats at every resolution they share. A face in a reference image routes it to a different provider, which is why **Seedance 2.0 \u00B7 Face** is its own row in the model picker at its own price.\r\n\r\n- **Seedance 2.5 is a second seat, not an upgrade.** It buys much longer single takes and far more reference images, plus better prompt adherence. What it gives up is 4K, and it costs more than 2.0 at every resolution the two share \u2014 so 2.0 stays the model for 4K, and for the same resolution at a lower price. Because 2.5 runs longer, a long clip on 2.5 can cost more than a shorter, higher-resolution one on 2.0 \u2014 read the Generate button, not the resolution. **Seedance 2.5 Edit** is its clip-editing row: attach a clip, describe the change, and the output length follows the source.\r\n- **Kling** is the cost-effective workhorse and the most flexible family: strong start-frame adherence for identity, layout and text, acting, dialogue, multi-shot (up to 6 cuts), and the widest range of clip lengths. **Kling V3.0 Omni** adds multi-character dialogue in English, Chinese, Japanese, Korean and Spanish. Standard and Pro are the same model at two fidelity and price tiers. **Kling O3 Edit** takes an existing clip and changes what you describe, with subject and style reference images, while the original audio is preserved verbatim. Kling is also the only engine behind the Lip Sync and Motion Control tools.\r\n- **MiniMax H3** is the seat to pick when the SOUND is part of what you are writing. Every other video model treats audio as a switch; H3 takes it as three separate instructions in one prompt \u2014 the lines and action sounds tied to a moment, the ambience running underneath, and a score only the audience hears \u2014 and generates all of it with the picture in a single pass. It is also the only model where you say how much of a reference should survive, including moving one subject's characteristic onto a different subject. It runs 5-15 seconds at 480p, 768p, 2K or 4K, and takes up to nine reference images plus reference video and audio. Two things to watch: **the first five reference images are free and every one after that costs extra**, so attach what the shot needs rather than the maximum; and 2K and 4K are upscales of a 768p render rather than larger generations \u2014 in our own testing the 2K pass showed more artifacting than the 768p original it was built from, at more than twice the price. Generate and judge at 768p; step up only when a delivery spec demands the pixels.\r\n- **MiniMax H3 Max** is the same model post-trained by fal for SPEED, and it is the more expensive seat, not the cheaper one. It is dramatically faster: on the same 5-second 768p prompt it finished in about 5 seconds against about 57 seconds for H3 \u2014 roughly 12x (measured 2026-08-27). It stops at 768p and costs more per second than H3 at the resolution they share. It still animates a start frame and an end frame, so image-to-video works normally; what it does not have is the reference set \u2014 the extra identity, style and environment images plus reference video and audio that base H3 reads. Pick it when a fast turnaround on a text-to-video or start-frame shot is worth paying for; pick H3 for resolution, references, or the same tier at a lower price.\r\n- **LTX-2.5** is the VOLUME seat \u2014 the cheapest native 1080p second in the catalogue, with synchronised audio included free at every resolution, so it is the model to reach for when the job is many takes rather than one hero shot. Two things are unique to it. It makes the LONGEST clips of anything here, up to 20 seconds, and it is the only model that reaches 1440p. It is also the only one with native MULTISHOT: a single generation can carry two to four connected shots that hold the character, lighting and voice across the cuts, which everywhere else means generating separate clips and watching identity drift between them. Its constraints are unusually sharp, though. Durations are EVEN NUMBERS ONLY starting at six \u2014 6, 8, 10, 12, 14, 16, 18, 20, with no 5-second or 7-second clip \u2014 and above 1080p that ceiling drops to 10 seconds. Aspect ratios are 16:9 and 9:16 only. And it takes FRAMES, not references: a start frame and an optional end frame that generates a transition between them, but no identity, style or environment reference images at all, so cross-shot character consistency belongs on MiniMax H3 or Kling. Because sound is generated in the same pass, write the audio into the prompt and anchor every cue to something visible \u2014 anything unanchored gets invented.\r\n- **LTX-2.5 Pro** is the fidelity seat of that pair, and it is NOT simply a better LTX. It renders the picture with more compute on busy frames, but on a narrower envelope than the base row: 720p and 1080p only (no 1440p, no 4K) and 6, 8 or 10 seconds only, for about a third more per second. Reaching for it because the name says Pro costs more AND takes away the reach. Pick it when a specific shot needs the extra fidelity and fits inside 1080p and ten seconds; pick base LTX for length, resolution and volume.\r\n- **Gemini Omni Flash** is the cheap 720p seat with native synced audio included in one pass. **Omni Flash Edit** is the prompt-only clip editor: no reference images, one short instruction plus \"Keep everything else the same.\" Long descriptive prompts destroy it.\r\n- **Veo 3.1** is niche and is never a default. Pick it only when you specifically want Google's audio pass. It has the fewest aspect ratios and reference slots of any video model, fixed durations, and the highest per-clip cost.\r\n\r\nBoth edit models take an existing clip as their canvas, so their output length follows the source clip rather than a duration you choose.\r\n\r\n### Audio\r\n\r\nAudio is a third media type alongside images and video \u2014 generated as its own asset, shown in the gallery's **Audio** tab, and dragged onto an audio track in the timeline. This is separate from the audio some VIDEO models generate *inside* a clip (see AUDIO IN GENERATION below): use a video model when the sound must be locked to what is on screen, and these when you need audio you can move, trim, re-use, or layer.\r\n\r\n**Seed Audio 1.0 is the default.** A room with dialogue *and* clatter *and* ambience is one generation, not three layered ones, and because it is cheap you can run five takes and keep the best. It makes a whole audio SCENE from one plain sentence. **It has no length setting of its own** \u2014 Slates writes your chosen duration into the prompt, and that is what you are charged. You describe the voice in words; there is no voice list to pick from. Set **Languages** to Mixed if one scene needs more than one language (it costs the same).\r\n\r\n**Sound Effects** makes one effect, or a seamless loop. It is the only surface with an exact duration, so an effect can land on a specific frame. Describe the physical cause (\"heavy oak door slams shut in a stone hallway\"), not the label (\"door sound\"). **Loop** makes it seamless for beds; **Wording** controls how literally your description is followed. Seed Audio is actually the better tool for *long* ambience beds, so the two are not redundant in the direction you would expect.\r\n\r\nKling's `SFX:` / `Ambient noise:` prompt syntax belongs to video prompts and makes Seed Audio results *worse* \u2014 write plain sentences there instead.\r\n\r\n**Inworld TTS-2 is the voice seat** \u2014 type the words, pick a voice, press Generate. In the prompt box's Audio lane pick **Voice** in the model picker; the prompt is the exact text that gets spoken (nothing is added or rewritten \u2014 open \"See what gets sent\" to confirm), and the **Voice** control on the bar opens the voice picker: **Presets** (ready-made voices with gender, accent and age filters \u2014 every one plays the same audition line, so you compare voices rather than scripts), **Clips** (any character's voice, or any audio clip in the project, cloned for the take), or **Describe** (a voice in words). The character counter beside the bar is the bill: the generated audio table above gives the text cap and billing buckets, and the Generate button shows the price. Direction goes in square brackets (`[whispering] \u2026`) \u2014 anything in parentheses is read aloud. Cloning a real person's voice needs their permission. Right-click any audio clip in the Audio tab \u2192 **Use as voice** lands you on the Voice lane with that clip as the voice. Studio Agent and the MCP/CLI do the same through `slates_generate_audio` (a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId`, or a `voiceDescription`).\r\n\r\n**There is no music generation.** For a song, use an external tool and import the audio (see PROJECTS \u2192 Supported File Formats). For spoken lines inside a scene, let Seed Audio perform them, put them in the video prompt on a model with native audio (Seedance, Kling Omni, Omni Flash, Veo), or use Kling Lip-Sync's text-to-speech against a shot.\r\n\r\n### Tools\r\n\r\n**Tools** is not a model family. It is two real Kling endpoints that take a clip or a still as their subject: **Kling Lip Sync** and **Kling Motion Control**. Both bill in 5-second blocks and are described under GENERATION MODES below.\r\n\r\n**How pricing works:** every generation is priced in Slates Credits, and the exact cost is shown on the Generate button before you commit. Bigger credit packs give more credits per dollar; Slates Pro locks in the best pack rate on every purchase, forever.\r\n\r\n---\r\n\r\n## GENERATION MODES\r\n\r\n### Create Image\r\nPrompt \u2192 select image model \u2192 set aspect ratio + resolution \u2192 generate. Batch grids available for quick iteration.\r\n\r\n### Text-to-Video\r\nPrompt \u2192 select video model \u2192 set duration + aspect ratio + resolution \u2192 generate. Output: MP4.\r\n\r\n### Image-to-Video (I2V)\r\nAttach start image + prompt \u2192 select model \u2192 generate video from that image. Optional: attach end image (Veo) for guided transitions.\r\n\r\n### Ingredients / References\r\nUse @character_name, @environment_name, or #style_name in prompt to attach reference images for visual consistency. Kling: up to 4 total references. Veo: up to 3. Nano Banana 2: up to 14. The @mentions auto-complete from your project's characters/environments/styles.\r\n\r\n### Lip Sync\r\n**Kling only.** Pick **Kling Lip Sync** under the Tools family in the model picker. Source: video or still image.\r\n\r\n- **Audio source** \u2014 Text to speech (type the line; six English/UK voices plus a storyteller, with a speed control) OR Upload audio (bring your own recording, max 5MB).\r\n- **Avatar tier** \u2014 only appears for a still-image source: Avatar v2 Standard (the value tier) or Avatar v2 Pro (higher fidelity, higher rate).\r\n- 5s output blocks.\r\n\r\nWorks very well with human-like characters. Less reliable with animals or non-human characters.\r\n\r\n### Motion Transfer\r\n**Kling only.** Pick **Kling Motion Control** under the Tools family. Target: still image (your character). Source: reference video (the motion).\r\n\r\n- **Engine** \u2014 Kling MC Standard (value tier) or Kling MC Pro (higher fidelity, higher rate).\r\n- **Orientation** \u2014 *Match video* copies skeleton and depth from the clip (best for dancing, walking, full-body action; driving clips up to 30s). *Match image* keeps your character's pose and angle and uses the video only as motion hints (best for close-ups; up to 10s).\r\n- 5s output.\r\n\r\n> **Note:** these two tools used to offer a second \"Seedance 2.0\" engine. It was not a separate engine \u2014 picking it made Slates write a sentence into your prompt that you never saw, which is no longer allowed anywhere in the app (see \"What gets sent\" below). Both tools are now Kling endpoints only.\r\n\r\n### Edit Image\r\nRight-click any image asset \u2192 open viewer \u2192 switch to Edit Mode. Enter an edit prompt describing the changes you want. Select edit model: Nano Banana 2 (supports up to 14 reference images), FLUX.2 Max, or Seedream 5 Lite. Choose resolution and aspect ratio. The result saves as a new asset with the original preserved. Useful for refining generated images without starting from scratch.\r\n\r\n### Edit Video (Kling O3 Edit / Omni Flash Edit)\r\nRight-click any video clip (gallery or timeline) \u2192 \"Edit with AI\". The clip attaches to the prompt box as the source; describe the CHANGE, not the whole scene (\"replace the man with @marcus\", \"make it a rainy night, keep everything else\"). Two engines in the model picker:\r\n- **Kling O3 Edit (default):** attach subject images (role: Subject) to swap someone in, or style images (role: Style) for a look \u2014 max 4 combined refs. Clips 3-15s. Original audio preserved.\r\n- **Omni Flash Edit (cheapest):** prompt only \u2014 no reference images; keep instructions simple and add \"Keep everything else the same.\" Clips 3-10s, 720p output.\r\n\r\nOutput length follows the source clip; the credit cost (clip seconds, rounded up, at the per-second rate) shows on the Generate button. The edited clip saves as a NEW asset linked to the original \u2014 chain edits freely. Trim longer clips on the timeline first.\r\n\r\n### Multi-Shot (Kling V3.0/Omni)\r\nEnable multi-shot toggle \u2192 multiple scene prompts in one generation, each with different framing. 6-axis camera controls per shot. Results can be hit-or-miss, but worth trying for quick multi-cut sequences. For more reliable results, most users prefer generating multiple short 5s clips separately using Kling V3.0 Omni in ingredients mode and assembling them on the timeline.\r\n\r\n---\r\n\r\n### Generate Audio\r\n\r\nSwitch the prompt box's lane pill from Image/Video to **Audio**, pick a surface, and generate. The prompt box offers three: **Seed Audio 1.0**, **Voice** (Inworld TTS-2) and **Sound Effects**. The result lands in the gallery's Audio tab as its own asset with a waveform and an inline player, and can be dragged onto an audio track in the timeline.\r\n\r\n- **Scene (Seed Audio 1.0)** \u2014 one plain sentence describing the moment. Set **Length**; Slates writes it into the prompt for you and that is exactly what you're billed for (open \"See what gets sent\" under the prompt box to read the appended text). **Say the crowd/room size out loud** \u2014 \"applause\" returns a full auditorium when you meant three people at an open mic. Ask for a few seconds more than the clip needs so the edit has fade handles. Describe the voice you want in the sentence itself (\"a weary dock foreman in his fifties, gravel in his voice\") \u2014 there is no voice picker.\r\n- **Voice (Inworld TTS-2)** \u2014 the prompt is the words to be spoken, verbatim. Pick the voice with the **Voice** control on the bar (presets you can play first, any clip in the project, a character's voice, or a description); the character counter is the bill. See MODELS \u2192 Inworld TTS-2 for direction tags and the cloning rules.\r\n- **Sound Effect** \u2014 describe the physical cause and set the length to roughly the event (\u22481s for an impact, 2\u20134s for a whoosh, 8\u201322s + **Loop** for a bed). **Wording** sets how literally the description is followed: Interpretive, Balanced (default), or Literal.\n\n#### Use your own voice recording\n\nIn the bottom prompt box, choose **Audio**, then **Inworld TTS-2** in the model picker. Open **Voice \u2192 Clips \u2192 Import voice clip** and select your recording. The import adds an audio asset to this project without generating anything. Click its play button to audition it, then click the recording's name to choose it. Type the words you want spoken in the prompt box and press **Generate**, which shows the price. The new take appears in **Gallery \u2192 Audio**.\n\nUse a clean recording of one speaker whose voice you have permission to use. Slates clones the recording for each take; there is no separate training wizard or persistent vendor voice to manage. **Presets** lets you audition ready-made voices; **Describe** lets you write a voice description and choose **Use this description**. Choosing in the prompt box sets up the next take; only Generate spends credits.\n\nTo attach your recording or a generated take to a character, open **Gallery \u2192 Characters**, then **Add voice** (or **Change voice**) on that character's card. Choose **Clips** and click the clip's name. Attaching an existing clip is free. The card displays its waveform and player. That character's voice is also listed under **Voice \u2192 Clips \u2192 Characters** in the prompt box. Selecting a preset or description from a character card generates and attaches a take; read the cost shown in that picker before choosing.\n\nFor **Inworld TTS-2**, choose the character under **Voice \u2192 Clips \u2192 Characters**. Typing `@name` does not choose a TTS voice: the prompt is spoken verbatim. Automatic voice attachment from a character mention belongs to **Seed Audio** only.\n\r\n---\r\n\r\n## AUDIO IN GENERATION\r\n\r\nThis section is about audio generated **inside a video clip**. For audio as its own asset, see Generate Audio above.\r\n\r\n**Veo 3.1 native audio:** Generates audio WITH video. Prompt syntax: `\"Hello!\"` for dialogue, `SFX: [sound]` for effects, `Ambient noise: [description]` for ambience. Max 10s dialogue. Add `(no subtitles)` to suppress text overlays.\r\n\r\n**Kling V3.0 Omni dialogue:** Multi-character dialogue with distinct voices. Languages: EN, ZH, JA, KO, ES. `Background music: [description]` for music. Max 10s dialogue.\r\n\r\n**Kling V3.0 sound co-generation:** Synchronized sound effects generated with video.\r\n\r\n\u26A0\uFE0F **This prompt syntax is video-only.** `SFX:`, `Ambient noise:` and `Background music:` are Kling/Veo conventions \u2014 the audio models above have no parser for them and will treat them as words in the scene.\r\n\r\n---\r\n\r\n## PROMPT SYSTEM\r\n\r\n### Unified Create Surface (roles + model-on-button)\r\nThe old mode tabs (text-to-video / frames-to-video / ingredients / create-image) are ONE \"Create\" surface. An **Image | Video | Audio pill** on the prompt bar switches your output lane \u2014 it remembers and restores the last model you used in each lane (pick Seedance once and the Video lane stays Seedance until you change it). Only the active lane shows its name; the other two are icons. Every attachment in the reference tray carries a tappable ROLE badge \u2014 Reference / First frame / Last frame / Subject / Style \u2014 you say what each attachment is; nothing is inferred. Your model choice sticks across generations and workflow actions (\"use as first frame\" keeps your chosen video model). Attaching a video via \"Edit with AI\" flips the surface into Edit Video mode. Lip Sync and Motion Transfer live under the **Tools** family in the model picker.\r\n\r\n### Floating Prompt Box \u2014 the bar holds everything\r\nPersistent across all pages. **There is no settings panel and no gear button.** Everything sits on one bottom bar, left to right:\r\n\r\n1. **Media toggle** \u2014 Image / Video / Audio.\r\n2. **Model picker** \u2014 a searchable menu plus a detached submenu. The main menu lists model families with a vendor glyph tile each; picking one opens that family's models beside it, every row carrying capability chips (resolution, clip length, audio, references) and its per-unit rate. Type to search across every model. The submenu is anchored to the row you opened it from, so it never travels. The trigger on the bar shows the model name and nothing else \u2014 no chevron, no resolution appended. Everything listed is a real model or endpoint.\r\n3. **Parameter controls** \u2014 one per setting the chosen model actually has (resolution, aspect, duration, length, quality, count, grid, face-in-reference, audio, loop, and so on). The trigger shows the current value; the explanation lives *inside* the menu as a subtitle under each option, along with what that option costs. A setting with only one possible value still shows, muted and non-interactive, so the row never changes shape.\r\n4. **Sliders for ranges.** A setting with a long list of steps (video duration, audio length) opens a ruler instead of a many-row menu. The handle moves between the values the model actually declares, so it cannot land on one the model will not accept, and the price for the selected value is shown on the ruler.\r\n5. **`More \u25BE`** \u2014 if the model has more parameters than fit the current window width, the extras are collected into a generated `More` dropdown automatically. Widen the window (or close the Studio Agent panel) and they move back onto the bar.\r\n6. **Generate** \u2014 reads `Generate \u00B7 <cost>`. **Cost only** \u2014 the model name is not repeated on the button (it's in the model picker) and there is no send arrow. A badge on the left of the button counts generations currently running.\r\n\r\nFor text-to-speech, the character counter is always visible because text length determines the price. On other surfaces it appears near the right of the bar after roughly three quarters of the model's prompt limit. It turns red over the limit.\r\n\r\n**Collapsing:** the chevron at the top-right of the box collapses it to a single arrow \u2014 nothing else. Click the arrow to bring it back.\r\n\r\nBelow the bar the queue shows pending/active generations with cost and progress.\r\n\r\n### What gets sent (prompt transparency)\r\nUnder the prompt box is a **\"See what gets sent\"** disclosure. Open it and you see the exact text that will be transmitted, produced by the same code that builds the request \u2014 so it can never disagree with what is actually sent. It shows:\r\n\r\n- **Reference numbering** \u2014 `@sarah` becomes `Sarah (image 1)` so the model knows which attached image is which.\r\n- **Key lines for attachments you did NOT mention** \u2014 one short neutral sentence per unmentioned attachment. Mention every reference in your own words and these never generate.\r\n- **The trailing style clause** when a `#style` is attached.\r\n- **The Seed Audio duration append** \u2014 the `\u2026 N seconds` Slates adds to the end of the prompt, which is also what you are billed for.\r\n- **Grid wrapping** when 2\u00D72 or 3\u00D73 is on.\r\n\r\nThe row stays hidden when the composed prompt is identical to what you typed, so it only appears when there is something to show.\r\n\r\n**Unresolved `#tags` and `@mentions` are named, not silently dropped.** If you type `#noir` and there is no saved style called \"noir\", the tag is removed from the text sent to the model (a raw tag confuses every model) \u2014 but the disclosure turns red and says so by name: *\"#noir matches nothing saved \u2014 removed from what gets sent.\"* Save the style, or reword it, and the warning clears.\r\n\r\n**Nothing is ever added that you cannot read here.** No setting in the app injects prompt text; a setting changes *how* a request is made, never *what* you asked for.\r\n\r\n### Prompting guide (on the web)\r\nPer-model prompting guidance lives at <https://slates.video/docs/prompting>, linked from the bottom of Settings. It covers every model Slates offers \u2014 Video, Image, Audio \u2014 with what that model reads, what it ignores, and its gotchas, all on one page so you can compare them. Markdown copy for pasting into an LLM: <https://slates.video/docs/prompting.md>.\r\n\r\nIt is generated from the same source the Slates CLI, the MCP server and Studio Agent are built on, so the guide and the app cannot disagree. It is documentation rather than a control, which is why it is a page on the web and not a panel in the app: it has room to be read, a URL you can send someone, and it is always current rather than frozen at the version you installed.\r\n\r\n### @Mentions\r\nType `@` \u2192 auto-complete shows project characters and environments. Type `#` \u2192 shows styles. Selecting inserts the reference image(s). At send time a mention is rewritten to a numbered citation (`@sarah` \u2192 `Sarah (image 1)`) so the model can tell your attachments apart \u2014 you can read the result in \"See what gets sent\". Nothing else about your wording is rewritten. Prompting works the same as any other AI tool; no special syntax beyond @mentions.\r\n\r\n### Writing the shot list with Studio Agent\r\nThere is no \"Enhance\" button and no \"Generate prompts\" button. **Studio Agent does this work**, because it reads the same shot list you do \u2014 every scene, every beat in order, with its references, its model and its price \u2014 and because you can steer it:\r\n\r\n> \"Write the beats for my current storyboard.\"\r\n> \"Now redo scene 3 handheld, and match its energy to scene 2.\"\r\n> \"SHOT-A4 runs long \u2014 split it after 'and then'.\"\r\n\r\nEvery field it writes is editable by hand, in place, in the storyboard's Document view. Open Studio Agent with **Ctrl+.**\r\n\r\n### Debug Panel (advanced)\r\nA developer panel showing the exact request body, with the ability to override the composed prompt before sending. **There is no toggle button for it on the prompt bar in any build** \u2014 open it with **Ctrl+Shift+D**. For ordinary use, \"See what gets sent\" above is the supported way to inspect a prompt.\r\n\r\n---\r\n\r\n## PROJECTS\r\n\r\n### Structure\r\nEach project = folder on your disk. Subdirectories: images/, videos/, audio/, references/, exports/. Location configurable in Settings \u2192 Projects Directory.\r\n\r\n### Assets\r\nEvery generated or imported file is an asset (image, video, audio). Metadata tracked: prompt, model, settings, cost, dimensions, timestamps. Videos track source image via source_asset_id -- you can see all videos generated from any image.\r\n\r\n### Supported File Formats\r\n- **Images:** PNG, JPEG, WEBP, GIF. Note: HEIC/HEIF (iPhone photos) NOT supported -- convert to JPEG/PNG first.\r\n- **Video:** MP4, MOV, WEBM, AVI, MKV.\r\n- **Audio:** MP3, WAV, OGG, M4A, AAC.\r\n- **Clipboard paste:** Any image format the OS clipboard provides (PNG, JPEG, WEBP, GIF). Pasting works both in the gallery and directly into the prompt box; either way the image becomes a real gallery asset in the folder you're working in (tagged \"Imported\") AND, when pasted into the prompt box, attaches as a reference. Anything generated from it links back to it as a source.\r\n- **Drag and drop:** Any file the browser recognizes as image/* or video/*.\r\n\r\n### Operations\r\nCreate/rename/delete projects. Import external files via drag-and-drop or file picker. Paste images from clipboard. Extract still frames from videos. Relocate project to different disk/folder (all paths auto-update). Cleanup orphaned assets.\r\n\r\n### Moving and copying assets between projects\r\nAssets (images, clips) can be sent to another project three ways: the selection band on the Images/Videos tabs, the right-click menu on any card, or by dragging cards and dropping on a project in the drop palette.\r\n\r\n- **Move** relocates the media files on disk into the destination project's folder. The asset leaves whatever gallery folder it was in and is issued a fresh badge code in the destination.\r\n- **Copy** duplicates it \u2014 new files, new thumbnails, new badge code \u2014 and changes nothing in the source project.\r\n\r\n**Why a move can be refused:** an image another project still builds with (a character/environment/style identity image, or a storyboard frame) cannot leave, because the entity left behind would point at a file it no longer owns. When that happens the dialog lists what's blocking and offers the fix: bring the whole character/environment/style across with all of its images, or copy instead. A storyboard frame is only ever offered a copy \u2014 moving its image out would empty the shot.\r\n\r\nRight-clicking a card that is part of a multi-selection acts on the whole selection (\"Move 5 to Project\u2026\"). Right-clicking a card outside the selection acts on that card alone.\r\n\r\n---\r\n\r\n## SHOTS \u2014 THE STORYBOARD IS THE SHOT LIST\r\n\r\nA **Shot** is the prompt bar, saved: the prompt, every reference with the job it carries, the model, every setting \u2014 and now the beat itself: who speaks, what they say, how it is said, what happens, the prop, the framing and the camera. It lives in the **storyboard**, which is the one place Shots are listed. It is never required: the prompt bar works exactly as it always has for anyone who never touches one.\r\n\r\n### Why it exists\r\nA generation's full recipe was already stored, but only once you had paid for it. A Shot can be written **before anything is generated**, so a whole piece can be planned, read, timed, priced and corrected while it is still free. That is the point of the thing: look at the entire ad or short film \u2014 every cheap asset lined up in the actual flow \u2014 before the videos exist.\r\n\r\n### What a Shot holds\r\nRaw prompt (@mentions intact); the model; aspect ratio / duration / resolution / negative prompt / sound and the rest of the bar's settings; every attachment with its ROLE (plain reference, subject, style, reference video, reference audio, first frame, last frame); and the script layer \u2014 `speaker`, `line`, `delivery`, `action`, `prop`, `shotSize`, `camera`, and a `continues` flag for one sentence running across two cuts. Characters, environments and styles are stored as the ENTITY, not a copied picture, so updating a character updates every Shot that names it.\r\n\r\n**The script fields are for reading and counting. Only the prompt is sent to a model.** Dialogue you want a model to perform still goes in the prompt, verbatim, with its delivery \u2014 writing it in `line` makes it readable and lets Slates check whether it fits the cut, not spoken.\r\n\r\n### Every Shot has an address\r\n`SHOT-A1`, `SHOT-A2` \u2026 per project, never reused \u2014 the same idea as the `IMG-A12` badge on a gallery card. Say it to Claude and you are both pointing at the same row. It is for **this session**, not for retrieval later: there is no shot search and no shot library, because a Shot is workspace state \u2014 alive while you build the piece, worthless once it ships.\r\n\r\n### Making one\r\n- **Save as Shot** on any generated image, clip or track's right-click menu restores that generation and keeps it \u2014 and that generation becomes the Shot's first take, so the row opens showing the result you kept it for. **Reuse Prompt** on those same menus does the restore WITHOUT saving anything.\r\n- Claude can write Shots directly (`slates_create_shot`), including for shots whose image does not exist yet \u2014 and can re-chop them with `slates_split_shot` / `slates_merge_shots`.\r\n- **It files itself.** A saved Shot lands in the scene you have open, else the last scene of the storyboard you were most recently working in; if the project has no storyboard, one appears named after the project. Nothing you save is ever somewhere you have to go and find.\r\n\r\n### There is no save button\r\nSelecting a Shot row **binds** the prompt bar to it. Edits write straight back to that row; `Clear` unbinds and returns the bar to free composing. There is no undo, and none is needed: the generations underneath a row are the permanent record of what actually fired, and the row itself is the working copy.\r\n\r\n### Two ways to look at it\r\nThe storyboard has one toggle and two jobs.\r\n\r\n- **Board** \u2014 arrange. One picture per Shot, dragged into the order you want. Drag one and the whole beat moves with it: the line, the references, the model, the settings, the takes. There is nothing else to drag, so there is never a question of what followed what.\r\n- **Document** \u2014 write. One continuous page: the script, the references beside the words that cite them, the prompts underneath. This is where you read the piece before paying for it.\r\n\r\nInside Document, choose **Script** to edit the words with scene headings and speakers. Delivery notes, model and pricing details, and warnings are hidden. **Shots** shows all layers. Open **Custom** to toggle Scene, Action, Character, Delivery, Dialogue, Shot, References, Prompt, Takes, and Warnings independently. Warnings covers missing models or references, unresolved mentions, model changes, and lines too long for their cut. Hiding a layer changes only what you see; it preserves your text and generation checks. There is a text-size slider and an independent toggle for shot numbers in the margin.\n\nDelivery is an optional performance note, not a required label for every line. TTS sends the authored prompt, or the dialogue when the prompt is empty, verbatim; separate Delivery notes are not added. For speech cues, use the selected TTS model's prompting guide and place supported tags in that spoken text. Script view does not strip inline speech cues from your words.\n\r\n### The header tells you what you are about to make\r\n`5 generations \u00B7 7 cuts \u00B7 54s \u00B7 84 credits`, and beneath it a variety strip like `6/7 wide \u00B7 5 push \u00B7 3 cuts in the loft`.\r\n\r\n**Two counts, because they measure different things.** Rhythm is counted in **cuts**; money is counted in **generations**. A multi-shot generation is several cuts inside one paid call, so mixing them would be wrong. A cut with no model chosen has no duration and shows as `\u2014` rather than `0s` \u2014 a runtime that invented seconds would be a lie about the one number this view exists to give.\r\n\r\n### Splitting and merging \u2014 the chop\r\nPut the caret mid-line and press Enter: the row becomes two, the second inheriting the model, settings and references, and marked as continuing the first if the split lands mid-sentence. Select two adjacent rows and press **Merge**: they become one, references combined, durations summed. **The price and the runtime move as you do it** \u2014 which is the whole reason to make the decision here rather than in a document somewhere else.\r\n\r\nSplit is also the move behind a voiceover that keeps talking while the picture hard-cuts to a new world: split at a word boundary and both rows carry one sentence, each with its own visuals.\n\nThe **Dialogue continues from previous shot** toggle is a planning note that the sentence spans a cut. It does not merge shots, join generated audio, or change generation settings. Its pressed state shows whether the note is set; toggling it does not move the script. **Merge these two shots** is the separate control between adjacent shots that actually combines them. In Document \u2192 Custom, **Scene** toggles scene headings and their controls; the dialogue remains visible even if a hidden heading belonged to a collapsed scene.\n\r\n### Variety, counted and never judged\r\nSlates counts what is in front of it \u2014 shot sizes, camera moves, cast, locations, durations, and any of them repeating three or more times in a row \u2014 and shows the counts. **It never changes anything, never suggests anything and never blocks.** `shotSize` and `camera` are free text: write `long-lens CU, other head blurred` if that is the shot. Anything unrecognised counts as \"other\", which is a fine answer.\r\n\r\nIf a spoken line cannot be read in its cut at any plausible pace, the row says so \u2014 and says it only when the line is genuinely impossible, never when it is merely long.\r\n\r\n### The animatic\r\nPress play on any beat and the storyboard plays as a rough cut: each picture held for **its own cut's duration**, with the line underneath. That tells you the rhythm of the finished piece before a single video exists. A multi-shot generation holds one picture across its internal cuts, and says so. A cut with no duration is held for 3 seconds and marked \u2014 the header leaves it out of the runtime for the same reason it is marked here.\r\n\r\n### Things that stay honest rather than being hidden\r\n- **Deleted references.** If an asset or character a Shot points at is gone, the Shot still loads and the row says how many items were left out of what gets sent \u2014 and editing the row does not quietly drop them.\r\n- **A swapped model.** Changing a Shot's model never rewrites your words \u2014 video models genuinely take different prompt grammars, so the row tells you which model the prompt was written for and leaves the sentence alone. The settings line shows what will actually be sent after the swap, and that is the value it prices.\r\n- **Image Shots carry no role badges.** Image generation sends every reference in one undifferentiated list, so an image Shot can remember that a picture is a style reference but cannot tell the model.\r\n- **The thumbnail is never a question.** Slates picks it \u2014 the first frame, else the first reference, else the newest take \u2014 and you can override it from any reference in the gutter.\r\n\r\n### Firing several\r\nSelect Shots and press **Generate all**: one total, the largest single Shot stated separately, one approval. They run **one at a time**. If one of the selected Shots has been deleted, the whole batch is refused and names it \u2014 nothing fires and nothing is billed. If a generation fails mid-run, the rest still fire, the failure is reported per Shot, and **nothing is retried automatically**.\r\n\r\n### Deleting\r\nDeleting a storyboard tells you how many Shots are attached before it does anything, and deletes them with it. It never moves them somewhere else without asking. A Shot that also lives in another storyboard survives.\r\n\r\n---\r\n\r\n## STORYBOARDING\r\n\r\n### Hierarchy\r\nStoryboard \u2192 Scenes \u2192 **Shots**. A scene is an ordered list of Shots, and a Shot is one beat: its picture, its references and their roles, its model and settings, its prompt, its words, and the generations it has produced. See **SHOTS** above \u2014 that section is the storyboard.\r\n\r\n### What happened to frame types\r\nThere used to be a \"frame type\" on each picture \u2014 first / last / ingredient \u2014 plus a separate motion-prompt box. Both were a second, weaker way of saying what a Shot already says: **a reference's role lives on the Shot** (first frame, last frame, subject, style, plain reference), and the motion prompt was just the Shot's prompt under another name. Existing storyboards were converted automatically and nothing was lost. Pick a picture's job on the Shot's reference rail; write the motion in the Shot's prompt.\r\n\r\n### Grid Exploration\r\n2x2 grid: 4 prompt variations for quick iteration. 3x3 grid: 9 variations for deeper exploration. Select individual cells \u2192 extract to full-resolution images. Tip: 2x2 is usually sufficient and produces better quality. 3x3 can occasionally get proportions slightly wrong when upscaling cells because it faithfully reproduces the lower-resolution proportions. Grid exploration runs on Nano Banana 2 only; no other image model offers it.\r\n\r\n### Storyboard \u2192 Video\r\nSelect frames \u2192 generate video for each \u2192 clips auto-insert into timeline in order with source tracking maintained.\r\n\r\n### The animatic\r\nPlay the storyboard as a rough cut. Each beat is held for **its own duration**, with its line underneath, so what you are watching runs at the finished piece's real length. Space = play/pause. Arrow keys = navigate. Escape = exit.\r\n\r\n### Paste a script\r\nPaste a script into the storyboard and it becomes one row per paragraph \u2014 ALL-CAPS cues become speakers, parentheticals become delivery. It is a plain parse, not a model: nothing is invented, nothing is sent anywhere, and prose that is not screenplay-formatted lands as one row per paragraph for you (or Claude) to chop.\r\n\r\n---\r\n\r\n## VIDEO EDITOR (TIMELINE)\r\n\r\n### Tracks\r\nMulti-track: video tracks + audio tracks stacked vertically. Clips independent per track. Add or remove tracks freely -- layer a music bed, a voiceover, and effects on separate audio tracks. Video assets go on video tracks, audio assets on audio tracks. Overlapping video clips resolve top-track-wins.\r\n\r\n### Audio Mixing\r\nEach track has a volume fader, and the timeline has a master output fader for the final mix. Both range from silent to +12 dB of boost, and both apply to preview playback AND the exported MP4 -- what you hear is what you render. Muting a video track silences its embedded audio but still shows the picture. Use the master fader to prevent clipping when stacking loud tracks.\r\n\r\n### Timeline Settings\r\nResolution and frame rate (24/30/60) are auto-managed: the first video clip sets both, and a later higher-resolution clip raises the canvas. All clips are conformed to the timeline frame rate on export. Changing the frame rate after clips are placed retimes them.\r\n\r\n### Clip Properties\r\nSource asset, in/out points (frame-level precision), duration, scale (fit/fill/custom %), position (X/Y offset), opacity (0-100%).\r\n\r\n### Tools\r\n- **Select (V):** Click/drag clips, view/edit properties\r\n- **Razor (C):** Split clip at playhead into two clips\r\n- **Slip (S):** Adjust clip in/out points without moving its position\r\n- **Snap toggle:** Snap to playhead/clip boundaries\r\n\r\n### Markers\r\nColor-coded timeline markers (6+ colors) with optional labels. Use for scene breaks, cue points, notes.\r\n\r\n### Playback & Navigation\r\nSpace = play/pause. Left/Right arrows = frame-by-frame. Up/Down = +-1 second. Page Up/Down = jump by screen width. Home/End = start/end of timeline.\r\n\r\n### Zoom\r\nCtrl+Plus = zoom in (finer precision). Ctrl+Minus = zoom out (see more timeline).\r\n\r\n### Undo/Redo\r\n50-step history. Ctrl+Z = undo. Ctrl+Shift+Z = redo.\r\n\r\n---\r\n\r\n## EXPORT\r\n\r\n### Video Export (FFmpeg)\r\nExport timeline \u2192 MP4 (H.264). Configure: resolution, frame rate, bitrate, output location. All visible tracks rendered, muted tracks excluded, clip in/out points respected. FFmpeg is bundled -- no separate install needed.\r\n\r\n### DaVinci Resolve XML Export\r\nGenerates XML project file containing: clip references (paths to source videos), timeline structure (tracks, clips), clip properties (scale, position, opacity, in/out points), timeline markers.\r\n\r\n**Importing into DaVinci Resolve:** File \u2192 Import \u2192 Timeline. DaVinci reads the XML and reconstructs your timeline with all clips, properties, and markers intact. From there you can color grade and export your final master.\r\n\r\nExports saved to project's exports/ directory with timestamped filenames.\r\n\r\n---\r\n\r\n## CHARACTERS, ENVIRONMENTS & STYLES\r\n\r\n### Characters\r\nCreate character with name + description. Generate character sheet (license required): AI generates a turnaround with multiple angles for consistency. Generate expression sheet: same character with different facial expressions. Use `@character_name` in any prompt to attach reference images.\r\n\r\n**Tips for consistency:** Experiment with character sheet generation using both the existing project style and photorealistic style. Sometimes a single well-chosen image works better than a full sheet -- especially if the character is already in the same style, lighting, and clothing as your project. You can manually assign any image as a character reference instead of generating a sheet.\r\n\r\n**Voice.** A character can carry one voice clip, the same way it carries one identity image \u2014 a shortcut for reusing a voice, never a requirement for speaking in one. **Add voice** / **Change voice** on the card opens the same voice picker the prompt box uses (a menu off the button, not a pop-up; the other cards stay on screen): a clip from the project attaches as it is; a preset or a described voice renders the character speaking a fixed audition line on Inworld TTS-2 and attaches that clip, at the credit cost the picker states first. Right-click the voice card \u2192 **Remove voice** detaches it without deleting the clip. The character's voice then shows under **Clips** in the Voice lane's picker, and mentioning the character (`@name`) in a Seed Audio prompt attaches the clip as a reference, so the scene casts that voice.\r\n\r\n### Environments\r\nCreate environment with name + description. Generate environment grid (license required, 3x3): 9 variations. Extract individual cells to full-resolution images. Use `@environment_name` in prompts.\r\n\r\n**Tip:** Like characters, sometimes a single strong environment image gives better consistency than a grid of 9. Experiment with both approaches.\r\n\r\n### Styles\r\nCreate style with name + description + upload reference image. Use `#style_name` in prompts. Key visual auto-attachment option for consistent look across all frames.\r\n\r\n---\r\n\r\n## SETTINGS\r\n\r\n### Generation\r\nEvery generation runs on Slates Credits \u2014 there are no API keys to configure. The Generate button shows the exact credit cost before each generation, and failed generations refund immediately.\r\n\r\n### Other Settings\r\n- **Projects Directory:** Where project folders live on disk. Changeable anytime.\r\n- **Default Model:** Pre-selected model for new generations. Override per-generation.\r\n- **Default Quality/Resolution:** Pre-selected resolution. Override per-generation.\r\n- **Grid Size:** Default 2x2 or 3x3 for grid exploration.\r\n- **Auto Naming:** Automatically name generated assets.\r\n- **Prompting guide:** A link at the bottom of Settings to <https://slates.video/docs/prompting> \u2014 per-model prompting guidance for every model (see PROMPT SYSTEM above).\r\n\r\n---\r\n\r\n## ACCOUNT & BILLING\r\n\r\n### Login\r\nEmail-only, no password. Enter email \u2192 receive magic link \u2192 click to log in. First login creates account automatically. Session persists across restarts.\r\n\r\n### License\r\nUnlocks: character sheet generation and environment grid generation. Includes 12 months of updates (Slates Pro includes lifetime updates). Major upgrades discounted after.\r\n\r\n### Credits\r\n\r\n<!-- BEGIN:GENERATED credits -->\nCredits are what every generation is paid with. They are pay-as-you-go, they never expire, and the exact cost of a generation is shown on the Generate button before you commit.\n\n- A **Slates Standard** license ($149 one time) starts you with **1,000 credits**.\n- **Slates Pro** ($297 one time, or $97 to upgrade later) starts you with **3,000 credits**.\n\n| Pack | Credits (Standard) | Credits per dollar | Versus the smallest pack |\n|------|--------------------|--------------------|--------------------------|\n| $10 | 250 | 25.0 | standard rate |\n| $25 | 650 | 26.0 | +4% more credits |\n| $50 | 1,375 | 27.5 | +10% more credits |\n| $100 | 3,000 | 30.0 | +20% more credits |\n| $250 | 8,000 | 32.0 | +28% more credits |\n| $500 | 17,000 | 34.0 | +36% more credits |\n| $1,000 | 35,000 | 35.0 | +40% more credits |\n\nPacks up to $500 are open to everyone; the $1,000 pack is offered inside the app to licensed accounts. Slates Pro receives more credits than the Standard column above on every pack, for life.\n<!-- END:GENERATED credits -->\r\n\r\nCredits NEVER expire, there is no monthly reset, and failed generations refund immediately. You can also turn on auto-topup so your balance refills when it runs low.\r\n\r\n### Standard vs Pro\r\n- **Standard:** the app, every AI model, and pay-as-you-go credits that never expire, plus 12 months of updates.\r\n- **Slates Pro:** everything in Standard, plus our lowest credit rate on every pack, forever (buy the smallest pack and pay the largest pack's rate), **4K video generation**, a priority generation queue, early access to every new model on release day, and lifetime updates. The more you top up, the more the better rate adds up.\r\n\r\nEvery AI model is available on both tiers. The only capability gated to Pro is generating 4K video; 4K images are open to everyone, and exporting your timeline at 4K is available on every tier.\r\n\r\n### 30-Day Guarantee\r\nFull refund within 30 days, no questions asked.\r\n\r\n---\r\n\r\n## KEYBOARD SHORTCUTS\r\n\r\n| Key | Action |\r\n|-----|--------|\r\n| Space | Play/pause |\r\n| V | Select tool |\r\n| C | Razor tool |\r\n| S | Slip tool / snap toggle |\r\n| M | Add marker |\r\n| Left/Right | Frame-by-frame |\r\n| Up/Down | Seek +-1 second |\r\n| Ctrl+Z | Undo |\r\n| Ctrl+Shift+Z | Redo |\r\n| Ctrl+Plus/Minus | Zoom timeline |\r\n| Delete/Backspace | Delete selected clip |\r\n| Escape | Close modal/viewer/slideshow |\r\n| Ctrl+Enter | Submit generation |\r\n| Home/End | Jump to timeline start/end |\r\n| Page Up/Down | Jump by screen width |\r\n\r\n---\r\n\r\n## OFFLINE USAGE\r\n\r\nThe app launches and works offline for everything except AI generation and login. Specifically:\r\n\r\n**Works offline:** Opening projects, viewing all assets (images/videos), editing timeline (trim, reorder, split clips), adding markers, slideshow playback, FFmpeg export to MP4, DaVinci XML export.\r\n\r\n**Requires internet:** AI generation (all models), login/signup, credit purchases, credit balance sync, license validation (only checked on first generation attempt per session, then cached), auto-updater.\r\n\r\nIf you lose internet mid-session, you can keep editing and exporting. Generation will fail until connectivity returns.\r\n\r\n---\r\n\r\n## GENERATION RECOVERY\r\n\r\nIf the app closes during a generation: on restart, Slates detects in-flight jobs, polls the AI provider, and downloads completed results automatically. Nothing is lost. Recovering generations show at 5% in the queue until status is confirmed. Works for every model.\r\n\r\n---\r\n\r\n## TROUBLESHOOTING\r\n\r\n**\"Insufficient credits\"** \u2014 Your credit balance is too low for this generation. Buy more credits in the app (packs from $10 to $500) or turn on auto-topup. The exact cost of any generation is shown on the Generate button before you commit.\r\n\r\n**\"Input was rejected by Kling\"** \u2014 Image may not meet quality requirements (character visibility, proportions, content policy). Try a different image or prompt.\r\n\r\n**\"Failed to upload image to FAL CDN\"** \u2014 Network issue during reference image upload. Check internet connection, retry.\r\n\r\n**\"Generation failed\" / \"Proxy generation failed\"** \u2014 Generic error from the AI provider. Usually temporary. Retry. If persistent, try a different model.\r\n\r\n**\"Source asset not found\" / \"Source video asset not found\" / \"Target image asset not found\"** \u2014 The image or video you're trying to use was deleted or moved. Re-import or select a different asset.\r\n\r\n**\"Invalid audio source\"** \u2014 Lip sync: either enter TTS text or upload an audio file. One is required.\r\n\r\n**\"TTS response missing audio URL\"** \u2014 Text-to-speech failed during lip sync. Retry.\r\n\r\n**Generation stuck** \u2014 Restart app. Recovery system polls providers and picks up where it left off.\r\n\r\n**API rate limit** \u2014 Too many requests (limit: 20 generations/minute). Wait 1-2 minutes, retry.\r\n\r\n**Project files missing** \u2014 Project folder was moved/deleted outside the app. Use project relocation in Settings to re-point to the correct folder.\r\n\r\n**License shows \"revoked\"** \u2014 Contact support. Character sheets and environment grids unavailable until resolved.\r\n\r\n**Session expired** \u2014 Magic link session timed out. Log in again via Settings.\r\n\r\n**iPhone photos won't import** \u2014 iPhones save photos as HEIC/HEIF format, which Slates doesn't support. Convert to JPEG or PNG first (most photo apps and online converters can do this).\r\n\r\n---\r\n\r\n## PRIVACY & DATA\r\n\r\n- Generated files stay on YOUR machine. Slates servers never store your videos/images.\r\n- No prompts logged server-side.\r\n- File uploads go directly to the AI provider via pre-signed URLs. Slates servers never buffer your media.\r\n- Server stores only: email, license status, credit balance, transaction history, session tokens.\r\n- Stripe handles all payment data. Slates never sees your card number.\r\n\r\n---\r\n\r\n## SYSTEM REQUIREMENTS\r\n\r\n- Windows 10/11 or macOS 12+\r\n- Internet connection required for AI generation (not for editing/exporting)\r\n- Disk space for project files (AI videos are typically 5-50MB each)\r\n- FFmpeg bundled with app (no separate install needed)\r\n- No GPU required (all AI processing happens in the cloud)\r\n\r\n---\r\n\r\n## COMMON TASKS (STEP-BY-STEP)\r\n\r\n### Generate an Image\r\n1. Open the floating prompt box (visible on every page).\r\n2. Enter your prompt describing the image.\r\n3. Select an image model (Nano Banana 2 recommended).\r\n4. Choose aspect ratio and resolution.\r\n5. Press Ctrl+Enter or click Generate.\r\n\r\n### Generate Video From an Image\r\n1. In the prompt box, attach a start image.\r\n2. Write a prompt describing the desired motion/action.\r\n3. Select a video model (Kling V3.0 Omni in ingredients mode recommended).\r\n4. Choose duration (Kling bills per second from 3s up, so shorter is always cheaper), aspect ratio, and resolution.\r\n5. Click Generate.\r\n\r\n### Use a Character Reference for Consistency\r\n1. Create a character in your project (name + description).\r\n2. Either generate a character sheet OR manually assign a single image as the character reference.\r\n3. In the prompt box, type `@` and select your character from auto-complete.\r\n4. The reference image is attached automatically. Generate normally.\r\n\r\n### Export to DaVinci Resolve for Color Grading\r\n1. In the video editor, finalize your timeline (clips, markers, timing).\r\n2. Click Export \u2192 DaVinci Resolve XML.\r\n3. Choose output location. File saves to exports/ directory.\r\n4. In DaVinci Resolve: File \u2192 Import \u2192 Timeline. Select the XML file.\r\n5. Your timeline loads with all clips, properties, and markers intact. Grade and export.\r\n\r\n### Extract a Still Frame From a Video\r\n1. Hover over any video clip in the gallery.\r\n2. Camera icon = extract **current frame**. Dropdown arrow next to it = **First frame** or **Last frame**.\r\n3. Extracted image saves to your project gallery. Use as start/end image for I2V, character reference, or storyboard frame.\r\n\r\nKey workflow: extract a clip's last frame \u2192 use it as the start image for the next generation \u2192 seamless visual continuity between scenes.\r\n\r\n### Buy More Credits\r\n1. Open Settings \u2192 Credits, or the credit badge in the top nav.\r\n2. Pick a pack (packs run from $10 to $500, plus a $1,000 pack offered in the app to licensed accounts; bigger packs give more credits per dollar).\r\n3. Pay via Stripe. Credits are added to your balance instantly and never expire.\r\n4. Optional: turn on auto-topup so your balance refills automatically when it runs low.\r\n\r\n---\r\n\r\n## COMMON QUESTIONS\r\n\r\n**Q: Which model should I use for most videos?**\r\nA: Seedance 2.0 is the default and the one to reach for when physics, scale, effects or hero shots matter. For everyday shots built from a start image, Kling V3.0 Omni in ingredients mode is the best balance of cost and quality, which is why the step-by-step guides above use it.\r\n\r\n**Q: What's the best image model?**\r\nA: GPT Image 2.5 is the strongest image model in the app, and the best available anywhere right now. It holds a long instruction more faithfully than anything else here, and it is the only model that renders words in the picture reliably. The two seats cost the same, so the only trade is time: **Flare** is the fast one and is best for drafts and exploring, **Sunburst** is the higher-quality one and is what finals, hero frames and reference-heavy edits should end up on. The quality knob has five settings spanning about 36\u00D7 from cheapest to dearest \u2014 `max` is four times `high`, though the smaller steps are uneven: **`medium`** is for drafts and where iteration belongs, **`high` is the default** and where most finished work should sit, and **`max` is the best output you can get** \u2014 finished frames, client deliverables, anything with exact text. `xhigh` sits just under `max` for about half the price and is worth trying before you jump to the top. Going past `high` should be a deliberate choice rather than a habit. Go to **4K only once you already know your references and your prompt are solid** \u2014 it is the most expensive setting on the model and the worst place to discover the composition was wrong. Prove the shot at 3K first, then re-run the settled prompt at 4K. If you used GPT Image 2 before, note the quality names all shifted by one: its `medium` is now `high`, its `high` is now `max`. Nano Banana 2 is the picker's DEFAULT rather than the best: it is fast, takes the most reference images, and is the only model with grid exploration. Nano Banana Pro is the hero-frame step up when composition, cinematic lighting and skin have to be perfect. Prices for every tier are in the model table above.\r\n\r\n**Q: Do I need to set up API keys?**\r\nA: No. There are no API keys in Slates \u2014 every generation runs on Slates Credits, which come with your license and never expire.\r\n\r\n**Q: How much does a generation cost?**\r\nA: It depends on the model, resolution, and length. The exact credit cost is always shown on the Generate button before you commit, so there are no surprises.\r\n\r\n**Q: Can I use Slates offline?**\r\nA: Yes for viewing projects, editing timeline, and exporting. No for AI generation -- that requires internet.\r\n\r\n**Q: Do credits expire?**\r\nA: No. Credits never expire.\r\n\r\n**Q: What happens if I close the app during a generation?**\r\nA: Nothing is lost. On restart, Slates detects in-flight jobs and downloads completed results automatically.\r\n\r\n---\r\n\r\n## FEATURES NOT IN SLATES\r\n\r\nThe following are NOT available. Do not suggest them:\r\n\r\n- Bring-your-own API keys (BYOK) \u2014 every generation runs on Slates Credits; there is no key-entry option\r\n- Local/on-device GPU inference (all AI runs in the cloud)\r\n- Built-in music generation (use external tools like Suno, import audio)\r\n- A voice picker for Seed Audio \u2014 you describe the voice you want in words instead. The preset voice shelf belongs to the Voice lane (Inworld TTS-2), and Kling Lip-Sync keeps its own small fixed list: six English/UK voices plus a storyteller, with a speed control\r\n- Automatic video editing from a script\r\n- A prompt \"Enhance\" button \u2014 ask Studio Agent to rewrite a prompt instead\r\n- A settings/gear panel on the prompt box \u2014 every parameter is a dropdown on the bar\r\n- Cloud project storage (all files are local)\r\n- Real-time collaboration / multi-user editing\r\n- Mobile app (desktop only: Windows and macOS)\r\n- HEIC/HEIF image import (convert to JPEG/PNG first)\r\n- Storyboard JSON export (import only)\r\n\r\n---\r\n\r\n## VERSION\r\n\r\n<!-- BEGIN:GENERATED version -->\nSlates Reference Version: 1.5.6\nLast Updated: 2026-09-09\n\nThis document is generated. Its source of truth is `slate/docs/slates-llm-manual.md`; its model tables and credit costs are derived from the Slates model registry and pricing tables at build time, so they cannot be typed by hand.\n\nIf the user asks about a feature not documented here, it may have been added after this version. The current copy is always at <https://slates.video/slates-reference.md>.\n<!-- END:GENERATED version -->\r\n\r\nIf this document didn't answer your question, email hello@slates.video so we can help and improve the app.\r\n\r\n</slates_reference>\r\n";
|
|
2
2
|
//# sourceMappingURL=content.d.ts.map
|