@slatesvideo/shared 0.6.8 → 0.6.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -9,8 +9,8 @@ export const SKILLS = {
9
9
  "slates-dialogue-blocking": "---\nname: slates-dialogue-blocking\ndescription: Keep multiple characters spatially consistent across cuts — seating, screen direction, eyelines, the 180-degree rule — by blocking the scene in 3D first. Use for any multi-character dialogue scene, conversations around a table or in a car, or when generated characters swap seats, change sides, or look the wrong way between shots.\n---\n\n# Dialogue blocking — six people who stay where you put them\n\nThe hardest thing to generate, and the case where previs beats raw prompting by the widest margin.\n\n## Why this is hard\n\nEvery cut is an independent guess unless something forces agreement. Prompt a six-person conversation four times and you get four different seating charts: characters swap places, the 180-degree line breaks, and nothing cuts together. The failure is not aesthetic — the shots are simply unusable as an edit, and you find out only after paying for all four.\n\n**The blocking fixes it structurally.** Positions exist in 3D, so every camera sees the same arrangement, and consistency stops being something the model has to remember.\n\n## Build\n\nFollow `slates-previs-blocking` and add these.\n\n### Seated proxies, colour-coded\n\nSimple seated shapes. **Do not animate the heads** — a proxy head turning the wrong way is worse than one that never turns.\n\nGive each character a distinct viewport colour and write the mapping down. This is the identity channel:\n\n> red = the boss · green = the kid · blue = the driver · yellow = the fixer · purple = the cousin · cyan = the nephew\n\nThat mapping goes verbatim into the generation prompt. It is what lets the model bind a grey body to a character sheet across four cuts.\n\nGive every proxy a material and set `mat.diffuse_color` to its identity colour — the blocking render pins Workbench to `MATERIAL` shading, so **the material's `diffuse_color` is what reaches the clip**. Set `object.color` to the same value too, so the user's viewport matches what renders. See `slates-previs-blocking` for the snippet.\n\n### Fix the geography, then never move it\n\nPlace people once. Write down who sits where relative to the camera's opening position, in words, because that sentence is going into the prompt:\n\n> Across the table, facing camera: yellow dead centre, purple far left, blue and green to the right.\n\n### The camera plan\n\nPer `slates-camera-language`, with two things specific to dialogue:\n\n- **Below shoulder height, slow rail glides.** Eye-level-and-above reads as surveillance.\n- **Decide who owns the near foreground in each cut and honour it.** A shoulder in frame is a spatial anchor; a different shoulder in the next cut relocates the whole room.\n\nThe move that earns the most: **a gaze handoff without a cut** — the camera keeps gliding while the target hands off across the table, face to face, slowing on each but never stopping. Build it by keyframing the Track To target's position between subjects.\n\n### Crossing behind someone\n\nA head wiping frame during a move is a strong depth cue. It is also a spatial claim, so pick who gets crossed and say so — *the camera crosses directly behind cyan's back mid-shot and his head wipes the frame once.*\n\n## The prompt\n\nEverything in `slates-blocking-to-prompt`, plus these blocks.\n\n### Geography — restate it as a rule\n\n> TABLE GEOGRAPHY — do not deviate: the camera is never parked behind red. Only in the opening seconds does his dark shoulder hang at the near frame RIGHT edge, and it slides out as the camera travels LEFT. The true near-foreground of this shot is cyan: the camera crosses directly behind him mid-shot. After the opening seconds red is gone from the foreground, and the camera never travels behind anyone except cyan.\n\n### Screen direction, and the mirror that is not a swap\n\nThe 180-degree rule survives on its own in the blocking. What breaks is the model **\"correcting\" a legitimate mirror** — when the camera faces back through a scene, sides invert, and that inversion is correct. Say so explicitly or it gets flipped:\n\n> The driver's seat is on the LEFT for the entire timeline; this layout never mirrors or flips. When a camera faces BACKWARD into the car, screen sides mirror naturally: the driver reads on the RIGHT of frame, the passenger on the LEFT — that is correct left-hand drive, not a swap. They never swap seats or roles anywhere in the timeline.\n\nThen compress it into the HOLD block: *(backward camera mirrors them: he right of frame, she left)*.\n\n### Presence\n\n> A seated person stays drawn even when partially occluded — in every interior frame some part of each seat's owner is visible: a hand, an arm, a shoulder, a head above the bolster. Every occupied seat visibly holds its person.\n\n### Keep everyone alive\n\nThree orthogonal layers. Without them, whoever is not speaking freezes:\n\n- **ONGOING BUSINESS** — a small continuous physical action per character, running whether or not they are speaking. *turns his glass a quarter every few seconds · thumbs a lighter without lighting it.*\n- **BACKGROUND LIFE** — soft-focus, low contrast, never pulls attention, never crosses in front of a speaking face.\n- **SCENE EVENT** — the unnamed thing everyone is playing but nobody says. One line, repeated verbatim in every character's direction: *keep tomorrow sounding like a fishing trip.*\n\n### Acting, per character\n\nSame six slots each. Terse:\n\n```\nACTING TASK — <character>\n SCENE DIRECTION (shared, unspoken): <the same line for everyone>\n MOTIVE (his fuel): <what he wants underneath>\n GOAL: <what he wants in this scene>\n OBSTACLE: <what is in the way>\n TACTIC: <how he goes about it>\n Moment to moment: <2-3 beats keyed to timestamps>\n (Safety: gaze always engaged in the task — never a frozen, glassy,\n unfocused stare; natural blink cadence.)\n```\n\nThat safety line is not filler. Dead eyes are the characteristic failure of generated faces in dialogue, and naming it is what prevents it.\n\n### Split any strong emotion into phases\n\nThe other characteristic failure is a face that strikes one extreme expression and holds it for the whole shot — a mouth stuck open for three seconds. Give the beat two phases with a hinge, and name the failure you are excluding:\n\n> PHASE 1 (17.3-18.6s) — she SCREAMS at him, mouth wide, eyes huge, hand clamped on the grab handle. PHASE 2 (18.6-19.9s) — the scream breaks off: she shuts her eyes tight and CLOSES her mouth, both hands now on the handle, head ducked, braced. Scream, then brace — never one frozen open mouth held through the whole shot.\n\nThe hinge timestamp is what makes it a performance instead of a pose.\n\n### Dialogue must not restructure the edit\n\nBoth of these, verbatim, every time:\n\n> DIALOGUE NEVER CREATES SHOTS: spoken lines happen inside the reference's takes exactly as blocked — no cutaways to a speaker, no reverse shots, no added close-ups. If a line plays while the camera is elsewhere, the line stays off-screen audio.\n\n> OFF-SCREEN VOICES RULE: a line marked off-screen must STAY off-screen — never show the speaker, never move him into frame, never route the camera behind him because he spoke.\n\nA sentence may cross a cut. Say so where it does: *the sentence does not pause for the edit.*\n\n## Model routing\n\nDialogue directed as separate layers (voices, scene sound, score) is **minimax-h3**'s seat; it also takes declared reference relationships, which suits a colour-coded cast. Native synced audio is **Veo**'s niche. seedance-2.5 carries the reference-video capacity. Route per `slates-model-selection` and read the chosen model's prompting skill before writing the audio block.\n\n## Checklist\n\n- [ ] Colour→character mapping written down and pasted into the prompt\n- [ ] Heads not animated in the blocking\n- [ ] Seating stated as a geography rule\n- [ ] Foreground owner named per cut\n- [ ] Mirror-is-not-a-swap clause present if any camera faces back through the scene\n- [ ] Presence rule present\n- [ ] Ongoing business, background life and scene event all specified\n- [ ] Acting task per character, safety line included\n- [ ] Any strong emotion split into phases with a hinge timestamp\n- [ ] Dialogue-never-creates-shots and off-screen-voices rules present\n\n## Related\n\n`slates-previs-blocking` · `slates-camera-language` · `slates-blocking-to-prompt` · `slates-character-identity` (the sheets) · `slates-prompting-minimax-h3`\n",
10
10
  "slates-direct-response-ad": "---\nname: slates-direct-response-ad\ndescription: Build a 30-second hyper-motion direct-response ad in Slates from a product image and brief. Composes upload → storyboard → frame gen → motion gen → timeline → export. Use when the user drops a product image and asks for \"an ad\", \"a promo video\", \"a TikTok ad\", \"an Instagram ad\", a launch video, or any short-form direct-response video built around a product.\n---\n\n# Direct-response ad — Slates workflow\n\n🚨 **Wrong skill if a PERSON talks to camera.** This file builds a product-led hyper-motion spot — product hero frames, punchy cuts, no presenter. **A creator-style ad where a synthetic person speaks to the lens is a different discipline with an inverted rulebook (ugly on purpose, one shot per generation, plate-before-video): read `slates-ugc-influencer-ad`.** Route on the presence of a talking person, not on the platform.\n\nYou are building a 30-second hyper-motion direct-response ad. The user has handed you a product image (or product URL) and a short brief. Slates desktop is open on the second monitor; the user watches it populate as you work.\n\n**Hard rules**\n\n- Always estimate cost before generating. Use `slates_estimate_generation_cost` and surface the total.\n- All Slates generation routes through Slates Credits, period (BYOK is retired) — don't suggest \"use your own keys\" workarounds.\n- Default model: `nano-banana-2-2k`. For close-up product hero frames step up to `4k` only if the user asks.\n- Hyper-motion = punchy cuts, 4 frames in 30 seconds, ~7s each. Don't over-storyboard.\n\n## Workflow\n\n### 1. Set up the project\n- Create a project named for the product (`slates_create_project`).\n- If the user gave a product image as a file path, upload it (`slates_upload_reference_image`).\n- If they pasted base64 / a data URL, use the same op with `dataUrl`.\n\n### 2. Generate the storyboard frames\nBuild exactly 4 frames in this order:\n\n| # | Beat | Visual goal |\n|---|------|-------------|\n| 1 | Hook | Hyper-close-up of the product, dramatic light, motion blur edge |\n| 2 | Lifestyle | Real person using/wearing/holding the product, eye contact |\n| 3 | Problem→solution | The before/after moment that justifies the buy |\n| 4 | CTA | Clean product hero with mental room for an overlaid CTA in editing |\n\nFor each frame:\n1. Draft a tight 1-2 sentence prompt (visual only — no copy text in the image).\n2. Reference the product upload's URL or asset ID for visual fidelity.\n3. Call `slates_generate_image` with that prompt + reference. **You see the result inline — evaluate it.**\n4. If it's wrong: refine prompt, regenerate. If it's right: bind it as a frame in the storyboard (`slates_add_frame`).\n\n### 3. Build the storyboard\n- `slates_create_storyboard` named \"30s ad — v1\".\n- Default scene already exists. Add 3 more scenes (\"Hook\", \"Lifestyle\", \"Problem-Solution\", \"CTA\") via `slates_add_scene`, or just add all 4 frames to the default scene.\n- For each generated image, add a frame referencing the asset id (`slates_add_frame`).\n\n### 4. Hand back to the user\n- Surface estimated total credits spent.\n- Tell the user the storyboard is ready and they can either:\n - **In Slates desktop:** click each frame to generate motion (the existing UI handles motion generation).\n - **Continue here:** ask you to keep going.\n\n### 5. If they say keep going — motion, assembly, export\n- Generate motion per frame with `slates_generate_video` (`firstFrameAssetId` = the frame's asset, `background: true`), routed per `slates-model-selection` (Kling 3.0 std 8s by default; Seedance 2 for any physics-heavy beat like the hook). Submit all four, then poll `slates_get_generation_status` until each completes (1-5 min).\n- Assemble: `slates_add_clip_to_timeline` for each completed clip in beat order (Hook → Lifestyle → Problem-Solution → CTA). Verify with `slates_get_timeline`; fix order with `slates_reorder_clips`.\n- Export: `slates_export_video` to an absolute `.mp4` path (default `<slates_get_project_directory>/exports/<product>-ad.mp4`), then `slates_reveal_file` so the user sees the file.\n- Full pipeline doctrine (batch cost authorization, model mixing, multi-take selection): `slates-one-prompt-film`.\n\n## Anti-patterns\n\n- **Don't** generate text overlays in the image. Slates renders captions/CTAs at the editor stage.\n- **Don't** burn credits on slot-machine prompting. If the first generation is off, refine the prompt; don't just regenerate.\n- **Don't** skip the cost estimate. Confirm with the user above ~17 credits.\n- **Don't** invent visual specifics about the product (colors, textures, angles) that aren't in the reference image. Reference-anchored prompts only.\n\n## Voice\n\nThe ad lives or dies on the hook frame. Tight, sensory, no fluff. Match the user's brand. Default tone is \"scroll-stopping\" not \"informative.\"\n",
11
11
  "slates-edit-and-iterate": "---\nname: slates-edit-and-iterate\ndescription: Iterate on an existing Slates asset — re-evaluate, refine prompt, regenerate or edit. Use when the user has an existing generated image in Slates and wants to \"tweak it\", \"change one thing\", \"make it warmer\", \"remove the second person\", or any other surgical refinement instead of full regeneration.\n---\n\n# Edit and iterate — Slates workflow\n\nThe user already has a generated image in Slates and wants to refine it. The vision-feedback-loop skill defines the general pattern; this skill is the specific recipe for \"I have asset X, here's what's wrong with it.\"\n\n## 🔴 The master rule — an edit is a LEAF, not a node\n\n**Never re-edit an edit. Always go back and re-edit the master.**\n\nEvery edit model silently re-renders the **whole frame**, not just the region you named. So the parts you didn't ask to change come back slightly different every pass — softer texture, drifted colour, mushier fine detail. It is barely visible after one edit and obvious by the second. Chaining edits compounds the damage and there is no way to undo it, because each generation *is* the new source.\n\nThe fix is structural, not a matter of care:\n\n- **Want two changes?** Make them in ONE edit off the master, or make them as two separate edits **both taken from the master**, then keep whichever you prefer.\n- **An edit came back wrong?** Do NOT edit the result to fix it. Discard it and re-edit the master with a better instruction.\n- **Only the changed region is worth keeping?** That is a compositing job — the edit supplies the new region, the untouched master supplies everything else.\n\nSlates records this: an edit result carries `sourceAssetIds` pointing at the asset it was made from, so **you can tell whether the thing you are about to edit is itself an edit.** Check before you edit — `[Edit]`-prefixed prompts and a populated source lineage both say \"this is a leaf; go back to its parent.\"\n\n## Workflow\n\n### 1. Pull the current asset back into context\n- The user references an asset by id, frame number, or \"the latest one.\"\n- Resolve to an asset id (`slates_list_assets` if needed).\n- `slates_get_asset_image` with that id to load it inline. **You see the image.**\n\n### 2. Identify the delta\nThe user's request is one of:\n- **Surgical** — \"remove the second figure\", \"make the sword red\", \"swap the background to a bamboo forest\".\n- **Aesthetic** — \"warmer light\", \"more dramatic\", \"softer focus\".\n- **Compositional** — \"wider shot\", \"lower angle\", \"centered subject\".\n- **Wholesale** — \"actually let's try a totally different look.\"\n\n### 3. Pick the right tool\n| Delta type | Approach |\n|---|---|\n| Surgical | `slates_edit_image` — `sourceAssetId` = the original, `prompt` = the change only (\"remove the second figure\"), not a re-description of the whole image. |\n| Aesthetic / compositional | `slates_generate_image` with the original in `referenceAssetIds` + a refined prompt. Don't re-roll from scratch. |\n| Wholesale | New prompt, no reference, fresh generation. Treat as a new brief. |\n\n**`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Default model `nano-banana-2` — the only edit model that also takes extra `referenceAssetIds`; `flux-2-max` / `seedream-5-lite` use their own edit endpoints and ignore references. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost above ~17 credits gates on `confirm=true`.\n\n### 4. Generate, evaluate, decide\n- Estimate cost first.\n- After generation, the result is inline. Compare side-by-side with the original (`slates_get_asset_image` again).\n- If the delta is correct: bind to the same role (frame, character identity, etc.) the original was bound to.\n- If the delta missed: one focused refinement, then regenerate. Cap at 3 tries.\n\n### 5. Hand back\n- \"Asset updated. Frame 3 now uses {new_asset_id}.\"\n- Always note what changed and what didn't, so the user can see the surgery worked: \"Lighting shifted to warmer, composition unchanged.\"\n\n## Anti-patterns\n\n- **Don't** delete the original asset until the user confirms the new one. Slates keeps both; the user picks.\n- **Don't** mix surgical and wholesale changes in one regeneration. The user said \"make it warmer\" — don't also reframe the shot.\n- **Don't** re-generate when `slates_edit_image` would work. Edits preserve composition and identity; full regen rolls the dice.\n- **Don't** edit an edit — ever. Not once, not \"just a small one.\" Go back to the master (see the master rule above). Every attempt re-renders the full frame and the degradation is cumulative and permanent.\n- **Don't** keep re-rolling the same failed edit. If three tries off the master didn't land, the brief is wrong, not the model — check in with the user.\n",
12
- "slates-model-selection": "---\nname: slates-model-selection\ndescription: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Seedance 2.5 is a SECOND SEAT beside 2.0 (30s takes, 30 references and timestamp control, but no 4K and dearer at every shared resolution — never an upgrade); MiniMax H3 is the AUTHORED-AUDIO seat (three directable sound layers in one pass, declared reference relationships, 480p-4K) with MiniMax H3 Max beside it as a faster 768p-capped, reference-free premium; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 or 9:16, 4/6/8s) and never the default.\n---\n\n# Model selection — the routing doctrine\n\nPick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.\n\n## 🔑 The meta-rule — above the table\n\nThe tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2.5 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:\n\n> **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.\n>\n> **Inspect the output at its intended crop.** A frame that holds up as a thumbnail can fall apart at the size it will actually be watched. For a location, look at atmosphere, material texture, and anchor objects; for a character, identity, skin, pose, and gradients.\n>\n> **Choose the model that PROVES that requirement** and leaves only failures you can afford to rerun or mask.\n>\n> **When the roster changes, repeat the evidence test.** Do not carry today's ranking forward on reputation.\n\n## Video routing\n\n| Job | Model | Why |\n|---|---|---|\n| **General-purpose — the default for most shots** | **Kling 3.0 std** | Cost-effective workhorse. Strong image-to-video: preserves identity, layout, and text from the start frame. 16:9 / 9:16 / 1:1, 3–15s. |\n| Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |\n| Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |\n| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |\n| The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |\n| **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs, and the only Seedance seat that **acts on timestamps** (rules in `slates-prompting-seedance-2-5` § Timestamps) — 480p / 720p / 1080p, **no 4K**, and **dearer than 2.0 at every shared resolution** (720p $0.231/s vs $0.15/s, +54%). If you want 4K, or the same resolution cheaper, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) LENGTH is the price dial, not resolution — a 30s 720p face gen is 489 credits and a 30s 1080p faceless gen is 614, against a 1,000-credit welcome grant. Quote before any take over ~10s. |\n| **The SOUND has to be directed, not just present** — a specific line delivered a specific way, scene sound that has to sit under it, and score that must stay out of the characters' world | **MiniMax H3** | The only seat where audio is authored in three separate layers in ONE pass (synchronised events in the body, ambience in a soundscape section, audience-only score in its own) rather than toggled on. 5–15s, 480p / 768p / 2K / 4K, 24fps, 32kHz stereo, 11 languages. Rules in `slates-prompting-minimax-h3`. |\n| **A reference has to keep a DECLARED amount of itself** — especially moving one subject's characteristic onto a *different* subject | **MiniMax H3** | The only seat that understands a stated retention relationship (kept whole / kept in part / transferred onto another subject / loose echo). 9 images + 3 video + 3 audio, 12 files total. 🚨 The first 5 reference images are free and every one after that costs 4 credits — pass `referenceImages` to `slates_estimate_generation_cost` before a reference-heavy job. |\n| **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p/768p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at 768p. It still animates a start frame and an end frame — image-to-video is one of the two things it is for — and since 2026-09-09 it takes the full omni-reference set too (9 images + 3 video + 3 audio), so the seats now differ on ladder and price rather than on what they accept. Never the default; never reach for it to save money. |\n| Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | Narrow, and now narrower: if the sound needs DIRECTING rather than merely existing, MiniMax H3 is the better seat. |\n\n### Named Seedance escalation triggers\n\n\"Physics matter\" is an abstract category and it under-fires. These are the beats Seedance is **observably** good at — if the shot contains one, escalate without deliberating:\n\n- **Real-time → slow-motion contrast.** The signature beat; nearly every strong clip rides it.\n- **The camera moving while debris, meteors, sparks or particles crash around the subject.** Distinctly a feature of this model, not just a thing it survives.\n- **Massive scale that has to read as genuinely huge** — not \"a big thing\", a thing whose size is the point of the shot.\n- **One continuous unbroken take.**\n\nConcrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).\n\n## Video EDIT routing (changing an existing clip)\n\n| Job | Tool | Why |\n|---|---|---|\n| **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + \"Keep everything else the same\"; long prompts destroy it (see below). |\n| **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |\n| **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but \"near-identical\" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |\n| Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. (2.5's relocate lane reaches 1080p too as of 2026-08-24, at $0.2457/s of combined input+output.) Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |\n| **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p/1080p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |\n| AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |\n\n- **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.\n- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.\n- **One change per pass, short prompts.** On Omni Flash this is documented law (\"overly descriptive prompts can lead to unintended changes\" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.\n- Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.\n\n## Motion Transfer & Lip Sync routing (Kling-only tools)\n\nBoth tools are **Kling-only**. Every entry in them is a real Kling endpoint that bolts motion or lip movement onto a finished source as a dedicated post-process.\n\n| Job | Tool | Why |\n|---|---|---|\n| Motion retarget onto a still character | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |\n| Re-voice a clip, or animate a still portrait | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |\n\n**Want the Seedance version of either?** It is not a switch on these tools — it is a normal `slates_generate_video` on `seedance-2` with the clip attached as a **video reference** and the motion or dialogue written into the prompt (\"the character from image 1 performs the exact motion from video 1\"). That routes to the same endpoint the tool would have called, with the prompt visible and editable instead of ghost-written. Single-pass conditioning genuinely beats post-hoc retargeting on fast choreography, contact, cloth and hair — and it carries native audio — so escalate there whenever fidelity matters.\n\n- Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane.\n- Faces on that route go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person (premium realface pricing).\n\n**Rules:**\n\n- **Default video = Kling 3.0 std.** Escalate to Seedance the moment the shot has physics/effects weight or is the hero moment — and say why in the plan (\"physics-heavy, routing to Seedance\").\n- **Veo is never the default.** 16:9 or 9:16 only, 4/6/8s only (and 8s only at 1080p/4K, or with reference images), and it is not the quality pick — treat it as a single-purpose tool for native-synced-audio shots. If audio can be added after (Kling lip-sync, edit stage), prefer Kling or Seedance + audio in post.\n- **9:16 vertical → Kling or Seedance by preference**, not by necessity: Veo does take 9:16 on the route Slates uses. Route away from it because it is the niche seat, not because it can't.\n- **Ratios and durations are enforced before submit.** `slates_generate_video` validates the aspect ratio, resolution and duration against the model you picked and refuses out-of-set values with the legal list — it will not silently ignore or downgrade them. The authoritative per-model sets are in the op's own param descriptions, which are generated from the capability SSOT; prefer those over any list written in prose here.\n- **Image-to-video from an NB2 start frame** (the standard pipeline) → Kling by default, Seedance when the motion is physics-heavy. Not Veo.\n- **User names a model explicitly → use it.** But if it's a mismatch for the job (crazy physics on Kling std, a 30s take on anything but Seedance 2.5, 4K on Seedance 2.5 which has none), say so in one line and offer the right route before generating.\n\n## Image routing\n\n**Video models (Kling, Seedance, Veo) cannot generate standalone images — ever.** A \"premium hero reference image\" is still an image job: it routes to an image model below, never to Seedance.\n\n- **Default: Nano Banana 2** — strongest reference HANDLING (14 refs; GPT Image now takes more, at 16, but Banana is still the one that holds many subjects coherently), best legible text, the standard start-frame generator.\n- **NB2 Lite** — the fast/draft seat: ~half NB2's price, ~2.7× faster, 1K only. Route iteration volume and drafts here; finals go back to NB2 full (2K/4K).\n- **Nano Banana Pro** — the hero-frame/typography ceiling (~2× NB2). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — feed it a full subject library.\n- **GPT Image 2.5** — two seats, `gpt-image-2-5-flare` and `gpt-image-2-5-sunburst`, **same price**. Readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. **Also the photoreal front-runner (Eric, 2026-08-24)** — it beat both Nano Banana rails head-to-head on skin realism, which is why the AI-influencer ad lane generates every plate on this line. **The seat split is SPEED vs QUALITY, not generate vs edit** (OpenAI's own rule): Flare is the small, fast model with quality *comparable to* GPT Image 2 — drafts, exploration, volume; Sunburst is OpenAI's *most capable* image model, higher quality than GPT Image 2, deliberately slower — finals, hero frames, photoreal, and multi-reference edits, where its lead is widest. **Explore on Flare, finish on Sunburst.** Five quality tiers, cheapest first — `low` (layout checks only), `medium` (drafts), **`high` (the default)**, `xhigh`, `max` (the top). Uneven: `max` is 4× `high`, `xhigh` only ~1.8× it. **16 reference images**, the schema ceiling. **Transparent backgrounds** via `backgroundMode` — free, and the only image family that offers them.\n\n 🚨 **The tier names moved when 2.5 replaced GPT Image 2, and the strings did not.** GPT Image 2's `medium` is 2.5's `high`; its `high` is 2.5's `max` — same money, one rung of renaming. The 2026-08-24 photoreal result was measured at GPT Image 2 `high`, so **the tier that reproduces it is `max`**. Nobody has re-run it on 2.5; the ranking is inherited, not re-measured.\n- **FLUX.2 Max** — photoreal texture, hex-color binding, typography, less censored.\n- **Seedream 5 Lite** — uncensored + any-resolution flat price; volume exploration when the Gemini filter is in the way.\n\n**Split rule of thumb:** readable text / panels / UI → GPT Image 2.5 (Flare to explore, Sunburst to finish); **photoreal people, finals and hero frames → Sunburst at `max`** — the 2026-08-24 result was measured at GPT Image 2's `high`, which is `max` here, and Flare only *matches* GPT Image 2 while Sunburst exceeds it; multi-reference edits where several references must all survive into one frame → Sunburst; edit-heavy work → the Banana line; drafts → GPT Image 2.5 Flare at `medium`, which now undercuts NB2 Lite on both price and resolution; uncensored or odd resolutions → Seedream/FLUX.\n\n⚠️ **This line said the opposite until 2026-08-24** — it sent photoreal *away* from GPT Image on reputation, which is the exact failure § The meta-rule above warns about. Re-run the evidence test when the roster moves. It moved again on 2026-09-09, and the ranking was carried across rather than re-measured — exactly what the meta-rule says not to trust. Treat it as a starting hypothesis for 2.5, not a receipt. **The seat choice above is likewise reasoned from OpenAI's positioning, not measured:** run Flare-`max` against Sunburst-`max` on one plate and write the answer into `slates-prompting-gpt-image-2-5`.\n\n## Audio routing\n\n**Image and video models cannot generate standalone audio, and neither audio model can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.\n\n| Job | Model | Why |\n|---|---|---|\n| **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines inside a scene | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. The continuity-bed workhorse; dialogue is performed inside the room, not cast. |\n| **One named voice saying one line** — a character's own voice, a narrator, a clean VO to lip-sync against | **Inworld TTS-2** (`inworld-tts-2`) | The prompt IS the words, spoken verbatim and billed per character. Voice = the character's clip (cloned for the take), a description, or a preset. No room tone — mix it on the timeline. |\n| **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control and a real loop mode. |\n\n**There is no music model.** A song is imported (Slates reads audio files and puts them on the timeline), not generated. A line that has to be spoken in a SPECIFIC voice is generated on Inworld TTS-2 and lip-synced against; a line that belongs to a scene is performed by Seed Audio inside it.\n\n### Named audio escalation triggers\n\n- **\"It needs to sound like a place\"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.\n- **\"Read this line\"** → Seed Audio, with the line in quotes inside the scene sentence. Re-roll until the take is right, then lip-sync against it.\n- **\"That needs a thump right there\"** → Sound Effects, with the duration set to roughly the length of the event.\n- **\"Give it a track\"** → there is no music generation. Say so and offer to lay an imported track on an audio track.\n\n**Rules:**\n\n- **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.\n- **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.\n- **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — and remember those extra seconds are billed on both surfaces.\n- **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.\n- Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`.\n\n## Cost is a tiebreaker, not the router\n\nRoute by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.\n",
13
- "slates-one-prompt-film": "---\nname: slates-one-prompt-film\ndescription: Use when the user gives ONE idea and wants a finished video out the other end — \"make me a video about X\", \"turn this idea into an ad\", \"make a short film from this\". The full pipeline: script, project, characters, storyboard, frame images, video generation, timeline assembly, MP4 export. This is the master recipe; the other Slates skills are its sub-steps.\n---\n\n# One prompt → finished film — Slates master pipeline\n\nThe user gives an idea. You hand back an MP4 on disk. Everything in between is yours, with exactly TWO mandatory user checkpoints: the creative plan, and ONE aggregated cost approval.\n\n## The pipeline\n\n### 1. Script the beats\nTurn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.\n\n🚨 **Before you fire the set, read its variety counts.** `slates_list_shots` returns the distribution with every listing — shot sizes, camera moves, durations, and any bucket repeating three or more times in a row. Read the table as a COLUMN, not as rows: if push-in is the plurality or every row says wide, the batch is wrong before a credit is spent. The craft is `slates-shot-variety`.\n\n**Surface a decision log with the plan.**\n\n<!-- @inject:decision-log -->\nWhen you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify **and that no row already records**:\n\n```\nsource phrase or declared default → what you wrote → what it resolves\n\"in a diner\" → warm, and the light is the reason → why the anchor was chosen, not what it is\n(no time of day) → late afternoon, low warm key → default; say the word and it changes\n```\n\n🚨 **Keep it to what is NOT already data — and almost everything now IS.** A Shot holds the references and their roles, the model, every param, the shot size, the camera, the prop, the action and the spoken line, and `slates_list_shots` reads the whole board back in order with its variety counts. Narrating any of those is retelling a row the user can open. **Write the Shot, and let the log carry only the judgement no field holds** — why this world, why this light, why this register.\n\n**Hard rule: never silently add weather, props, style, or camera movement.** Four of those are now FIELDS: put the value on the Shot (`prop`, `camera`, `shotSize`, `action`) so the user can read and change it, and put the *reason* in the log only when you invented it rather than being told it. The rule has not softened — it moved from narration into data, which is stronger, because a field can be corrected and a sentence in chat cannot.\n\n> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.\n<!-- @end:decision-log -->\n\nA 4-10 shot script is where you invent the most on the user's behalf — time of day, wardrobe, weather, lens feel, camera moves the brief never mentioned. The log is what makes those visible while they are still free to change.\n\n### 2. Set up the project\n- `slates_create_project` named for the piece.\n- Recurring character? Build it properly — `slates_create_character` + the `slates-character-identity` recipe — so every frame references the same identity.\n- Recurring location? `slates_create_environment`.\n- One-off shots don't need character/environment records; skip the ceremony.\n\n### 3. Storyboard skeleton and the Shots (no generation yet)\n- `slates_create_storyboard`, `slates_add_scene` per script scene.\n- `slates_create_shot` per beat — the prompt, the model, the params and the references, with the roles they carry. **A Shot needs no image**, so the entire film exists as rows before anything is paid for.\n- `slates_get_shot` reads one back COMPOSED: the prompt the model will actually receive, its numbered references, and its exact quote. Audit your own work there — you cannot approve something the request will not contain.\n- Structure first, spend second — the user catches script problems on the free skeleton, not on burned credits.\n\n### 4. ONE aggregated cost approval — then hands-off\nThe Shots ARE the quote. `slates_generate_from_shots` without `confirm` returns one itemised total for the set plus the largest single item — no hand arithmetic, no `slates_estimate_generation_cost` per call:\n\n> Plan: 6 frames at 1k 16:9 + 5 × 8s Kling 3.0 std + 1 × 8s Seedance 2 hero shot ≈ N credits total, largest single N. Proceed with the batch?\n\nPer `slates-cost-discipline` 3b: that single OK authorizes `confirm=true` for **every enumerated call in the batch** — no per-call re-asking. Re-confirm only if a call's price overruns the plan >25% or new calls get added (extra retakes, new shots).\n\n### 5. Generate frame images\nFire the image Shots with `slates_generate_from_shots` (`confirm: true` — step 4 authorized it). Slates names each reference inline as \"image N\"; you never hand-write a role label or a number. Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`, then `slates_update_shot` with `attachFrameId` so the recipe travels with the picture.\n\n**Multi-take where it matters:** for the hook shot and any shot the whole film hangs on, generate 2-4 variants (cheap model or 1k), pull them back with `slates_get_assets_batch`, pick the strongest on composition + identity, discard the rest. Don't multi-take filler shots.\n\n### 6. Generate video per Shot\nFork each bound frame's image Shot with `slates_duplicate_shot` (`model:` the video model — that is the A/B lever the op takes inline), then `slates_update_shot` the copy with `firstFrameAssetId` = the bound frame. Two calls, because `slates_duplicate_shot` forks the prompt, the model and the params; **attachments are changed with `slates_update_shot`.** Then fire the set with `slates_generate_from_shots`.\n\n⚠️ **It runs SEQUENTIALLY and blocks until the last clip lands** — a 6-shot film is one long wait, and it will usually outlast the HTTP timeout while the run keeps going. When that happens, poll `slates_get_shot` for each Shot's `generationIds` and then `slates_get_generation_status`; **never re-fire, that double-spends.** (Concurrent batch firing needs a real queue — concurrency limiting, per-item failure isolation, partial-billing semantics — and is deliberately not built yet.)\n\n**Model mixing — route per `slates-model-selection`** (details in the per-model guides):\n- **Kling V3** (`slates-prompting-kling-v3`): the DEFAULT for most shots — 16:9 / 9:16 / 1:1, 3-15s, strong start-frame adherence; std is the workhorse, Omni for multi-character dialogue.\n- **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K (4K video is Pro-only).\n- **MiniMax H3** (`slates-prompting-minimax-h3`): route here when a shot's SOUND is part of the writing — a line delivered a particular way, scene sound under it, score that must stay outside the characters' world. It authors all three in one pass, which **collapses a shot's audio pass into its video pass** and removes the separate `slates_generate_audio` step for that shot. 5-15s, 480p/768p/2K/4K. Its sibling `minimax-h3-max` is faster, tops out at 1080p, takes the same references, and costs MORE at 768p — a deliberate speed pick, never a saving.\n- **Veo 3.1** (`slates-prompting-veo-3`): niche, never the default — only when native synced audio must generate WITH the video in one gen; 16:9 or 9:16, 4/6/8s (8s only at 1080p/4K or with reference images).\n\nFailed gen? The run continues past it and **nothing is retried automatically**. Read the per-Shot error in the result, fix that Shot with `slates_update_shot`, and re-fire only it (a retry beyond the plan = announce the delta cost).\n\n### 7. Assemble the timeline\n- `slates_get_timeline` once to get the lay of the land.\n- `slates_add_clip_to_timeline` for each completed video asset **in story order** — defaults append back-to-back on the first video track, which is exactly an assembly cut.\n- Order wrong? `slates_reorder_clips` with the full clip-id list. Dropped a shot? `slates_remove_clip`, then reorder to close the gap.\n\n### 8. Export + deliver\n- Output path: ask the user, or default to `<slates_get_project_directory>/exports/<name>.mp4`.\n- `slates_export_video` (absolute path, `.mp4`; blocks while ffmpeg renders — minutes for long timelines).\n- `slates_reveal_file` so the file is literally in front of them.\n- Offer the finishing path: `slates_export_timeline_xml` → DaVinci Resolve (File → Import → Timeline) for grading, sound, and titles.\n\n### 9. Report\nShots delivered, total spent vs. approved plan, the export path, and the single best next lever (\"re-take shot 3 with a tighter prompt\" / \"add a CTA end-card\").\n\n## Hard rules\n\n- **Two checkpoints only.** Creative plan (step 1) and total cost (step 4). Everything else runs without asking — that's the product promise.\n- **Skeleton before spend.** Project + storyboard structure are free; generation isn't.\n- **Look at everything.** Every image inline, every video via `slates_get_asset_video_frames` if a clip seems off. Never assemble a timeline from clips you haven't evaluated.\n- **3-strike rule per shot.** Three failed takes on one shot = stop, show the user what you tried, ask.\n- **Consistency comes from references, not luck.** Same identity asset on every character frame; same environment refs across a location's shots.\n- **Plan in Shots, not in chat.** Every decision that ends up in a sentence you have to remember is a decision the user cannot see, price, fork or re-fire. A Shot is a row: it survives the conversation, and the user can open it in the app and fix one reference without you.\n",
12
+ "slates-model-selection": "---\r\nname: slates-model-selection\r\ndescription: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Seedance 2.5 is a SECOND SEAT beside 2.0 (30s takes, 30 references and timestamp control, but no 4K and dearer at every shared resolution — never an upgrade); MiniMax H3 is the AUTHORED-AUDIO seat (three directable sound layers in one pass, declared reference relationships, 480p-4K) with MiniMax H3 Max beside it as a faster 768p-capped premium with omni-references; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 or 9:16, 4/6/8s) and never the default.\r\n---\r\n\r\n# Model selection — the routing doctrine\r\n\r\nPick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.\r\n\r\n## 🔑 The meta-rule — above the table\r\n\r\nThe tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2.5 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:\r\n\r\n> **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.\r\n>\r\n> **Inspect the output at its intended crop.** A frame that holds up as a thumbnail can fall apart at the size it will actually be watched. For a location, look at atmosphere, material texture, and anchor objects; for a character, identity, skin, pose, and gradients.\r\n>\r\n> **Choose the model that PROVES that requirement** and leaves only failures you can afford to rerun or mask.\r\n>\r\n> **When the roster changes, repeat the evidence test.** Do not carry today's ranking forward on reputation.\r\n\r\n## Video routing\r\n\r\n| Job | Model | Why |\r\n|---|---|---|\r\n| **General-purpose — the default for most shots** | **Kling 3.0 std** | Cost-effective workhorse. Strong image-to-video: preserves identity, layout, and text from the start frame. 16:9 / 9:16 / 1:1, 3–15s. |\r\n| Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |\r\n| Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |\r\n| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |\r\n| The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |\r\n| **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs, and the only Seedance seat that **acts on timestamps** (rules in `slates-prompting-seedance-2-5` § Timestamps) — 480p / 720p / 1080p, **no 4K**, and **dearer than 2.0 at every shared resolution** (720p $0.231/s vs $0.15/s, +54%). If you want 4K, or the same resolution cheaper, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) LENGTH is the price dial, not resolution — a 30s 720p face gen is 489 credits and a 30s 1080p faceless gen is 614, against a 1,000-credit welcome grant. Quote before any take over ~10s. |\r\n| **The SOUND has to be directed, not just present** — a specific line delivered a specific way, scene sound that has to sit under it, and score that must stay out of the characters' world | **MiniMax H3** | The only seat where audio is authored in three separate layers in ONE pass (synchronised events in the body, ambience in a soundscape section, audience-only score in its own) rather than toggled on. 5–15s, 480p / 768p / 2K / 4K, 24fps, 32kHz stereo, 11 languages. Rules in `slates-prompting-minimax-h3`. |\r\n| **A reference has to keep a DECLARED amount of itself** — especially moving one subject's characteristic onto a *different* subject | **MiniMax H3** | The only seat that understands a stated retention relationship (kept whole / kept in part / transferred onto another subject / loose echo). 9 images + 3 video + 3 audio, 12 files total. 🚨 The first 5 reference images are free and every one after that costs 4 credits — pass `referenceImages` to `slates_estimate_generation_cost` before a reference-heavy job. |\r\n| **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p/768p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at 768p. It still animates a start frame and an end frame — image-to-video is one of the two things it is for — and since 2026-09-09 it takes the full omni-reference set too (9 images + 3 video + 3 audio), so the seats now differ on ladder and price rather than on what they accept. Never the default; never reach for it to save money. |\r\n| Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | Narrow, and now narrower: if the sound needs DIRECTING rather than merely existing, MiniMax H3 is the better seat. |\r\n\r\n### Named Seedance escalation triggers\r\n\r\n\"Physics matter\" is an abstract category and it under-fires. These are the beats Seedance is **observably** good at — if the shot contains one, escalate without deliberating:\r\n\r\n- **Real-time → slow-motion contrast.** The signature beat; nearly every strong clip rides it.\r\n- **The camera moving while debris, meteors, sparks or particles crash around the subject.** Distinctly a feature of this model, not just a thing it survives.\r\n- **Massive scale that has to read as genuinely huge** — not \"a big thing\", a thing whose size is the point of the shot.\r\n- **One continuous unbroken take.**\r\n\r\nConcrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).\r\n\r\n## Video EDIT routing (changing an existing clip)\r\n\r\n| Job | Tool | Why |\r\n|---|---|---|\r\n| **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + \"Keep everything else the same\"; long prompts destroy it (see below). |\r\n| **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |\r\n| **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but \"near-identical\" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |\r\n| Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. (2.5's relocate lane reaches 1080p too as of 2026-08-24, at $0.2457/s of combined input+output.) Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |\r\n| **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p/1080p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |\r\n| AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |\r\n\r\n- **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.\r\n- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.\r\n- **One change per pass, short prompts.** On Omni Flash this is documented law (\"overly descriptive prompts can lead to unintended changes\" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.\r\n- Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.\r\n\r\n## Motion Transfer & Lip Sync routing (Kling-only tools)\r\n\r\nBoth tools are **Kling-only**. Every entry in them is a real Kling endpoint that bolts motion or lip movement onto a finished source as a dedicated post-process.\r\n\r\n| Job | Tool | Why |\r\n|---|---|---|\r\n| Motion retarget onto a still character | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |\r\n| Re-voice a clip, or animate a still portrait | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |\r\n\r\n**Want the Seedance version of either?** It is not a switch on these tools — it is a normal `slates_generate_video` on `seedance-2` with the clip attached as a **video reference** and the motion or dialogue written into the prompt (\"the character from image 1 performs the exact motion from video 1\"). That routes to the same endpoint the tool would have called, with the prompt visible and editable instead of ghost-written. Single-pass conditioning genuinely beats post-hoc retargeting on fast choreography, contact, cloth and hair — and it carries native audio — so escalate there whenever fidelity matters.\r\n\r\n- Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane.\r\n- Faces on that route go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person (premium realface pricing).\r\n\r\n**Rules:**\r\n\r\n- **Default video = Kling 3.0 std.** Escalate to Seedance the moment the shot has physics/effects weight or is the hero moment — and say why in the plan (\"physics-heavy, routing to Seedance\").\r\n- **Veo is never the default.** 16:9 or 9:16 only, 4/6/8s only (and 8s only at 1080p/4K, or with reference images), and it is not the quality pick — treat it as a single-purpose tool for native-synced-audio shots. If audio can be added after (Kling lip-sync, edit stage), prefer Kling or Seedance + audio in post.\r\n- **9:16 vertical → Kling or Seedance by preference**, not by necessity: Veo does take 9:16 on the route Slates uses. Route away from it because it is the niche seat, not because it can't.\r\n- **Ratios and durations are enforced before submit.** `slates_generate_video` validates the aspect ratio, resolution and duration against the model you picked and refuses out-of-set values with the legal list — it will not silently ignore or downgrade them. The authoritative per-model sets are in the op's own param descriptions, which are generated from the capability SSOT; prefer those over any list written in prose here.\r\n- **Image-to-video from an NB2 start frame** (the standard pipeline) → Kling by default, Seedance when the motion is physics-heavy. Not Veo.\r\n- **User names a model explicitly → use it.** But if it's a mismatch for the job (crazy physics on Kling std, a 30s take on anything but Seedance 2.5, 4K on Seedance 2.5 which has none), say so in one line and offer the right route before generating.\r\n\r\n## Image routing\r\n\r\n**Video models (Kling, Seedance, Veo) cannot generate standalone images — ever.** A \"premium hero reference image\" is still an image job: it routes to an image model below, never to Seedance.\r\n\r\n- **Default: Nano Banana 2** — strongest reference HANDLING (14 refs; GPT Image now takes more, at 16, but Banana is still the one that holds many subjects coherently), best legible text, the standard start-frame generator.\r\n- **NB2 Lite** — the fast/draft seat: ~half NB2's price, ~2.7× faster, 1K only. Route iteration volume and drafts here; finals go back to NB2 full (2K/4K).\r\n- **Nano Banana Pro** — the hero-frame/typography ceiling (~2× NB2). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — feed it a full subject library.\r\n- **GPT Image 2.5** — two seats, `gpt-image-2-5-flare` and `gpt-image-2-5-sunburst`, **same price**. Readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. **Also the photoreal front-runner (Eric, 2026-08-24)** — it beat both Nano Banana rails head-to-head on skin realism, which is why the AI-influencer ad lane generates every plate on this line. **The seat split is SPEED vs QUALITY, not generate vs edit** (OpenAI's own rule): Flare is the small, fast model with quality *comparable to* GPT Image 2 — drafts, exploration, volume; Sunburst is OpenAI's *most capable* image model, higher quality than GPT Image 2, deliberately slower — finals, hero frames, photoreal, and multi-reference edits, where its lead is widest. **Explore on Flare, finish on Sunburst.** Five quality tiers, cheapest first — `low` (layout checks only), `medium` (drafts), **`high` (the default)**, `xhigh`, `max` (the top). Uneven: `max` is 4× `high`, `xhigh` only ~1.8× it. **16 reference images**, the schema ceiling. **Transparent backgrounds** via `backgroundMode` — free, and the only image family that offers them.\r\n\r\n 🚨 **The tier names moved when 2.5 replaced GPT Image 2, and the strings did not.** GPT Image 2's `medium` is 2.5's `high`; its `high` is 2.5's `max` — same money, one rung of renaming. The 2026-08-24 photoreal result was measured at GPT Image 2 `high`, so **the tier that reproduces it is `max`**. Nobody has re-run it on 2.5; the ranking is inherited, not re-measured.\r\n- **FLUX.2 Max** — photoreal texture, hex-color binding, typography, less censored.\r\n- **Seedream 5 Lite** — uncensored + any-resolution flat price; volume exploration when the Gemini filter is in the way.\r\n\r\n**Split rule of thumb:** readable text / panels / UI → GPT Image 2.5 (Flare to explore, Sunburst to finish); **photoreal people, finals and hero frames → Sunburst at `max`** — the 2026-08-24 result was measured at GPT Image 2's `high`, which is `max` here, and Flare only *matches* GPT Image 2 while Sunburst exceeds it; multi-reference edits where several references must all survive into one frame → Sunburst; edit-heavy work → the Banana line; drafts → GPT Image 2.5 Flare at `medium`, which now undercuts NB2 Lite on both price and resolution; uncensored or odd resolutions → Seedream/FLUX.\r\n\r\n⚠️ **This line said the opposite until 2026-08-24** — it sent photoreal *away* from GPT Image on reputation, which is the exact failure § The meta-rule above warns about. Re-run the evidence test when the roster moves. It moved again on 2026-09-09, and the ranking was carried across rather than re-measured — exactly what the meta-rule says not to trust. Treat it as a starting hypothesis for 2.5, not a receipt. **The seat choice above is likewise reasoned from OpenAI's positioning, not measured:** run Flare-`max` against Sunburst-`max` on one plate and write the answer into `slates-prompting-gpt-image-2-5`.\r\n\r\n## Audio routing\r\n\r\n**Image and video models cannot generate standalone audio, and neither audio model can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.\r\n\r\n| Job | Model | Why |\r\n|---|---|---|\r\n| **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines inside a scene | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. The continuity-bed workhorse; dialogue is performed inside the room, not cast. |\r\n| **One named voice saying one line** — a character's own voice, a narrator, a clean VO to lip-sync against | **Inworld TTS-2** (`inworld-tts-2`) | The prompt IS the words, spoken verbatim and billed per character. Voice = the character's clip (cloned for the take), a description, or a preset. No room tone — mix it on the timeline. |\r\n| **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control and a real loop mode. |\r\n\r\n**There is no music model.** A song is imported (Slates reads audio files and puts them on the timeline), not generated. A line that has to be spoken in a SPECIFIC voice is generated on Inworld TTS-2 and lip-synced against; a line that belongs to a scene is performed by Seed Audio inside it.\r\n\r\n### Named audio escalation triggers\r\n\r\n- **\"It needs to sound like a place\"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.\r\n- **\"Read this line\"** → Seed Audio, with the line in quotes inside the scene sentence. Re-roll until the take is right, then lip-sync against it.\r\n- **\"That needs a thump right there\"** → Sound Effects, with the duration set to roughly the length of the event.\r\n- **\"Give it a track\"** → there is no music generation. Say so and offer to lay an imported track on an audio track.\r\n\r\n**Rules:**\r\n\r\n- **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.\r\n- **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.\r\n- **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — and remember those extra seconds are billed on both surfaces.\r\n- **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.\r\n- Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`.\r\n\r\n## Cost is a tiebreaker, not the router\r\n\r\nRoute by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.\r\n",
13
+ "slates-one-prompt-film": "---\r\nname: slates-one-prompt-film\r\ndescription: Use when the user gives ONE idea and wants a finished video out the other end — \"make me a video about X\", \"turn this idea into an ad\", \"make a short film from this\". The full pipeline: script, project, characters, storyboard, frame images, video generation, timeline assembly, MP4 export. This is the master recipe; the other Slates skills are its sub-steps.\r\n---\r\n\r\n# One prompt → finished film — Slates master pipeline\r\n\r\nThe user gives an idea. You hand back an MP4 on disk. Everything in between is yours, with exactly TWO mandatory user checkpoints: the creative plan, and ONE aggregated cost approval.\r\n\r\n## The pipeline\r\n\r\n### 1. Script the beats\r\nTurn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.\r\n\r\n🚨 **Before you fire the set, read its variety counts.** `slates_list_shots` returns the distribution with every listing — shot sizes, camera moves, durations, and any bucket repeating three or more times in a row. Read the table as a COLUMN, not as rows: if push-in is the plurality or every row says wide, the batch is wrong before a credit is spent. The craft is `slates-shot-variety`.\r\n\r\n**Surface a decision log with the plan.**\r\n\r\n<!-- @inject:decision-log -->\r\nWhen you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify **and that no row already records**:\r\n\r\n```\r\nsource phrase or declared default → what you wrote → what it resolves\r\n\"in a diner\" → warm, and the light is the reason → why the anchor was chosen, not what it is\r\n(no time of day) → late afternoon, low warm key → default; say the word and it changes\r\n```\r\n\r\n🚨 **Keep it to what is NOT already data — and almost everything now IS.** A Shot holds the references and their roles, the model, every param, the shot size, the camera, the prop, the action and the spoken line, and `slates_list_shots` reads the whole board back in order with its variety counts. Narrating any of those is retelling a row the user can open. **Write the Shot, and let the log carry only the judgement no field holds** — why this world, why this light, why this register.\r\n\r\n**Hard rule: never silently add weather, props, style, or camera movement.** Four of those are now FIELDS: put the value on the Shot (`prop`, `camera`, `shotSize`, `action`) so the user can read and change it, and put the *reason* in the log only when you invented it rather than being told it. The rule has not softened — it moved from narration into data, which is stronger, because a field can be corrected and a sentence in chat cannot.\r\n\r\n> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.\r\n<!-- @end:decision-log -->\r\n\r\nA 4-10 shot script is where you invent the most on the user's behalf — time of day, wardrobe, weather, lens feel, camera moves the brief never mentioned. The log is what makes those visible while they are still free to change.\r\n\r\n### 2. Set up the project\r\n- `slates_create_project` named for the piece.\r\n- Recurring character? Build it properly — `slates_create_character` + the `slates-character-identity` recipe — so every frame references the same identity.\r\n- Recurring location? `slates_create_environment`.\r\n- One-off shots don't need character/environment records; skip the ceremony.\r\n\r\n### 3. Storyboard skeleton and the Shots (no generation yet)\r\n- `slates_create_storyboard`, `slates_add_scene` per script scene.\r\n- `slates_create_shot` per beat — the prompt, the model, the params and the references, with the roles they carry. **A Shot needs no image**, so the entire film exists as rows before anything is paid for.\r\n- `slates_get_shot` reads one back COMPOSED: the prompt the model will actually receive, its numbered references, and its exact quote. Audit your own work there — you cannot approve something the request will not contain.\r\n- Structure first, spend second — the user catches script problems on the free skeleton, not on burned credits.\r\n\r\n### 4. ONE aggregated cost approval — then hands-off\r\nThe Shots ARE the quote. `slates_generate_from_shots` without `confirm` returns one itemised total for the set plus the largest single item — no hand arithmetic, no `slates_estimate_generation_cost` per call:\r\n\r\n> Plan: 6 frames at 1k 16:9 + 5 × 8s Kling 3.0 std + 1 × 8s Seedance 2 hero shot ≈ N credits total, largest single N. Proceed with the batch?\r\n\r\nPer `slates-cost-discipline` 3b: that single OK authorizes `confirm=true` for **every enumerated call in the batch** — no per-call re-asking. Re-confirm only if a call's price overruns the plan >25% or new calls get added (extra retakes, new shots).\r\n\r\n### 5. Generate frame images\r\nFire the image Shots with `slates_generate_from_shots` (`confirm: true` — step 4 authorized it). Slates names each reference inline as \"image N\"; you never hand-write a role label or a number. Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`, then `slates_update_shot` with `attachFrameId` so the recipe travels with the picture.\r\n\r\n**Multi-take where it matters:** for the hook shot and any shot the whole film hangs on, generate 2-4 variants (cheap model or 1k), pull them back with `slates_get_assets_batch`, pick the strongest on composition + identity, discard the rest. Don't multi-take filler shots.\r\n\r\n### 6. Generate video per Shot\r\nFork each bound frame's image Shot with `slates_duplicate_shot` (`model:` the video model — that is the A/B lever the op takes inline), then `slates_update_shot` the copy with `firstFrameAssetId` = the bound frame. Two calls, because `slates_duplicate_shot` forks the prompt, the model and the params; **attachments are changed with `slates_update_shot`.** Then fire the set with `slates_generate_from_shots`.\r\n\r\n⚠️ **It runs SEQUENTIALLY and blocks until the last clip lands** — a 6-shot film is one long wait, and it will usually outlast the HTTP timeout while the run keeps going. When that happens, poll `slates_get_shot` for each Shot's `generationIds` and then `slates_get_generation_status`; **never re-fire, that double-spends.** (Concurrent batch firing needs a real queue — concurrency limiting, per-item failure isolation, partial-billing semantics — and is deliberately not built yet.)\r\n\r\n**Model mixing — route per `slates-model-selection`** (details in the per-model guides):\r\n- **Kling V3** (`slates-prompting-kling-v3`): the DEFAULT for most shots — 16:9 / 9:16 / 1:1, 3-15s, strong start-frame adherence; std is the workhorse, Omni for multi-character dialogue.\r\n- **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K (4K video is Pro-only).\r\n- **MiniMax H3** (`slates-prompting-minimax-h3`): route here when a shot's SOUND is part of the writing — a line delivered a particular way, scene sound under it, score that must stay outside the characters' world. It authors all three in one pass, which **collapses a shot's audio pass into its video pass** and removes the separate `slates_generate_audio` step for that shot. 5-15s, 480p/768p/2K/4K. Its sibling `minimax-h3-max` is faster, tops out at 768p, takes the same references, and costs MORE at 768p — a deliberate speed pick, never a saving.\r\n- **Veo 3.1** (`slates-prompting-veo-3`): niche, never the default — only when native synced audio must generate WITH the video in one gen; 16:9 or 9:16, 4/6/8s (8s only at 1080p/4K or with reference images).\r\n\r\nFailed gen? The run continues past it and **nothing is retried automatically**. Read the per-Shot error in the result, fix that Shot with `slates_update_shot`, and re-fire only it (a retry beyond the plan = announce the delta cost).\r\n\r\n### 7. Assemble the timeline\r\n- `slates_get_timeline` once to get the lay of the land.\r\n- `slates_add_clip_to_timeline` for each completed video asset **in story order** — defaults append back-to-back on the first video track, which is exactly an assembly cut.\r\n- Order wrong? `slates_reorder_clips` with the full clip-id list. Dropped a shot? `slates_remove_clip`, then reorder to close the gap.\r\n\r\n### 8. Export + deliver\r\n- Output path: ask the user, or default to `<slates_get_project_directory>/exports/<name>.mp4`.\r\n- `slates_export_video` (absolute path, `.mp4`; blocks while ffmpeg renders — minutes for long timelines).\r\n- `slates_reveal_file` so the file is literally in front of them.\r\n- Offer the finishing path: `slates_export_timeline_xml` → DaVinci Resolve (File → Import → Timeline) for grading, sound, and titles.\r\n\r\n### 9. Report\r\nShots delivered, total spent vs. approved plan, the export path, and the single best next lever (\"re-take shot 3 with a tighter prompt\" / \"add a CTA end-card\").\r\n\r\n## Hard rules\r\n\r\n- **Two checkpoints only.** Creative plan (step 1) and total cost (step 4). Everything else runs without asking — that's the product promise.\r\n- **Skeleton before spend.** Project + storyboard structure are free; generation isn't.\r\n- **Look at everything.** Every image inline, every video via `slates_get_asset_video_frames` if a clip seems off. Never assemble a timeline from clips you haven't evaluated.\r\n- **3-strike rule per shot.** Three failed takes on one shot = stop, show the user what you tried, ask.\r\n- **Consistency comes from references, not luck.** Same identity asset on every character frame; same environment refs across a location's shots.\r\n- **Plan in Shots, not in chat.** Every decision that ends up in a sentence you have to remember is a decision the user cannot see, price, fork or re-fire. A Shot is a row: it survives the conversation, and the user can open it in the app and fix one reference without you.\r\n",
14
14
  "slates-previs-blocking": "---\nname: slates-previs-blocking\ndescription: Build a 3D blocking pass in Blender, render it grey-box, and use it as a reference video so the generated shot follows a camera path you designed instead of one the model invented. Use when the user wants precise camera control, a multi-cut sequence, a one-take move, spatial consistency across shots, or says the camera keeps drifting / they keep burning credits re-rolling.\n---\n\n# Previs blocking — design the shot, then generate it\n\nThe spine of the whole workflow. Read this first; the other four previs skills are branches off it.\n\n## The mechanism (why this works at all)\n\nA text prompt asks the model to *invent* camera motion, so it invents differently every roll. You cannot iterate on a variable you do not control, so you re-roll and pay again.\n\nA **reference video** removes the invention. You build the shot in Blender as untextured proxies — a neutral grey set with colour-coded figures, free, instant, deterministic — render the camera's path to mp4, and hand the model that clip alongside the prompt. **Blender locks the motion; the model builds the world.** Iteration moves to the free half, and the paid half usually lands first try.\n\nTwo halves, and keeping them separate is the whole discipline:\n\n| Half | Lives in | Changes when |\n|---|---|---|\n| **Structure** — cuts, camera, timing, who is where | the blocking clip | you re-block |\n| **Style** — what any of it looks like | references + prompt text | you restyle (see `slates-restyle-from-blocking`) |\n\n## Before you start\n\n1. `slates_blender_status` — confirms the bridge is up and returns fps, frame range, existing camera. If it reports `connected: false`, relay its hint and stop; nothing else here works.\n2. Settle **format first**, because the blocking render *is* the film's format: fps, aspect, duration. 24fps is the default and makes cut times land on clean frames. Duration ≤ 30s (seedance-2.5's reference-video ceiling; 15s on the others).\n3. Know the shot count. \"One take\" and \"19 cuts\" are different builds.\n\n## Build order\n\nDo these in order. Each stage is verifiable on its own, and a camera built before the geometry has nothing to frame.\n\n### 1. Set the format\n\n```python\nscene = bpy.context.scene\nscene.render.fps = 24\nscene.render.fps_base = 1.0\nscene.render.resolution_x, scene.render.resolution_y = 1920, 1080\nscene.frame_start, scene.frame_end = 1, 720 # 30s at 24fps\nresult = {\"seconds\": 720 / 24}\n```\n\nFrame maths, stated once so you never redo it in your head: **frame = seconds × fps + 1**. A cut at 7.79s is frame 188.\n\n### 2. Geometry and light — grey set, coded figures, named\n\nProxies only. A person is a box or a capsule with a sphere head. A car is a stretched cube. A can is a cylinder. **The SET is neutral grey — one light, a floor and enough wall that the space reads.** Colour is reserved for the figures, where it carries meaning (below); a grey set is what makes those few colours legible as notation rather than décor. Anything you spend on materials here you pay for twice, because the model repaints every surface anyway.\n\n**Name every object for what it *is* in the story**, not `Cube.003`. The name is how you refer to it later, and it is how you keep your own timeline honest.\n\nTwo conventions that cost nothing now and save a re-roll later:\n\n- **Colour is identity.** Give each character a distinct viewport colour and *write the mapping down* — `red = the boss, green = the kid, blue = the driver`. The generation prompt will restate that mapping so the model knows which grey body is which person across cuts. Without it, characters swap.\n- **Encode facing on featureless proxies.** A box has no front. Mark one face red, the back black, the sides green, and say so in the prompt: `RED face = the direction he faces`. Otherwise the model guesses which way people are looking.\n- **Checker a surface when SCALE or SPEED has to read.** Flat grey gives a model no parallax cue, so a fast move over a featureless floor reads as slow, and a big room reads as a small one. A black-and-white checker on the ground (or the wall a camera races past) gives it something to measure against. ⚠️ **Build it as GEOMETRY, never as a Checker Texture node.** The blocking render is Workbench, which draws one flat colour per material and never evaluates a shader node tree — a `TEX_CHECKER` comes out flat grey and you lose the cue without being told. Subdivide the plane and alternate `material_index` per face. Like every other colour here it is notation, so it goes in the translation list and gets dressed over.\n\n```python\n# Two materials, alternated per face. `TILE` is the square size in metres.\ndark = bpy.data.materials.new(\"Checker_Dark\")\ndark.diffuse_color = (0.05, 0.05, 0.05, 1.0)\nlight = bpy.data.materials.new(\"Checker_Light\")\nlight.diffuse_color = (0.80, 0.80, 0.80, 1.0)\nfloor.data.materials.append(dark) # material_index 0\nfloor.data.materials.append(light) # material_index 1\n# Subdivide first (edit mode or a Subdivide modifier applied) so there ARE\n# faces to alternate — a 2-triangle plane can only ever be one colour.\nfor face in floor.data.polygons:\n cx, cy = face.center.x, face.center.y\n face.material_index = (int(cx // TILE) + int(cy // TILE)) % 2\n```\n\nAnd the identity colour on each proxy:\n\n```python\nmat = bpy.data.materials.new(\"ID_Red\")\nmat.diffuse_color = (0.8, 0.1, 0.1, 1.0) # what the blocking render draws\nobj.data.materials.append(mat)\nobj.color = (0.8, 0.1, 0.1, 1.0) # same value, for viewport parity\n```\n\nThe blocking render pins Workbench to `MATERIAL` shading, so **`mat.diffuse_color` is the value that reaches the clip** — and an object with no material at all falls back to a neutral grey, which is why an unpainted set still reads correctly. Set `obj.color` to the same value anyway: it costs one line, it makes the user's viewport match what renders, and keeping the two equal means you never have to remember which one is authoritative.\n\n### 3. Camera\n\nThe whole of `slates-camera-language`. Build the rig, then keyframe it. Then **read back what you built** with `slates_blender_scene` — its `cutSeconds` is your cut list, and it is the number you will write timings against. That field is the authoritative one on EITHER rig — marker frames when cameras are bound to markers, the active camera's own keyframes when they are not. `camera.keyframeSeconds` is empty on a marker-bound edit, which is the rig `slates-camera-language` recommends for anything past a handful of cuts.\n\n### 4. Handheld, last\n\nAdd it after the moves are right, never before — noise on top of a wrong path just hides the wrong path.\n\n### 5. Verify the cuts\n\nThe one check that catches the most damage: on a multi-cut blocking, camera position, target and focal length must all change **exactly on the cut frame, with no transition frame between**. One interpolated frame reads as a whip-pan the model will faithfully reproduce.\n\n```python\n# Every camera f-curve keyframe on a cut frame must be CONSTANT out of the\n# previous key, or the cut smears.\nfor fc in cam.animation_data.action.fcurves:\n for kp in fc.keyframe_points:\n if int(kp.co[0]) in CUT_FRAMES:\n kp.interpolation = 'CONSTANT'\n```\n\nAlso check nothing interpenetrates — proxies through floors, clones through the hero object, letters through each other. The model renders intersections as faithfully as it renders everything else.\n\n### 6. Save a backup after every stage\n\nCheap, and blocking is iterative by nature.\n\n```python\nbpy.ops.wm.save_as_mainfile(filepath=path, copy=True)\n```\n\n## Render and generate\n\n```\nslates_blender_render_blocking { projectId, fps: 24 }\n```\n\nRenders the **scene camera** through scene settings — never the user's viewport, so the result does not depend on where they left their mouse — imports the mp4 into the project, and returns `assetId` + `durationSeconds`.\n\nThen:\n\n```\nslates_generate_video {\n model: \"seedance-2.5\",\n videoReferenceAssetIds: [<the blocking asset>],\n videoReferenceSecondsEach: [<durationSeconds>],\n characterAssetIds: [...], environmentAssetIds: [...], styleAssetIds: [...],\n prompt: <written per slates-blocking-to-prompt>\n}\n```\n\n**Four inputs, and that is the entire stack:** a character sheet each, one location/style reference, the blocking clip, and a prompt written against the blocking. Resist adding a fifth.\n\nModel note: seedance-2.5 is the seat for this — 10 reference videos at up to 30s each. seedance-2 and minimax-h3 take 3 at 15s. Route per `slates-model-selection`.\n\n## Leaving holes on purpose\n\nWhere the model outperforms any blockout you could build — liquid, smoke, fire, cloth — **block a black gap instead** and say so in the prompt: `CUT 7 (14.5-17.0, black gap in the reference)`. You are reserving a slot, not forgetting one.\n\n## What not to do\n\n- **Don't texture, light or material the blocking.** Grey is the specification. The reference supplies motion; the references supply look.\n- **Don't animate what you don't need.** Heads especially — a proxy head turning wrong is worse than one that never turns.\n- **Don't build the camera before the geometry.** It has nothing to aim at, and every value you set gets redone.\n- **Don't skip reading the scene back.** Write timings from `slates_blender_scene`'s `cutSeconds`, never from what you intended to build.\n- **Don't exceed the model's reference-video ceiling.** A 40s blocking against a 30s cap silently truncates.\n\n## Related\n\n`slates-camera-language` (rigs and moves) · `slates-blocking-to-prompt` (writing the prompt against the clip) · `slates-dialogue-blocking` (multi-character continuity) · `slates-restyle-from-blocking` (one blocking, many worlds) · `slates-model-selection` (routing)\n",
15
15
  "slates-project-organization": "---\nname: slates-project-organization\ndescription: Use when the user names an asset by code (\"use IMG-A36\"), asks what a code means, or is organizing or navigating a project. Covers the asset short-code system (IMG-A12 / VID-V3 / AUD-S1 badges on every gallery card), folders for film STRUCTURE, and the typed tabs for reusable references.\n---\n\n# Organizing a Slates project\n\nSlates already gives each REUSABLE reference type its own home — the **Characters**, **Environments**, and **Styles** tabs, each with its own generation + `@mention`/`#ref` behavior. Do NOT recreate those as folders. Folders are for **structure**, never type.\n\n**Folders = where an asset sits in the FILM**, and they mirror to real subfolders on disk (`projects/<id>/…`), so a human can open the project in Resolve/Finder and navigate it like an edit. Use them for work product, not references.\n\nCreate with `slates_create_folder`; file assets with `slates_move_assets_to_folder`. Generations land in the project's active folder, so set it before a batch.\n\nConventions by project type:\n- **Short film / narrative:** `Shots` (scene stills) · `Clips` (generated video) · `Final` (the export). Use one folder per scene (`Scene 1`, `Scene 2`, …) instead when the piece has distinct locations/beats.\n- **Ad / UGC:** `Hooks` · `B-roll` · `Talking-head` · `Final`.\n\nRules of thumb:\n- Reusable cast / sets / look → leave in the Characters/Environments/Styles tabs. Don't fold them.\n- Scene stills, clips, and the final cut → file into the structural folder they belong to, as you make them.\n- One folder per asset (folders are structure). Cross-cutting status (hero take, reject, variant) is a tag concern, not a folder.\n- Keep the gallery legible: work product lives in folders; the reference scaffolding (sheets, plates, style images) stays in its tabs.\n\n## Asset codes — the shared vocabulary (IMG-A12 / VID-V3 / AUD-S1)\n\nEvery asset gets a short, stable code the moment it lands in a project, and the user sees it as the badge in the **top-left corner of every image and video card** in the gallery. This is the shared vocabulary between you and the user — it exists so neither of you ever has to quote a UUID.\n\n**The scheme:**\n- `IMG-A{n}` = images · `VID-V{n}` = videos · `AUD-S{n}` = audio.\n- Numbering is **per project, per type**, counts up from 1, and **numbers are never reused** — deleting IMG-A12 doesn't renumber anything, so a code always means the same asset forever.\n- Each asset also carries a **label**: the first ~4 meaningful words of its prompt, title-cased. Chat format is code + label: `IMG-A12 — Beach Sunset`.\n\n**How to use it:**\n- **User names a code** (\"use IMG-A36 as the reference\", \"animate VID-V3's last frame\") → resolve it via `slates_list_assets` (match the `code` field) to get the assetId, confirm back in the same vocabulary: \"Got it — IMG-A36 — Marcus Rooftop Close-Up as the first frame.\"\n- **You name assets** → ALWAYS code + label, never UUID, never \"the beach one\" (which of three?). The user matches your words to the badge by eye.\n- **User seems confused** about what a code is or how to point you at an image → explain it in one line: \"Every image and video in your gallery has a code badge in its top-left corner — like IMG-A36. Just say that code and I'll know exactly which one you mean.\"\n- **Ambiguity** (\"the sunset image\" when several exist) → pull candidates with `slates_get_assets_batch` and offer the codes: \"I see IMG-A12, IMG-A19, and IMG-A24 with sunsets — which one?\"\n",
16
16
  "slates-prompting-elevenlabs": "---\nname: slates-prompting-elevenlabs\ndescription: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before calling slates_generate_audio with model eleven-sfx — ONE short effect with an EXACT duration, or a seamless loop, billed per second. Covers describing an effect by its physical cause, the one-sound-per-generation rule, picking a duration, loops, prompt_influence, and when to use Seed Audio instead.\n---\n\n# ElevenLabs Sound Effects v2 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — ElevenLabs Sound Effects v2.** ONE short sound with an exact length, or a seamless loop. The only Slates audio surface with a real duration control and a real loop mode.\n\n**The five levers**\n1. **Describe the physical CAUSE, not the label** — `heavy oak door slams shut`, `boot scuffs on grit`, `a latch drops home`.\n2. **Name the material and the space.** The material decides the timbre and the space decides the tail: `on wet concrete`, `in a tiled stairwell`, `across an empty warehouse`.\n3. **One sound per generation.** A room with dialogue AND clatter AND ambience is one Seed Audio pass, not three effects.\n4. **Pick the duration from the cut**, not from a feeling: roughly 0.5-1s for an `impact`, 2-4s for a `whoosh`, 8-22s for a `loopable bed`.\n5. **Ask for a loop explicitly** — `seamless loop` — when the sound has to lie under a whole scene, and keep it featureless enough to survive the seam.\n\n**Examples**\n- `A heavy oak door slams shut in a stone hallway, brief reverberant tail.` (1.5s)\n- `Steady rain on a tin awning, no thunder, no wind gusts, seamless loop.` (18s)\n\n**Hard constraint:** it is billed per second and the duration is never left for the model to pick — that would make the charge non-deterministic. It is NOT a speech surface: a line in a specific voice is `inworld-tts-2`, and dialogue inside a scene is Seed Audio, which casts and performs the line in the room.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** — a label is not a sound; describe the physical cause:\n- `door sound`, `whoosh`, `footsteps`, `impact`, `ambience` standing alone\n<!-- @banned:end -->\n\nOne short sound with an exact length, carried on fal (`fal-ai/elevenlabs/sound-effects/v2`). This is the only Slates audio surface with a real duration control and a real loop mode.\n\n## Where it routes\n\n- **A single hit that has to land on a known frame** — door slam, whoosh, impact, UI blip, riser.\n- **A seamless loop** you can lay under a whole scene — rain, engine hum, crowd murmur, machine noise.\n- **NOT** layered scenes. A room with dialogue *and* clatter *and* ambience is one `seed-audio` pass, not three SFX generations.\n- **NOT** speech. Dialogue, narration and scratch VO are `seed-audio` — it casts and performs the line inside the scene.\n- **AUDIO-ONLY.** It cannot produce images or video.\n\n## THE RULES\n\n### 1. Describe the physical CAUSE, not the label\n\n```\n✗ door sound\n✓ heavy oak door slams shut in a stone hallway\n\n✗ whoosh\n✓ a thick rope swung fast past a microphone, low air displacement\n\n✗ footsteps\n✓ boots on wet gravel, slow, one person\n```\n\nMaterial + weight + surface + room. Naming all four is the difference between a usable effect and a stock-library shrug. Cap is 450 characters — you will not need them.\n\n### 2. One sound per generation\n\nThis surface makes a single event. A door, then footsteps, then a siren is three generations layered on the timeline — or one `seed-audio` scene, which is usually cheaper and always more coherent.\n\n### 3. Duration is always explicit, and it is the price\n\nSlates **always sends** `durationSeconds`. (Left null the model picks, which makes the charge non-deterministic — so it is never left null.) The window it must fall in:\n\n<!-- @inject:thresholds -->\n<!-- GENERATED from @slatesvideo/shared — do not edit between the markers.\n Source: CONFIRM_CREDITS, DEVIATION_FACTOR and the audio bounds in\n packages/shared/src/operations/index.ts. Every number here is REFUSED by an\n op when a prompt gets it wrong, which is why none of them is typed by hand\n any more: this block replaced four claims that contradicted the code. -->\n\n**The thresholds, from the code that enforces them:**\n\n- **Confirm gate:** above **17 credits** an op returns `requires_confirm` and will not\n proceed until you re-call with `confirm: true`. Below it, announce the cost once and go.\n- **Deviation pause:** the desktop Studio Agent stops and re-asks when projected generation spend\n exceeds the approved plan by more than **20%**. You do not trigger this; the app does.\n- **Seed Audio duration:** **3–120 seconds.** There is no duration\n parameter on the model — the number you pass is written into the prompt AND is what the user is\n billed. Outside that range the op refuses rather than clamping.\n- **Sound Effects duration:** **1–22 seconds**, billed per second, never left for the\n model to pick.\n\nNever quote a credit figure from memory: `slates_estimate_generation_cost` returns the real one.\n<!-- @end:thresholds -->\n\n| Kind of sound | Ask for |\n|---|---|\n| impact, hit, click | 0.5–1s |\n| whoosh, riser, transition | 2–4s |\n| loopable bed | 8–22s + `loop: true` |\n\nOver-asking pads the tail with room tone you then trim. Under-asking clips the decay.\n\n### 4. Loops\n\n`loop: true` tiles without a seam — rain, engine hum, crowd murmur, machine noise. Combine with a longer duration so the loop point is not obvious.\n\nFor a bed longer than 22s, this is the wrong surface: `seed-audio` runs to 120s in one pass.\n\n### 5. Prompt influence\n\n`promptInfluence` 0–1, default 0.3. Higher hugs your wording with less variation between takes; lower explores. Raise it when a re-roll keeps wandering off the brief; lower it when every take sounds like the same take.\n\n## Iterating\n\n- Re-rolls that keep missing = the prompt named a **label** instead of a **cause**. Rewrite it as a physical event.\n- A hit that lands but sounds wrong in the scene is usually a *room* problem — name the space (\"in a stone hallway\", \"in a padded studio\", \"outdoors, no reflections\").\n- Three failed takes means the prompt is wrong, not the seed.\n\n## Content notes\n\nElevenLabs applies its own moderation. See slates-content-policy.\n",
@@ -20,7 +20,7 @@ export const SKILLS = {
20
20
  "slates-prompting-kling-v3": "---\nname: slates-prompting-kling-v3\ndescription: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_generate_video with kling-v3.0-std, kling-v3.0-pro, or kling-v3.0-omni. Kling has dialogue + SFX + ambient native syntax (Omni adds multi-character dialogue and language codes). Multi-shot rules differ from Seedance/Veo — don't cross syntaxes.\n---\n\n# Kling V3.0 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Kling V3.0.** The general default. Define the core subjects clearly at the START and keep those descriptions identical across shots. Up to 15s, up to 6 cuts, and the strongest image-to-video identity hold in the catalogue.\n\n**The five levers**\n1. **Dialogue in quotes** — `Character says, \"exact words here\"`. On Omni, direct the voice with `Gender + Age + Voice quality + Speech rate + Emotional tone + Language`: `[Character A: Detective, mid-40s, raspy, slow cadence, weary]: \"I've seen this before.\"`\n2. **Unique speaker labels, no pronouns after the introduction.** `he`, `the agent`, any synonym causes voice drift.\n3. **Sound has real syntax** — `SFX: heavy boots on wet pavement, distant siren wailing`, `Ambient noise: city traffic`, `Background music: low cello`. Always physical-cause specific; `SFX: footsteps` is not enough.\n4. **Motion adverbs modulate energy directly** — `slowly`, `rapidly`, `gently`, `explosively`. One primary camera move per shot, never stacked.\n5. **On image-to-video, do NOT re-describe the image.** It is an anchor; prompt how the scene EVOLVES from it — movement, camera, environmental change.\n\n**Examples**\n- `A detective in a wet grey overcoat stands under a stairwell light. He steps forward slowly as the light flickers. [Character A: Detective, mid-40s, raspy voice, slow cadence, weary]: \"I've seen this before.\" SFX: heavy boots on wet concrete, distant siren wailing. Ambient noise: rain on metal.`\n- `Camera tracks right alongside a cyclist crossing a bridge at dusk. She rises out of the saddle rapidly as the grade steepens. Ambient noise: wind, tyres on wet asphalt, distant traffic.`\n\n**Hard constraint:** `Immediately` (Omni only) removes the natural conversational beat between speakers — use it when timing matters and leave it out when it does not. Kling has a real `negativePrompt` field, unlike Seedance; start from the standard block and layer scene-specific suppressions.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- `SFX: footsteps` and any label-only effect — physical-cause specificity or nothing\n- a pronoun or synonym for a speaker after the first introduction (`he`, `the agent`) — it causes voice drift; repeat the full label\n- `single continuous take` — Seedance's phrase, and it fights Kling's multi-shot\n<!-- @banned:end -->\n\nKuaishou's video model. Three tiers: `kling-v3.0-std` (general use, no audio), `kling-v3.0-pro` (higher visual quality, no audio), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).\n\nUp to 15s. Multi-shot supported (up to 6 cuts in 15s total). Strong on image-to-video — preserves identity, layout, and text from the input image well.\n\n## Subject definition rule (verbatim, fal blog)\n\n> \"Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots.\"\n\n## Dialogue syntax\n\n```\nCharacter says, \"exact words here\"\n```\n\nUse quotation marks for precise speech. Languages (Omni only): EN, ZH, JA, KO, ES.\n\n## Voice direction formula (Omni)\n\n```\nGender + Age Range + Voice Quality + Speech Rate + Emotional Tone + Language\n```\n\nExample:\n```\n[Character A: Detective, mid-40s, raspy voice, slow cadence, weary]: \"I've seen this before.\"\n```\n\nTone phrases that fire:\n- `speaking in a hushed, trembling whisper`\n- `shouting with commanding authority`\n- `clear, fearful voice`\n- `with a trembling voice, \"I'm scared\"`\n\n## The `Immediately` keyword (Omni only)\n\nWithout `Immediately`, Kling adds a natural conversational beat between speakers. With it, dialogue is back-to-back. Use when timing matters.\n\n```\n[Alice]: \"Get down!\" Immediately, [Bob]: \"Where?\"\n```\n\n## Speaker label discipline\n\nUnique labels per character. **No pronouns or synonyms after first introduction** — they cause voice drift.\n\n✅ `[Character A: Black-suited Agent]` ... `[Character A: Black-suited Agent]: \"Stop.\"`\n❌ `[Agent]... then he says...`\n\n## Multi-character dialogue (Omni)\n\n```\nAlice says in English, \"Hello!\" Then Bob replies in Spanish, \"¡Hola!\"\n```\n\n## Sound effects, ambient noise, music\n\n```\nSFX: thunder cracks, footsteps approaching\nAmbient noise: city traffic, birds chirping, ocean waves\nBackground music: tense orchestral strings, low cello\n```\n\nSFX accepts physical-cause specificity:\n- ✅ `SFX: heavy boots on wet pavement, distant siren wailing`\n- ❌ `SFX: footsteps`\n\n## Image-to-video guidance\n\n**Verbatim (fal blog):**\n> \"Treat the input image as an anchor. Kling 3.0 excels at preserving the identity, layout, and text details. Focus prompts on how the scene evolves *from* the image: subtle movements, camera motion, or environmental changes.\"\n\n**Don't re-describe what's already in the image.** Focus on motion, changes, evolution.\n\n## Multi-shot — what makes them hit\n\n**Hard cap: total duration ≤ 15s across all shots. Max 6 cuts.**\n\nHit conditions:\n- Shot labels are explicit: `Shot 1:`, `Shot 2:`\n- One primary action per shot\n- Subject described identically in each shot block\n- Camera move per shot is **one verb**, not a chain\n- Per-shot blocks: 30-60 words\n\nMiss conditions:\n- Compressing narrative into one paragraph\n- Pronoun-only references after the first shot\n- Mixing camera moves within a shot (\"pan then orbit then push in\")\n- Extreme wide → extreme close in adjacent shots without reference images\n\n## Element references (Omni)\n\nUpload 2-4 multi-angle reference photos per character/object. Tag inline:\n\n```\n@element1 is the protagonist (refs: front, side, back angles).\n@element2 is the antagonist.\n```\n\n## Reference discipline (character / environment refs)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Kling specifically\n\n- **Kling's consistency lever is \"lock the subject with a fixed label reused verbatim.\"** That is Kling's phrasing for rules 2 and 3, and it is stricter than the others: **pronoun and synonym drift breaks it**, so the exact same label must appear on every single mention — not \"he\", not \"the detective\" after you named him. Reusing the label verbatim is the whole game. Slates composes this for you from `@mentions`.\n- **Element references are the transport for rule 1** — 2-4 multi-angle photos per character/object, tagged `@element1` / `@element2` (see Element references above). The cap is 4 combined refs on the edit path.\n\n## Negative prompting — has a real field\n\nKling exposes `negative_prompt` on the fal endpoint (different from Seedance which has none). Default block to start from:\n\n```\nblurry, low quality, watermark, text overlay, distorted hands, extra fingers,\nduplicate limbs, unnatural skin texture, overly saturated colors, lens flare,\nfloating objects, inconsistent shadows, jittery, flickering, morphing face\n```\n\nLayer scene-specific suppressions on top.\n\n## Cinematic tactics\n\n- **Motion adverb precision** modulates motion energy directly: `slowly`, `rapidly`, `gently`, `explosively`\n- **Camera vocabulary that registers as instructions:** profile shot, tracking, following, freezing, panning, \"moving in sync with the subject\"\n- **One primary camera move per shot** — never stack\n\n## Tier choice\n\n- **Standard**: general use, no audio\n- **Pro**: higher visual quality, no audio\n- **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references\n\nPick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.\n\n## Benchmark prompt structure\n\n```\n[Character A: <role>, <voice quality>]: \"<line>.\" Immediately, [Character B: <role>, <voice quality>]: \"<reply>.\"\nAmbient noise: <soundscape>.\nCamera <single move>.\n```\n\nCinematic example (paraphrasing fal blog patterns):\n> \"Shot 1: Wide establishing shot of a neon-lit alleyway in heavy rain, steam rising from grates. Camera slowly tracks forward.\n> Shot 2: Medium shot of a detective in a trench coat ducking under an awning, water dripping from his hat brim. [Detective: weary, raspy]: 'I knew she'd come back.' Ambient noise: distant traffic, rain on metal.\n> Shot 3: Close-up on his eyes, narrowing as headlights flash across his face.\"\n\n<!-- slates-only -->\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with `firstFrameAssetId` or `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside cost + `requires_confirm: true`. Look at them, revise prompt if needed, then re-call with `confirm=true`. Kling Omni multi-character with several ingredient images especially benefits — confirm each character image lands cleanly before spending.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Detective Closeup`. The user sees that code as a gallery badge.\n\n- ✅ \"I'm anchoring on **IMG-A12** as the detective and **IMG-A18** as the alleyway environment — Omni will handle the line delivery in EN.\"\n- ❌ \"I'm using the detective image and the alley one...\" (which alley? Three exist.)\n<!-- /slates-only -->\n\n## Video-to-video EDIT<!-- slates-only --> (`slates_edit_video`)<!-- /slates-only --> — @Video1 / @ElementN / @ImageN\n\nKling O3 edit takes an EXISTING 3-15s clip and changes only what the prompt names — character swap, environment change, style transfer — in one pass, no masking. Original motion, camera, and audio are preserved by default. Its notation is Kling's own, different from the \"image N\" naming used everywhere else:\n\n- **`@Video1`** — the source clip (always; the transport anchors the instruction to it).\n- **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images<!-- slates-only --> (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically)<!-- /slates-only -->.\n- **`@Image1..`** — style/appearance references<!-- slates-only --> (pass as `styleAssetIds`)<!-- /slates-only -->.\n- Max **4 combined** element + image refs per edit.\n\n**Prompt shape — the change, not the whole scene:**\n\n```\nReplace the man in @Video1 with @Element1, keeping his walk cycle, the camera move, and the rain unchanged.\n```\n\n```\nEdit @Video1: turn the daytime street into a neon-lit Tokyo alley at night, wet asphalt reflections. Apply the visual style of @Image1. Keep the subject and camera motion exactly as they are.\n```\n\nRules:\n- Name what CHANGES; explicitly state what stays (\"keep the motion / camera / everything else unchanged\") — the model preserves better when told to.\n- One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).\n- Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.\n- Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.\n- Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.\n\n## Sources\n\n- [fal.ai — Kling 3.0 Prompting Guide](https://blog.fal.ai/kling-3-0-prompting-guide/)\n- [Vidguru — Kling 3.0 Omni Guide](https://www.vidguru.ai/blog/kling-3.0-omni-guide.html)\n- [AcceptPrompt — Kling 3 Prompt Guide](https://www.acceptprompt.com/blog/kling-3-prompt-guide)\n- [DataCamp — Kling 3.0 Tutorial](https://www.datacamp.com/tutorial/kling-3-0)\n",
21
21
  "slates-prompting-lip-sync": "---\nname: slates-prompting-lip-sync\ndescription: How to set up lip-sync — Kling-only (dedicated lip-sync and avatar endpoints, 5-second outputs). Read before calling slates_generate_lip_sync. Two flows — video→video re-dub and image→video avatar — with different inputs, pricing, and gotchas. Voice catalog, framing rules, audio file constraints, and which tier to pick. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.\n---\n\n# Lip-sync — setup guide\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Lip-sync (Kling only).** Two different flows with different inputs and different prices; every output is 5 seconds.\n\n**The five levers**\n1. **Pick `sourceType` deliberately** — `video` re-dubs an existing talking head (cheapest); `image` animates a still portrait (avatar-standard, then avatar-pro only on the final selected take).\n2. **The `prompt` on the avatar flows is SCENE CONTEXT, not motion direction.** Ambience, lighting, micro-expression: `Soft rim light`, `warm office`, `cool blue evening light through a window`, `gentle confident smile between sentences`, `focused intent expression`.\n3. **Clean the audio before uploading** — `noise-reduced`, `levelled`. Lip detection is sensitive, and a raw recording is the most common cause of a bad take.\n4. **Iterate on the SOURCE or the AUDIO, never on a refinement prompt** — there is not one. If the output is wrong, change the input.\n5. **Use avatar-standard for first-pass dialogue takes**, and switch to pro only once the line is locked. Facial fidelity is not visible until then.\n\n**Examples**\n- `Soft rim light, warm office, gentle confident smile between sentences.`\n- `Cool blue evening light through a window, focused intent expression.` (Or `.` — an empty prompt is fine when you have nothing to add.)\n\n**Hard constraint:** it is Kling-only and always 5 seconds. For a generated PERFORMANCE instead — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal — that is a normal Seedance video generation with the clip attached as a video reference, not a mode of this tool. A real recording, or a cloned/cast voice rendered on `inworld-tts-2`, for production; this tool's built-in TTS is for scratch.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** — the avatar prompt is scene context and motion verbs are ignored:\n- `turns her head`, `raises an eyebrow`, `hand gestures`, `nods`, `walks`\n- `reader_en_m-v1` — listed in fal's docs, returns \"Voice id not found\" in production\n<!-- @banned:end -->\n\n**This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; every entry is a real endpoint and every output is 5 seconds.\n\n| Flow | Source | Model | Cost | Use case |\n|------|--------|-------|-----------|----------|\n| Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |\n| Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |\n| Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |\n\nPick `sourceType` deliberately — it decides the pricing tier and the underlying endpoint.\n\n## Want Seedance instead? That is a video generation, not a mode here\n\nSeedance can generate the performance rather than bolting a mouth onto finished pixels — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal, and a video source keeps its own voice. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the clip (or portrait) attached as a video/ingredient reference and the dialogue written into the prompt yourself.\n\nThat is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a \"video 1\" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.\n\n- Driving clips must be 2–15s; output duration is whatever you set (4–15s).\n- Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming.\n- Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.\n\nEverything below is about the Kling tool.\n\n## Choosing video vs avatar\n\nUse **video** (re-dub) when:\n- A talking-head clip already exists (Slates-generated, recorded, or imported)\n- The mouth/face is already moving and only the audio needs to change\n- ~4 credits is hard to beat for short dialogue replacement\n\nUse **avatar** when:\n- Only a still portrait exists\n- The character needs to come alive from a single image\n- Identity + face fidelity matter (avatar-pro for hero shots, standard for everything else)\n\n## Source asset constraints\n\n### Video flow (`sourceType: 'video'`)\n- Format: mp4 or mov\n- Duration: 2–10s (lip-sync output is always 5s — long videos get trimmed)\n- Resolution: 720p or 1080p (480p will be rejected)\n- Max file size: 100MB\n- Face must be visible and roughly facing camera. Profile shots fail.\n- Existing audio is replaced.\n\n### Avatar flow (`sourceType: 'image'`)\n- Min 512×512, PNG/JPG/WebP\n- **Face occupies 60–70% of frame.** This is the single biggest avatar quality lever.\n- Eyes open, mouth neutral, looking near-camera. Side profile = bad output.\n- Single subject, clean background. Group photos confuse the face anchor.\n\n## Audio source\n\nTwo ways to drive the lips:\n\n### TTS (`audioMethod: 'tts'`)\n- Pass `ttsText` (the words spoken)\n- Optional: `ttsVoice` (default `oversea_male1`), `ttsLanguage` (default EN), `ttsSpeed` (default 1.0)\n- **Hard cap: 120 characters of text.** Longer = silently truncated.\n- Languages: EN, ZH, JA, KO, ES\n\n### Upload (`audioMethod: 'upload'`)\n- Pass `audioFilePath` — absolute path to an audio file on the user's machine\n- Format: mp3, wav, m4a, ogg, aac\n- Max 5MB\n- Duration: 2–60s (output is 5s — longer audio gets trimmed)\n- Single clean voice. Music underneath, multiple speakers, or noisy mics produce garbage lips.\n\nPrefer upload for production-quality voice. TTS for fast iteration / placeholder dialogue.\n\n## Voice catalog (TTS)\n\nReliable English voices (verified working on the fal endpoint as of 2026):\n\n| Voice ID | Description |\n|----------|-------------|\n| `oversea_male1` | Male, English — default, stable |\n| `commercial_lady_en_f-v1` | Female commercial English |\n| `uk_boy1` | Young man, UK accent |\n| `uk_man2` | Man, UK accent |\n| `uk_oldman3` | Older man, UK accent |\n| `calm_story1` | Storyteller / narrator |\n\nAvoid `reader_en_m-v1` — listed in fal.ai docs but returns \"Voice id not found\" in production.\n\nFull 48-voice list (ZH, JA, KO included): https://fal.ai/models/fal-ai/kling-video/lipsync/text-to-video/api\n\n## Speech-rate notes\n\n`ttsSpeed` range 0.5–2.0:\n- 0.8–1.0: natural conversational\n- 1.1–1.3: punchy ad delivery\n- 1.4+: rushed, clips consonants\n- 0.6–0.7: slow, weighty (good for dramatic lines)\n\nDefault 1.0 unless the line specifically calls for slower or faster cadence.\n\n## Avatar prompt usage\n\nThe `prompt` parameter on avatar-v2 (standard + pro) is **scene context**, not motion direction. The mouth animation comes from the audio — the prompt sets ambiance, lighting, micro-expression.\n\nGood:\n- `Soft rim light, warm office, gentle confident smile between sentences.`\n- `Cool blue evening light through a window, focused intent expression.`\n\nBad (the model ignores motion verbs):\n- ❌ `She turns her head, raises an eyebrow, then speaks.`\n- ❌ `Hand gestures while talking.`\n\nDefault `\".\"` is fine if you have nothing useful to add.\n\n## Tier selection — avatar standard vs pro\n\n**Use standard** when:\n- Drafts, A/B testing voices, internal review reels\n- Wide / medium shots where face isn't the focal point\n- Cost matters more than micro-expression fidelity\n\n**Use pro** when:\n- Final ads where the avatar's face fills the screen\n- The character is named / branded — identity drift kills the take\n- You're already paying tens of credits for the surrounding video pipeline\n\nDon't default to pro. The ~15-credit delta per take adds up across iteration.\n\n## Common failure modes\n\n| Symptom | Likely cause | Fix |\n|---------|--------------|-----|\n| Lip movement looks \"rubber\" / disconnected | Source face <60% of frame | Re-crop the still tighter |\n| Voice doesn't match character age/gender | Default voice id used | Pick from voice catalog |\n| Output truncated mid-word | TTS text >120 chars | Shorten or chain two takes |\n| Garbled mouth on uploaded audio | Background music / multi-voice | Use clean dialogue-only audio |\n| \"Voice id not found\" 422 | Hit `reader_en_m-v1` | Switch to `oversea_male1` |\n| Avatar eyes drift / cross | Source had closed/angled eyes | Pick a frame with neutral open eyes |\n| Generation completes but lips don't move | Profile shot / face >70° off-axis | Use a near-frontal portrait |\n\n## Cost discipline\n\n- Video re-dub at ~4 credits is the cheapest dialogue iteration in the entire Slates stack — use it for voice A/B testing\n- Avatar standard at ~14 credits is fine for medium use\n- Avatar pro at ~29 credits trips the confirm gate — explicit user OK required every time\n- All 5s. There is no shorter option.\n\n## Workflow patterns\n\n**Voice A/B test (cheap):**\n1. Generate one base talking-head video clip with Veo or Seedance (~40 credits)\n2. Run `slates_generate_lip_sync` with `sourceType: 'video'` against 3–5 different `ttsVoice` values\n3. Total cost: ~40 + (5 × ~4) ≈ 60 credits to compare voices\n\n**Brand avatar from a single portrait:**\n1. Generate or upload the hero portrait (face fills frame, eyes open, neutral mouth)\n2. Avatar standard for first-pass dialogue takes\n3. Avatar pro only on the final selected take\n\n**Avoid:**\n- Avatar pro on first iteration (waste — facial fidelity isn't visible until you've locked the line)\n- TTS for final ads (production should use real voice or cloned voice — the upload flow)\n- Uploading raw recordings — clean noise + level the file first, lip detection is sensitive\n\n## Confirm gate: cost + codes, no inline preview\n\nLip-sync is mechanical — the model re-syncs the chosen source to the chosen audio. The confirm response carries the source asset's code so you can announce it in chat.\n\n- ✅ \"Lip-syncing **IMG-A12 — Founder Portrait** to the new line. ~29 credits on avatar-pro. Confirm?\"\n- ❌ \"Using the founder image...\" (which? Three exist.)\n\nDon't second-guess the source. If the output is wrong, iterate on source choice or audio, not on a refinement prompt (there isn't one).\n\n## Sources\n\n- [fal.ai — Kling LipSync API](https://fal.ai/models/fal-ai/kling-video/lipsync/text-to-video/api)\n- [fal.ai — AI Avatar v2 Standard](https://fal.ai/models/fal-ai/kling-video/ai-avatar/v2/standard/api)\n- [fal.ai — AI Avatar v2 Pro](https://fal.ai/models/fal-ai/kling-video/ai-avatar/v2/pro/api)\n",
22
22
  "slates-prompting-ltx-2-5": "---\nname: slates-prompting-ltx-2-5\ndescription: How to prompt LTX-2.5 and LTX-2.5 Pro. Read before calling slates_generate_video with model ltx-2-5 or ltx-2-5-pro. LTX scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE — Lightricks ranks the prompt sound, camera, character detail, shot type and scene, then scene dressing, all in one flowing paragraph. It is also the catalogue's native MULTISHOT seat: one generation carries two to four connected shots holding character, light and voice across the cuts. Base ltx-2-5 is the distilled build — 720p/1080p/1440p/4K, clips of 6 to 20 seconds in EVEN steps, and the cheapest native 1080p second in Slates; ltx-2-5-pro is the full diffusion build and is NOT a superset, reaching only 1080p and 10 seconds for about a third more money. Three hazards live here: durations are even numbers only from six (there is no 5s or 7s clip), the model has NO reference endpoint at all so identity references are unavailable, and any sound not anchored to something in frame gets invented for you.\n---\n\n# LTX-2.5 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — LTX-2.5.** It scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE. Lightricks' own priority order: sound, camera, character detail, shot type and scene, then scene dressing. One flowing paragraph, not labelled sections. When a prompt sprawls, cut from the bottom.\n\n**The five levers**\n1. **Lead with sound, and anchor every sound to something in frame.** The test is \"visible, or at least locatable\" — `the rope creaks against the cleat`, `the hull knocking hollow against the fenders`, `rain on the awning`. A distant whistle is fine IF you have named the post it comes from.\n2. **Camera second**, because framing decides the visual weight of the shot — `low camera at the gunwale`, `slow drift right`, `static medium behind the counter`.\n3. **Character detail as physical ACTION**, not as adjectives about a person — `she braces a boot on the rail and hauls`, `his hands counting notes`.\n4. **It is the native MULTISHOT seat** — one generation carries two to four connected shots holding character, light and voice across the cuts. Write the cuts.\n5. **Quote dialogue and name the language and accent** — `in English with a slight German accent` — `\"We should not have come back,\" in English with a slight German accent.`\n\n**Examples**\n- `The rope creaks against the cleat as she leans back, gulls calling somewhere off the port bow, the hull knocking hollow against the fenders. Low camera at the gunwale, slow drift right. She braces a boot on the rail and hauls, twice, then stops.`\n- `A till drawer bangs shut, a fan ticks against its cage, rain on the awning outside. Static medium behind the counter, then cut to a close-up of his hands counting notes, then cut wide as he looks up at the door.`\n\n**Hard constraint:** durations are EVEN numbers from six — there is no 5s or 7s clip. There is NO reference endpoint at all, so identity references are unavailable; use MiniMax H3 or Kling when a character must hold across shots. Any sound not anchored to something in frame gets invented for you. And never write mood adjectives as sound: \"tense atmosphere\", \"a sense of dread\" and \"ominous ambience\" produce nothing usable — the fix is one more moving object in frame with a sound attached to it.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** — mood adjectives standing in for sound produce nothing usable:\n- `tense atmosphere`, `a sense of dread`, `ominous ambience`, `eerie silence`\n- an unanchored sound: name the thing in frame it comes from, or cut it\n<!-- @banned:end -->\n\nLTX-2.5 generates picture and sound **in a single pass**, with a Gemma-4 12B text encoder reading\none flowing paragraph. That single fact drives everything below: the prompt is not a shot\ndescription with audio bolted on, it is **a scene where the sound is load-bearing** — and\nLightricks' own priority order puts sound first, ahead of the camera.\n\nTwo seats, and the naming is a trap:\n\n| | `ltx-2-5` (base) | `ltx-2-5-pro` |\n|---|---|---|\n| Build | Distilled, 8-step | Full diffusion (\"Diffusion Fidelity Rendering\") |\n| Resolutions | 720p / 1080p / **1440p** / 4K | 720p / 1080p |\n| Durations | 6–20s, even steps | 6 / 8 / 10s |\n| Price | $0.09–$0.30 per second | $0.12–$0.17 per second |\n| Reach for it when | iterating, long takes, 4K delivery, batch volume | one dense final render inside 1080p and 10s |\n\n**Pro is not \"base plus more.\"** It buys picture quality on a *narrower* envelope — it cannot make\na 1440p frame and it cannot make a 12-second clip. Reaching for it out of habit costs a third more\n*and* takes away the reach.\n\n---\n\n## 1. The six parts, in priority order, in one paragraph\n\nLightricks ranks the elements of an LTX prompt like this. When a prompt sprawls, **cut from the\nbottom.**\n\n1. **Sound** — highest priority; the model scores the picture as it draws it.\n2. **Camera** — framing decides visual weight and the feel of the shot.\n3. **Character detail** — expressed as physical action.\n4. **Shot type and scene** — the action itself.\n5. **Scene dressing** — the first thing to trim.\n\nWrite it as **one flowing paragraph**, not a list of labelled sections. LTX is not Seedance (eight\nengineering slots) and not H3 (three separate audio layers) — it wants continuous prose.\n\n---\n\n## 2. Sound: anchor it or it gets invented\n\n**Write the audio line last, then go back and check every cue has a source you could point at.**\nAnything unanchored, the model invents for you.\n\nThe test is **\"visible, or at least locatable.\"** A distant whistle is fine *if* you have named the\nmarshal's post it comes from. A \"distant whistle\" with nothing to attach to is a coin flip.\n\n> the rope creaks against the cleat as she leans back, gulls calling somewhere off the port bow,\n> the hull knocking hollow against the fenders\n\n**Never write mood adjectives as sound.** \"Tense atmosphere\", \"a sense of dread\" and \"ominous\nambience\" produce nothing usable. If a scene feels thin, the fix is **one more moving object in\nframe with a sound attached to it** — never another adjective.\n\n### Dialogue\n\nQuote it, and name the language and accent:\n\n> \"We should not have come back,\" in English with a slight German accent.\n\nTwo rules that decide whether the lip sync lands:\n\n- **Give the character a beat of stillness before they speak.** The sync needs something to lock\n against; a character already mid-motion when the line starts drifts.\n- **Describe the beat structure** — when they look, how long they wait, when they speak, where they\n look afterwards.\n\nSlates pins the frame rate at 25fps, which is also what Lightricks recommends for dialogue: at 50fps\nthe performance \"pulls toward a video look.\"\n\n---\n\n## 3. Character emotion is physical\n\nThe model renders actions. It does not render adjectives.\n\n| Instead of | Write |\n|---|---|\n| she looks anxious | her jaw sets, she turns the ring on her finger twice |\n| he seems exhausted | he blinks slowly and lets his shoulder take the doorframe |\n| a tense standoff | neither moves; his thumb finds the strap and stays there |\n\n---\n\n## 4. Multishot — the thing this model is uniquely for\n\n**One LTX generation can carry several connected shots**, holding character, environment, lighting,\nvoice and style across every cut. Nothing else in the catalogue does this natively; everywhere else\nyou generate separate clips and stitch them, and identity drifts between them.\n\n**Working range is two to four shots.** Three is the comfortable stopping point.\n\nAt **every** transition you must supply four things:\n\n1. **Name the edit in the prose** — \"hard cut\", \"dissolve\", \"match cut\".\n2. **Re-establish the shot completely** — scale, angle, lens and light all reset at a cut. A cut is\n not a continuation.\n3. **Re-identify recurring characters by their original descriptor.** \"The woman in the bronze\n gown\", never \"she\". Pronouns lose the character across a cut — this is the single most common\n multishot failure.\n4. **State what the sound does at the cut.** Silence is not assumed; if the room tone should drop\n out, say so.\n\nA shape that works:\n\n> Wide establishing shot of the workshop, dust in the window light, a lathe turning somewhere off\n> frame — hard cut — macro close-up of the brass fitting as it seats, the turning noise gone,\n> replaced by a single dry click — match cut — medium shot of the woman in the bronze gown stepping\n> back, the room tone returning underneath her.\n\n---\n\n## 5. Camera: write it, don't enumerate it\n\nfal exposes a `camera_motion` enum (dolly in/out/left/right, jib up/down, static, focus shift).\n**Slates does not surface it, deliberately** — and prose is the better instrument anyway:\n\n- **A written move can be tied to a specific moment.** \"A slow push-in that settles as she reaches\n the door, then holds\" is not expressible as an enum value.\n- **For multishot it would be actively wrong** — one enum value would impose a single camera\n behaviour on three shots that each want their own.\n\nSo name the lens, the framing, the move, and **the moment the move resolves**.\n\n---\n\n## 6. The hard constraints\n\n### Durations are even numbers only, starting at six\n\n**6, 8, 10, 12, 14, 16, 18, 20.** There is no 5-second LTX clip and no odd duration of any length.\nAsking for 7s is not a rounding matter — that generation does not exist.\n\n**And the long end is 1080p-and-below only.** At 1440p and 4K the ceiling drops to **6, 8 or 10**.\n\nfal's own default is `auto`, which lets the model pick the length from the described action.\n**Slates always sends an explicit length instead**, so what you choose is what you are billed for.\nChoose the length the beat needs.\n\n### Aspect ratios: 16:9 and 9:16, and nothing else\n\nThe narrowest set in the catalogue alongside Veo. Square, 4:5 and 21:9 are not available on this\nmodel at any resolution.\n\n### Frames, not references\n\nLTX takes a **start frame** and an **optional end frame** (which generates a transition between the\ntwo). It has **no reference-to-video endpoint at all** — no identity references, no style\nreferences, no environment references, no reference video, no reference audio.\n\n**For character consistency across separate shots, use MiniMax H3 or Kling.** Within a single LTX\ngeneration, use multishot instead — that is precisely the gap it fills.\n\nIn image-to-video, **do not cut away from the opening frame too early.** You have paid for that\nframe; let it play before the first move.\n\n### Do not ask for text on screen\n\nNeither the spelling nor its stability from frame to frame can be relied on. Signage, labels,\ncaptions and lower-thirds belong in post.\n\n---\n\n## 7. Audio is free here, and that changes the routing\n\nNative synchronised audio is **included at every resolution on both seats**, with no surcharge and\nno toggle that costs money — unlike Kling, where sound is a paid dimension. A 6-second 1080p LTX\nclip **with sound** is 39 credits.\n\nCombined with 1080p at $0.13/s — the cheapest native 1080p second in Slates — this makes LTX **the\ncoverage seat**: the one to reach for when the job is many takes rather than one hero shot, when a\nsequence needs its own sound, or when the credit budget is the binding constraint.\n\nRoute away from it when you need identity references (H3, Kling), a ratio other than 16:9 or 9:16\n(Seedance, Kling), or authored multi-layer audio direction (H3).\n",
23
- "slates-prompting-minimax-h3": "---\nname: slates-prompting-minimax-h3\ndescription: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09), so the seats differ on ladder and price, not on what they accept. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; 4 free then +1 on Max), and audio written into the wrong section is dropped or duplicated.\n---\n\n# MiniMax H3 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — MiniMax H3.** The only seat where audio is AUTHORED rather than toggled: dialogue, scene sound and score are three separate sections of the prompt, generated in one pass, and putting a sound in the wrong section drops or doubles it.\n\n**The five levers**\n1. **Write the three audio layers separately** — `Scene sound:` for what is in the room, `Score:` for what only the audience hears, and the dialogue quoted inline. Section decides attribution.\n2. **Quote dialogue and name the language** — `says in English`, `speaks in Spanish`. Eleven languages are stably supported; the language is part of the instruction, not an afterthought.\n3. **Declare the reference RELATIONSHIP**, which no other seat has: `kept whole`, `partly kept`, `transferred`, or `a loose echo`. An undeclared reference is a guess.\n4. **Give a beat of stillness before a line** — `sits still for a beat, then looks up`. The sync needs something to lock against; a character already mid-motion when the line starts drifts.\n5. **Describe the beat structure** — `waits`, `then speaks`, `under the last three seconds`. H3 is a timeline, so write one.\n\n**Examples**\n- `A woman sits still at a kitchen table for a beat, then looks up. She says in English, \"You said Tuesday.\" Scene sound: a fridge hum, a spoon set down on formica. Score: none.`\n- `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, \"No es el alternador.\" Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`\n\n**Hard constraint:** the two seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p rather than 4K, takes the same 9+3+3 references, and costs MORE at the tier they share — it is a speed pick, never the cheap one. H3's top two resolution tiers are UPSCALES of the native render: judge at native. Reference images past the free allowance are a paid dimension of the cost key (5 free on base, 4 on Max) — declare the count when quoting.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- a sound written into the wrong audio section — it is dropped, doubled, or attributed to the wrong layer\n- `background music` as a bare instruction: the score is its own authored layer, audience-only, and it is named as such\n- an undeclared reference relationship — say kept whole, partly kept, transferred, or a loose echo\n<!-- @banned:end -->\n\nH3 is an **omni transformer**: it generates picture and sound in the same pass, at 24fps with\n32kHz stereo, 5–15 seconds, in 11 stably-supported languages (Arabic, Chinese, English, French,\nGerman, Italian, Japanese, Korean, Portuguese, Russian, Spanish). That single fact drives\neverything below — the prompt is not a shot description with sound bolted on, it is a **timeline\nwith three audio layers you author separately**.\n\n**Two seats, one grammar.** Everything in this file applies to both. They differ only in what the\nendpoint accepts:\n\n| | `minimax-h3` | `minimax-h3-max` |\n|---|---|---|\n| Resolution | 480p / 768p / **2K / 4K** | 480p / 768p / **1080p** |\n| References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) |\n| Frames | start and/or end | start and/or end |\n| Price at 768p | **$0.060/s** | $0.080/s |\n| Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) |\n\n**Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it\ntops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth\npaying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.\n\n**The speed is measured, not claimed** (2026-08-27, same prompt and params on both rows): a 5-second\n768p text-to-video finished in **4.8 seconds** on Max against **57 seconds** on base H3 — roughly\n**12x**, queue to finished file. fal advertises \"under 3 seconds\"; the literal claim did not hold at\n4.8s wall-clock, but the order of magnitude did. For iteration loops and client-present work that gap\nis the entire reason the seat exists.\n\n---\n\n## The one thing that makes H3 different: audio is a THREE-LAYER instruction\n\nEvery other video seat treats sound as on or off. H3 splits it, and the split is enforced by where\nyou write each thing. Get the section wrong and the sound is dropped, doubled, or attributed to the\nwrong source.\n\n| Layer | What belongs in it | Where it goes |\n|---|---|---|\n| **Synchronised events** | dialogue, singing, and any sound tied to a specific shot or action | the **body** of the prompt, on the beat it lands |\n| **Scene sound** | ambience and physical sounds that run across the whole clip — room tone, rain, traffic, a ventilation hum | the **soundscape** section |\n| **Score** | music the characters cannot hear; audience-only | the **music** section |\n\n**Three rules, all from MiniMax's own guide:**\n\n1. **Dialogue and singing NEVER go in the soundscape section.** They are synchronised events; they\n belong in the body, at the moment they happen.\n2. **Diegetic music — music the characters can hear** (a radio in the scene, a busker) — also\n belongs in the **body**, not in the score section. The score section is audience-only.\n3. **Write the score in instrumental terms, not mood words.** Name the instruments, the tempo, and\n how it develops. *\"A restrained solo-piano score at a slow tempo, sustained low cello underneath,\n no swell\"* — not *\"emotional music\"*.\n\nUse **N/A** for a section only when silence or absence is genuinely what the shot wants. An empty\nscore section is a real choice; a vague one is a wasted layer.\n\n### The shape, in the one prompt field\n\nSlates sends one prompt string, so write the three layers as labelled paragraphs in this order:\n\n```\n[Shot 1] Live-action, cinematic. A medium-wide shot frames a baker opening the shutters of a\nsmall street bakery before sunrise. The camera pushes in with small amplitude at slow speed as\nthe middle-aged baker with a calm, slightly raspy voice places a fresh loaf on the counter and\nsays: \"First batch of the morning.\" [Shot 2] At 00:05.000, the camera cuts to a close-up of\nsteam rising from the sliced bread while his final words carry over from the previous shot.\n\nSoundscape: wooden shutters scrape open over a quiet street, trays clink softly inside, a\ndoorbell rings once, then light footsteps and the crisp sound of bread being sliced.\n\nScore: a soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes,\ngentle fade at the end.\n```\n\n**Body target: 350–500 words** for a reference-carrying shot. Dialogue-heavy content prioritises\nfitting the complete spoken timeline over hitting a word count.\n\n🚨 **Slates disables the provider's prompt expander.** H3's API can rewrite your prompt before\ngeneration; Slates turns that off, because a model rewriting the user's words invisibly is banned\noutright (prompt transparency: what the composer shows is what the model gets). The practical\nconsequence is on you: **nothing will pad a thin prompt.** Write the whole body.\n\n---\n\n## Shots and timing\n\nThe first shot carries **no timestamp**. Every later shot opens with the bracket and a cut time\nthat increases and stays inside the clip length:\n\n```\n[Shot 2] At 00:03.500, the camera cuts to ...\n```\n\nTransition verbs the model knows: **cuts to · transitions to · changes to · switches to**.\n\n**Dialogue that continues across a cut** needs the continuity said out loud — *\"his final words\ncarry over from the previous shot\"* — or the line restarts. **Speech that ends abruptly** should be\ndescribed as cut off rather than trailed off.\n\n---\n\n## Camera — write the move into the sentence\n\nThe model has a named motion vocabulary:\n\n> Zoom In / Zoom Out · Push In / Pull Out · Pan Left / Pan Right · Truck Left / Truck Right ·\n> Tilt Up / Tilt Down · Pedestal Up / Pedestal Down · Arc Shot · Tracking Shot · Static Shot ·\n> Shake Slightly / Shake Strongly · POV · Roll Clockwise / Roll Counterclockwise\n\nModify with **amplitude** (`with small amplitude` / `with large amplitude`) and **speed**\n(`at slow speed` / `at fast speed`).\n\n🚨 **Integrate the motion into the sentence — never stack labels.** MiniMax's own example:\n*\"The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.\"*\nNot *\"Push In. Small amplitude. Slow.\"*\n\n---\n\n## Speakers and dialogue\n\nGive each speaking character a stable identity in the prose and keep it: describe the voice once\n(*\"a young woman with a quiet, breathy voice\"*), then refer back to the same description at every\nline. Identification, delivery and action sit **outside** the quoted line; the line itself is only\nthe words.\n\n```\nThe young woman with a quiet, breathy voice says: \"I get off at the next station.\"\n```\n\n**Voiceover** needs two things — the phrase *\"says in an off-screen voiceover\"* **and** an explicit\nstatement that the lips stay closed. Without the second half the model animates a mouth.\n\n```\nThe man says in an off-screen voiceover: \"I still remember that road.\" — his lips remain\ncompletely closed.\n```\n\n**On-screen text** — signs, banners, labels, subtitles, neon — goes in double quotes with the\noriginal wording preserved exactly: *A red neon sign reading \"Open Late\" glows above the doorway.*\n\n---\n\n## References — H3's real differentiator is the declared RELATIONSHIP\n\n*(BOTH rows. `minimax-h3-max` gained the reference set on 2026-09-09; its free allowance is\nFOUR images rather than the base row's five.)*\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### Cite references by number — Slates already does it for you\n\nH3 on fal takes references as **typed slots** and expects the prompt to name them by modality and\norder: **`image 1`, `image 2`, `video 1`, `audio 1`**. That is exactly what the Slates composer\nemits from your `@mentions` and `#tags` (`Marcus (image 1) in the workshop (image 2)`), in the\nexact order it sends them.\n\n🚨 **Do NOT hand-write angle-bracket reference tags.** MiniMax's own model-card grammar uses\n`<Subject N>` / `<Picture N>` / `<Video N>` / `<Audio N>` labels; the fal endpoints Slates calls do\nnot — they build the binding from the typed slots and ask for plain numbered prose. Typing the tags\nyourself puts literal angle brackets in the prompt the model reads.\n\n### State how much of each reference survives\n\nThis is the lever no other model in the catalogue gives you. Say, in plain words, what each\nreference is FOR and how much of it should carry through:\n\n| Intent | Say something like |\n|---|---|\n| Keep it whole | *\"Keep the woman in image 1 exactly as she appears — hair, cardigan, necklace.\"* |\n| Keep part of it | *\"Use the café in image 2 for the brick wall and the sofa; the lighting is late evening, not daylight.\"* |\n| **Move a trait onto someone else** | *\"Give the man in image 3 the weathered leather texture of the jacket in image 4.\"* |\n| Loose echo | *\"Match the general palette and grain of image 5; nothing else from it.\"* |\n\nThe third row is the one with no equivalent anywhere else in Slates: **transferring a characteristic\nonto a different subject** is a first-class thing H3 understands. Reach for H3 when that is the job.\n\n**Audio references** bind a voice or a texture without copying the words. Say which speaker an\naudio reference is for (*\"the woman in image 1 speaks in the voice timbre of audio 1\"*), and when\nyou are referencing only the timbre, **do not carry the reference clip's original dialogue into your\nprompt** — write the new line. When you genuinely want the same words re-performed, quote them\nexactly and say so.\n\n**An audio reference cannot travel alone** — H3 refuses a reference set that is audio only. Pair it\nwith at least one image or video reference.\n\n### 💸 Reference images past the free allowance are billed — and the two rows differ\n\nOn `minimax-h3` the first **5** are free and each additional image adds **4 credits**. On\n`minimax-h3-max` the first **4** are free and each additional image adds **1 credit** — fal prices\nMax's references by token rather than per image, and Slates normalises every Max reference to\n1024x1024 so that per-image number is exact. Both rows take **9** images, at every resolution and\nevery length. Four extra images on a 10s\n768p clip add 16 credits to a 30-credit generation: **more than half again**, for references that\noften make the output worse rather than better (see the 2–4 rule above).\n\nAttach the references the shot needs, not the ceiling. Call\n`slates_estimate_generation_cost` with `referenceImages` set to the real count before a\nreference-heavy job — a quote that omits it under-reports the bill.\n\n---\n\n## Frames\n\n`minimax-h3` and `minimax-h3-max` both take a **start frame**, an **end frame**, or both. With an\nend frame, land it explicitly: describe the final pose, spacing and composition as the thing the\nshot **settles into** at the end, rather than hoping the model finds it.\n\n> *\"…she rotates the handle into the final angle and settles into the pose, spacing and composition\n> of image 2 at the end of the shot.\"*\n\n**Frames and references are mutually exclusive** on both rows — they are different endpoints, and\nthe reference endpoint has no frame slots at all. Slates refuses the combination rather than\ndropping one side.\n\n---\n\n## Cost discipline\n\n| Combination | Credits |\n|---|---:|\n| `minimax-h3` · 768p · 5s | 15 |\n| `minimax-h3` · 768p · 10s | 30 |\n| `minimax-h3` · 2K · 10s | 65 |\n| `minimax-h3` · 4K · 10s | 80 |\n| `minimax-h3-max` · 768p · 10s | 40 |\n| `minimax-h3-max` · 1080p · 10s | 80 |\n| `minimax-h3` — every reference image past the **fifth** | **+4** |\n| `minimax-h3-max` — every reference image past the **fourth** | **+1** |\n\n**768p is the default for a reason.** It is the tier the model natively generates.\n\n🚨 **2K and 4K are UPSCALES of a 768p render, not larger generations.** fal's own schema says so:\n*\"480P and 768P are native generation modes; 2K and 4K upscale a 768P base result.\"* The upscaler\n(H3-Regenerate-2K) is a separate stage bolted onto a finished take — it can enlarge detail but it\ncannot add information.\n\n**In our own test (2026-08-27, same prompt, same seed) the 2K pass came back with MORE artifacting\nthan the 768p original it was built from**, while costing 33 credits for a 5-second take against 15,\nand taking nearly twice as long to return. One shot, so treat it as a warning rather than a law —\nbut the mechanism explains it, and the burden of proof is on 2K.\n\n**So: generate at 768p and judge it at 768p.** Reach for 2K or 4K only when a delivery spec demands\nthe pixels, and expect to be paying for size rather than quality — a post-production upscale from a\nclean 768p master is very often the better result. **4K video is Pro-only** (the server returns\n`PRO_REQUIRED` for a base account); 2K is open to every tier.\n\n---\n\n## Quick checklist\n\n- Body written as a timeline, first shot untimestamped, later shots on `[Shot N] At MM:SS.mmm`.\n- Camera motion written **into** a sentence with amplitude and speed.\n- Dialogue and diegetic music in the body; ambience in the soundscape section; audience-only score\n in the score section, described by instrument and tempo.\n- Voiceover carries both the off-screen phrase and the closed-lips statement.\n- References cited as `image 1` / `video 1` / `audio 1`, each with a stated job and a stated degree\n of retention. No angle-bracket tags.\n- Reference count is deliberate — you are paying 4 credits for each one past the fifth.\n- Frames **or** references, never both.\n- The prompt is the prompt: no expander will fill it out for you.\n",
23
+ "slates-prompting-minimax-h3": "---\nname: slates-prompting-minimax-h3\ndescription: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09), so the seats differ on ladder and price, not on what they accept. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; pooled media tokens on Max), and audio written into the wrong section is dropped or duplicated.\n---\n\n# MiniMax H3 — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — MiniMax H3.** The only seat where audio is AUTHORED rather than toggled: dialogue, scene sound and score are three separate sections of the prompt, generated in one pass, and putting a sound in the wrong section drops or doubles it.\n\n**The five levers**\n1. **Write the three audio layers separately** — `Scene sound:` for what is in the room, `Score:` for what only the audience hears, and the dialogue quoted inline. Section decides attribution.\n2. **Quote dialogue and name the language** — `says in English`, `speaks in Spanish`. Eleven languages are stably supported; the language is part of the instruction, not an afterthought.\n3. **Declare the reference RELATIONSHIP**, which no other seat has: `kept whole`, `partly kept`, `transferred`, or `a loose echo`. An undeclared reference is a guess.\n4. **Give a beat of stillness before a line** — `sits still for a beat, then looks up`. The sync needs something to lock against; a character already mid-motion when the line starts drifts.\n5. **Describe the beat structure** — `waits`, `then speaks`, `under the last three seconds`. H3 is a timeline, so write one.\n\n**Examples**\n- `A woman sits still at a kitchen table for a beat, then looks up. She says in English, \"You said Tuesday.\" Scene sound: a fridge hum, a spoon set down on formica. Score: none.`\n- `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, \"No es el alternador.\" Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`\n\n**Hard constraint:** the two seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 768p rather than 4K, takes the same 9+3+3 references, and costs MORE at the tier they share — it is a speed pick, never the cheap one. H3's top two resolution tiers are UPSCALES of the native render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- a sound written into the wrong audio section — it is dropped, doubled, or attributed to the wrong layer\n- `background music` as a bare instruction: the score is its own authored layer, audience-only, and it is named as such\n- an undeclared reference relationship — say kept whole, partly kept, transferred, or a loose echo\n<!-- @banned:end -->\n\nH3 is an **omni transformer**: it generates picture and sound in the same pass, at 24fps with\n32kHz stereo, 5–15 seconds, in 11 stably-supported languages (Arabic, Chinese, English, French,\nGerman, Italian, Japanese, Korean, Portuguese, Russian, Spanish). That single fact drives\neverything below — the prompt is not a shot description with sound bolted on, it is a **timeline\nwith three audio layers you author separately**.\n\n**Two seats, one grammar.** Everything in this file applies to both. They differ only in what the\nendpoint accepts:\n\n| | `minimax-h3` | `minimax-h3-max` |\n|---|---|---|\n| Resolution | 480p / 768p / **2K / 4K** | 480p / 768p |\n| References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) |\n| Frames | start and/or end | start and/or end |\n| Price at 768p | **$0.060/s** | $0.080/s |\n| Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) |\n\n**Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it\ntops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth\npaying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.\n\n**The speed is measured, not claimed** (2026-08-27, same prompt and params on both rows): a 5-second\n768p text-to-video finished in **4.8 seconds** on Max against **57 seconds** on base H3 — roughly\n**12x**, queue to finished file. fal advertises \"under 3 seconds\"; the literal claim did not hold at\n4.8s wall-clock, but the order of magnitude did. For iteration loops and client-present work that gap\nis the entire reason the seat exists.\n\n🚨 **Max's known weakness: colour banding in low light (Eric, 2026-09-09).** Certain shots —\nespecially dark or low-key ones — come back with low-bitrate-looking banding across gradients (skies,\nwalls, shadow falloff). It is the one place the seat visibly gives something up. If a shot is dark\nand gradient-heavy, either light it up in the prompt or route to base H3 at 768p; do not fix it by\nreaching for 2K, which adds its own artifacting on top.\n\n---\n\n## The one thing that makes H3 different: audio is a THREE-LAYER instruction\n\nEvery other video seat treats sound as on or off. H3 splits it, and the split is enforced by where\nyou write each thing. Get the section wrong and the sound is dropped, doubled, or attributed to the\nwrong source.\n\n| Layer | What belongs in it | Where it goes |\n|---|---|---|\n| **Synchronised events** | dialogue, singing, and any sound tied to a specific shot or action | the **body** of the prompt, on the beat it lands |\n| **Scene sound** | ambience and physical sounds that run across the whole clip — room tone, rain, traffic, a ventilation hum | the **soundscape** section |\n| **Score** | music the characters cannot hear; audience-only | the **music** section |\n\n**Three rules, all from MiniMax's own guide:**\n\n1. **Dialogue and singing NEVER go in the soundscape section.** They are synchronised events; they\n belong in the body, at the moment they happen.\n2. **Diegetic music — music the characters can hear** (a radio in the scene, a busker) — also\n belongs in the **body**, not in the score section. The score section is audience-only.\n3. **Write the score in instrumental terms, not mood words.** Name the instruments, the tempo, and\n how it develops. *\"A restrained solo-piano score at a slow tempo, sustained low cello underneath,\n no swell\"* — not *\"emotional music\"*.\n\nUse **N/A** for a section only when silence or absence is genuinely what the shot wants. An empty\nscore section is a real choice; a vague one is a wasted layer.\n\n### The shape, in the one prompt field\n\nSlates sends one prompt string, so write the three layers as labelled paragraphs in this order:\n\n```\n[Shot 1] Live-action, cinematic. A medium-wide shot frames a baker opening the shutters of a\nsmall street bakery before sunrise. The camera pushes in with small amplitude at slow speed as\nthe middle-aged baker with a calm, slightly raspy voice places a fresh loaf on the counter and\nsays: \"First batch of the morning.\" [Shot 2] At 00:05.000, the camera cuts to a close-up of\nsteam rising from the sliced bread while his final words carry over from the previous shot.\n\nSoundscape: wooden shutters scrape open over a quiet street, trays clink softly inside, a\ndoorbell rings once, then light footsteps and the crisp sound of bread being sliced.\n\nScore: a soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes,\ngentle fade at the end.\n```\n\n**Body target: 350–500 words** for a reference-carrying shot. Dialogue-heavy content prioritises\nfitting the complete spoken timeline over hitting a word count.\n\n🚨 **Slates disables the provider's prompt expander.** H3's API can rewrite your prompt before\ngeneration; Slates turns that off, because a model rewriting the user's words invisibly is banned\noutright (prompt transparency: what the composer shows is what the model gets). The practical\nconsequence is on you: **nothing will pad a thin prompt.** Write the whole body.\n\n---\n\n## Shots and timing\n\nThe first shot carries **no timestamp**. Every later shot opens with the bracket and a cut time\nthat increases and stays inside the clip length:\n\n```\n[Shot 2] At 00:03.500, the camera cuts to ...\n```\n\nTransition verbs the model knows: **cuts to · transitions to · changes to · switches to**.\n\n**Dialogue that continues across a cut** needs the continuity said out loud — *\"his final words\ncarry over from the previous shot\"* — or the line restarts. **Speech that ends abruptly** should be\ndescribed as cut off rather than trailed off.\n\n---\n\n## Camera — write the move into the sentence\n\nThe model has a named motion vocabulary:\n\n> Zoom In / Zoom Out · Push In / Pull Out · Pan Left / Pan Right · Truck Left / Truck Right ·\n> Tilt Up / Tilt Down · Pedestal Up / Pedestal Down · Arc Shot · Tracking Shot · Static Shot ·\n> Shake Slightly / Shake Strongly · POV · Roll Clockwise / Roll Counterclockwise\n\nModify with **amplitude** (`with small amplitude` / `with large amplitude`) and **speed**\n(`at slow speed` / `at fast speed`).\n\n🚨 **Integrate the motion into the sentence — never stack labels.** MiniMax's own example:\n*\"The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.\"*\nNot *\"Push In. Small amplitude. Slow.\"*\n\n---\n\n## Speakers and dialogue\n\nGive each speaking character a stable identity in the prose and keep it: describe the voice once\n(*\"a young woman with a quiet, breathy voice\"*), then refer back to the same description at every\nline. Identification, delivery and action sit **outside** the quoted line; the line itself is only\nthe words.\n\n```\nThe young woman with a quiet, breathy voice says: \"I get off at the next station.\"\n```\n\n**Voiceover** needs two things — the phrase *\"says in an off-screen voiceover\"* **and** an explicit\nstatement that the lips stay closed. Without the second half the model animates a mouth.\n\n```\nThe man says in an off-screen voiceover: \"I still remember that road.\" — his lips remain\ncompletely closed.\n```\n\n**On-screen text** — signs, banners, labels, subtitles, neon — goes in double quotes with the\noriginal wording preserved exactly: *A red neon sign reading \"Open Late\" glows above the doorway.*\n\n---\n\n## References — H3's real differentiator is the declared RELATIONSHIP\n\n*(BOTH rows. `minimax-h3-max` gained the reference set on 2026-09-09; its free allowance is\nFOUR images rather than the base row's five.)*\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### Cite references by number — Slates already does it for you\n\nH3 on fal takes references as **typed slots** and expects the prompt to name them by modality and\norder: **`image 1`, `image 2`, `video 1`, `audio 1`**. That is exactly what the Slates composer\nemits from your `@mentions` and `#tags` (`Marcus (image 1) in the workshop (image 2)`), in the\nexact order it sends them.\n\n🚨 **Do NOT hand-write angle-bracket reference tags.** MiniMax's own model-card grammar uses\n`<Subject N>` / `<Picture N>` / `<Video N>` / `<Audio N>` labels; the fal endpoints Slates calls do\nnot — they build the binding from the typed slots and ask for plain numbered prose. Typing the tags\nyourself puts literal angle brackets in the prompt the model reads.\n\n### State how much of each reference survives\n\nThis is the lever no other model in the catalogue gives you. Say, in plain words, what each\nreference is FOR and how much of it should carry through:\n\n| Intent | Say something like |\n|---|---|\n| Keep it whole | *\"Keep the woman in image 1 exactly as she appears — hair, cardigan, necklace.\"* |\n| Keep part of it | *\"Use the café in image 2 for the brick wall and the sofa; the lighting is late evening, not daylight.\"* |\n| **Move a trait onto someone else** | *\"Give the man in image 3 the weathered leather texture of the jacket in image 4.\"* |\n| Loose echo | *\"Match the general palette and grain of image 5; nothing else from it.\"* |\n\nThe third row is the one with no equivalent anywhere else in Slates: **transferring a characteristic\nonto a different subject** is a first-class thing H3 understands. Reach for H3 when that is the job.\n\n**Audio references** bind a voice or a texture without copying the words. Say which speaker an\naudio reference is for (*\"the woman in image 1 speaks in the voice timbre of audio 1\"*), and when\nyou are referencing only the timbre, **do not carry the reference clip's original dialogue into your\nprompt** — write the new line. When you genuinely want the same words re-performed, quote them\nexactly and say so.\n\n**An audio reference cannot travel alone** — H3 refuses a reference set that is audio only. Pair it\nwith at least one image or video reference.\n\n### 💸 Reference images past the free allowance are billed — and the two rows differ\n\nOn `minimax-h3` the first **5** are free and each additional image adds **4 credits**.\nMax pools image pixels, reference-video seconds and reference-audio seconds into one token\nallowance. Include `referenceImages`, `videoRefSeconds` and `audioRefSeconds` when estimating;\ncharacter voices count as audio. The generation preflight resolves the actual attached media.\n\nFour extra images on a 10s\n768p clip add 16 credits to a 30-credit generation: **more than half again**, for references that\noften make the output worse rather than better (see the 2–4 rule above).\n\nAttach the references the shot needs, not the ceiling. Call\n`slates_estimate_generation_cost` with `referenceImages` set to the real count before a\nreference-heavy job — a quote that omits it under-reports the bill.\n\n---\n\n## Frames\n\n`minimax-h3` and `minimax-h3-max` both take a **start frame**, an **end frame**, or both. With an\nend frame, land it explicitly: describe the final pose, spacing and composition as the thing the\nshot **settles into** at the end, rather than hoping the model finds it.\n\n> *\"…she rotates the handle into the final angle and settles into the pose, spacing and composition\n> of image 2 at the end of the shot.\"*\n\n**Frames and references are mutually exclusive** on both rows — they are different endpoints, and\nthe reference endpoint has no frame slots at all. Slates refuses the combination rather than\ndropping one side.\n\n---\n\n## Cost discipline\n\n| Combination | Credits |\n|---|---:|\n| `minimax-h3` · 768p · 5s | 15 |\n| `minimax-h3` · 768p · 10s | 30 |\n| `minimax-h3` · 2K · 10s | 65 |\n| `minimax-h3` · 4K · 10s | 80 |\n| `minimax-h3-max` · 768p · 10s | 40 |\n| `minimax-h3` — every reference image past the **fifth** | **+4** |\n\n**768p is the default for a reason.** It is the tier the model natively generates.\n\n🚨 **2K and 4K are UPSCALES of a 768p render, not larger generations.** fal's own schema says so:\n*\"480P and 768P are native generation modes; 2K and 4K upscale a 768P base result.\"* The upscaler\n(H3-Regenerate-2K) is a separate stage bolted onto a finished take — it can enlarge detail but it\ncannot add information.\n\n**In our own test (2026-08-27, same prompt, same seed) the 2K pass came back with MORE artifacting\nthan the 768p original it was built from**, while costing 33 credits for a 5-second take against 15,\nand taking nearly twice as long to return.\n\n🚨 **Confirmed independently (Eric, 2026-09-09): 2K and 4K carry visible AI noise artifacting and\n\"just look bad\".** That is now TWO separate observations, months apart, pointing the same way — it is\nno longer a single-shot warning. The tiers stay available because a delivery spec sometimes demands\nthe pixels, but **do not route to 2K/4K for quality**: you are paying more, waiting longer, and\nadding artifacts to a 768p render. Upscale in post from a clean 768p master instead.\n\n**So: generate at 768p and judge it at 768p.** Reach for 2K or 4K only when a delivery spec demands\nthe pixels, and expect to be paying for size rather than quality — a post-production upscale from a\nclean 768p master is very often the better result. **4K video is Pro-only** (the server returns\n`PRO_REQUIRED` for a base account); 2K is open to every tier.\n\n---\n\n## Quick checklist\n\n- Body written as a timeline, first shot untimestamped, later shots on `[Shot N] At MM:SS.mmm`.\n- Camera motion written **into** a sentence with amplitude and speed.\n- Dialogue and diegetic music in the body; ambience in the soundscape section; audience-only score\n in the score section, described by instrument and tempo.\n- Voiceover carries both the off-screen phrase and the closed-lips statement.\n- References cited as `image 1` / `video 1` / `audio 1`, each with a stated job and a stated degree\n of retention. No angle-bracket tags.\n- Reference count is deliberate — you are paying 4 credits for each one past the fifth.\n- Frames **or** references, never both.\n- The prompt is the prompt: no expander will fill it out for you.\n",
24
24
  "slates-prompting-motion-transfer": "---\nname: slates-prompting-motion-transfer\ndescription: How to set up motion transfer — Kling Motion Control only (std and pro tiers, 5-second outputs). Read before calling slates_generate_motion_transfer. Reference image (character) + driving video (motion source) → new video of the character performing the motion. Asset selection rules, character_orientation, tiers, and prompt usage. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.\n---\n\n# Motion transfer — setup guide\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Motion transfer (Kling Motion Control only).** A target IMAGE (your character) plus a source VIDEO (the motion) produces your character performing that motion. Always 5 seconds.\n\n**The five levers**\n1. **The target image must show body proportions clearly** and the character must occupy more than about 5% of the frame. A tiny figure in a wide shot has nothing to drive.\n2. **Single character in the target.** A group image breaks the identity anchor.\n3. **Choose `characterOrientation` on purpose** — `video` takes the source clip's framing, `image` preserves the portrait's. It is the most-missed choice here.\n\n4. **The prompt is atmosphere only** — `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.` Motion verbs are ignored; the motion is already in the driving video.\n5. **Pick the best 5 seconds of the source up front**, and write only atmosphere: `soft afternoon sunlight`, `vintage warm color grade`, `clean studio backdrop`. The output is 5s regardless, so a long driving clip just wastes the choice.\n\n**Examples**\n- `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.`\n- `Clean studio backdrop, sharp focus on the character.` (Or leave it empty.)\n\n**Hard constraint:** cartoon driving videos fail, and a cropped or partial target character drifts. std is fine while the motion-and-framing combination is still moving; switch to pro once it is locked. For a REGENERATED shot instead — physical contact, cloth and hair, camera motion, native audio — that is a normal Seedance video generation with the driving clip as a video reference, not a mode of this tool.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** — the motion is already in the driving video, so motion verbs are ignored:\n- `spins faster`, `jumps higher`, `add more energy`, `moves quicker`\n<!-- @banned:end -->\n\nTake a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion. **This tool is Kling-only** — it wraps Kling Motion Control and nothing else.\n\n| Tier | Cost | Use case |\n|------|-----------|----------|\n| Kling std (`kling-mc-std-5s`) | ~32 credits / 5s | General motion transfer, budget lane |\n| Kling pro (`kling-mc-pro-5s`) | ~42 credits / 5s | Cleaner anatomy, better identity preservation |\n\nBoth tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)\n\n## Want Seedance instead? That is a video generation, not a mode here\n\nKling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion, and the output carries native audio. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the driving clip attached as a video reference and the character image as an ingredient, then write the prompt yourself:\n\n```\nThe character from image 1 performs the exact motion, choreography, and camera\nmovement from video 1. Preserve the character's identity, appearance, and outfit.\n```\n\nThat is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.\n\n- **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).\n- **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.\n- **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).\n- `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.\n\nEverything below is about the Kling tool.\n\n## Inputs\n\n- `sourceVideoAssetId` — driving video. **Must be a realistic human** with clear proportions. Anime/cartoon/CG driving videos fail.\n- `targetImageAssetId` — character to be animated. Can be any style (cartoon, anime, realistic, painted).\n- Both must already exist as assets in the project. Use `slates_list_assets` to find them or upload first.\n\n## Source video constraints\n\n- Realistic human (not animated, not CG)\n- Entire body OR upper body visible — head must not be obstructed\n- Subject occupies a clear share of the frame\n- Single primary subject. Multi-person driving videos confuse the motion anchor.\n- Clean motion — choppy / cut-edited driving videos produce jittery output\n\nGood driving video sources:\n- Reference dance footage with one subject\n- Walking / gesture / posing clips\n- Talking-head footage when paired with character_orientation: 'video'\n\nBad driving video sources:\n- Music videos with multi-shot edits\n- Anime / animation clips\n- Heavily stylized footage with smoke / particles obscuring the body\n- Footage where the subject's head leaves frame mid-clip\n\n## Target image constraints\n\n- Character body proportions clearly visible\n- Character occupies >5% of image area (not a tiny figure in a wide shot)\n- Single character. Group images break the identity anchor.\n- Any artistic style works — cartoon, anime, painted, realistic, 3D render\n\nAvoid:\n- Extreme close-up of just the face (no body to drive)\n- Character partially cropped at the waist when the driving video is full-body\n- Multiple characters\n\n## character_orientation — the most-missed choice\n\nThis single parameter changes the output dramatically. Pick deliberately.\n\n| Value | Output framing | Max source duration | Best for |\n|-------|----------------|---------------------|----------|\n| `video` | Matches driving video framing | Up to 30s source | Complex full-body motion (dance, action, athletics) |\n| `image` | Matches target image framing | Up to 10s source | Camera moves, simpler motion, preserving original composition |\n\n**Default `video`** when the driving video has the look you want (most cases).\n\nSwitch to `image` when the target image's composition is the brand asset and the motion is secondary (e.g., a hero shot of a character that needs subtle gesture, not a full performance).\n\n## Tier choice — std vs pro\n\n**std (~32 credits)** for:\n- Drafts, motion exploration, blocking\n- Group scenes where the character isn't a hero shot\n- When the budget is tight and the motion is the focus\n\n**pro (~42 credits)** for:\n- Final hero takes\n- Branded characters where identity drift = unacceptable\n- Anatomically complex motion (limbs crossing, fast direction changes)\n- Anime / cartoon target images — pro handles non-realistic styles better\n\nDon't default to pro. The ~10-credit delta compounds fast across iteration.\n\n## Prompt usage (optional)\n\nThe `prompt` field is **scene/style refinement**, not motion direction. The motion comes from the driving video — the prompt sets ambiance, lighting, additional detail.\n\nGood:\n- `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.`\n- `Clean studio backdrop, sharp focus on the character.`\n\nBad (model ignores motion verbs — they're already in the driving video):\n- ❌ `She spins faster and jumps higher.`\n- ❌ `Add more energy to the dance.`\n\nLeave it empty if you don't have a specific atmospheric note.\n\n## Common failure modes\n\n| Symptom | Likely cause | Fix |\n|---------|--------------|-----|\n| Limbs distort / extra fingers | std tier, complex motion | Switch to pro |\n| Character identity drifts | Target image cropped too tight | Use a fuller-body target |\n| Output looks \"stuck\" / minimal motion | Driving video subject too small in frame | Pick a driving video where the subject fills more of the frame |\n| Cartoon target turns realistic | std tier on stylized art | Switch to pro — handles non-realistic styles better |\n| Garbled output entirely | Anime / CG driving video | Use realistic human driving footage |\n| Wrong framing on output | character_orientation set wrong | Try the other value |\n| Background bleeds through character | Target image had complex background | Use a target with cleaner background separation |\n\n## Workflow patterns\n\n**Reference dance to brand character:**\n1. Generate or upload the brand character as a still image (clean background, full body, single subject)\n2. Find driving footage — a clean reference video of the dance you want\n3. Upload both as project assets\n4. Run motion transfer with `motionModel: 'kling-mc-pro'`, `characterOrientation: 'video'`\n5. Total cost: ~42 credits per 5s take\n\n**Subtle motion on a hero portrait:**\n1. Use the locked hero portrait as the target image\n2. Pick a driving video with subtle gesture (head turn, slight posture shift)\n3. `characterOrientation: 'image'` to preserve the portrait's framing\n4. std tier is fine for this case — motion isn't dramatic\n\n**Avoid:**\n- Pro tier on first iteration — waste, switch to it once the motion + framing combo is locked\n- Cartoon driving videos — guaranteed failure\n- Cropped or partial target characters — identity will drift\n- Long driving videos when output is 5s — pick the best 5s of the source upfront\n\n## Cost discipline\n\n- 5 seconds, no shorter option\n- Both tiers trip the confirm gate — every call needs explicit user OK\n- Iteration is expensive: 4 takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.\n- Always run a single std take first to validate the motion + framing combo before committing to pro\n\n## Confirm gate: cost + codes, no inline preview\n\nMotion transfer is mechanical — the model deterministically applies source motion to target image. Both tiers trip the confirm gate; the response includes the asset codes for source and target so you can announce them in chat.\n\n- ✅ \"Transferring motion from **VID-V3** onto **IMG-A12 — Detective Closeup**. ~42 credits, confirm?\"\n- ❌ \"Using the walk video and the detective image...\" (multiple of each in the project.)\n\nDon't second-guess the assets the user picked — the model executes the transfer. If the output is wrong, iterate on motion source or target choice, not on a refinement prompt.\n\n## Sources\n\n- [fal.ai — Kling Motion Control V3 Standard](https://fal.ai/models/fal-ai/kling-video/v3/standard/motion-control)\n- [fal.ai — Kling Motion Control V3 Pro](https://fal.ai/models/fal-ai/kling-video/v3/pro/motion-control)\n",
25
25
  "slates-prompting-nano-banana-2": "---\nname: slates-prompting-nano-banana-2\ndescription: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3.1 Flash Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.\n---\n\n# Nano Banana 2 — cinematic & photorealistic prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Nano Banana 2 (Gemini 3.1 Flash Image).** Brief it like a creative director, not a tag list. Structure: `Film still from [director] [genre]. Shot on [camera] with [lens]. [Subject and action]. [3-5 specific visual details]. [Lighting — direction + quality]. [Color palette]. [Film stock]. [1-2 word tone].`\n\n**The five levers**\n1. **Named lens + aperture** beats \"shallow depth of field\" — `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin), `Panavision anamorphic`, `400mm telephoto`.\n2. **Light by direction and quality**, never \"good lighting\" — `hard sidelight from a single window, deep falloff`, `overcast north light`, `practical tungsten spill`.\n3. **A named film stock or sensor** carries a whole palette — `Kodak Portra 400`, `Cinestill 800T`, `ARRI Alexa 65`.\n4. **Composition as a shot** — `low angle`, `aerial view`, `rule of thirds with the subject camera-left`, `foreground occlusion`.\n5. **Positive framing only.** Describe what is there. \"Empty street\", never \"no cars\"; \"unstaged documentary photography\", never \"not anime\".\n\n**Examples**\n- `Film still from a Denis Villeneuve thriller. Shot on ARRI Alexa 65, 85mm f/1.4. A woman in a charcoal wool coat stands at a rain-slick bus stop, breath visible. Hard sodium light from a single overhead lamp, deep falloff into blue night. Kodak Vision3 500T. Isolated.`\n- `Editorial still life on seamless bone paper. 100mm macro, f/8. A cracked ceramic bowl holding three figs. Soft north light from camera-left, one gentle shadow. Muted earth palette. Portra 400 grain. Quiet.`\n\n**Hard constraint:** there is no `negativePrompt` field. Suppress by reframing positively, or inline `without` / `free of`. Knowledge cutoff January 2025 — anything later needs reference images.\n<!-- @card:end -->\n\nNano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the default model behind `slates_generate_image` — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.\n\nKnowledge cutoff: January 2025. Anything after needs explicit reference images.\n\n## Google's 4 official rules (verbatim)\n\n1. **Be specific.** Provide concrete details on subject, lighting, and composition.\n2. **Use positive framing.** Describe what you want, not what you don't want.\n3. **Control the camera.** Use photographic and cinematic terms like \"low angle\" and \"aerial view.\"\n4. **Iterate.** Refine images with follow-up prompts in a conversational manner.\n\n## Official prompt formula\n\n```\n[Subject] + [Action] + [Location/context] + [Composition] + [Style]\n```\n\nFor the cinematic / photoreal use case, expand to:\n\n```\nFilm still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and action]. [3-5 specific visual details]. [LIGHTING — direction + quality]. [COLOR PALETTE]. [FILM STOCK or sensor language]. [1-2 word emotional tone].\n```\n\n## Photorealism positives — what consistently works\n\n> ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern — do not carry it across.**\n> Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.\n> The leak happens in one specific way — you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `slates-prompting-seedance` (Part 3, \"Don't cross-pollinate image-model syntax\").\n\n**Named lenses + apertures** beat generic \"shallow depth of field\":\n- `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`\n- `Panavision anamorphic` for horizontal flares + cinematic width\n- `400mm telephoto` for compression + isolation\n- `24mm` for environmental interiors\n\n**Named cameras / sensors:**\n- `ARRI Alexa 65`, `Hasselblad X2D`, `Canon EOS R5`, `Sony A7III`, `Fujifilm X-T5`\n- \"Specific gear\" beats \"DSLR\"\n\n**Named film stocks** (one per prompt — never mix):\n- `Kodak Portra 400` — natural skin, warm\n- `Fuji Velvia 50` — saturated, landscape\n- `Ilford HP5 Plus` — black and white, gritty grain\n- `CineStill 800T` — tungsten night, halation\n\n**Physics-based lighting** (direction + quality):\n- `Single key light at 45 degrees from upper left`\n- `Late afternoon sun at 15 degrees above horizon`\n- `Color temperature 4500K` beats `slightly warm`\n- `Practicals only — no fill` for Deakins-style realism\n\n**Imperfection vocabulary** (forces away from AI-clean):\n- `visible pores`, `natural skin grain`, `peach fuzz`, `slight hyperpigmentation`\n- `unretouched raw photography`, `ISO noise`, `sweat beading`\n- `crisp catchlights in the eyes`, `skin micro-detail`\n\n**Director references** (use when locking style):\n| Director | Tone | Visual signature |\n|---|---|---|\n| Denis Villeneuve | Cold, vast, existential | Desaturated, overwhelming scale |\n| Roger Deakins | Precise motivated light | Single source, deep shadows, practicals |\n| Emmanuel Lubezki | Natural, spiritual | Available light, golden hour |\n| Bradford Young | Warm darkness | Underexposed, rich shadows, skin tones |\n\n**Genre cues that move the model:**\n- `unstaged documentary photography style`\n- `fashion magazine editorial, shot on medium-format analog film, pronounced grain`\n- `Film still from [Director] [genre]`\n\n## The anti-list — phrases that DEGRADE realism\n\nThese are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is extracted\n by src/prompts/banned-tokens.ts, inlined verbatim into the slates_generate_image\n op description (always in context on both surfaces), and matched against every\n submitted prompt. Editing this list changes what the agent is told AND what it\n is warned about — keep every entry backticked, and keep prose outside the\n backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- `8k`, `4k` (as a quality token)\n- `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone\n- `masterpiece`, `best quality`, `highly detailed`, `ultra-detailed`\n- `trending on ArtStation`, `award-winning`\n- `perfect skin`, `flawless`, `airbrushed`, `smooth skin`\n- `cinematic` standing alone — always specify *which cinema* (director, lens, era, stock)\n- `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue\n<!-- @banned:end -->\n\n## Negative prompting — there is no field\n\nNano Banana 2 has **no `negativePrompt` parameter**. Three patterns to suppress unwanted content:\n\n1. **Positive reframing (preferred):** \"empty street\" not \"no cars\". \"Unstaged documentary photography\" not \"not anime.\"\n2. **Inline `without` / `free of`:** \"without any people, vehicles, or man-made structures\", \"free of text overlays, logos, or watermarks.\"\n3. **Constraint clauses for anatomy/quality:** \"accurate anatomy with five fingers per hand, symmetrical features, natural proportions\"; \"sharp, well-exposed, free of blur or JPEG artifacts.\"\n\nDefault to #1. Reach for #2 only when positive framing can't suppress the unwanted element.\n\n## Reference images\n\n- **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.\n- **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as \"image N\" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **\"assign a distinct name to each character/object\"**. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render the scene's expression\") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n\n### Reference rules (the verified ones)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Nano Banana 2 specifically\n\n- **NB2's own consistency lever is \"assign a distinct name to each character/object.\"** That is Google's phrasing for rule 3 — cite each canonical identity inline by name.\n- **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.\n- **Character consistency is officially \"not 100% perfect\"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.\n- **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.\n\n## Common failure modes + fixes\n\n**Hands:** Append `accurate anatomy with five fingers per hand, symmetrical features, natural proportions, relaxed open palm`. Avoid heavy jewelry, props intersecting fingers, motion blur in references.\n\n**Text in images:** Quote-wrap target text. Specify font (`Century Gothic, 12pt`). Long phrases work; small text degrades. Two-step works best — generate text concepts conversationally first, then ask for the image.\n\n**Left/right confusion:** Default is **viewer's perspective**, not subject's. Append `left and right are from the character's perspective, NOT the camera's` when scene-blocking matters.\n\n**Surreal / absurd prompts trip uncanny valley:** The model drags toward realism. If you want surrealism, lean hard into stylization keywords (`painted`, `illustrated`, `stop-motion`).\n\n**Soft faces / dead eyes:** Add `crisp catchlights in the eyes`, `skin micro-detail`, `peach fuzz visible`. Don't stack quality enhancers — single clean prompt beats multiple re-interpretations.\n\n**Post-cutoff content (anything after Jan 2025):** Use reference images. The model has no knowledge of recent franchises, products, events.\n\n## Resolution tactics\n\n- Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — check current numbers<!-- slates-only --> by calling `slates_estimate_generation_cost`<!-- /slates-only -->. Pick the cheapest resolution that serves the use case.\n- **At 2K and above, the model allocates more tokens to surface detail** — explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.\n- 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.\n- 2K generations vary 20-60s+. Don't time-budget tightly.\n\n## Boring vs cinema — examples\n\n❌ **Boring:** \"Wide shot of a man on a dock looking at the forest.\"\n\n✅ **Cinema:** \"Direct overhead drone shot on weathered dock surface. Single figure standing center frame, climbing up from frame bottom. Boot prints leading away from him toward shore. Pale winter light. Anamorphic lens flare from low sun. Desaturated blue and slate grey palette. Kodak Portra 400 grain. The path already walked by someone else. Map of threat.\"\n\n❌ **Boring:** \"Close up of a woman looking scared.\"\n\n✅ **Cinema:** \"Extreme close on subject's mouth and nose, 135mm f/2.8, shallow depth of field. Breath pluming out, catching cold light from upper-left key. Lips slightly parted, peach fuzz visible. The breath holds. CineStill 800T halation around catchlights. Waiting.\"\n\n## The 3-strike rule\n\nIf three iterations on the same prompt haven't produced what the user wants, stop. Hand back to the user with what you tried and what isn't working. The slot machine doesn't converge — the prompt structure is wrong, not the seed.\n\n## Family variants — Lite and Pro\n\nEverything in this skill applies to the whole Nano Banana family; two variants trade speed/ceiling around NB2 full:\n\n- **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.\n- **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.\n\n<!-- slates-only -->\nRouting between them (and vs GPT Image 2.5 / FLUX / Seedream): `slates-model-selection`.\n<!-- /slates-only -->\n",
26
26
  "slates-prompting-omni-flash": "---\nname: slates-prompting-omni-flash\ndescription: How to prompt Gemini Omni Flash (Google, via fal). Read before calling slates_generate_video with omni-flash or slates_edit_video with omni-flash-edit. Cheap 720p tier with native synced audio included — 3-10s, 16:9/9:16 only; t2v, single-start-frame i2v, or reference-to-video with up to 7 reference images. The edit variant is the EDIT-FIDELITY WINNER for footage-synced VFX (receipt 2026-07-09) — but ONLY with short prompts: one change + \"Keep everything else the same.\" Long descriptive prompts destroy fidelity.\n---\n\n# Gemini Omni Flash — prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card — Gemini Omni Flash.** Two different jobs with OPPOSITE prompt rules, and getting them the wrong way round is the whole failure mode.\n\n**The five levers**\n1. **Editing: short prompt, ONE change, nothing else.** Google's own doc says so and a 2026-07-09 receipt confirms it — a long \"keep every frame identical\" preamble produced WORSE drift than two sentences.\n2. **Editing: always end with `Keep everything else the same.`** — the one documented preservation lever.\n\n3. **Editing: describe the EFFECT, never a real object as a metaphor.** \"Candle-like flame\" rendered a literal candle in the subject's hand.\n4. **Editing: no conditional timing cues.** \"…when he calls it, as he walks…\" hard-fails with `invalid_request`. Collapse to one continuous action; the model syncs to the footage's own motion.\n5. **Generation: the opposite — describe fully.** Subject, action, setting, `camera tracking alongside`, `overcast flat light`, tone. Audio is prompt-driven with no parameters: dialogue in quotes, sound in plain language — `rain patters on the tin roof`, `spray from tyres`, `a horn somewhere behind`.\n\n**Examples**\n- Edit: `Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.`\n- Generate: `A courier in a yellow shell jacket weaves between stalled cars on a wet arterial road, camera tracking alongside at shoulder height. Overcast flat light, spray from tyres. Rain patters on car roofs, a horn somewhere behind.`\n\n**Hard constraint:** it is a CHEAP DRAFT seat for generation and the EDIT-fidelity winner for footage-synced VFX — never a hero generation shot. Expect a possible jitter or doubled speech beat in the last half second of an edit: trim the tail rather than burning a re-roll.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use in an EDIT prompt** (each one has a receipt above):\n- a long preservation preamble — it produces WORSE drift than `Keep everything else the same.`\n- a real object as a metaphor: `candle-like`, `flame-like`, `laser-like`\n- a conditional timing cue: `when he`, `as she`, `once they` — these hard-fail, they do not merely drift\n- harm-to-person framing: `ignite`, `catch fire`, `on fire` applied to a person trips the safety filter\n<!-- @banned:end -->\n\nGoogle's fast video generation + editing model (\"Nano Banana Pro for video\" in creator slang — a nickname; it is NOT the NB Pro image model). Carried on fal (`google/gemini-omni-flash*`). 720p only, 24fps, 3–10 second clips, 16:9 or 9:16. **Audio is native and included** — dialogue, SFX, and ambient generate WITH the video at no extra cost.\n\n## Where it routes\n\n- **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.\n- **Cheap drafts and iteration volume** — lowest-cost audio-native video seat (~6.4 cr/s at 720p).\n- **NOT hero GENERATION shots** — Kling 3.0 stays the general gen default, Seedance 2.0 the premium tier; Omni Flash's *generation* quality seat is still unproven.\n\n## Editing (`slates_edit_video`, model `omni-flash-edit`) — THE RULES (receipts, not theory)\n\n1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *\"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes.\"* Live receipt 2026-07-09: a long \"keep every frame/word/movement identical…\" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *\"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.\"*\n2. **Always end with \"Keep everything else the same.\"** — the one documented preservation lever.\n3. **Never name a real-world object as a metaphor.** \"Candle-like flame\" rendered a literal candle in his hand. Describe the effect itself (\"small magical flames on his fingertips\").\n3b. **No conditional timing cues — they HARD-FAIL, not drift.** Receipt 2026-07-09: \"a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…\" → deterministic `invalid_request` (2×, \"could not generate with the given inputs\"); collapsing to one continuous action — \"A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke.\" — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video.\n4. **Safety filter (Google's, strict about harm-to-person):** \"fingertips ignite / catch fire\" → `content_policy_violation`. Frame effects as magical/harmless VFX: \"small magical flames appear on his fingertips\" passed. See slates-content-policy §Gemini for the substitution patterns.\n5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.\n6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.\n7. Source clip 3–10s (trim longer clips first). Output length follows the source; billing per output second, rounded up. Voice editing unsupported — never ask it to change dialogue.\n8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).\n9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.\n\n## Generation (`slates_generate_video`, model `omni-flash`)\n\n- **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.\n- Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.\n- **Name references inline** the standard Slates way (\"Marcus (image 1) walks…\"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.\n- **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language (\"rain patters on the tin roof\"). Negative direction as plain instructions (\"Do not show text\").\n- Duration is an explicit 3–10s integer param; cost scales linearly per second.\n\n## Input conditioning (Slates handles this — know it exists)\n\nPhone footage stores rotation as a metadata flag; models ignore it and edit the raw sideways pixels. Clips must be rotation-normalized (and oversized sources downscaled) before upload — receipt 2026-07-09: a portrait Pixel clip came back sideways until conditioned. If an edit output comes back rotated, the source wasn't normalized.\n\n## Content notes\n\n- Google applies its own safety filters to input images/clips and output. Uploads containing recognizable real people are restricted by Google's policy — though own-footage editing of the uploader passed on our route 2026-07-09. See slates-content-policy.\n- Output carries an invisible SynthID watermark (Google-side, programmatic detection only).\n",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@slatesvideo/shared",
3
- "version": "0.6.8",
3
+ "version": "0.6.10",
4
4
  "description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
5
5
  "license": "MIT",
6
6
  "type": "module",