@slatesvideo/shared 0.5.3 → 0.5.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +31 -0
- package/dist/operations/index.d.ts +4 -8
- package/dist/operations/index.js +26 -38
- package/dist/prompts/character-sheet.d.ts +19 -18
- package/dist/prompts/character-sheet.js +84 -38
- package/dist/prompts/environment-sheet.d.ts +9 -1
- package/dist/prompts/environment-sheet.js +17 -3
- package/dist/prompts/model-facts.js +4 -1
- package/dist/prompts/partials.generated.d.ts +2 -0
- package/dist/prompts/partials.generated.js +14 -0
- package/dist/prompts/prompting-tips.js +50 -19
- package/dist/prompts/reference-composer.d.ts +1 -1
- package/dist/prompts/reference-composer.js +3 -4
- package/dist/prompts/reference-rules.d.ts +43 -14
- package/dist/prompts/reference-rules.js +51 -27
- package/dist/skills/content.js +14 -14
- package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
- package/exports/slates-prompt-builder/generated/reference-character.md +78 -0
- package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
- package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
- package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +8 -3
- package/skills/_partials/decision-log.md +12 -0
- package/skills/_partials/reference-rules-core.md +12 -0
- package/skills/_partials/reference-tips-short.md +2 -0
- package/skills/_partials/references-read-literally.md +11 -0
- package/skills/_partials/still-gate.md +3 -0
- package/skills/slates-character-identity.md +100 -0
- package/skills/slates-cost-discipline.md +10 -0
- package/skills/slates-edit-and-iterate.md +17 -2
- package/skills/slates-model-selection.md +24 -1
- package/skills/slates-one-prompt-film.md +22 -3
- package/skills/slates-prompting-flux-2-max.md +36 -5
- package/skills/slates-prompting-gpt-image-2.md +1 -1
- package/skills/slates-prompting-kling-v3.md +40 -9
- package/skills/slates-prompting-nano-banana-2.md +44 -12
- package/skills/slates-prompting-omni-flash.md +1 -1
- package/skills/slates-prompting-seedance.md +295 -90
- package/skills/slates-prompting-veo-3.md +33 -4
- package/skills/slates-storyboard-from-script.md +19 -0
- package/skills/slates-vision-feedback-loop.md +49 -2
- package/skills/slates-character-turnaround.md +0 -55
package/dist/skills/content.js
CHANGED
|
@@ -1,26 +1,26 @@
|
|
|
1
1
|
// GENERATED — do not edit. Source: packages/shared/skills/*.md
|
|
2
2
|
// Regenerated by scripts/embed-skills.mjs on every build.
|
|
3
3
|
export const SKILLS = {
|
|
4
|
-
"slates-character-
|
|
4
|
+
"slates-character-identity": "---\nname: slates-character-identity\ndescription: Build a Slates character from a reference image — generate one identity sheet and bind it to the character so the card updates live. Use when the user wants to create a character, build a character from an image, or starts a storyboard flow that needs consistent character references.\n---\n\n# Character identity sheet — Slates workflow\n\nA character's identity sheet is attached to **every** downstream generation that mentions it, so a flaw in the sheet becomes a flaw in every shot made from it. Building it well is the highest-leverage thing you can do for a project.\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n## The shape: ONE sheet, three panels\n\nSlates generates **one identity sheet per character**, bound as the character's canonical reference:\n\n| Panel | What it carries |\n|---|---|\n| **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |\n| **Full-body front, relaxed A-pose — framed from the collarbone down, head not shown** | Build, proportion, wardrobe. Headless on purpose: a front-facing body panel renders a ~40px face that can't match the portrait's, so the sheet would carry two competing identities and the model averages them. |\n| **Full-body back, head and hair visible** | Hair fall and the back of the outfit — the only panel where either reads. Keeps its head because there's no face to compete with. |\n\nThe rule is **kill every competing rendering of the FACE, not every head** — which is why exactly one body panel is headless.\n\nOn a deep neutral-grey plate (`#3a3a3c`), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look. Quadrupeds and non-bipedal characters are carved out — natural standing stance, head shown on both body panels.\n\n**The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.\n\n**Why one sheet.** Every `@character` mention attaches that character's canonical identity image, so each character costs one reference slot. It also reduces competing facial renderings to **one** — with the front panel headless and the back panel turned away, the portrait is the only face on the sheet, so there is nothing left to average.\n\n## Workflow\n\n### Get the reference\nThe user has either:\n- Pasted/uploaded an image of the character (real person, drawing, AI render).\n- Described the character in text only.\n\nIf image: upload it as a reference<!-- slates-only --> (`slates_upload_reference_image`)<!-- /slates-only -->.\nIf text only: generate from prompt-only — less consistent, so warn the user.\n\n<!-- slates-only -->\n### Create the character record\n`slates_create_character` with:\n- `name` (ask if not given)\n- `description` — 1-2 sentences, *visual* only (\"tall, dark hair, scar over left eye\"), not personality.\n- `style` — leave as the source's own medium by default. Only name a transform if the user wants one (e.g. anime → realistic).\n<!-- /slates-only -->\n\n### Generate the sheet\n<!-- slates-only -->\n`slates_generate_character_identity` with `characterId`, `projectId`, and `baseAssetId` (the source portrait).\n\n**Do not hand-write the sheet prompt.** Slates builds it from the canonical template in `@slatesvideo/shared/prompts` (`buildCharacterIdentityPrompt`) — panels, plate, lighting and craft clauses included — and appends your `userNotes`. Use `userNotes` for what the template can't know: *\"use the woman on the left\"*, *\"keep the scar on the right cheek\"*. A hand-written prompt is a fork of the template and will drift from it.\n\n- Estimate cost first with `slates_estimate_generation_cost` and announce in **credits** — never quote a price from memory.\n<!-- /slates-only -->\n\n- Default to Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted spend.\n- When the result returns inline, **evaluate it before binding**:\n - Is the portrait clearly the largest panel, and is it off-frontal?\n - **Is the front body panel cleanly headless** — an empty collar with the garment holding its shape, no partial face, no floating jaw, no smeared neck stump? A botched crop is worse than no crop.\n - Do the body panels read as the same build, wardrobe and hair as the portrait?\n - Catchlights present, irises readable rather than black holes?\n - Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?\n - Plate a flat deep grey, not white and not black?\n- If off: one focused refinement, then regenerate. The sheet is upstream of everything — it is worth a re-roll that a scene frame is not.\n<!-- slates-only -->\n- The op binds the result as the canonical identity automatically.\n\n### Hand back\n> \"Character {name} ready — identity sheet bound. Use `@{name}` in any prompt and Slates attaches it and names it inline, so the face stays consistent.\"\n<!-- /slates-only -->\n\n## How the reference gets used at scene time\n\nSlates cites the sheet inline under the character's name — `{name} (image N)` — in the exact order it sends references. That **name** is the anti-averaging lever, and it is each model's own official mechanism (NB2: \"assign a distinct name\"; Seedance: `Reference <Subject_N> in <Image_N>`; Kling: reuse a fixed label verbatim).\n\nCritically, the app injects **no** wardrobe, expression, or lighting directive. The user's scene prompt owns all of that — which is why `@{name}` dropped into a movie-still injection keeps the still's own clothing and lighting instead of dragging the sheet's.\n\n## Anti-patterns\n\n- **Don't** studio-light, white-background, or black-background the sheet. White bleeds into the video and washes out the location; black eats edge detail. Flat, even, shadowless light on a deep neutral grey.\n- **Don't** hand-write the sheet prompt when the op will build it — that is how the template and the shipped prompt fork.\n- **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.\n- **Don't** skip binding. An unbound asset doesn't help downstream.\n- **Don't** invent character details. Stick to what's in the reference image and the user's description.\n- **Don't** describe the headless front panel as removal or decapitation — in `userNotes` or any hand-written variant. The template asks for it as *framing* — \"cropped at the collarbone, head not shown, invisible-mannequin presentation\" — which is a standard e-commerce genre with deep training data. Removal phrasing is untested and invites a refusal.\n- **Don't** use 4K — wastes credits, no quality gain at sheet scale.\n- **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `slates-prompting-seedance`.\n",
|
|
5
5
|
"slates-content-policy": "---\nname: slates-content-policy\ndescription: Content-policy-safe construction — build any scene safe from the first word so it hits full cinematic impact without depicting prohibited content (and without getting silently rejected or degraded by the model's filter). Read this before writing any prompt that involves conflict, creatures, crowds, destruction, weapons, or young characters. Mirror of @slatesvideo/shared/prompts content-policy fragment — SSOT: second-brain business/projects/slates/product/prompting-ssot.md.\n---\n\n# Content-policy-safe construction — read before any risk-surface prompt\n\nDon't depict the harm — depict the energy, the aftermath, the threat, or the scale. Build the scene safe from the first word.\n\nWrite scenes that hit full cinematic impact without ever *needing* to depict prohibited content. This is a craft move, not a compromise — the substitutions below usually read as more cinematic, not less, and they keep your generation from getting silently rejected or degraded by the model's filter. A standoff is more tense than a massacre; an evacuated city is eerier than a crowd in panic; a roar lands harder than a kill. Load this whenever a prompt involves conflict, creatures, crowds, destruction, weapons, or young characters.\n\n## Substitution table\n\n| Avoid | Use instead |\n|---|---|\n| Civilians in panic, crowds fleeing under debris | An evacuated / empty city; abandoned streets; a lone figure for scale |\n| Weapons firing into buildings or at people | Energy-discharge standoffs, searchlights sweeping, charged auras, shockwaves with no muzzle fire |\n| Creatures tearing into each other, gore | A grapple / standoff — roars, near-misses, circling, an energy clash; combat that stays contained (in/on the water, never lifting into the air) |\n| Destruction with people in harm's way | Destruction in uninhabited terrain — glaciers, deserts, ruins, open sea, evacuated zones |\n| Realistic guns as the focus | Stylized / fantasy implements, weapons slung-not-fired, the weapon as silhouette or prop only |\n| Blood, wounds, death | Impact light, dust, debris, buckling and collapse, a silhouette dropping out of frame |\n| Real, named public figures | Original / anonymous characters |\n| Real brand logos | Original or generalized branding — except the user's own product, which is the whole point of a brand film |\n\n## The safe benchmark\n\nWhen a scene starts drifting risky, pull it back toward this shape: **one original creature, in a generalized monument or amphitheatre, in daylight, no weapons present, performing expressive action** (rising, roaring, spreading wings). Original design, generalized location, daylight, expressive rather than violent. That's the confirmed-safe envelope — most epic ideas re-stage into it without losing the punch.\n\n## Containment rule — it doubles as a physics win\n\nGive any creature or combat scene a **containment rule** that grounds the physics at the same time:\n- \"the fight STAYS at the sea surface — they breach, dive, grapple, submerge, but never fly or get carried into the air\"\n- \"boss scale locked ~2.5 human-heights, NOT kaiju-giant\"\n- \"destruction stays in the evacuated valley\"\n\nThis improves coherence (the model isn't inventing absurd escalation) AND keeps the scene inside policy — same clause buys both.\n\n## Scale and stakes without harm\n\nEpic stakes come from environmental danger and reaction, not depicted victims: tiny figures diving clear of *collapsing* terrain (not being crushed), a war-horn over an *empty* field, an army *scrambling* across a frozen valley as a titan tears free of a glacier. The danger is the environment; the figures are reacting, not dying. Snow plumes, glowing runes, splintering ice, shockwaves, and dust carry the chaos.\n\n## Minors — hard rule\n\nNever write romantic, sexual, or suggestive content involving or directed at minors, and never anything that sexualizes a young-presenting character. Any scene with children stays wholesome and age-appropriate. Non-negotiable — it overrides every stylistic goal.\n\n## Pre-flight (run before delivering any risk-surface prompt)\n\n- [ ] No civilians depicted in panic/harm; crowds are evacuated or absent.\n- [ ] No weapons firing at people/buildings; threat is energy / searchlight / silhouette.\n- [ ] No creature-on-creature or creature-on-person gore; combat is grapple / standoff / roar, contained.\n- [ ] Destruction is in uninhabited / evacuated terrain.\n- [ ] Creatures are original (\"not based on any franchise\"); no real public figures; no real brand logos except the user's own product.\n- [ ] Anything with children is wholesome and age-appropriate.\n\nIf a box fails, apply the substitution table before writing the prompt.\n\n## Editing real footage (Kling O3 edit / Omni Flash edit) — real people in the SOURCE\n\nVideo edit takes the user's own footage, which often contains real people. Rules:\n\n- The user must hold rights/consent for any real person's likeness in footage they edit — ask once when it's clearly someone other than the user, then proceed.\n- Kling's video-to-video filter behavior on real faces is **not yet verified** (unlike Seedance, where the consent-gated real-face route is confirmed). If an edit of real-person footage is rejected by the provider, do NOT retry-spam variations — tell the user the filter blocked it and offer a no-face crop/segment or an AI-character swap instead.\n- **Omni Flash: own-footage editing of the uploader's own face PASSED live 2026-07-09** (real talking-head clip, edited on our fal route) despite Google's documented \"recognizable people\" restriction — treat that restriction as aimed at third-party/public figures, but expect probabilistic refusals and never promise passage.\n- Never use edit to put a real, named public figure into a scene, or to make someone appear to say/do something they didn't. Faceless b-roll (hands, products, landscapes, crowds-from-behind) edits freely.\n\n## Gemini / Omni Flash filter regime (video gen + edit) — receipts 2026-07-09\n\nGoogle's filter is its own regime (stricter than fal-hosted Kling about harm-to-a-person, looser than BytePlus about faces). Live receipts:\n\n| Blocked (`content_policy_violation`) | Passed |\n|---|---|\n| \"his fingertips **ignite** with a small real flame\" (fire ON a body part = harm) | \"small **magical** flames appear on his fingertips … vanish when he blows on them\" |\n\n- **Harm-to-person framing is the tripwire**, not the effect itself. Reframe body-contact effects as magical / supernatural / harmless VFX: \"magical flames\", \"a glowing aura\", \"sparks of light dance on\". Avoid ignite / burn / on fire / catch fire applied to a person.\n- **Never use a real object as a metaphor** — \"candle-like flame\" rendered a literal candle in the subject's hand. Describe the effect, not an object that resembles it.\n- The block is a 422 refund (no credits lost) and arrives mid-generation — one reframe per the substitution mindset above, don't retry-spam.\n",
|
|
6
|
-
"slates-cost-discipline": "---\nname: slates-cost-discipline\ndescription: Mandatory pre-flight discipline before ANY generation call (image or video) — estimate cost, announce in credits, get confirmation, aggregate batches. Read this every time before calling slates_generate_image or any future slates_generate_* op. Skipping this risks burning the user's credits on guesses.\n---\n\n# Slates cost discipline — read before every generation\n\nGeneration costs real money. Every call is on the user's credits. The user can't see what you're about to spend until you tell them. **Tell them first, generate second.**\n\n## The 4 rules\n\n### 1. Pre-flight estimate — never call generate without one\n\nBefore ANY `slates_generate_*` call, run `slates_estimate_generation_cost` first. Inputs you must lock before estimating:\n\n- **Model** — derived from the op (`slates_generate_image` → `nano-banana-2-{resolution}`)\n- **Resolution** — never let the op default. Pick deliberately. Drafts → 1k. Hero → 2k. Print → 4k.\n- **Aspect ratio** — never let the op default to 1:1. Pick from the use case (cinematic → 16:9, mobile vertical → 9:16, square feed → 1:1).\n- **Count** — explicit. Don't generate 4 when 1 will tell you if the prompt works.\n\nIf aspect ratio or resolution isn't obvious from the user's request, **ask before estimating**. Don't guess.\n\n### 2. Announce in credits, plainly, before spending\n\nSlates bills abstract **credits** (they never expire). Announce the credit total the estimate returns — never dollars.\n\nFormat: `About to spend N credits on M image(s) at [resolution] [aspect ratio]. Proceed?`\n\nExamples:\n- `About to spend 4 credits on 1 image at 1k 16:9. Proceed?`\n- `About to spend 24 credits on 4 images at 2k 9:16 (variants). Proceed?`\n\nBelow ~7 credits you can proceed silently after announcing once. Above ~7 credits wait for explicit confirmation. Above ~17 credits the server itself will gate with `requires_confirm` — pass `confirm: true` only after the user explicitly OKs.\n\n### 3. Aggregate batches into ONE upfront announcement\n\nIf you're planning a multi-call workflow (5 storyboard frames, 3 character variants, a grid of options), **announce the total before the first call**, not five small announcements after the fact.\n\nFormat: `Plan: N generations totaling C credits. [Brief description of the sequence.] Proceed with the batch?`\n\nExample: `Plan: 6 frame generations at 1k 16:9 totaling 24 credits — establishing wide, push-in, two-shot, reverse, OTS, insert. Proceed?`\n\n### 3b. Batch authorization — one approval covers the enumerated batch\n\nWhen the user approves a batch plan with one aggregated cost total up front (\"8 scenes, ~$X total — go\"), that single approval authorizes `confirm=true` on **each enumerated call in that batch** — and nothing beyond it. You do not need to re-ask per call; that's the point of the upfront announcement. Hands-off multi-scene runs depend on this.\n\nBoundaries that re-trigger confirmation:\n\n- Any call's actual estimate exceeds what the announced plan implied for it by **>25%** → stop, surface the delta, get a fresh OK.\n- New calls are added that weren't in the enumerated plan (extra variants, retries beyond the plan, a new scene) → those are NOT covered. Announce and confirm separately.\n- The batch scope changes (different model, resolution, or duration than announced) → re-announce, re-confirm.\n\nOne approval = that plan, as enumerated, at those prices. Nothing else.\n\n### 4. Track the running total\n\nAfter each generation completes, the response includes `cost_credits` (when available). Keep a running tally in your context. Surface it every 3 generations or whenever the user asks \"how much have we spent?\"\n\n## Resolution decision rules\n\n| Use case | Resolution |\n|---|---|\n| First draft of a new prompt | 1k |\n| Storyboard frame (will likely regenerate) | 1k |\n| Hero shot, locked composition | 2k |\n| Print, marketing asset, final delivery | 4k |\n| Iterating to refine | match the previous resolution |\n\nResolution is a price lever, not a free choice: on Nano Banana 2 and FLUX.2 Max, 4k costs roughly 2x 1k (Seedream 5 Lite is flat-priced regardless of resolution). Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models` for current numbers instead of assuming. Pick the cheapest resolution that serves the use case.\n\n**4K VIDEO is Pro-only (2026-07-07).** The ladder above is for IMAGES (open at every tier). For VIDEO — Kling, Seedance, Veo — 4K requires a Slates Pro account; a base-tier 4K video gen is rejected server-side with `PRO_REQUIRED`. Default video to 1080p or lower and only reach for 4K when the user is on Pro and explicitly asks. 4K *images* are never gated.\n\n## Aspect ratio decision rules\n\nAsk the user when ambiguous. Otherwise:\n\n| Context cue | Aspect ratio |\n|---|---|\n| \"cinematic\", \"film\", \"movie\", \"wide\" | 16:9 |\n| \"TikTok\", \"Reels\", \"Story\", \"mobile vertical\", \"phone\" | 9:16 |\n| \"square\", \"Instagram feed\", \"thumbnail\" | 1:1 |\n| \"ultra-wide\", \"anamorphic\", \"cinemascope\" | 21:9 |\n| \"portrait\", \"magazine cover\", \"vertical\" | 4:5 or 2:3 |\n| \"landscape photo\", \"horizontal\" | 3:2 or 4:3 |\n\nIf the user prompt mixes signals (e.g. \"cinematic Instagram post\"), ask. Don't guess.\n\n## When the gate fires\n\nThe server returns `requires_clarification` when aspect ratio or resolution is missing. The server returns `requires_confirm` when total spend exceeds ~17 credits. In both cases:\n\n1. Surface the gate response to the user\n2. Get a clean answer\n3. Re-call with the explicit values + `confirm: true` if applicable\n\nDon't bypass the gate by silently filling in defaults. The gates exist because defaults waste money.\n\n## Video is slow + async — a timeout is NOT a failure\n\nVideo gens take minutes (Seedance 4K can run far longer). A client/CLI timeout or a slow, empty-looking response is **not** a failed generation — the job is still running on the provider.\n\n- **Never re-submit a video gen because it \"timed out.\"** That double-charges the user for one video. Re-rolling a slow gen is the single most expensive mistake here.\n- **Poll, don't re-roll.** Use `background: true` on `slates_generate_video`, then poll `slates_get_generation_status` (free, read-only) until it reports `completed` or `failed`. In-flight jobs survive app restarts and are recovered.\n- A gen has only failed when the status comes back `failed` — and a provider *rejection* **refunds** the credits, so failed isolation tests are ~free. Until you see a terminal status, the job is in flight. Wait.\n\n## The 3-strike rule\n\nStop after 3 iterations on the same prompt. Hand back to the user with what you tried and what's not working. The slot machine doesn't converge — if it's not landing, the prompt structure is wrong, not the seed.\n",
|
|
6
|
+
"slates-cost-discipline": "---\nname: slates-cost-discipline\ndescription: Mandatory pre-flight discipline before ANY generation call (image or video) — estimate cost, announce in credits, get confirmation, aggregate batches. Read this every time before calling slates_generate_image or any future slates_generate_* op. Skipping this risks burning the user's credits on guesses.\n---\n\n# Slates cost discipline — read before every generation\n\nGeneration costs real money. Every call is on the user's credits. The user can't see what you're about to spend until you tell them. **Tell them first, generate second.**\n\n## The 4 rules\n\n### 1. Pre-flight estimate — never call generate without one\n\nBefore ANY `slates_generate_*` call, run `slates_estimate_generation_cost` first. Inputs you must lock before estimating:\n\n- **Model** — derived from the op (`slates_generate_image` → `nano-banana-2-{resolution}`)\n- **Resolution** — never let the op default. Pick deliberately. Drafts → 1k. Hero → 2k. Print → 4k.\n- **Aspect ratio** — never let the op default to 1:1. Pick from the use case (cinematic → 16:9, mobile vertical → 9:16, square feed → 1:1).\n- **Count** — explicit. Don't generate 4 when 1 will tell you if the prompt works.\n\nIf aspect ratio or resolution isn't obvious from the user's request, **ask before estimating**. Don't guess.\n\n### 2. Announce in credits, plainly, before spending\n\nSlates bills abstract **credits** (they never expire). Announce the credit total the estimate returns — never dollars.\n\nFormat: `About to spend N credits on M image(s) at [resolution] [aspect ratio]. Proceed?`\n\nExamples:\n- `About to spend 4 credits on 1 image at 1k 16:9. Proceed?`\n- `About to spend 24 credits on 4 images at 2k 9:16 (variants). Proceed?`\n\nBelow ~7 credits you can proceed silently after announcing once. Above ~7 credits wait for explicit confirmation. Above ~17 credits the server itself will gate with `requires_confirm` — pass `confirm: true` only after the user explicitly OKs.\n\n### 3. Aggregate batches into ONE upfront announcement\n\nIf you're planning a multi-call workflow (5 storyboard frames, 3 character variants, a grid of options), **announce the total before the first call**, not five small announcements after the fact.\n\nFormat: `Plan: N generations totaling C credits. [Brief description of the sequence.] Proceed with the batch?`\n\nExample: `Plan: 6 frame generations at 1k 16:9 totaling 24 credits — establishing wide, push-in, two-shot, reverse, OTS, insert. Proceed?`\n\n### 3b. Batch authorization — one approval covers the enumerated batch\n\nWhen the user approves a batch plan with one aggregated cost total up front (\"8 scenes, ~$X total — go\"), that single approval authorizes `confirm=true` on **each enumerated call in that batch** — and nothing beyond it. You do not need to re-ask per call; that's the point of the upfront announcement. Hands-off multi-scene runs depend on this.\n\nBoundaries that re-trigger confirmation:\n\n- Any call's actual estimate exceeds what the announced plan implied for it by **>25%** → stop, surface the delta, get a fresh OK.\n- New calls are added that weren't in the enumerated plan (extra variants, retries beyond the plan, a new scene) → those are NOT covered. Announce and confirm separately.\n- The batch scope changes (different model, resolution, or duration than announced) → re-announce, re-confirm.\n\nOne approval = that plan, as enumerated, at those prices. Nothing else.\n\n### 4. Track the running total\n\nAfter each generation completes, the response includes `cost_credits` (when available). Keep a running tally in your context. Surface it every 3 generations or whenever the user asks \"how much have we spent?\"\n\n## Resolution decision rules\n\n| Use case | Resolution |\n|---|---|\n| First draft of a new prompt | 1k |\n| Storyboard frame (will likely regenerate) | 1k |\n| Hero shot, locked composition | 2k |\n| Print, marketing asset, final delivery | 4k |\n| Iterating to refine | match the previous resolution |\n\nResolution is a price lever, not a free choice: on Nano Banana 2 and FLUX.2 Max, 4k costs roughly 2x 1k (Seedream 5 Lite is flat-priced regardless of resolution). Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models` for current numbers instead of assuming. Pick the cheapest resolution that serves the use case.\n\n**4K VIDEO is Pro-only (2026-07-07).** The ladder above is for IMAGES (open at every tier). For VIDEO — Kling, Seedance, Veo — 4K requires a Slates Pro account; a base-tier 4K video gen is rejected server-side with `PRO_REQUIRED`. Default video to 1080p or lower and only reach for 4K when the user is on Pro and explicitly asks. 4K *images* are never gated.\n\n## Aspect ratio decision rules\n\nAsk the user when ambiguous. Otherwise:\n\n| Context cue | Aspect ratio |\n|---|---|\n| \"cinematic\", \"film\", \"movie\", \"wide\" | 16:9 |\n| \"TikTok\", \"Reels\", \"Story\", \"mobile vertical\", \"phone\" | 9:16 |\n| \"square\", \"Instagram feed\", \"thumbnail\" | 1:1 |\n| \"ultra-wide\", \"anamorphic\", \"cinemascope\" | 21:9 |\n| \"portrait\", \"magazine cover\", \"vertical\" | 4:5 or 2:3 |\n| \"landscape photo\", \"horizontal\" | 3:2 or 4:3 |\n\nIf the user prompt mixes signals (e.g. \"cinematic Instagram post\"), ask. Don't guess.\n\n## When the gate fires\n\nThe server returns `requires_clarification` when aspect ratio or resolution is missing. The server returns `requires_confirm` when total spend exceeds ~17 credits. In both cases:\n\n1. Surface the gate response to the user\n2. Get a clean answer\n3. Re-call with the explicit values + `confirm: true` if applicable\n\nDon't bypass the gate by silently filling in defaults. The gates exist because defaults waste money.\n\n## Video is slow + async — a timeout is NOT a failure\n\nVideo gens take minutes (Seedance 4K can run far longer). A client/CLI timeout or a slow, empty-looking response is **not** a failed generation — the job is still running on the provider.\n\n- **Never re-submit a video gen because it \"timed out.\"** That double-charges the user for one video. Re-rolling a slow gen is the single most expensive mistake here.\n- **Poll, don't re-roll.** Use `background: true` on `slates_generate_video`, then poll `slates_get_generation_status` (free, read-only) until it reports `completed` or `failed`. In-flight jobs survive app restarts and are recovered.\n- A gen has only failed when the status comes back `failed` — and a provider *rejection* **refunds** the credits, so failed isolation tests are ~free. Until you see a terminal status, the job is in flight. Wait.\n\n## 🔴 The still-gate — the most expensive mistake in the pipeline\n\n<!-- @inject:still-gate -->\n**A visible defect in the still is already a STOP.** Do not animate it. Fix the frame first, then move to motion — and go to motion only when the crop passes the still scan and you genuinely need movement to confirm an uncertain edge, reflection, or object.\n\nThis is a **cost** rule as much as a craft rule: a 1080p/10s premium video generation costs many multiples of an image re-roll, and video is where a defect stops being fixable. Anything wrong in the still gets worse in motion — soft geometry mushes, broken-but-plausible objects fall apart, oily textures start crawling. **Animating a known-bad frame is the single most expensive mistake in the pipeline.** Re-rolling the image is the cheap move; re-rolling the video is not.\n<!-- @end:still-gate -->\n\nThe check itself lives in `slates-vision-feedback-loop` (the four slop tells and the per-model accents). The **stop** is a cost rule and belongs here: before every image→video call, confirm the source frame passed the still scan. If it didn't, spending video credits on it is not iteration — it is buying a more expensive copy of a defect you already found.\n\n## The 3-strike rule\n\nStop after 3 iterations on the same prompt. Hand back to the user with what you tried and what's not working. The slot machine doesn't converge — if it's not landing, the prompt structure is wrong, not the seed.\n",
|
|
7
7
|
"slates-direct-response-ad": "---\nname: slates-direct-response-ad\ndescription: Build a 30-second hyper-motion direct-response ad in Slates from a product image and brief. Composes upload → storyboard → frame gen → motion gen → timeline → export. Use when the user drops a product image and asks for \"an ad\", \"a promo video\", \"a TikTok ad\", \"an Instagram ad\", a launch video, or any short-form direct-response video built around a product.\n---\n\n# Direct-response ad — Slates workflow\n\nYou are building a 30-second hyper-motion direct-response ad. The user has handed you a product image (or product URL) and a short brief. Slates desktop is open on the second monitor; the user watches it populate as you work.\n\n**Hard rules**\n\n- Always estimate cost before generating. Use `slates_estimate_generation_cost` and surface the total.\n- All Slates generation routes through Slates Credits, period (BYOK is retired) — don't suggest \"use your own keys\" workarounds.\n- Default model: `nano-banana-2-2k`. For close-up product hero frames step up to `4k` only if the user asks.\n- Hyper-motion = punchy cuts, 4 frames in 30 seconds, ~7s each. Don't over-storyboard.\n\n## Workflow\n\n### 1. Set up the project\n- Create a project named for the product (`slates_create_project`).\n- If the user gave a product image as a file path, upload it (`slates_upload_reference_image`).\n- If they pasted base64 / a data URL, use the same op with `dataUrl`.\n\n### 2. Generate the storyboard frames\nBuild exactly 4 frames in this order:\n\n| # | Beat | Visual goal |\n|---|------|-------------|\n| 1 | Hook | Hyper-close-up of the product, dramatic light, motion blur edge |\n| 2 | Lifestyle | Real person using/wearing/holding the product, eye contact |\n| 3 | Problem→solution | The before/after moment that justifies the buy |\n| 4 | CTA | Clean product hero with mental room for an overlaid CTA in editing |\n\nFor each frame:\n1. Draft a tight 1-2 sentence prompt (visual only — no copy text in the image).\n2. Reference the product upload's URL or asset ID for visual fidelity.\n3. Call `slates_generate_image` with that prompt + reference. **You see the result inline — evaluate it.**\n4. If it's wrong: refine prompt, regenerate. If it's right: bind it as a frame in the storyboard (`slates_add_frame`).\n\n### 3. Build the storyboard\n- `slates_create_storyboard` named \"30s ad — v1\".\n- Default scene already exists. Add 3 more scenes (\"Hook\", \"Lifestyle\", \"Problem-Solution\", \"CTA\") via `slates_add_scene`, or just add all 4 frames to the default scene.\n- For each generated image, add a frame referencing the asset id (`slates_add_frame`).\n\n### 4. Hand back to the user\n- Surface estimated total credits spent.\n- Tell the user the storyboard is ready and they can either:\n - **In Slates desktop:** click each frame to generate motion (the existing UI handles motion generation).\n - **Continue here:** ask you to keep going.\n\n### 5. If they say keep going — motion, assembly, export\n- Generate motion per frame with `slates_generate_video` (`firstFrameAssetId` = the frame's asset, `background: true`), routed per `slates-model-selection` (Kling 3.0 std 8s by default; Seedance 2 for any physics-heavy beat like the hook). Submit all four, then poll `slates_get_generation_status` until each completes (1-5 min).\n- Assemble: `slates_add_clip_to_timeline` for each completed clip in beat order (Hook → Lifestyle → Problem-Solution → CTA). Verify with `slates_get_timeline`; fix order with `slates_reorder_clips`.\n- Export: `slates_export_video` to an absolute `.mp4` path (default `<slates_get_project_directory>/exports/<product>-ad.mp4`), then `slates_reveal_file` so the user sees the file.\n- Full pipeline doctrine (batch cost authorization, model mixing, multi-take selection): `slates-one-prompt-film`.\n\n## Anti-patterns\n\n- **Don't** generate text overlays in the image. Slates renders captions/CTAs at the editor stage.\n- **Don't** burn credits on slot-machine prompting. If the first generation is off, refine the prompt; don't just regenerate.\n- **Don't** skip the cost estimate. Confirm with the user above ~17 credits.\n- **Don't** invent visual specifics about the product (colors, textures, angles) that aren't in the reference image. Reference-anchored prompts only.\n\n## Voice\n\nThe ad lives or dies on the hook frame. Tight, sensory, no fluff. Match the user's brand. Default tone is \"scroll-stopping\" not \"informative.\"\n",
|
|
8
|
-
"slates-edit-and-iterate": "---\nname: slates-edit-and-iterate\ndescription: Iterate on an existing Slates asset — re-evaluate, refine prompt, regenerate or edit. Use when the user has an existing generated image in Slates and wants to \"tweak it\", \"change one thing\", \"make it warmer\", \"remove the second person\", or any other surgical refinement instead of full regeneration.\n---\n\n# Edit and iterate — Slates workflow\n\nThe user already has a generated image in Slates and wants to refine it. The vision-feedback-loop skill defines the general pattern; this skill is the specific recipe for \"I have asset X, here's what's wrong with it.\"\n\n## Workflow\n\n### 1. Pull the current asset back into context\n- The user references an asset by id, frame number, or \"the latest one.\"\n- Resolve to an asset id (`slates_list_assets` if needed).\n- `slates_get_asset_image` with that id to load it inline. **You see the image.**\n\n### 2. Identify the delta\nThe user's request is one of:\n- **Surgical** — \"remove the second figure\", \"make the sword red\", \"swap the background to a bamboo forest\".\n- **Aesthetic** — \"warmer light\", \"more dramatic\", \"softer focus\".\n- **Compositional** — \"wider shot\", \"lower angle\", \"centered subject\".\n- **Wholesale** — \"actually let's try a totally different look.\"\n\n### 3. Pick the right tool\n| Delta type | Approach |\n|---|---|\n| Surgical | `slates_edit_image` — `sourceAssetId` = the original, `prompt` = the change only (\"remove the second figure\"), not a re-description of the whole image. |\n| Aesthetic / compositional | `slates_generate_image` with the original in `referenceAssetIds` + a refined prompt. Don't re-roll from scratch. |\n| Wholesale | New prompt, no reference, fresh generation. Treat as a new brief. |\n\n**`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Default model `nano-banana-2` — the only edit model that also takes extra `referenceAssetIds`; `flux-2-max` / `seedream-5-lite` use their own edit endpoints and ignore references. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost above ~17 credits gates on `confirm=true`.\n\n### 4. Generate, evaluate, decide\n- Estimate cost first.\n- After generation, the result is inline. Compare side-by-side with the original (`slates_get_asset_image` again).\n- If the delta is correct: bind to the same
|
|
9
|
-
"slates-model-selection": "---\nname: slates-model-selection\ndescription: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 only) and never the default.\n---\n\n# Model selection — the routing doctrine\n\nPick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.\n\n## Video routing\n\n| Job | Model | Why |\n|---|---|---|\n| **General-purpose — the default for most shots** | **Kling 3.0 std** | Cost-effective workhorse. Strong image-to-video: preserves identity, layout, and text from the start frame. Any aspect ratio, 5–15s. |\n| Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |\n| Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |\n| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |\n| The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |\n| Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | The only job Veo wins. |\n\n## Video EDIT routing (changing an existing clip)\n\n| Job | Tool | Why |\n|---|---|---|\n| **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + \"Keep everything else the same\"; long prompts destroy it (see below). |\n| **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |\n| **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but \"near-identical\" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |\n| Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |\n| AI-edit the user's OWN footage | Omni Flash Edit (3–10s) or Kling O3 Edit (3–15s, 720–3840px) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |\n\n- **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.\n- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the pro demos (e.g. Higgsfield's split-screen short) actually work, plus gesture-only beats with voiceover laid over in post.\n- **One change per pass, short prompts.** On Omni Flash this is documented law (\"overly descriptive prompts can lead to unintended changes\" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.\n- Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.\n\n## Motion Transfer & Lip Sync routing (two engines per tool)\n\nBoth tools have a cheap Kling utility lane and a premium Seedance lane. The capability is the same; the execution model differs: Kling bolts motion/lip onto the source as a dedicated post-process; Seedance generates in a single pass with the driving clip / dialogue as native conditioning signals — better motion fidelity, natural speech delivery, audio included.\n\n| Job | Engine | Why |\n|---|---|---|\n| Quick motion retarget, budget lane, or driving clip >15s | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |\n| **Motion transfer where fidelity or audio matters** — dance, choreography, cinematic action | **Seedance 2.0** (`motionModel=seedance-2`) | Single-pass conditioning beats post-hoc retargeting; prompt-driven; native audio. Driving clip 2–15s; bills input+output seconds (vref keys). |\n| Cheap lip-sync utility (re-voice a clip, simple avatar) | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |\n| **Natural speech, voice cloned from the source clip, premium delivery** | **Seedance 2.0** (`engine=seedance-2`) | The line is spoken IN the generation (no TTS layer); a video source keeps its own voice; uploaded ≤15s audio can drive it. |\n\n- Faces: Seedance tool gens default `seedanceFace=true` (sources are people). A REAL person triggers the consent cascade (`[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent`, premium realface pricing).\n- Billing: any Seedance gen with a video reference bills COMBINED input+output seconds (`seedance-2*-vref-*` keys) — always pass the clip duration and quote before confirming.\n\n**Rules:**\n\n- **Default video = Kling 3.0 std.** Escalate to Seedance the moment the shot has physics/effects weight or is the hero moment — and say why in the plan (\"physics-heavy, routing to Seedance\").\n- **Veo is never the default.** 16:9 only, 4/6/8s only, and it is not the quality pick — treat it as a single-purpose tool for native-synced-audio shots. If audio can be added after (Kling lip-sync, edit stage), prefer Kling or Seedance + audio in post.\n- **9:16 vertical → Kling or Seedance.** Veo can't.\n- **Image-to-video from an NB2 start frame** (the standard pipeline) → Kling by default, Seedance when the motion is physics-heavy. Not Veo.\n- **User names a model explicitly → use it.** But if it's a mismatch for the job (crazy physics on Kling std, vertical on Veo), say so in one line and offer the right route before generating.\n\n## Image routing\n\n**Video models (Kling, Seedance, Veo) cannot generate standalone images — ever.** A \"premium hero reference image\" is still an image job: it routes to an image model below, never to Seedance.\n\n- **Default: Nano Banana 2** — best reference handling (14 refs), best legible text, the standard start-frame generator.\n- **NB2 Lite** — the fast/draft seat: ~half NB2's price, ~2.7× faster, 1K only. Route iteration volume and drafts here; finals go back to NB2 full (2K/4K).\n- **Nano Banana Pro** — the hero-frame/typography ceiling (~2× NB2). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — feed it a full subject library.\n- **GPT Image 2** — readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. Medium quality is the default (half NB2's price at 1080p); high (~4×) only when text precision is the whole job. 4K at both tiers is API-only — even paid ChatGPT can't render it.\n- **FLUX.2 Max** — photoreal texture, hex-color binding, typography, less censored.\n- **Seedream 5 Lite** — uncensored + any-resolution flat price; volume exploration when the Gemini filter is in the way.\n\n**Split rule of thumb:** readable text / panels / UI → GPT Image 2; photoreal, character-locked, widescreen, or edit-heavy → the Banana line; drafts → NB2 Lite; uncensored or odd resolutions → Seedream/FLUX.\n\n## Cost is a tiebreaker, not the router\n\nRoute by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.\n",
|
|
10
|
-
"slates-one-prompt-film": "---\nname: slates-one-prompt-film\ndescription: The full one-prompt-to-finished-film pipeline in Slates — script, project, characters, storyboard, frame images, video generation, timeline assembly, MP4 export. Use when the user gives one idea and wants a finished video out the other end - \"make me a video about X\", \"turn this idea into an ad\", \"make a short film from this\", \"one prompt, finished film\". This is the master recipe; other Slates skills are its sub-steps.\n---\n\n# One prompt → finished film — Slates master pipeline\n\nThe user gives an idea. You hand back an MP4 on disk. Everything in between is yours, with exactly TWO mandatory user checkpoints: the creative plan, and ONE aggregated cost approval.\n\n## The pipeline\n\n### 1. Script the beats\nTurn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.\n\n### 2. Set up the project\n- `slates_create_project` named for the piece.\n- Recurring character? Build it properly — `slates_create_character` + the `slates-character-
|
|
8
|
+
"slates-edit-and-iterate": "---\nname: slates-edit-and-iterate\ndescription: Iterate on an existing Slates asset — re-evaluate, refine prompt, regenerate or edit. Use when the user has an existing generated image in Slates and wants to \"tweak it\", \"change one thing\", \"make it warmer\", \"remove the second person\", or any other surgical refinement instead of full regeneration.\n---\n\n# Edit and iterate — Slates workflow\n\nThe user already has a generated image in Slates and wants to refine it. The vision-feedback-loop skill defines the general pattern; this skill is the specific recipe for \"I have asset X, here's what's wrong with it.\"\n\n## 🔴 The master rule — an edit is a LEAF, not a node\n\n**Never re-edit an edit. Always go back and re-edit the master.**\n\nEvery edit model silently re-renders the **whole frame**, not just the region you named. So the parts you didn't ask to change come back slightly different every pass — softer texture, drifted colour, mushier fine detail. It is barely visible after one edit and obvious by the second. Chaining edits compounds the damage and there is no way to undo it, because each generation *is* the new source.\n\nThe fix is structural, not a matter of care:\n\n- **Want two changes?** Make them in ONE edit off the master, or make them as two separate edits **both taken from the master**, then keep whichever you prefer.\n- **An edit came back wrong?** Do NOT edit the result to fix it. Discard it and re-edit the master with a better instruction.\n- **Only the changed region is worth keeping?** That is a compositing job — the edit supplies the new region, the untouched master supplies everything else.\n\nSlates records this: an edit result carries `sourceAssetIds` pointing at the asset it was made from, so **you can tell whether the thing you are about to edit is itself an edit.** Check before you edit — `[Edit]`-prefixed prompts and a populated source lineage both say \"this is a leaf; go back to its parent.\"\n\n## Workflow\n\n### 1. Pull the current asset back into context\n- The user references an asset by id, frame number, or \"the latest one.\"\n- Resolve to an asset id (`slates_list_assets` if needed).\n- `slates_get_asset_image` with that id to load it inline. **You see the image.**\n\n### 2. Identify the delta\nThe user's request is one of:\n- **Surgical** — \"remove the second figure\", \"make the sword red\", \"swap the background to a bamboo forest\".\n- **Aesthetic** — \"warmer light\", \"more dramatic\", \"softer focus\".\n- **Compositional** — \"wider shot\", \"lower angle\", \"centered subject\".\n- **Wholesale** — \"actually let's try a totally different look.\"\n\n### 3. Pick the right tool\n| Delta type | Approach |\n|---|---|\n| Surgical | `slates_edit_image` — `sourceAssetId` = the original, `prompt` = the change only (\"remove the second figure\"), not a re-description of the whole image. |\n| Aesthetic / compositional | `slates_generate_image` with the original in `referenceAssetIds` + a refined prompt. Don't re-roll from scratch. |\n| Wholesale | New prompt, no reference, fresh generation. Treat as a new brief. |\n\n**`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Default model `nano-banana-2` — the only edit model that also takes extra `referenceAssetIds`; `flux-2-max` / `seedream-5-lite` use their own edit endpoints and ignore references. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost above ~17 credits gates on `confirm=true`.\n\n### 4. Generate, evaluate, decide\n- Estimate cost first.\n- After generation, the result is inline. Compare side-by-side with the original (`slates_get_asset_image` again).\n- If the delta is correct: bind to the same role (frame, character identity, etc.) the original was bound to.\n- If the delta missed: one focused refinement, then regenerate. Cap at 3 tries.\n\n### 5. Hand back\n- \"Asset updated. Frame 3 now uses {new_asset_id}.\"\n- Always note what changed and what didn't, so the user can see the surgery worked: \"Lighting shifted to warmer, composition unchanged.\"\n\n## Anti-patterns\n\n- **Don't** delete the original asset until the user confirms the new one. Slates keeps both; the user picks.\n- **Don't** mix surgical and wholesale changes in one regeneration. The user said \"make it warmer\" — don't also reframe the shot.\n- **Don't** re-generate when `slates_edit_image` would work. Edits preserve composition and identity; full regen rolls the dice.\n- **Don't** edit an edit — ever. Not once, not \"just a small one.\" Go back to the master (see the master rule above). Every attempt re-renders the full frame and the degradation is cumulative and permanent.\n- **Don't** keep re-rolling the same failed edit. If three tries off the master didn't land, the brief is wrong, not the model — check in with the user.\n",
|
|
9
|
+
"slates-model-selection": "---\nname: slates-model-selection\ndescription: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 only) and never the default.\n---\n\n# Model selection — the routing doctrine\n\nPick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.\n\n## 🔑 The meta-rule — above the table\n\nThe tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:\n\n> **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.\n>\n> **Inspect the output at its intended crop.** A frame that holds up as a thumbnail can fall apart at the size it will actually be watched. For a location, look at atmosphere, material texture, and anchor objects; for a character, identity, skin, pose, and gradients.\n>\n> **Choose the model that PROVES that requirement** and leaves only failures you can afford to rerun or mask.\n>\n> **When the roster changes, repeat the evidence test.** Do not carry today's ranking forward on reputation.\n\n## Video routing\n\n| Job | Model | Why |\n|---|---|---|\n| **General-purpose — the default for most shots** | **Kling 3.0 std** | Cost-effective workhorse. Strong image-to-video: preserves identity, layout, and text from the start frame. Any aspect ratio, 5–15s. |\n| Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |\n| Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |\n| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |\n| The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |\n\n### Named Seedance escalation triggers\n\n\"Physics matter\" is an abstract category and it under-fires. These are the beats Seedance is **observably** good at — if the shot contains one, escalate without deliberating:\n\n- **Real-time → slow-motion contrast.** The signature beat; nearly every strong clip rides it.\n- **The camera moving while debris, meteors, sparks or particles crash around the subject.** Distinctly a feature of this model, not just a thing it survives.\n- **Massive scale that has to read as genuinely huge** — not \"a big thing\", a thing whose size is the point of the shot.\n- **One continuous unbroken take.**\n\nConcrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).\n| Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | The only job Veo wins. |\n\n## Video EDIT routing (changing an existing clip)\n\n| Job | Tool | Why |\n|---|---|---|\n| **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + \"Keep everything else the same\"; long prompts destroy it (see below). |\n| **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |\n| **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but \"near-identical\" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |\n| Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |\n| AI-edit the user's OWN footage | Omni Flash Edit (3–10s) or Kling O3 Edit (3–15s, 720–3840px) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |\n\n- **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.\n- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.\n- **One change per pass, short prompts.** On Omni Flash this is documented law (\"overly descriptive prompts can lead to unintended changes\" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.\n- Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.\n\n## Motion Transfer & Lip Sync routing (two engines per tool)\n\nBoth tools have a cheap Kling utility lane and a premium Seedance lane. The capability is the same; the execution model differs: Kling bolts motion/lip onto the source as a dedicated post-process; Seedance generates in a single pass with the driving clip / dialogue as native conditioning signals — better motion fidelity, natural speech delivery, audio included.\n\n| Job | Engine | Why |\n|---|---|---|\n| Quick motion retarget, budget lane, or driving clip >15s | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |\n| **Motion transfer where fidelity or audio matters** — dance, choreography, cinematic action | **Seedance 2.0** (`motionModel=seedance-2`) | Single-pass conditioning beats post-hoc retargeting; prompt-driven; native audio. Driving clip 2–15s; bills input+output seconds (vref keys). |\n| Cheap lip-sync utility (re-voice a clip, simple avatar) | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |\n| **Natural speech, voice cloned from the source clip, premium delivery** | **Seedance 2.0** (`engine=seedance-2`) | The line is spoken IN the generation (no TTS layer); a video source keeps its own voice; uploaded ≤15s audio can drive it. |\n\n- Faces: Seedance tool gens default `seedanceFace=true` (sources are people). A REAL person triggers the consent cascade (`[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent`, premium realface pricing).\n- Billing: any Seedance gen with a video reference bills COMBINED input+output seconds (`seedance-2*-vref-*` keys) — always pass the clip duration and quote before confirming.\n\n**Rules:**\n\n- **Default video = Kling 3.0 std.** Escalate to Seedance the moment the shot has physics/effects weight or is the hero moment — and say why in the plan (\"physics-heavy, routing to Seedance\").\n- **Veo is never the default.** 16:9 only, 4/6/8s only, and it is not the quality pick — treat it as a single-purpose tool for native-synced-audio shots. If audio can be added after (Kling lip-sync, edit stage), prefer Kling or Seedance + audio in post.\n- **9:16 vertical → Kling or Seedance.** Veo can't.\n- **Image-to-video from an NB2 start frame** (the standard pipeline) → Kling by default, Seedance when the motion is physics-heavy. Not Veo.\n- **User names a model explicitly → use it.** But if it's a mismatch for the job (crazy physics on Kling std, vertical on Veo), say so in one line and offer the right route before generating.\n\n## Image routing\n\n**Video models (Kling, Seedance, Veo) cannot generate standalone images — ever.** A \"premium hero reference image\" is still an image job: it routes to an image model below, never to Seedance.\n\n- **Default: Nano Banana 2** — best reference handling (14 refs), best legible text, the standard start-frame generator.\n- **NB2 Lite** — the fast/draft seat: ~half NB2's price, ~2.7× faster, 1K only. Route iteration volume and drafts here; finals go back to NB2 full (2K/4K).\n- **Nano Banana Pro** — the hero-frame/typography ceiling (~2× NB2). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — feed it a full subject library.\n- **GPT Image 2** — readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. Medium quality is the default (half NB2's price at 1080p); high (~4×) only when text precision is the whole job. 4K at both tiers is API-only — even paid ChatGPT can't render it.\n- **FLUX.2 Max** — photoreal texture, hex-color binding, typography, less censored.\n- **Seedream 5 Lite** — uncensored + any-resolution flat price; volume exploration when the Gemini filter is in the way.\n\n**Split rule of thumb:** readable text / panels / UI → GPT Image 2; photoreal, character-locked, widescreen, or edit-heavy → the Banana line; drafts → NB2 Lite; uncensored or odd resolutions → Seedream/FLUX.\n\n## Cost is a tiebreaker, not the router\n\nRoute by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.\n",
|
|
10
|
+
"slates-one-prompt-film": "---\nname: slates-one-prompt-film\ndescription: The full one-prompt-to-finished-film pipeline in Slates — script, project, characters, storyboard, frame images, video generation, timeline assembly, MP4 export. Use when the user gives one idea and wants a finished video out the other end - \"make me a video about X\", \"turn this idea into an ad\", \"make a short film from this\", \"one prompt, finished film\". This is the master recipe; other Slates skills are its sub-steps.\n---\n\n# One prompt → finished film — Slates master pipeline\n\nThe user gives an idea. You hand back an MP4 on disk. Everything in between is yours, with exactly TWO mandatory user checkpoints: the creative plan, and ONE aggregated cost approval.\n\n## The pipeline\n\n### 1. Script the beats\nTurn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.\n\n**Surface a decision log with the plan.**\n\n<!-- @inject:decision-log -->\nWhen you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:\n\n```\nsource phrase or declared default → what you wrote → what it resolves\n\"in a diner\" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable\n(no time of day) → late afternoon, low warm key → default; say the word and it changes\n(no camera) → slow push-in, single move → one move per shot; stacking increases instability\n```\n\n**Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the \"why did you add that?\" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.\n\n> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.\n<!-- @end:decision-log -->\n\nA 4-10 shot script is where you invent the most on the user's behalf — time of day, wardrobe, weather, lens feel, camera moves the brief never mentioned. The log is what makes those visible while they are still free to change.\n\n### 2. Set up the project\n- `slates_create_project` named for the piece.\n- Recurring character? Build it properly — `slates_create_character` + the `slates-character-identity` recipe — so every frame references the same identity.\n- Recurring location? `slates_create_environment`.\n- One-off shots don't need character/environment records; skip the ceremony.\n\n### 3. Storyboard skeleton (no generation yet)\n- `slates_create_storyboard`, `slates_add_scene` per script scene.\n- Structure first, spend second — the user catches script problems on the free skeleton, not on burned credits.\n\n### 4. ONE aggregated cost approval — then hands-off\nPrice the whole batch before the first generation: frame images (count × model — `slates_estimate_generation_cost`) + video gens (count × model × duration). Present a single total:\n\n> Plan: 6 frames at 1k 16:9 + 5 × 8s Kling 3.0 std + 1 × 8s Seedance 2 hero shot ≈ $X.XX total. Proceed with the batch?\n\nPer `slates-cost-discipline` 3b: that single OK authorizes `confirm=true` for **every enumerated call in the batch** — no per-call re-asking. Re-confirm only if a call's price overruns the plan >25% or new calls get added (extra retakes, new shots).\n\n### 5. Generate frame images\nPer shot: `slates_generate_image` with `referenceAssetIds` pointing at the character identity / environment / prior frames for consistency (Slates names each reference inline as \"image N\" — you don't hand-write role labels; reuse the same subject name across shots). Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`.\n\n**Multi-take where it matters:** for the hook shot and any shot the whole film hangs on, generate 2-4 variants (cheap model or 1k), pull them back with `slates_get_assets_batch`, pick the strongest on composition + identity, discard the rest. Don't multi-take filler shots.\n\n### 6. Generate video per frame — background mode\n`slates_generate_video` with `firstFrameAssetId` = the bound frame, `background: true`. Submit ALL shots, collect the generationIds, then poll `slates_get_generation_status` every 10-15s (1-5 min per gen; they survive app restarts). This parallelizes a 6-shot film into one wait instead of six.\n\n**Model mixing — route per `slates-model-selection`** (details in the per-model guides):\n- **Kling V3** (`slates-prompting-kling-v3`): the DEFAULT for most shots — any aspect ratio, 5-15s, strong start-frame adherence; std is the workhorse, Omni for multi-character dialogue.\n- **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K (4K video is Pro-only).\n- **Veo 3.1** (`slates-prompting-veo-3`): niche, never the default — only when native synced audio must generate WITH the video in one gen; 16:9 only, 4/6/8s.\n\nFailed gen? Check the error via `slates_get_generation_status`, fix the prompt, resubmit that one shot (a retry beyond the plan = announce the delta cost).\n\n### 7. Assemble the timeline\n- `slates_get_timeline` once to get the lay of the land.\n- `slates_add_clip_to_timeline` for each completed video asset **in story order** — defaults append back-to-back on the first video track, which is exactly an assembly cut.\n- Order wrong? `slates_reorder_clips` with the full clip-id list. Dropped a shot? `slates_remove_clip`, then reorder to close the gap.\n\n### 8. Export + deliver\n- Output path: ask the user, or default to `<slates_get_project_directory>/exports/<name>.mp4`.\n- `slates_export_video` (absolute path, `.mp4`; blocks while ffmpeg renders — minutes for long timelines).\n- `slates_reveal_file` so the file is literally in front of them.\n- Offer the finishing path: `slates_export_timeline_xml` → DaVinci Resolve (File → Import → Timeline) for grading, sound, and titles.\n\n### 9. Report\nShots delivered, total spent vs. approved plan, the export path, and the single best next lever (\"re-take shot 3 with a tighter prompt\" / \"add a CTA end-card\").\n\n## Hard rules\n\n- **Two checkpoints only.** Creative plan (step 1) and total cost (step 4). Everything else runs without asking — that's the product promise.\n- **Skeleton before spend.** Project + storyboard structure are free; generation isn't.\n- **Look at everything.** Every image inline, every video via `slates_get_asset_video_frames` if a clip seems off. Never assemble a timeline from clips you haven't evaluated.\n- **3-strike rule per shot.** Three failed takes on one shot = stop, show the user what you tried, ask.\n- **Consistency comes from references, not luck.** Same identity asset on every character frame; same environment refs across a location's shots.\n",
|
|
11
11
|
"slates-project-organization": "---\nname: slates-project-organization\ndescription: How a Slates project's assets are organized AND named — the asset short-code system (IMG-A12 / VID-V3 / AUD-S1 badges on every gallery card), folders for film STRUCTURE, the typed tabs for reusable references. Read when the user refers to an asset by code, asks what a code like IMG-A36 means, or when organizing/navigating a project.\n---\n\n# Organizing a Slates project\n\nSlates already gives each REUSABLE reference type its own home — the **Characters**, **Environments**, and **Styles** tabs, each with its own generation + `@mention`/`#ref` behavior. Do NOT recreate those as folders. Folders are for **structure**, never type.\n\n**Folders = where an asset sits in the FILM**, and they mirror to real subfolders on disk (`projects/<id>/…`), so a human can open the project in Resolve/Finder and navigate it like an edit. Use them for work product, not references.\n\nCreate with `slates_create_folder`; file assets with `slates_move_assets_to_folder`. Generations land in the project's active folder, so set it before a batch.\n\nConventions by project type:\n- **Short film / narrative:** `Shots` (scene stills) · `Clips` (generated video) · `Final` (the export). Use one folder per scene (`Scene 1`, `Scene 2`, …) instead when the piece has distinct locations/beats.\n- **Ad / UGC:** `Hooks` · `B-roll` · `Talking-head` · `Final`.\n\nRules of thumb:\n- Reusable cast / sets / look → leave in the Characters/Environments/Styles tabs. Don't fold them.\n- Scene stills, clips, and the final cut → file into the structural folder they belong to, as you make them.\n- One folder per asset (folders are structure). Cross-cutting status (hero take, reject, variant) is a tag concern, not a folder.\n- Keep the gallery legible: work product lives in folders; the reference scaffolding (sheets, plates, style images) stays in its tabs.\n\n## Asset codes — the shared vocabulary (IMG-A12 / VID-V3 / AUD-S1)\n\nEvery asset gets a short, stable code the moment it lands in a project, and the user sees it as the badge in the **top-left corner of every image and video card** in the gallery. This is the shared vocabulary between you and the user — it exists so neither of you ever has to quote a UUID.\n\n**The scheme:**\n- `IMG-A{n}` = images · `VID-V{n}` = videos · `AUD-S{n}` = audio.\n- Numbering is **per project, per type**, counts up from 1, and **numbers are never reused** — deleting IMG-A12 doesn't renumber anything, so a code always means the same asset forever.\n- Each asset also carries a **label**: the first ~4 meaningful words of its prompt, title-cased. Chat format is code + label: `IMG-A12 — Beach Sunset`.\n\n**How to use it:**\n- **User names a code** (\"use IMG-A36 as the reference\", \"animate VID-V3's last frame\") → resolve it via `slates_list_assets` (match the `code` field) to get the assetId, confirm back in the same vocabulary: \"Got it — IMG-A36 — Marcus Rooftop Close-Up as the first frame.\"\n- **You name assets** → ALWAYS code + label, never UUID, never \"the beach one\" (which of three?). The user matches your words to the badge by eye.\n- **User seems confused** about what a code is or how to point you at an image → explain it in one line: \"Every image and video in your gallery has a code badge in its top-left corner — like IMG-A36. Just say that code and I'll know exactly which one you mean.\"\n- **Ambiguity** (\"the sunset image\" when several exist) → pull candidates with `slates_get_assets_batch` and offer the codes: \"I see IMG-A12, IMG-A19, and IMG-A24 with sunsets — which one?\"\n",
|
|
12
|
-
"slates-prompting-flux-2-max": "---\nname: slates-prompting-flux-2-max\ndescription: How to prompt FLUX.2 Max (Black Forest Labs image model). Read before calling slates_generate_image with model flux-2-max, or slates_edit_image with editModel flux-2-max. FLUX.2 wants front-loaded structure, real camera vocabulary, and positive-only phrasing — no negative prompts, no tag soup.\n---\n\n# FLUX.2 Max — prompting\n\nBlack Forest Labs' top image model, routed via fal.ai. In Slates: `slates_generate_image` with `model: flux-2-max` (REQUIRES projectId — no headless path), priced per resolution (1k/2k/4k — call `slates_estimate_generation_cost` for current numbers, never quote from memory). Strengths vs Nano Banana 2: photoreal texture, less censored, precise hex-color control, strong typography. Reference images route through FLUX's edit endpoint and carry a lower per-model cap than NB2's 14.\n\n## Core structure — front-load what matters\n\n```\nSubject + Action + Style + Context\n```\n\nWord order is weight. FLUX.2 attends hardest to the start of the prompt: main subject → key action → critical style → essential context → secondary details.\n\n**Length:** 10-30 words for concept tests, 30-80 words for most work, 80+ only for genuinely complex scenes.\n\n## Photorealism: name real gear, not \"professional photo\"\n\nThe single biggest realism lever is concrete camera vocabulary:\n\n```\nShot on Hasselblad X2D, 80mm lens, f/2.8, natural lighting\nShot on Sony A7IV, 35mm, golden hour, shallow depth of field\nKodak Portra 400, natural grain, organic colors\n```\n\nEra cues work the same way: \"early digital camera, slight noise, flash photography, candid\" reads 2000s digicam; \"film grain, warm color cast, soft focus\" reads 80s.\n\nFor portraits add: natural skin texture, realistic pores, subtle imperfections, soft diffused lighting.\n\n## No negative prompts — reframe positively\n\nFLUX.2 has no negative prompt support. Describe the presence you want, not the absence:\n\n- ❌ \"no blur\" → ✅ \"sharp focus throughout\"\n- ❌ \"no people\" → ✅ \"empty scene\"\n- ❌ \"no harsh shadows\" → ✅ \"soft, diffused lighting\"\n\n## Hex colors — bind them to objects\n\nFLUX.2 matches hex codes, but only when each code is attached to a specific object:\n\n```\nwalls in hex #C4725A, sofa in #1B6B6F, accent pillows #E8A847\ngradient starting with color #02eb3c and finishing with color #edfa3c\n```\n\n❌ \"use #FF0000 somewhere\" — unbound colors land inconsistently.\n\n## Text rendering\n\nQuote the exact text, then place and style it:\n\n```\nThe text 'OPEN' appears in red neon letters above the door\nLogo text 'ACME' in color #FF5733, ultra-bold decorative serif, centered\n```\n\nSpecify placement relative to other elements, font family feel (serif / sans / script), and relative size (\"large headline,\" \"small body copy\").\n\n## JSON prompting for production work\n\nFor multi-element scenes that must come out exactly right (product shots, infographics, brand work), FLUX.2 parses structured JSON prompts:\n\n```json\n{\n \"scene\": \"Professional studio product photography on polished concrete\",\n \"subjects\": [{ \"description\": \"matte black ceramic mug with steam\", \"position\": \"center foreground\" }],\n \"style\": \"commercial product photography\",\n \"color_palette\": [\"#1B1B1B\", \"#E8A847\"],\n \"lighting\": \"three-point softbox, soft diffused highlights\",\n \"camera\": { \"lens-mm\": 85, \"f-number\": \"f/5.6\" }\n}\n```\n\nUse natural language for exploration, JSON when the layout is locked and you're matching a spec.\n\n## Reference images (edit path)\n\nIn Slates, pass `referenceAssetIds` on `slates_generate_image` — FLUX routes them through its edit endpoint. Slates names each reference inline in the prompt (\"the subject (image 1), the style (image 2)\") in the order it sends them, so you don't hand-write role labels; the name carries the role and unnamed-by-position blending is avoided. For surgical changes to one existing image use `slates_edit_image` with `editModel: flux-2-max` (note: FLUX edits ignore extra referenceAssetIds — that's NB2-only).\n\nReference discipline (FLUX caps refs lower than NB2's 14, so be deliberate):\n- **2-4 strong refs**, one per role, named — not 1 (warps), not many (blends).\n- **Flat-lit identity refs** — a studio-lit / scene-lit character sheet bleeds its lighting into the output.\n- **Attach both character sheets, named as one entity** — turnaround (body/proportion/outfit) + close-up expression sheet (face detail), cited under the same name; the shared name keeps the expressions from averaging the face. Don't write a role essay or \"render neutral\" instruction — the user's prompt owns the expression, wardrobe, and lighting.\n- **Environment: describe it, don't feed a multi-panel grid** — reserve a ref for a hard exact-match, then use ONE clean establishing image.\n\n## Character consistency across a series\n\nDefine the character exhaustively once, then repeat those exact descriptors verbatim in every subsequent prompt. FLUX has no memory between generations — the repeated description IS the consistency mechanism.\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Generic \"AI look\" on photoreal | Name a camera body + lens + f-stop instead of \"professional photo\" |\n| Colors drift from brand spec | Bind each hex code to a named object |\n| Text garbled | Quote the exact string, specify font feel + placement + size |\n| Multi-reference blend chaos | Name each reference inline (Slates does this from your @mentions/referenceAssetIds) — the same name for one entity, distinct names per role |\n| Wanted element missing | Move it earlier in the prompt — order is weight |\n\n## Pre-flight: references arrive inline, refer by code\n\nWhen you pass `referenceAssetIds`, the first call returns the references **inline as image content blocks** with a cost estimate and `requires_confirm: true`. Look at them — revise the prompt if they suggest a different composition or style — then re-call with `confirm=true`. Refer to each asset by its short code (`IMG-A12 — Beach Sunset`) when talking to the user; it matches the badge on their gallery thumbnail.\n\n## Sources\n\n- [Black Forest Labs — FLUX.2 Prompting Guide](https://docs.bfl.ml/guides/prompting_guide_flux2)\n- [fal.ai — FLUX.2 [max] Prompt Guide](https://fal.ai/learn/devs/flux-2-max-prompt-guide)\n",
|
|
13
|
-
"slates-prompting-gpt-image-2": "---\nname: slates-prompting-gpt-image-2\ndescription: Prompting GPT Image 2 — the readable-text / character-sheet / shot-grid engine. Read before calling slates_generate_image with model gpt-image-2. Covers the quality tiers (medium default, high for max text precision), resolution classes (1k/2k=1080p/3k=1440p/4k), text-accuracy prompting, panel/grid layout direction, and when to route to the Banana line instead.\n---\n\n# GPT Image 2 — sheets, grids, and text that actually reads\n\nGPT Image 2's edge is **character-level text accuracy** (~99% on English), ordered panels, and exact element placement — the jobs where every other model garbles a word or shuffles a layout. It is NOT the photoreal or character-locked pick: route those to the Banana line (`slates-model-selection` has the split).\n\n## Quality tiers — always set explicitly\n\n- **medium** (default) — sharp text, fast, the value seat: half NB2's price at the 1080p class. Blind benchmarks put it within a hair of high at a quarter of the cost. Start here.\n- **high** — ~4× the price; max text precision + reasoning. A deliberate premium pick when tiny type, dense diagrams, or many labeled elements ARE the job.\n\nNever rely on the provider default (it's high — the priciest tier). The Slates ops send medium unless you say otherwise.\n\n## Resolution classes\n\n`1k` = 1024²-class · `2k` = 1920×1080-class · `3k` = 2560×1440-class · `4k` = 3840×2160-class. Pick 2k for most sheets/panels; 4k for print-density grids. 4K exists at BOTH tiers and is API-only — even paid ChatGPT can't render it.\n\n## Prompting for text accuracy\n\n- **Quote every string that must render verbatim**: `the sign reads \"OPEN 24 HOURS\"` — quoted strings render most reliably.\n- Specify font *feel*, not font names: \"clean geometric sans, high contrast\", \"hand-painted brush lettering\".\n- For dense text (posters, UI mocks), list the copy as ordered lines: `Line 1: \"...\" Line 2: \"...\"` — GPT Image 2 respects ordering.\n- Keep total on-image text under ~30 words for perfect accuracy; beyond that, accuracy degrades gracefully but degrades.\n\n## Panels, sheets, and grids\n\n- State the grid explicitly and number the cells: \"a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom\".\n- Give each cell ONE content clause: \"Panel 3: the character mid-jump, side view\".\n- Character sheets:
|
|
14
|
-
"slates-prompting-kling-v3": "---\nname: slates-prompting-kling-v3\ndescription: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_generate_video with kling-v3.0-std, kling-v3.0-pro, or kling-v3.0-omni. Kling has dialogue + SFX + ambient native syntax (Omni adds multi-character dialogue and language codes). Multi-shot rules differ from Seedance/Veo — don't cross syntaxes.\n---\n\n# Kling V3.0 — prompting\n\nKuaishou's video model. Three tiers: `kling-v3.0-std` (general use, no audio), `kling-v3.0-pro` (higher visual quality, no audio), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).\n\nUp to 15s. Multi-shot supported (up to 6 cuts in 15s total). Strong on image-to-video — preserves identity, layout, and text from the input image well.\n\n## Subject definition rule (verbatim, fal blog)\n\n> \"Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots.\"\n\n## Dialogue syntax\n\n```\nCharacter says, \"exact words here\"\n```\n\nUse quotation marks for precise speech. Languages (Omni only): EN, ZH, JA, KO, ES.\n\n## Voice direction formula (Omni)\n\n```\nGender + Age Range + Voice Quality + Speech Rate + Emotional Tone + Language\n```\n\nExample:\n```\n[Character A: Detective, mid-40s, raspy voice, slow cadence, weary]: \"I've seen this before.\"\n```\n\nTone phrases that fire:\n- `speaking in a hushed, trembling whisper`\n- `shouting with commanding authority`\n- `clear, fearful voice`\n- `with a trembling voice, \"I'm scared\"`\n\n## The `Immediately` keyword (Omni only)\n\nWithout `Immediately`, Kling adds a natural conversational beat between speakers. With it, dialogue is back-to-back. Use when timing matters.\n\n```\n[Alice]: \"Get down!\" Immediately, [Bob]: \"Where?\"\n```\n\n## Speaker label discipline\n\nUnique labels per character. **No pronouns or synonyms after first introduction** — they cause voice drift.\n\n✅ `[Character A: Black-suited Agent]` ... `[Character A: Black-suited Agent]: \"Stop.\"`\n❌ `[Agent]... then he says...`\n\n## Multi-character dialogue (Omni)\n\n```\nAlice says in English, \"Hello!\" Then Bob replies in Spanish, \"¡Hola!\"\n```\n\n## Sound effects, ambient noise, music\n\n```\nSFX: thunder cracks, footsteps approaching\nAmbient noise: city traffic, birds chirping, ocean waves\nBackground music: tense orchestral strings, low cello\n```\n\nSFX accepts physical-cause specificity:\n- ✅ `SFX: heavy boots on wet pavement, distant siren wailing`\n- ❌ `SFX: footsteps`\n\n## Image-to-video guidance\n\n**Verbatim (fal blog):**\n> \"Treat the input image as an anchor. Kling 3.0 excels at preserving the identity, layout, and text details. Focus prompts on how the scene evolves *from* the image: subtle movements, camera motion, or environmental changes.\"\n\n**Don't re-describe what's already in the image.** Focus on motion, changes, evolution.\n\n## Multi-shot — what makes them hit\n\n**Hard cap: total duration ≤ 15s across all shots. Max 6 cuts.**\n\nHit conditions:\n- Shot labels are explicit: `Shot 1:`, `Shot 2:`\n- One primary action per shot\n- Subject described identically in each shot block\n- Camera move per shot is **one verb**, not a chain\n- Per-shot blocks: 30-60 words\n\nMiss conditions:\n- Compressing narrative into one paragraph\n- Pronoun-only references after the first shot\n- Mixing camera moves within a shot (\"pan then orbit then push in\")\n- Extreme wide → extreme close in adjacent shots without reference images\n\n## Element references (Omni)\n\nUpload 2-4 multi-angle reference photos per character/object. Tag inline:\n\n```\n@element1 is the protagonist (refs: front, side, back angles).\n@element2 is the antagonist.\n```\n\n## Reference discipline (character / environment refs)\n\n- **2-4 strong refs per role**, named (the same fixed label reused verbatim) and reused across every shot — swapping mid-sequence drifts. Kling's consistency lever is **\"lock the subject with a fixed label reused verbatim\"** (pronoun/synonym drift breaks it), so reusing the exact name on every mention is the whole game. Slates composes this for you from `@mentions`.\n- **Flat-lit identity refs.** A studio-lit / scene-lit character sheet bleeds its lighting into the clip. Prep refs flat and plain.\n- **Attach both character sheets, named as one entity** — the turnaround (body/proportion/outfit) and the close-up expression sheet (face detail), cited under the same name. The shared name keeps the varied expressions from averaging the face; don't write a role essay or tell it to \"render neutral\" — the user's prompt owns the expression, wardrobe, and lighting.\n- **Environment: describe it, don't feed a multi-panel grid.** Reserve an environment ref for a hard exact-match, then use ONE clean establishing image.\n\n## Negative prompting — has a real field\n\nKling exposes `negative_prompt` on the fal endpoint (different from Seedance which has none). Default block to start from:\n\n```\nblurry, low quality, watermark, text overlay, distorted hands, extra fingers,\nduplicate limbs, unnatural skin texture, overly saturated colors, lens flare,\nfloating objects, inconsistent shadows, jittery, flickering, morphing face\n```\n\nLayer scene-specific suppressions on top.\n\n## Cinematic tactics\n\n- **Motion adverb precision** modulates motion energy directly: `slowly`, `rapidly`, `gently`, `explosively`\n- **Camera vocabulary that registers as instructions:** profile shot, tracking, following, freezing, panning, \"moving in sync with the subject\"\n- **One primary camera move per shot** — never stack\n\n## Tier choice\n\n- **Standard**: general use, no audio\n- **Pro**: higher visual quality, no audio\n- **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references\n\nPick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models` for current numbers before choosing a tier.\n\n## Benchmark prompt structure\n\n```\n[Character A: <role>, <voice quality>]: \"<line>.\" Immediately, [Character B: <role>, <voice quality>]: \"<reply>.\"\nAmbient noise: <soundscape>.\nCamera <single move>.\n```\n\nCinematic example (paraphrasing fal blog patterns):\n> \"Shot 1: Wide establishing shot of a neon-lit alleyway in heavy rain, steam rising from grates. Camera slowly tracks forward.\n> Shot 2: Medium shot of a detective in a trench coat ducking under an awning, water dripping from his hat brim. [Detective: weary, raspy]: 'I knew she'd come back.' Ambient noise: distant traffic, rain on metal.\n> Shot 3: Close-up on his eyes, narrowing as headlights flash across his face.\"\n\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with `firstFrameAssetId` or `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside cost + `requires_confirm: true`. Look at them, revise prompt if needed, then re-call with `confirm=true`. Kling Omni multi-character with several ingredient images especially benefits — confirm each character image lands cleanly before spending.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Detective Closeup`. The user sees that code as a gallery badge.\n\n- ✅ \"I'm anchoring on **IMG-A12** as the detective and **IMG-A18** as the alleyway environment — Omni will handle the line delivery in EN.\"\n- ❌ \"I'm using the detective image and the alley one...\" (which alley? Three exist.)\n\n## Video-to-video EDIT (`slates_edit_video`) — @Video1 / @ElementN / @ImageN\n\nKling O3 edit takes an EXISTING 3-15s clip and changes only what the prompt names — character swap, environment change, style transfer — in one pass, no masking. Original motion, camera, and audio are preserved by default. Its notation is Kling's own, different from the \"image N\" naming used everywhere else:\n\n- **`@Video1`** — the source clip (always; the transport anchors the instruction to it).\n- **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically).\n- **`@Image1..`** — style/appearance references (pass as `styleAssetIds`).\n- Max **4 combined** element + image refs per edit.\n\n**Prompt shape — the change, not the whole scene:**\n\n```\nReplace the man in @Video1 with @Element1, keeping his walk cycle, the camera move, and the rain unchanged.\n```\n\n```\nEdit @Video1: turn the daytime street into a neon-lit Tokyo alley at night, wet asphalt reflections. Apply the visual style of @Image1. Keep the subject and camera motion exactly as they are.\n```\n\nRules:\n- Name what CHANGES; explicitly state what stays (\"keep the motion / camera / everything else unchanged\") — the model preserves better when told to.\n- One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).\n- Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.\n- Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.\n- Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings — see `slates-model-selection`.\n\n## Sources\n\n- [fal.ai — Kling 3.0 Prompting Guide](https://blog.fal.ai/kling-3-0-prompting-guide/)\n- [Vidguru — Kling 3.0 Omni Guide](https://www.vidguru.ai/blog/kling-3.0-omni-guide.html)\n- [AcceptPrompt — Kling 3 Prompt Guide](https://www.acceptprompt.com/blog/kling-3-prompt-guide)\n- [DataCamp — Kling 3.0 Tutorial](https://www.datacamp.com/tutorial/kling-3-0)\n",
|
|
12
|
+
"slates-prompting-flux-2-max": "---\nname: slates-prompting-flux-2-max\ndescription: How to prompt FLUX.2 Max (Black Forest Labs image model). Read before calling slates_generate_image with model flux-2-max, or slates_edit_image with editModel flux-2-max. FLUX.2 wants front-loaded structure, real camera vocabulary, and positive-only phrasing — no negative prompts, no tag soup.\n---\n\n# FLUX.2 Max — prompting\n\nBlack Forest Labs' top image model, routed via fal.ai. In Slates: `slates_generate_image` with `model: flux-2-max` (REQUIRES projectId — no headless path), priced per resolution (1k/2k/4k — call `slates_estimate_generation_cost` for current numbers, never quote from memory). Strengths vs Nano Banana 2: photoreal texture, less censored, precise hex-color control, strong typography. Reference images route through FLUX's edit endpoint and carry a lower per-model cap than NB2's 14.\n\n## Core structure — front-load what matters\n\n```\nSubject + Action + Style + Context\n```\n\nWord order is weight. FLUX.2 attends hardest to the start of the prompt: main subject → key action → critical style → essential context → secondary details.\n\n**Length:** 10-30 words for concept tests, 30-80 words for most work, 80+ only for genuinely complex scenes.\n\n## Photorealism: name real gear, not \"professional photo\"\n\nThe single biggest realism lever is concrete camera vocabulary:\n\n```\nShot on Hasselblad X2D, 80mm lens, f/2.8, natural lighting\nShot on Sony A7IV, 35mm, golden hour, shallow depth of field\nKodak Portra 400, natural grain, organic colors\n```\n\nEra cues work the same way: \"early digital camera, slight noise, flash photography, candid\" reads 2000s digicam; \"film grain, warm color cast, soft focus\" reads 80s.\n\nFor portraits add: natural skin texture, realistic pores, subtle imperfections, soft diffused lighting.\n\n## No negative prompts — reframe positively\n\nFLUX.2 has no negative prompt support. Describe the presence you want, not the absence:\n\n- ❌ \"no blur\" → ✅ \"sharp focus throughout\"\n- ❌ \"no people\" → ✅ \"empty scene\"\n- ❌ \"no harsh shadows\" → ✅ \"soft, diffused lighting\"\n\n## Hex colors — bind them to objects\n\nFLUX.2 matches hex codes, but only when each code is attached to a specific object:\n\n```\nwalls in hex #C4725A, sofa in #1B6B6F, accent pillows #E8A847\ngradient starting with color #02eb3c and finishing with color #edfa3c\n```\n\n❌ \"use #FF0000 somewhere\" — unbound colors land inconsistently.\n\n## Text rendering\n\nQuote the exact text, then place and style it:\n\n```\nThe text 'OPEN' appears in red neon letters above the door\nLogo text 'ACME' in color #FF5733, ultra-bold decorative serif, centered\n```\n\nSpecify placement relative to other elements, font family feel (serif / sans / script), and relative size (\"large headline,\" \"small body copy\").\n\n## JSON prompting for production work\n\nFor multi-element scenes that must come out exactly right (product shots, infographics, brand work), FLUX.2 parses structured JSON prompts:\n\n```json\n{\n \"scene\": \"Professional studio product photography on polished concrete\",\n \"subjects\": [{ \"description\": \"matte black ceramic mug with steam\", \"position\": \"center foreground\" }],\n \"style\": \"commercial product photography\",\n \"color_palette\": [\"#1B1B1B\", \"#E8A847\"],\n \"lighting\": \"three-point softbox, soft diffused highlights\",\n \"camera\": { \"lens-mm\": 85, \"f-number\": \"f/5.6\" }\n}\n```\n\nUse natural language for exploration, JSON when the layout is locked and you're matching a spec.\n\n## Reference images (edit path)\n\nIn Slates, pass `referenceAssetIds` on `slates_generate_image` — FLUX routes them through its edit endpoint. Slates names each reference inline in the prompt (\"the subject (image 1), the style (image 2)\") in the order it sends them, so you don't hand-write role labels; the name carries the role and unnamed-by-position blending is avoided. For surgical changes to one existing image use `slates_edit_image` with `editModel: flux-2-max` (note: FLUX edits ignore extra referenceAssetIds — that's NB2-only).\n\n### Reference rules (the verified ones)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For FLUX.2 Max specifically\n\n- **FLUX caps references well below NB2's 14, so rule 1's \"2-4\" is a ceiling here, not a starting point.** Be deliberate about which roles earn a slot.\n- **Rule 9 has a hard edge on this model:** `slates_edit_image` with `editModel: flux-2-max` ignores extra `referenceAssetIds` — that is NB2-only. A FLUX edit sees the source image and the prompt, nothing else.\n- **FLUX has no memory between generations, so rule 7 is enforced by repetition.** Define the character exhaustively once and repeat those exact descriptors verbatim in every subsequent prompt — see Character consistency across a series below.\n\n## Character consistency across a series\n\nDefine the character exhaustively once, then repeat those exact descriptors verbatim in every subsequent prompt. FLUX has no memory between generations — the repeated description IS the consistency mechanism.\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Generic \"AI look\" on photoreal | Name a camera body + lens + f-stop instead of \"professional photo\" |\n| Colors drift from brand spec | Bind each hex code to a named object |\n| Text garbled | Quote the exact string, specify font feel + placement + size |\n| Multi-reference blend chaos | Name each reference inline (Slates does this from your @mentions/referenceAssetIds) — the same name for one entity, distinct names per role |\n| Wanted element missing | Move it earlier in the prompt — order is weight |\n\n## Pre-flight: references arrive inline, refer by code\n\nWhen you pass `referenceAssetIds`, the first call returns the references **inline as image content blocks** with a cost estimate and `requires_confirm: true`. Look at them — revise the prompt if they suggest a different composition or style — then re-call with `confirm=true`. Refer to each asset by its short code (`IMG-A12 — Beach Sunset`) when talking to the user; it matches the badge on their gallery thumbnail.\n\n## Sources\n\n- [Black Forest Labs — FLUX.2 Prompting Guide](https://docs.bfl.ml/guides/prompting_guide_flux2)\n- [fal.ai — FLUX.2 [max] Prompt Guide](https://fal.ai/learn/devs/flux-2-max-prompt-guide)\n",
|
|
13
|
+
"slates-prompting-gpt-image-2": "---\nname: slates-prompting-gpt-image-2\ndescription: Prompting GPT Image 2 — the readable-text / character-sheet / shot-grid engine. Read before calling slates_generate_image with model gpt-image-2. Covers the quality tiers (medium default, high for max text precision), resolution classes (1k/2k=1080p/3k=1440p/4k), text-accuracy prompting, panel/grid layout direction, and when to route to the Banana line instead.\n---\n\n# GPT Image 2 — sheets, grids, and text that actually reads\n\nGPT Image 2's edge is **character-level text accuracy** (~99% on English), ordered panels, and exact element placement — the jobs where every other model garbles a word or shuffles a layout. It is NOT the photoreal or character-locked pick: route those to the Banana line (`slates-model-selection` has the split).\n\n## Quality tiers — always set explicitly\n\n- **medium** (default) — sharp text, fast, the value seat: half NB2's price at the 1080p class. Blind benchmarks put it within a hair of high at a quarter of the cost. Start here.\n- **high** — ~4× the price; max text precision + reasoning. A deliberate premium pick when tiny type, dense diagrams, or many labeled elements ARE the job.\n\nNever rely on the provider default (it's high — the priciest tier). The Slates ops send medium unless you say otherwise.\n\n## Resolution classes\n\n`1k` = 1024²-class · `2k` = 1920×1080-class · `3k` = 2560×1440-class · `4k` = 3840×2160-class. Pick 2k for most sheets/panels; 4k for print-density grids. 4K exists at BOTH tiers and is API-only — even paid ChatGPT can't render it.\n\n## Prompting for text accuracy\n\n- **Quote every string that must render verbatim**: `the sign reads \"OPEN 24 HOURS\"` — quoted strings render most reliably.\n- Specify font *feel*, not font names: \"clean geometric sans, high contrast\", \"hand-painted brush lettering\".\n- For dense text (posters, UI mocks), list the copy as ordered lines: `Line 1: \"...\" Line 2: \"...\"` — GPT Image 2 respects ordering.\n- Keep total on-image text under ~30 words for perfect accuracy; beyond that, accuracy degrades gracefully but degrades.\n\n## Panels, sheets, and grids\n\n- State the grid explicitly and number the cells: \"a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom\".\n- Give each cell ONE content clause: \"Panel 3: the character mid-jump, side view\".\n- Character identity sheets: GPT Image 2 holds structured panel layouts; the Banana line holds the *face* better. Prefer NB2/NB Pro for identity-critical sheets and GPT Image 2 when labels or annotations are the main requirement.\n\n## References & editing\n\nReference images route through the edit endpoint (up to ~10). The composed \"image N\" naming applies as everywhere else. Mask-based inpainting exists at the API level but isn't surfaced — describe the change instead.\n\n## Filter regime\n\nOpenAI moderate — a third regime distinct from Gemini (NB family) and ByteDance (Seedream). Real-face references pass more readily than Gemini; violence/brand rules are similar. `slates-content-policy` applies unchanged.\n",
|
|
14
|
+
"slates-prompting-kling-v3": "---\nname: slates-prompting-kling-v3\ndescription: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_generate_video with kling-v3.0-std, kling-v3.0-pro, or kling-v3.0-omni. Kling has dialogue + SFX + ambient native syntax (Omni adds multi-character dialogue and language codes). Multi-shot rules differ from Seedance/Veo — don't cross syntaxes.\n---\n\n# Kling V3.0 — prompting\n\nKuaishou's video model. Three tiers: `kling-v3.0-std` (general use, no audio), `kling-v3.0-pro` (higher visual quality, no audio), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).\n\nUp to 15s. Multi-shot supported (up to 6 cuts in 15s total). Strong on image-to-video — preserves identity, layout, and text from the input image well.\n\n## Subject definition rule (verbatim, fal blog)\n\n> \"Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots.\"\n\n## Dialogue syntax\n\n```\nCharacter says, \"exact words here\"\n```\n\nUse quotation marks for precise speech. Languages (Omni only): EN, ZH, JA, KO, ES.\n\n## Voice direction formula (Omni)\n\n```\nGender + Age Range + Voice Quality + Speech Rate + Emotional Tone + Language\n```\n\nExample:\n```\n[Character A: Detective, mid-40s, raspy voice, slow cadence, weary]: \"I've seen this before.\"\n```\n\nTone phrases that fire:\n- `speaking in a hushed, trembling whisper`\n- `shouting with commanding authority`\n- `clear, fearful voice`\n- `with a trembling voice, \"I'm scared\"`\n\n## The `Immediately` keyword (Omni only)\n\nWithout `Immediately`, Kling adds a natural conversational beat between speakers. With it, dialogue is back-to-back. Use when timing matters.\n\n```\n[Alice]: \"Get down!\" Immediately, [Bob]: \"Where?\"\n```\n\n## Speaker label discipline\n\nUnique labels per character. **No pronouns or synonyms after first introduction** — they cause voice drift.\n\n✅ `[Character A: Black-suited Agent]` ... `[Character A: Black-suited Agent]: \"Stop.\"`\n❌ `[Agent]... then he says...`\n\n## Multi-character dialogue (Omni)\n\n```\nAlice says in English, \"Hello!\" Then Bob replies in Spanish, \"¡Hola!\"\n```\n\n## Sound effects, ambient noise, music\n\n```\nSFX: thunder cracks, footsteps approaching\nAmbient noise: city traffic, birds chirping, ocean waves\nBackground music: tense orchestral strings, low cello\n```\n\nSFX accepts physical-cause specificity:\n- ✅ `SFX: heavy boots on wet pavement, distant siren wailing`\n- ❌ `SFX: footsteps`\n\n## Image-to-video guidance\n\n**Verbatim (fal blog):**\n> \"Treat the input image as an anchor. Kling 3.0 excels at preserving the identity, layout, and text details. Focus prompts on how the scene evolves *from* the image: subtle movements, camera motion, or environmental changes.\"\n\n**Don't re-describe what's already in the image.** Focus on motion, changes, evolution.\n\n## Multi-shot — what makes them hit\n\n**Hard cap: total duration ≤ 15s across all shots. Max 6 cuts.**\n\nHit conditions:\n- Shot labels are explicit: `Shot 1:`, `Shot 2:`\n- One primary action per shot\n- Subject described identically in each shot block\n- Camera move per shot is **one verb**, not a chain\n- Per-shot blocks: 30-60 words\n\nMiss conditions:\n- Compressing narrative into one paragraph\n- Pronoun-only references after the first shot\n- Mixing camera moves within a shot (\"pan then orbit then push in\")\n- Extreme wide → extreme close in adjacent shots without reference images\n\n## Element references (Omni)\n\nUpload 2-4 multi-angle reference photos per character/object. Tag inline:\n\n```\n@element1 is the protagonist (refs: front, side, back angles).\n@element2 is the antagonist.\n```\n\n## Reference discipline (character / environment refs)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Kling specifically\n\n- **Kling's consistency lever is \"lock the subject with a fixed label reused verbatim.\"** That is Kling's phrasing for rules 2 and 3, and it is stricter than the others: **pronoun and synonym drift breaks it**, so the exact same label must appear on every single mention — not \"he\", not \"the detective\" after you named him. Reusing the label verbatim is the whole game. Slates composes this for you from `@mentions`.\n- **Element references are the transport for rule 1** — 2-4 multi-angle photos per character/object, tagged `@element1` / `@element2` (see Element references above). The cap is 4 combined refs on the edit path.\n\n## Negative prompting — has a real field\n\nKling exposes `negative_prompt` on the fal endpoint (different from Seedance which has none). Default block to start from:\n\n```\nblurry, low quality, watermark, text overlay, distorted hands, extra fingers,\nduplicate limbs, unnatural skin texture, overly saturated colors, lens flare,\nfloating objects, inconsistent shadows, jittery, flickering, morphing face\n```\n\nLayer scene-specific suppressions on top.\n\n## Cinematic tactics\n\n- **Motion adverb precision** modulates motion energy directly: `slowly`, `rapidly`, `gently`, `explosively`\n- **Camera vocabulary that registers as instructions:** profile shot, tracking, following, freezing, panning, \"moving in sync with the subject\"\n- **One primary camera move per shot** — never stack\n\n## Tier choice\n\n- **Standard**: general use, no audio\n- **Pro**: higher visual quality, no audio\n- **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references\n\nPick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.\n\n## Benchmark prompt structure\n\n```\n[Character A: <role>, <voice quality>]: \"<line>.\" Immediately, [Character B: <role>, <voice quality>]: \"<reply>.\"\nAmbient noise: <soundscape>.\nCamera <single move>.\n```\n\nCinematic example (paraphrasing fal blog patterns):\n> \"Shot 1: Wide establishing shot of a neon-lit alleyway in heavy rain, steam rising from grates. Camera slowly tracks forward.\n> Shot 2: Medium shot of a detective in a trench coat ducking under an awning, water dripping from his hat brim. [Detective: weary, raspy]: 'I knew she'd come back.' Ambient noise: distant traffic, rain on metal.\n> Shot 3: Close-up on his eyes, narrowing as headlights flash across his face.\"\n\n<!-- slates-only -->\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with `firstFrameAssetId` or `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside cost + `requires_confirm: true`. Look at them, revise prompt if needed, then re-call with `confirm=true`. Kling Omni multi-character with several ingredient images especially benefits — confirm each character image lands cleanly before spending.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Detective Closeup`. The user sees that code as a gallery badge.\n\n- ✅ \"I'm anchoring on **IMG-A12** as the detective and **IMG-A18** as the alleyway environment — Omni will handle the line delivery in EN.\"\n- ❌ \"I'm using the detective image and the alley one...\" (which alley? Three exist.)\n<!-- /slates-only -->\n\n## Video-to-video EDIT<!-- slates-only --> (`slates_edit_video`)<!-- /slates-only --> — @Video1 / @ElementN / @ImageN\n\nKling O3 edit takes an EXISTING 3-15s clip and changes only what the prompt names — character swap, environment change, style transfer — in one pass, no masking. Original motion, camera, and audio are preserved by default. Its notation is Kling's own, different from the \"image N\" naming used everywhere else:\n\n- **`@Video1`** — the source clip (always; the transport anchors the instruction to it).\n- **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images<!-- slates-only --> (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically)<!-- /slates-only -->.\n- **`@Image1..`** — style/appearance references<!-- slates-only --> (pass as `styleAssetIds`)<!-- /slates-only -->.\n- Max **4 combined** element + image refs per edit.\n\n**Prompt shape — the change, not the whole scene:**\n\n```\nReplace the man in @Video1 with @Element1, keeping his walk cycle, the camera move, and the rain unchanged.\n```\n\n```\nEdit @Video1: turn the daytime street into a neon-lit Tokyo alley at night, wet asphalt reflections. Apply the visual style of @Image1. Keep the subject and camera motion exactly as they are.\n```\n\nRules:\n- Name what CHANGES; explicitly state what stays (\"keep the motion / camera / everything else unchanged\") — the model preserves better when told to.\n- One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).\n- Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.\n- Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.\n- Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.\n\n## Sources\n\n- [fal.ai — Kling 3.0 Prompting Guide](https://blog.fal.ai/kling-3-0-prompting-guide/)\n- [Vidguru — Kling 3.0 Omni Guide](https://www.vidguru.ai/blog/kling-3.0-omni-guide.html)\n- [AcceptPrompt — Kling 3 Prompt Guide](https://www.acceptprompt.com/blog/kling-3-prompt-guide)\n- [DataCamp — Kling 3.0 Tutorial](https://www.datacamp.com/tutorial/kling-3-0)\n",
|
|
15
15
|
"slates-prompting-lip-sync": "---\nname: slates-prompting-lip-sync\ndescription: How to set up lip-sync — Kling (cheap utility lane) or Seedance 2.0 (premium single-pass lane). Read before calling slates_generate_lip_sync. Flows — video→video re-dub, image→video avatar, and Seedance native speech — with different inputs, pricing, and gotchas. Voice catalog, framing rules, audio file constraints, and which engine/tier to pick.\n---\n\n# Lip-sync — setup guide\n\nTwo engines. Kling is the cheap utility lane (dedicated lip-sync endpoints, 5-second outputs); Seedance is the premium lane (speech generated IN the video itself, single pass):\n\n| Flow | Source | Engine/Model | Cost | Use case |\n|------|--------|-------|-----------|----------|\n| Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |\n| Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |\n| Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |\n| **Seedance native** | image or video | `engine=seedance-2` | per second (`seedance-2-face-*`; video sources bill input+output seconds) | **Premium**: natural delivery, whole-body performance, voice cloned from a video source, audio included |\n\nPick engine + `sourceType` deliberately — they decide the pricing tier and the underlying endpoint.\n\n## Seedance engine (premium single-pass)\n\nKling lip-sync moves the mouth on finished pixels; Seedance *generates* the performance — head movement, gesture, delivery energy — with the dialogue as a native conditioning signal. Key facts:\n\n- **`ttsText` becomes the spoken line, natively.** No TTS voice/speed params — a VIDEO source keeps its **own voice** (the model clones it from the clip's audio track); for an image source the voice follows the character's look, or describe it in the line's context.\n- **`audioMethod=upload`** drives the speech from a ≤15s audio file instead (a reference-audio input, no billing surcharge).\n- **Sources:** image (any style; same framing rules as the avatar flow below) or a 2–15s video clip. Output duration follows the source/audio/line length (4–15s), not a fixed 5s.\n- **Billing:** image sources bill the normal `seedance-2-face-{res}-{N}s` keys; video sources bill combined input+output seconds (`-vref-` keys) — pass `sourceSeconds` and quote via the confirm gate.\n- **Faces:** `seedanceFace` defaults true. A REAL person → `[REAL_FACE_DETECTED]` → confirm consent → retry with `seedanceRealFace=true, realFaceConsent=true` (premium realface pricing).\n- When to pick it: hero dialogue shots, natural delivery, \"make this clip's person say X in their own voice\". Stay on Kling for cheap utility re-dubs and long clips.\n\nEverything below applies to the **Kling** engine.\n\n## Choosing video vs avatar\n\nUse **video** (re-dub) when:\n- A talking-head clip already exists (Slates-generated, recorded, or imported)\n- The mouth/face is already moving and only the audio needs to change\n- ~4 credits is hard to beat for short dialogue replacement\n\nUse **avatar** when:\n- Only a still portrait exists\n- The character needs to come alive from a single image\n- Identity + face fidelity matter (avatar-pro for hero shots, standard for everything else)\n\n## Source asset constraints\n\n### Video flow (`sourceType: 'video'`)\n- Format: mp4 or mov\n- Duration: 2–10s (lip-sync output is always 5s — long videos get trimmed)\n- Resolution: 720p or 1080p (480p will be rejected)\n- Max file size: 100MB\n- Face must be visible and roughly facing camera. Profile shots fail.\n- Existing audio is replaced.\n\n### Avatar flow (`sourceType: 'image'`)\n- Min 512×512, PNG/JPG/WebP\n- **Face occupies 60–70% of frame.** This is the single biggest avatar quality lever.\n- Eyes open, mouth neutral, looking near-camera. Side profile = bad output.\n- Single subject, clean background. Group photos confuse the face anchor.\n\n## Audio source\n\nTwo ways to drive the lips:\n\n### TTS (`audioMethod: 'tts'`)\n- Pass `ttsText` (the words spoken)\n- Optional: `ttsVoice` (default `oversea_male1`), `ttsLanguage` (default EN), `ttsSpeed` (default 1.0)\n- **Hard cap: 120 characters of text.** Longer = silently truncated.\n- Languages: EN, ZH, JA, KO, ES\n\n### Upload (`audioMethod: 'upload'`)\n- Pass `audioFilePath` — absolute path to an audio file on the user's machine\n- Format: mp3, wav, m4a, ogg, aac\n- Max 5MB\n- Duration: 2–60s (output is 5s — longer audio gets trimmed)\n- Single clean voice. Music underneath, multiple speakers, or noisy mics produce garbage lips.\n\nPrefer upload for production-quality voice. TTS for fast iteration / placeholder dialogue.\n\n## Voice catalog (TTS)\n\nReliable English voices (verified working on the fal endpoint as of 2026):\n\n| Voice ID | Description |\n|----------|-------------|\n| `oversea_male1` | Male, English — default, stable |\n| `commercial_lady_en_f-v1` | Female commercial English |\n| `uk_boy1` | Young man, UK accent |\n| `uk_man2` | Man, UK accent |\n| `uk_oldman3` | Older man, UK accent |\n| `calm_story1` | Storyteller / narrator |\n\nAvoid `reader_en_m-v1` — listed in fal.ai docs but returns \"Voice id not found\" in production.\n\nFull 48-voice list (ZH, JA, KO included): https://fal.ai/models/fal-ai/kling-video/lipsync/text-to-video/api\n\n## Speech-rate notes\n\n`ttsSpeed` range 0.5–2.0:\n- 0.8–1.0: natural conversational\n- 1.1–1.3: punchy ad delivery\n- 1.4+: rushed, clips consonants\n- 0.6–0.7: slow, weighty (good for dramatic lines)\n\nDefault 1.0 unless the line specifically calls for slower or faster cadence.\n\n## Avatar prompt usage\n\nThe `prompt` parameter on avatar-v2 (standard + pro) is **scene context**, not motion direction. The mouth animation comes from the audio — the prompt sets ambiance, lighting, micro-expression.\n\nGood:\n- `Soft rim light, warm office, gentle confident smile between sentences.`\n- `Cool blue evening light through a window, focused intent expression.`\n\nBad (the model ignores motion verbs):\n- ❌ `She turns her head, raises an eyebrow, then speaks.`\n- ❌ `Hand gestures while talking.`\n\nDefault `\".\"` is fine if you have nothing useful to add.\n\n## Tier selection — avatar standard vs pro\n\n**Use standard** when:\n- Drafts, A/B testing voices, internal review reels\n- Wide / medium shots where face isn't the focal point\n- Cost matters more than micro-expression fidelity\n\n**Use pro** when:\n- Final ads where the avatar's face fills the screen\n- The character is named / branded — identity drift kills the take\n- You're already paying tens of credits for the surrounding video pipeline\n\nDon't default to pro. The ~15-credit delta per take adds up across iteration.\n\n## Common failure modes\n\n| Symptom | Likely cause | Fix |\n|---------|--------------|-----|\n| Lip movement looks \"rubber\" / disconnected | Source face <60% of frame | Re-crop the still tighter |\n| Voice doesn't match character age/gender | Default voice id used | Pick from voice catalog |\n| Output truncated mid-word | TTS text >120 chars | Shorten or chain two takes |\n| Garbled mouth on uploaded audio | Background music / multi-voice | Use clean dialogue-only audio |\n| \"Voice id not found\" 422 | Hit `reader_en_m-v1` | Switch to `oversea_male1` |\n| Avatar eyes drift / cross | Source had closed/angled eyes | Pick a frame with neutral open eyes |\n| Generation completes but lips don't move | Profile shot / face >70° off-axis | Use a near-frontal portrait |\n\n## Cost discipline\n\n- Video re-dub at ~4 credits is the cheapest dialogue iteration in the entire Slates stack — use it for voice A/B testing\n- Avatar standard at ~14 credits is fine for medium use\n- Avatar pro at ~29 credits trips the confirm gate — explicit user OK required every time\n- All 5s. There is no shorter option.\n\n## Workflow patterns\n\n**Voice A/B test (cheap):**\n1. Generate one base talking-head video clip with Veo or Seedance (~40 credits)\n2. Run `slates_generate_lip_sync` with `sourceType: 'video'` against 3–5 different `ttsVoice` values\n3. Total cost: ~40 + (5 × ~4) ≈ 60 credits to compare voices\n\n**Brand avatar from a single portrait:**\n1. Generate or upload the hero portrait (face fills frame, eyes open, neutral mouth)\n2. Avatar standard for first-pass dialogue takes\n3. Avatar pro only on the final selected take\n\n**Avoid:**\n- Avatar pro on first iteration (waste — facial fidelity isn't visible until you've locked the line)\n- TTS for final ads (production should use real voice or cloned voice — the upload flow)\n- Uploading raw recordings — clean noise + level the file first, lip detection is sensitive\n\n## Confirm gate: cost + codes, no inline preview\n\nLip-sync is mechanical — the model re-syncs the chosen source to the chosen audio. The confirm response carries the source asset's code so you can announce it in chat.\n\n- ✅ \"Lip-syncing **IMG-A12 — Founder Portrait** to the new line. ~29 credits on avatar-pro. Confirm?\"\n- ❌ \"Using the founder image...\" (which? Three exist.)\n\nDon't second-guess the source. If the output is wrong, iterate on source choice or audio, not on a refinement prompt (there isn't one).\n\n## Sources\n\n- [fal.ai — Kling LipSync API](https://fal.ai/models/fal-ai/kling-video/lipsync/text-to-video/api)\n- [fal.ai — AI Avatar v2 Standard](https://fal.ai/models/fal-ai/kling-video/ai-avatar/v2/standard/api)\n- [fal.ai — AI Avatar v2 Pro](https://fal.ai/models/fal-ai/kling-video/ai-avatar/v2/pro/api)\n",
|
|
16
16
|
"slates-prompting-motion-transfer": "---\nname: slates-prompting-motion-transfer\ndescription: How to set up motion transfer — Kling Motion Control (cheap utility lane) or Seedance 2.0 (premium single-pass lane). Read before calling slates_generate_motion_transfer. Reference image (character) + driving video (motion source) → new video of the character performing the motion. Asset selection rules, engine choice, character_orientation, tiers, and prompt usage.\n---\n\n# Motion transfer — setup guide\n\nTake a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion. Two engines:\n\n| Engine | Cost | Use case |\n|------|-----------|----------|\n| Kling std (`kling-mc-std-5s`) | ~32 credits / 5s | General motion transfer, budget lane |\n| Kling pro (`kling-mc-pro-5s`) | ~42 credits / 5s | Cleaner anatomy, better identity preservation |\n| **Seedance 2.0** (`motionModel=seedance-2`) | per second of input+output (`seedance-2-face-vref-*`) | **Premium lane** — single-pass generation with the driving clip as a native conditioning signal: better motion fidelity, native audio, prompt-directed |\n\nAll tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)\n\n## Seedance engine (premium single-pass)\n\nKling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion. Key differences from Kling:\n\n- **Prompt-driven.** The prompt is the primary control, using ordinal references: `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.` Add style/setting/camera direction freely — Seedance re-generates the whole shot.\n- **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`character_orientation: 'video'` takes up to 30s).\n- **Billing = combined input+output seconds** (the vref keys). Pass `sourceVideoSeconds`; output `duration` defaults to the clip length (4–15s). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.\n- **Faces route through the face cascade**: `seedanceFace` defaults true (driving clips contain people). A REAL person → `[REAL_FACE_DETECTED]` → confirm consent with the user → retry `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).\n- **Native audio included** — the output can carry sound from the prompt (or the driving clip's vibe); no audio surcharge.\n- `characterOrientation` is Kling-only; Seedance framing follows the prompt + `aspectRatio`.\n\nEverything below applies to the **Kling** engine.\n\n## Inputs\n\n- `sourceVideoAssetId` — driving video. **Must be a realistic human** with clear proportions. Anime/cartoon/CG driving videos fail.\n- `targetImageAssetId` — character to be animated. Can be any style (cartoon, anime, realistic, painted).\n- Both must already exist as assets in the project. Use `slates_list_assets` to find them or upload first.\n\n## Source video constraints\n\n- Realistic human (not animated, not CG)\n- Entire body OR upper body visible — head must not be obstructed\n- Subject occupies a clear share of the frame\n- Single primary subject. Multi-person driving videos confuse the motion anchor.\n- Clean motion — choppy / cut-edited driving videos produce jittery output\n\nGood driving video sources:\n- Reference dance footage with one subject\n- Walking / gesture / posing clips\n- Talking-head footage when paired with character_orientation: 'video'\n\nBad driving video sources:\n- Music videos with multi-shot edits\n- Anime / animation clips\n- Heavily stylized footage with smoke / particles obscuring the body\n- Footage where the subject's head leaves frame mid-clip\n\n## Target image constraints\n\n- Character body proportions clearly visible\n- Character occupies >5% of image area (not a tiny figure in a wide shot)\n- Single character. Group images break the identity anchor.\n- Any artistic style works — cartoon, anime, painted, realistic, 3D render\n\nAvoid:\n- Extreme close-up of just the face (no body to drive)\n- Character partially cropped at the waist when the driving video is full-body\n- Multiple characters\n\n## character_orientation — the most-missed choice\n\nThis single parameter changes the output dramatically. Pick deliberately.\n\n| Value | Output framing | Max source duration | Best for |\n|-------|----------------|---------------------|----------|\n| `video` | Matches driving video framing | Up to 30s source | Complex full-body motion (dance, action, athletics) |\n| `image` | Matches target image framing | Up to 10s source | Camera moves, simpler motion, preserving original composition |\n\n**Default `video`** when the driving video has the look you want (most cases).\n\nSwitch to `image` when the target image's composition is the brand asset and the motion is secondary (e.g., a hero shot of a character that needs subtle gesture, not a full performance).\n\n## Tier choice — std vs pro\n\n**std (~32 credits)** for:\n- Drafts, motion exploration, blocking\n- Group scenes where the character isn't a hero shot\n- When the budget is tight and the motion is the focus\n\n**pro (~42 credits)** for:\n- Final hero takes\n- Branded characters where identity drift = unacceptable\n- Anatomically complex motion (limbs crossing, fast direction changes)\n- Anime / cartoon target images — pro handles non-realistic styles better\n\nDon't default to pro. The ~10-credit delta compounds fast across iteration.\n\n## Prompt usage (optional)\n\nThe `prompt` field is **scene/style refinement**, not motion direction. The motion comes from the driving video — the prompt sets ambiance, lighting, additional detail.\n\nGood:\n- `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.`\n- `Clean studio backdrop, sharp focus on the character.`\n\nBad (model ignores motion verbs — they're already in the driving video):\n- ❌ `She spins faster and jumps higher.`\n- ❌ `Add more energy to the dance.`\n\nLeave it empty if you don't have a specific atmospheric note.\n\n## Common failure modes\n\n| Symptom | Likely cause | Fix |\n|---------|--------------|-----|\n| Limbs distort / extra fingers | std tier, complex motion | Switch to pro |\n| Character identity drifts | Target image cropped too tight | Use a fuller-body target |\n| Output looks \"stuck\" / minimal motion | Driving video subject too small in frame | Pick a driving video where the subject fills more of the frame |\n| Cartoon target turns realistic | std tier on stylized art | Switch to pro — handles non-realistic styles better |\n| Garbled output entirely | Anime / CG driving video | Use realistic human driving footage |\n| Wrong framing on output | character_orientation set wrong | Try the other value |\n| Background bleeds through character | Target image had complex background | Use a target with cleaner background separation |\n\n## Workflow patterns\n\n**Reference dance to brand character:**\n1. Generate or upload the brand character as a still image (clean background, full body, single subject)\n2. Find driving footage — a clean reference video of the dance you want\n3. Upload both as project assets\n4. Run motion transfer with `motionModel: 'kling-mc-pro'`, `characterOrientation: 'video'`\n5. Total cost: ~42 credits per 5s take\n\n**Subtle motion on a hero portrait:**\n1. Use the locked hero portrait as the target image\n2. Pick a driving video with subtle gesture (head turn, slight posture shift)\n3. `characterOrientation: 'image'` to preserve the portrait's framing\n4. std tier is fine for this case — motion isn't dramatic\n\n**Avoid:**\n- Pro tier on first iteration — waste, switch to it once the motion + framing combo is locked\n- Cartoon driving videos — guaranteed failure\n- Cropped or partial target characters — identity will drift\n- Long driving videos when output is 5s — pick the best 5s of the source upfront\n\n## Cost discipline\n\n- 5 seconds, no shorter option\n- Both tiers trip the confirm gate — every call needs explicit user OK\n- Iteration is expensive: 4 takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.\n- Always run a single std take first to validate the motion + framing combo before committing to pro\n\n## Confirm gate: cost + codes, no inline preview\n\nMotion transfer is mechanical — the model deterministically applies source motion to target image. Both tiers trip the confirm gate; the response includes the asset codes for source and target so you can announce them in chat.\n\n- ✅ \"Transferring motion from **VID-V3** onto **IMG-A12 — Detective Closeup**. ~42 credits, confirm?\"\n- ❌ \"Using the walk video and the detective image...\" (multiple of each in the project.)\n\nDon't second-guess the assets the user picked — the model executes the transfer. If the output is wrong, iterate on motion source or target choice, not on a refinement prompt.\n\n## Sources\n\n- [fal.ai — Kling Motion Control V3 Standard](https://fal.ai/models/fal-ai/kling-video/v3/standard/motion-control)\n- [fal.ai — Kling Motion Control V3 Pro](https://fal.ai/models/fal-ai/kling-video/v3/pro/motion-control)\n",
|
|
17
|
-
"slates-prompting-nano-banana-2": "---\nname: slates-prompting-nano-banana-2\ndescription: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3 Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.\n---\n\n# Nano Banana 2 — cinematic & photorealistic prompting\n\nThe **default** model behind `slates_generate_image` is **Gemini 3 Image** (Nano Banana 2 / Flash) — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill. NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.\n\nKnowledge cutoff: January 2025. Anything after needs explicit reference images.\n\n## Google's 4 official rules (verbatim)\n\n1. **Be specific.** Provide concrete details on subject, lighting, and composition.\n2. **Use positive framing.** Describe what you want, not what you don't want.\n3. **Control the camera.** Use photographic and cinematic terms like \"low angle\" and \"aerial view.\"\n4. **Iterate.** Refine images with follow-up prompts in a conversational manner.\n\n## Official prompt formula\n\n```\n[Subject] + [Action] + [Location/context] + [Composition] + [Style]\n```\n\nFor the cinematic / photoreal use case, expand to:\n\n```\nFilm still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and action]. [3-5 specific visual details]. [LIGHTING — direction + quality]. [COLOR PALETTE]. [FILM STOCK or sensor language]. [1-2 word emotional tone].\n```\n\n## Photorealism positives — what consistently works\n\n**Named lenses + apertures** beat generic \"shallow depth of field\":\n- `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`\n- `Panavision anamorphic` for horizontal flares + cinematic width\n- `400mm telephoto` for compression + isolation\n- `24mm` for environmental interiors\n\n**Named cameras / sensors:**\n- `ARRI Alexa 65`, `Hasselblad X2D`, `Canon EOS R5`, `Sony A7III`, `Fujifilm X-T5`\n- \"Specific gear\" beats \"DSLR\"\n\n**Named film stocks** (one per prompt — never mix):\n- `Kodak Portra 400` — natural skin, warm\n- `Fuji Velvia 50` — saturated, landscape\n- `Ilford HP5 Plus` — black and white, gritty grain\n- `CineStill 800T` — tungsten night, halation\n\n**Physics-based lighting** (direction + quality):\n- `Single key light at 45 degrees from upper left`\n- `Late afternoon sun at 15 degrees above horizon`\n- `Color temperature 4500K` beats `slightly warm`\n- `Practicals only — no fill` for Deakins-style realism\n\n**Imperfection vocabulary** (forces away from AI-clean):\n- `visible pores`, `natural skin grain`, `peach fuzz`, `slight hyperpigmentation`\n- `unretouched raw photography`, `ISO noise`, `sweat beading`\n- `crisp catchlights in the eyes`, `skin micro-detail`\n\n**Director references** (use when locking style):\n| Director | Tone | Visual signature |\n|---|---|---|\n| Denis Villeneuve | Cold, vast, existential | Desaturated, overwhelming scale |\n| Roger Deakins | Precise motivated light | Single source, deep shadows, practicals |\n| Emmanuel Lubezki | Natural, spiritual | Available light, golden hour |\n| Bradford Young | Warm darkness | Underexposed, rich shadows, skin tones |\n\n**Genre cues that move the model:**\n- `unstaged documentary photography style`\n- `fashion magazine editorial, shot on medium-format analog film, pronounced grain`\n- `Film still from [Director] [genre]`\n\n## The anti-list — phrases that DEGRADE realism\n\nThese are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.\n\n**Never use:**\n- `8k`, `4k` (as a quality token)\n- `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone\n- `masterpiece`, `best quality`, `highly detailed`, `ultra-detailed`\n- `trending on ArtStation`, `award-winning`\n- `perfect skin`, `flawless`, `airbrushed`, `smooth skin`\n- `cinematic` standing alone — always specify *which cinema* (director, lens, era, stock)\n- `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue\n\n## Negative prompting — there is no field\n\nNano Banana 2 has **no `negativePrompt` parameter**. Three patterns to suppress unwanted content:\n\n1. **Positive reframing (preferred):** \"empty street\" not \"no cars\". \"Unstaged documentary photography\" not \"not anime.\"\n2. **Inline `without` / `free of`:** \"without any people, vehicles, or man-made structures\", \"free of text overlays, logos, or watermarks.\"\n3. **Constraint clauses for anatomy/quality:** \"accurate anatomy with five fingers per hand, symmetrical features, natural proportions\"; \"sharp, well-exposed, free of blur or JPEG artifacts.\"\n\nDefault to #1. Reach for #2 only when positive framing can't suppress the unwanted element.\n\n## Reference images\n\n- **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.\n- **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style (or pass `referenceAssetIds`), Slates composes the prompt so each reference is named inline as \"image N\" — e.g. `Marcus (images 1 and 2) sits across from the woman (images 3 and 4) in the cafe (image 5)`, with a trailing `Render in the visual style of image 6.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **\"assign a distinct name to each character/object\"**, so citing both of a subject's sheets under the SAME name (\"Marcus\") is what tells the model they are ONE person and stops the face averaging. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render the scene's expression\") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n\n### Reference rules (the verified ones)\n1. **2-4 strong refs beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each adds context AND variables to balance.\n2. **One reference per ROLE, named** (identity / style-grade / environment). Same-role competitors drift. The model doesn't infer roles from order — the inline name does it.\n3. **Identity refs: attach both sheets, named as one entity — don't gate them.** A character's turnaround (body/proportion/outfit) AND its close-up expression sheet (high-res face: eyes, skin, teeth) both go in, cited under the SAME name (\"Marcus (images 1 and 2)\"). That shared name — not a role essay — is what stops the varied expressions from averaging the face. An *unnamed* expression sheet hurts; named as one entity, the close-ups are a fidelity win.\n4. **Flat-light identity refs.** Prep them with flat, even, shadowless lighting on a plain neutral background. Studio-lit / scene-lit sheets bleed their lighting into the generation (\"green-screen pasted in front of mountains\").\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words. Reserve an environment ref for a mandatory exact-match, and then use ONE clean establishing image — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions, then pick a cell. Never feed a grid back in as a reference — cells share a split detail budget, so flaws propagate.\n7. **Reuse the same refs across all shots.** Swapping mid-sequence causes drift.\n8. **Legible in-shot text → bake it into the NB2 start frame**, then animate from it. Never trust text-to-video to render clean text.\n- **Character consistency is officially \"not 100% perfect\"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.\n- **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.\n\n## Common failure modes + fixes\n\n**Hands:** Append `accurate anatomy with five fingers per hand, symmetrical features, natural proportions, relaxed open palm`. Avoid heavy jewelry, props intersecting fingers, motion blur in references.\n\n**Text in images:** Quote-wrap target text. Specify font (`Century Gothic, 12pt`). Long phrases work; small text degrades. Two-step works best — generate text concepts conversationally first, then ask for the image.\n\n**Left/right confusion:** Default is **viewer's perspective**, not subject's. Append `left and right are from the character's perspective, NOT the camera's` when scene-blocking matters.\n\n**Surreal / absurd prompts trip uncanny valley:** The model drags toward realism. If you want surrealism, lean hard into stylization keywords (`painted`, `illustrated`, `stop-motion`).\n\n**Soft faces / dead eyes:** Add `crisp catchlights in the eyes`, `skin micro-detail`, `peach fuzz visible`. Don't stack quality enhancers — single clean prompt beats multiple re-interpretations.\n\n**Post-cutoff content (anything after Jan 2025):** Use reference images. The model has no knowledge of recent franchises, products, events.\n\n## Resolution tactics\n\n- Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — call `slates_estimate_generation_cost` for current numbers. Pick the cheapest resolution that serves the use case.\n- **At 2K and above, the model allocates more tokens to surface detail** — explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.\n- 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.\n- 2K generations vary 20-60s+. Don't time-budget tightly.\n\n## Boring vs cinema — examples\n\n❌ **Boring:** \"Wide shot of a man on a dock looking at the forest.\"\n\n✅ **Cinema:** \"Direct overhead drone shot on weathered dock surface. Single figure standing center frame, climbing up from frame bottom. Boot prints leading away from him toward shore. Pale winter light. Anamorphic lens flare from low sun. Desaturated blue and slate grey palette. Kodak Portra 400 grain. The path already walked by someone else. Map of threat.\"\n\n❌ **Boring:** \"Close up of a woman looking scared.\"\n\n✅ **Cinema:** \"Extreme close on subject's mouth and nose, 135mm f/2.8, shallow depth of field. Breath pluming out, catching cold light from upper-left key. Lips slightly parted, peach fuzz visible. The breath holds. CineStill 800T halation around catchlights. Waiting.\"\n\n## The 3-strike rule\n\nIf three iterations on the same prompt haven't produced what the user wants, stop. Hand back to the user with what you tried and what isn't working. The slot machine doesn't converge — the prompt structure is wrong, not the seed.\n\n## Family variants — Lite and Pro\n\nEverything in this skill applies to the whole Nano Banana family; two variants trade speed/ceiling around NB2 full:\n\n- **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.\n- **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.\n\nRouting between them (and vs GPT Image 2 / FLUX / Seedream): `slates-model-selection`.\n",
|
|
18
|
-
"slates-prompting-omni-flash": "---\nname: slates-prompting-omni-flash\ndescription: How to prompt Gemini Omni Flash (Google, via fal). Read before calling slates_generate_video with omni-flash or slates_edit_video with omni-flash-edit. Cheap 720p tier with native synced audio included — 3-10s, 16:9/9:16 only; t2v, single-start-frame i2v, or reference-to-video with up to 7 reference images. The edit variant is the EDIT-FIDELITY WINNER for footage-synced VFX (receipt 2026-07-09) — but ONLY with short prompts: one change + \"Keep everything else the same.\" Long descriptive prompts destroy fidelity.\n---\n\n# Gemini Omni Flash — prompting\n\nGoogle's fast video generation + editing model (\"Nano Banana Pro for video\" in creator slang — a nickname; it is NOT the NB Pro image model). Carried on fal (`google/gemini-omni-flash*`). 720p only, 24fps, 3–10 second clips, 16:9 or 9:16. **Audio is native and included** — dialogue, SFX, and ambient generate WITH the video at no extra cost.\n\n## Where it routes\n\n- **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.\n- **Cheap drafts and iteration volume** — lowest-cost audio-native video seat (~6.4 cr/s at 720p).\n- **NOT hero GENERATION shots** — Kling 3.0 stays the general gen default, Seedance 2.0 the premium tier; Omni Flash's *generation* quality seat is still unproven.\n\n## Editing (`slates_edit_video`, model `omni-flash-edit`) — THE RULES (receipts, not theory)\n\n1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *\"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes.\"* Live receipt 2026-07-09: a long \"keep every frame/word/movement identical…\" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *\"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.\"*\n2. **Always end with \"Keep everything else the same.\"** — the one documented preservation lever.\n3. **Never name a real-world object as a metaphor.** \"Candle-like flame\" rendered a literal candle in his hand. Describe the effect itself (\"small magical flames on his fingertips\").\n3b. **No conditional timing cues — they HARD-FAIL, not drift.** Receipt 2026-07-09: \"a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…\" → deterministic `invalid_request` (2×, \"could not generate with the given inputs\"); collapsing to one continuous action — \"A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke.\" — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video.\n4. **Safety filter (Google's, strict about harm-to-person):** \"fingertips ignite / catch fire\" → `content_policy_violation`. Frame effects as magical/harmless VFX: \"small magical flames appear on his fingertips\" passed. See slates-content-policy §Gemini for the substitution patterns.\n5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.\n6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.\n7. Source clip 3–10s (trim longer clips first). Output length follows the source; billing per output second, rounded up. Voice editing unsupported — never ask it to change dialogue.\n8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).\n9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.\n\n## Generation (`slates_generate_video`, model `omni-flash`)\n\n- **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.\n- Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.\n- **Name references inline** the standard Slates way (\"Marcus (
|
|
19
|
-
"slates-prompting-seedance": "---\nname: slates-prompting-seedance\ndescription: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance prompts have very specific structure (6-step formula + narrative timing beats) that differs from Kling and Veo — don't cross-pollinate the syntax.\n---\n\n# Seedance 2.0 — prompting\n\nByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K — 4K video is Pro-only, default 1080p), 4–15s, first+last frame + up to 9 reference images.\n\n## Official 6-step formula\n\n```\nSubject + Action + Environment + Camera + Style + Constraints\n```\n\n**Sweet spot length:** 60-150 words (not 150-300 — that's the upper bound). Multi-shot can run longer.\n\n## Pin the subject in the first 20-30 words\n\nThe opening sentence is the **identity anchor**. If the subject isn't locked early, the model hallucinates new subjects mid-clip.\n\n```\nA matte black earbud case sits on a polished obsidian surface...\n```\n\n## Narrative timing beats — \"At N seconds\"\n\nUse natural-language time markers, NOT shot brackets or labels.\n\n```\nAt 2 seconds, the camera begins a slow dolly forward.\nAt 4 seconds, the lid opens in slow-motion...\n```\n\nBeat count: **2 for 5s, 3 for 10s, 4-5 for 15s**.\n\n## Camera moves — exact terms only\n\n8 supported: `push-in`, `pull-out`, `pan`, `tracking`, `orbit`, `aerial`, `handheld`, `fixed`.\n\n- \"Dolly in\" not \"zoom in\"\n- \"Orbit\" not \"circle\"\n- **One primary camera move per beat** — never stack them\n\n## Lighting is the #1 quality lever\n\nByteDance says lighting has the biggest impact of any prompt element. Describe before or alongside the subject.\n\n```\nA cool-white diagonal beam from upper left, dust particles drifting through.\nSoft golden hour lighting from low west angle.\nDramatic rim light against dark background.\n```\n\n## Camera and subject motion — separate sentences\n\nMixing them is the #1 cause of glitchy / shaky output.\n\n❌ \"The camera speed ramps as the earbud rises.\"\n✅ \"The earbud rises smoothly. The camera tracks upward.\"\n\n## Slow-motion works. \"Fast\" doesn't.\n\n`fast` is ByteDance's #1 quality-degrading keyword. Speed ramps and slow-motion are supported in natural language.\n\n```\nthe lid opens in slow-motion · the blade whips through the air\n```\n\nOther dangerous tags (treated as slop): `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.\n\n## Style block at the end\n\nOne primary anchor + 2-3 supporting details. End with `Single continuous take` if you want one shot with no cuts. **Never** write `no cut` or `seamless transition` — not in the training vocabulary.\n\n## Reference media — `@Image1` / `@Video1` / `@Audio1` syntax\n\nReference-to-video endpoint accepts up to **9 reference images, 3 reference videos, 3 audio clips**. Tag inline:\n\n```\n@Image1 is the character. @Image2 is the environment. @Audio1 is the foley.\n```\n\n**Mutually exclusive:** First-frame/last-frame mode CANNOT be combined with reference images. The error reads `\"first/last frame content cannot be mixed with reference media content.\"` Pick one or the other.\n\n### Motion transfer & lip-sync recipes (reference video / audio)\n\nThese aren't separate Seedance features — they're prompting strategies over reference media, and the Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:\n\n- **Motion transfer:** subject image as a reference + the driving clip via `videoReferenceAssetId` (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`\n- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: \"…\"` — with `generate_audio` on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (`audioReferenceAssetId`, ≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`\n- **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script → their likeness AND voice deliver the new line.\n- **Billing:** a reference VIDEO switches the cost key to `seedance-2*-vref-{res}-{T}s` where T = clip seconds + output seconds — quote before confirming. Audio references are free (audio is included on every route).\n\n## Faces — set `seedanceFace` for AI-character faces\n\nSeedance routes through **three tiers** depending on the face in the reference, exposed as the \"Face in Reference\" toggle plus the real-face params on `slates_generate_video`:\n\n- **Faceless / object / environment refs → default route (cheapest).** Leave `seedanceFace` off.\n- **An AI-character's FACE in a reference → `seedanceFace: true`.** The default route's baseline moderation rejects or degrades faces, so this reroutes to the face-capable provider. It costs **~45% more** — the cost key becomes `seedance-2-face-{res}-{N}s`, so the pre-flight quote already reflects it. Announce the face-route price, not the faceless one.\n- **A REAL person's photo (the user themselves, an actor) → the consent-gated premium route.** If a `seedanceFace` gen fails with `[REAL_FACE_DETECTED]`, the provider classified the reference as a real person: confirm with the user that (a) they hold the rights/consent to the likeness and (b) they accept the higher price (cost key `seedance-2-realface-{res}-{N}s`, roughly 2× the AI-face rate — quote via `slates_estimate_generation_cost`), then retry with `seedanceRealFace: true` + `realFaceConsent: true`. Never set `realFaceConsent` without the user's explicit confirmation.\n\nRules:\n- **The real-vs-AI call is the PROVIDER'S, not yours.** ByteDance's classifier is probabilistic — some real photos pass the standard face route (billed at the cheap rate; fine), others get rejected with `[REAL_FACE_DETECTED]` (auto-refunded). Don't preemptively route to the real-face tier just because a photo looks real; try `seedanceFace: true` first and escalate only on the marked rejection. Public figures / celebrities fail on every route.\n- It's about the **reference, not the output.** If your character refs (turnaround, expression sheet, a generated portrait) show a face, turn it on. A product shot with no person stays off.\n- Don't toggle it on \"just in case\" — a faceless gen on the face route burns ~45% extra for nothing.\n\n## Reference rules (the verified ones)\n\n- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say \"still / scene / from a movie / from the image.\" The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one — see slates-cost-discipline).\n- **2-4 strong refs beat both extremes** — not 1 (warps), not 12 (averages worse). Start focused.\n- **One reference per ROLE, named in the prompt.** Seedance's official idiom is **\"Reference \\<Subject_N\\> in \\<Image_N\\>\"** — `Image_N` indexes the order the refs are attached, so the name + index carries the role; the model doesn't infer it from order alone. Slates composes this for you from `@mentions`: it cites each reference inline as \"image N\" in the exact order it sends them. You don't hand-write role labels.\n- **Character identity: attach the turnaround AND the close-up expression sheet, named as one entity** — cite both under the same name. The shared name — not a role essay — is what keeps the varied expressions from averaging the face; don't gate the expression sheet, and don't tell it to \"render neutral / ignore the outfit\" (the user's prompt owns expression, wardrobe, and lighting). The trend is MORE references (video/audio into Seedance), all addressed by name — lean into attaching rich refs and let the naming do the work.\n- **Flat-lit identity refs.** A studio-lit / scene-lit character sheet bleeds its lighting into the clip (\"green-screen pasted in front of mountains\"). Prep refs flat and plain.\n- **Environment: describe it, don't feed a grid.** Default to words and let the model build the space to fit; reserve an environment ref for a hard exact-match, and then use ONE clean establishing image — never a multi-panel grid.\n- **Reuse the same refs across every shot** in a sequence — swapping mid-sequence drifts.\n- **Legible on-screen text → bake it into an NB2 start frame** and animate from it; Seedance won't render clean text from scratch.\n- **Grids are for EXPLORING compositions, not for inputting** — pick a cell, don't feed the grid back as a reference.\n\n## Image-to-video / first-frame guidance\n\n**Describe motion, not image.** The model already sees the visual; tokens spent re-describing appearance are wasted.\n\nRequired stability phrases:\n- `preserve composition and colors`\n- `maintain exact appearance from reference image`\n- `consistent character throughout, no deformation or drift`\n\n**Cap I2V prompts under 60 words** when possible. Over 100 words frequently triggers silent generation failure.\n\n## Negative prompting — inline only\n\nSeedance has **no `negativePrompt` field**. Use the Constraints slot:\n\n```\navoid jitter and bent limbs\navoid temporal flicker\navoid identity drift\nno distortion, no stretching\n```\n\nAlso fine: positive reframing (\"empty street\" not \"no cars\").\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Hallucinated subject mid-clip | First 20-30 words = identity anchor |\n| Bent limbs / extra fingers | `avoid jitter and bent limbs` in Constraints |\n| Identity drift across multi-shot | Repeat subject anchor at each beat |\n| Silent generation failure on I2V | Cut prompt under 100 words, single primary camera move |\n| Speech / motion conflict | Limit dialogue to one line per action shot |\n\n## Benchmark prompts (verbatim from authoritative sources)\n\n**Single-shot (fal.ai):**\n> \"A golden retriever runs across a sandy beach at sunset, kicking up wet sand with each stride, the camera tracking alongside at ground level. Waves crash softly in the background.\"\n\n**Multi-shot commercial (fal.ai):**\n> \"Shot 1: extreme close-up of condensation dripping down a glass bottle, the sound of ice clinking. Shot 2: the bottle rises from a bed of crushed ice, camera tilting up slowly, bright backlight creating a halo effect. Shot 3: a hand grabs the bottle against a sunset rooftop backdrop, the city humming below.\"\n\n**Cinematic anchor (atlabs):**\n> \"Modern Rural Aesthetics, Cinematic Commercial quality, shot with Sony A7S3/cinema camera, 4K/8K ultra-clear, Extreme Macro, natural transparent lighting, healing ASMR, no historical costume drama feel.\"\n\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** — if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Beach Sunset`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.\n\n- ✅ \"I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame — the camera move is going to be a slow dolly forward through the gap.\"\n- ❌ \"I'm using the first beach image and the last one...\" (which? They have four.)\n\n## Sources\n\n- [fal.ai — How to Use Seedance 2.0](https://fal.ai/learn/tools/how-to-use-seedance-2-0)\n- [apiyi.com — Seedance 2.0 Prompt Guide](https://help.apiyi.com/en/seedance-2-0-prompt-guide-video-generation-camera-style-tips-en.html)\n- [atlabs.ai — Ultimate Seedance 2.0 Prompting Guide](https://www.atlabs.ai/blog/the-ultimate-seedance-2.0-prompting-guide-47-prompts-2026)\n",
|
|
17
|
+
"slates-prompting-nano-banana-2": "---\nname: slates-prompting-nano-banana-2\ndescription: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3.1 Flash Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.\n---\n\n# Nano Banana 2 — cinematic & photorealistic prompting\n\nNano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the default model behind `slates_generate_image` — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.\n\nKnowledge cutoff: January 2025. Anything after needs explicit reference images.\n\n## Google's 4 official rules (verbatim)\n\n1. **Be specific.** Provide concrete details on subject, lighting, and composition.\n2. **Use positive framing.** Describe what you want, not what you don't want.\n3. **Control the camera.** Use photographic and cinematic terms like \"low angle\" and \"aerial view.\"\n4. **Iterate.** Refine images with follow-up prompts in a conversational manner.\n\n## Official prompt formula\n\n```\n[Subject] + [Action] + [Location/context] + [Composition] + [Style]\n```\n\nFor the cinematic / photoreal use case, expand to:\n\n```\nFilm still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and action]. [3-5 specific visual details]. [LIGHTING — direction + quality]. [COLOR PALETTE]. [FILM STOCK or sensor language]. [1-2 word emotional tone].\n```\n\n## Photorealism positives — what consistently works\n\n> ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern — do not carry it across.**\n> Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.\n> The leak happens in one specific way — you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `slates-prompting-seedance` (Part 3, \"Don't cross-pollinate image-model syntax\").\n\n**Named lenses + apertures** beat generic \"shallow depth of field\":\n- `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`\n- `Panavision anamorphic` for horizontal flares + cinematic width\n- `400mm telephoto` for compression + isolation\n- `24mm` for environmental interiors\n\n**Named cameras / sensors:**\n- `ARRI Alexa 65`, `Hasselblad X2D`, `Canon EOS R5`, `Sony A7III`, `Fujifilm X-T5`\n- \"Specific gear\" beats \"DSLR\"\n\n**Named film stocks** (one per prompt — never mix):\n- `Kodak Portra 400` — natural skin, warm\n- `Fuji Velvia 50` — saturated, landscape\n- `Ilford HP5 Plus` — black and white, gritty grain\n- `CineStill 800T` — tungsten night, halation\n\n**Physics-based lighting** (direction + quality):\n- `Single key light at 45 degrees from upper left`\n- `Late afternoon sun at 15 degrees above horizon`\n- `Color temperature 4500K` beats `slightly warm`\n- `Practicals only — no fill` for Deakins-style realism\n\n**Imperfection vocabulary** (forces away from AI-clean):\n- `visible pores`, `natural skin grain`, `peach fuzz`, `slight hyperpigmentation`\n- `unretouched raw photography`, `ISO noise`, `sweat beading`\n- `crisp catchlights in the eyes`, `skin micro-detail`\n\n**Director references** (use when locking style):\n| Director | Tone | Visual signature |\n|---|---|---|\n| Denis Villeneuve | Cold, vast, existential | Desaturated, overwhelming scale |\n| Roger Deakins | Precise motivated light | Single source, deep shadows, practicals |\n| Emmanuel Lubezki | Natural, spiritual | Available light, golden hour |\n| Bradford Young | Warm darkness | Underexposed, rich shadows, skin tones |\n\n**Genre cues that move the model:**\n- `unstaged documentary photography style`\n- `fashion magazine editorial, shot on medium-format analog film, pronounced grain`\n- `Film still from [Director] [genre]`\n\n## The anti-list — phrases that DEGRADE realism\n\nThese are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.\n\n**Never use:**\n- `8k`, `4k` (as a quality token)\n- `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone\n- `masterpiece`, `best quality`, `highly detailed`, `ultra-detailed`\n- `trending on ArtStation`, `award-winning`\n- `perfect skin`, `flawless`, `airbrushed`, `smooth skin`\n- `cinematic` standing alone — always specify *which cinema* (director, lens, era, stock)\n- `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue\n\n## Negative prompting — there is no field\n\nNano Banana 2 has **no `negativePrompt` parameter**. Three patterns to suppress unwanted content:\n\n1. **Positive reframing (preferred):** \"empty street\" not \"no cars\". \"Unstaged documentary photography\" not \"not anime.\"\n2. **Inline `without` / `free of`:** \"without any people, vehicles, or man-made structures\", \"free of text overlays, logos, or watermarks.\"\n3. **Constraint clauses for anatomy/quality:** \"accurate anatomy with five fingers per hand, symmetrical features, natural proportions\"; \"sharp, well-exposed, free of blur or JPEG artifacts.\"\n\nDefault to #1. Reach for #2 only when positive framing can't suppress the unwanted element.\n\n## Reference images\n\n- **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.\n- **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as \"image N\" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **\"assign a distinct name to each character/object\"**. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render the scene's expression\") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n\n### Reference rules (the verified ones)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Nano Banana 2 specifically\n\n- **NB2's own consistency lever is \"assign a distinct name to each character/object.\"** That is Google's phrasing for rule 3 — cite each canonical identity inline by name.\n- **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.\n- **Character consistency is officially \"not 100% perfect\"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.\n- **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.\n\n## Common failure modes + fixes\n\n**Hands:** Append `accurate anatomy with five fingers per hand, symmetrical features, natural proportions, relaxed open palm`. Avoid heavy jewelry, props intersecting fingers, motion blur in references.\n\n**Text in images:** Quote-wrap target text. Specify font (`Century Gothic, 12pt`). Long phrases work; small text degrades. Two-step works best — generate text concepts conversationally first, then ask for the image.\n\n**Left/right confusion:** Default is **viewer's perspective**, not subject's. Append `left and right are from the character's perspective, NOT the camera's` when scene-blocking matters.\n\n**Surreal / absurd prompts trip uncanny valley:** The model drags toward realism. If you want surrealism, lean hard into stylization keywords (`painted`, `illustrated`, `stop-motion`).\n\n**Soft faces / dead eyes:** Add `crisp catchlights in the eyes`, `skin micro-detail`, `peach fuzz visible`. Don't stack quality enhancers — single clean prompt beats multiple re-interpretations.\n\n**Post-cutoff content (anything after Jan 2025):** Use reference images. The model has no knowledge of recent franchises, products, events.\n\n## Resolution tactics\n\n- Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — check current numbers<!-- slates-only --> by calling `slates_estimate_generation_cost`<!-- /slates-only -->. Pick the cheapest resolution that serves the use case.\n- **At 2K and above, the model allocates more tokens to surface detail** — explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.\n- 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.\n- 2K generations vary 20-60s+. Don't time-budget tightly.\n\n## Boring vs cinema — examples\n\n❌ **Boring:** \"Wide shot of a man on a dock looking at the forest.\"\n\n✅ **Cinema:** \"Direct overhead drone shot on weathered dock surface. Single figure standing center frame, climbing up from frame bottom. Boot prints leading away from him toward shore. Pale winter light. Anamorphic lens flare from low sun. Desaturated blue and slate grey palette. Kodak Portra 400 grain. The path already walked by someone else. Map of threat.\"\n\n❌ **Boring:** \"Close up of a woman looking scared.\"\n\n✅ **Cinema:** \"Extreme close on subject's mouth and nose, 135mm f/2.8, shallow depth of field. Breath pluming out, catching cold light from upper-left key. Lips slightly parted, peach fuzz visible. The breath holds. CineStill 800T halation around catchlights. Waiting.\"\n\n## The 3-strike rule\n\nIf three iterations on the same prompt haven't produced what the user wants, stop. Hand back to the user with what you tried and what isn't working. The slot machine doesn't converge — the prompt structure is wrong, not the seed.\n\n## Family variants — Lite and Pro\n\nEverything in this skill applies to the whole Nano Banana family; two variants trade speed/ceiling around NB2 full:\n\n- **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.\n- **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.\n\n<!-- slates-only -->\nRouting between them (and vs GPT Image 2 / FLUX / Seedream): `slates-model-selection`.\n<!-- /slates-only -->\n",
|
|
18
|
+
"slates-prompting-omni-flash": "---\nname: slates-prompting-omni-flash\ndescription: How to prompt Gemini Omni Flash (Google, via fal). Read before calling slates_generate_video with omni-flash or slates_edit_video with omni-flash-edit. Cheap 720p tier with native synced audio included — 3-10s, 16:9/9:16 only; t2v, single-start-frame i2v, or reference-to-video with up to 7 reference images. The edit variant is the EDIT-FIDELITY WINNER for footage-synced VFX (receipt 2026-07-09) — but ONLY with short prompts: one change + \"Keep everything else the same.\" Long descriptive prompts destroy fidelity.\n---\n\n# Gemini Omni Flash — prompting\n\nGoogle's fast video generation + editing model (\"Nano Banana Pro for video\" in creator slang — a nickname; it is NOT the NB Pro image model). Carried on fal (`google/gemini-omni-flash*`). 720p only, 24fps, 3–10 second clips, 16:9 or 9:16. **Audio is native and included** — dialogue, SFX, and ambient generate WITH the video at no extra cost.\n\n## Where it routes\n\n- **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.\n- **Cheap drafts and iteration volume** — lowest-cost audio-native video seat (~6.4 cr/s at 720p).\n- **NOT hero GENERATION shots** — Kling 3.0 stays the general gen default, Seedance 2.0 the premium tier; Omni Flash's *generation* quality seat is still unproven.\n\n## Editing (`slates_edit_video`, model `omni-flash-edit`) — THE RULES (receipts, not theory)\n\n1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *\"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes.\"* Live receipt 2026-07-09: a long \"keep every frame/word/movement identical…\" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *\"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.\"*\n2. **Always end with \"Keep everything else the same.\"** — the one documented preservation lever.\n3. **Never name a real-world object as a metaphor.** \"Candle-like flame\" rendered a literal candle in his hand. Describe the effect itself (\"small magical flames on his fingertips\").\n3b. **No conditional timing cues — they HARD-FAIL, not drift.** Receipt 2026-07-09: \"a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…\" → deterministic `invalid_request` (2×, \"could not generate with the given inputs\"); collapsing to one continuous action — \"A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke.\" — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video.\n4. **Safety filter (Google's, strict about harm-to-person):** \"fingertips ignite / catch fire\" → `content_policy_violation`. Frame effects as magical/harmless VFX: \"small magical flames appear on his fingertips\" passed. See slates-content-policy §Gemini for the substitution patterns.\n5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.\n6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.\n7. Source clip 3–10s (trim longer clips first). Output length follows the source; billing per output second, rounded up. Voice editing unsupported — never ask it to change dialogue.\n8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).\n9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.\n\n## Generation (`slates_generate_video`, model `omni-flash`)\n\n- **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.\n- Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.\n- **Name references inline** the standard Slates way (\"Marcus (image 1) walks…\"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.\n- **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language (\"rain patters on the tin roof\"). Negative direction as plain instructions (\"Do not show text\").\n- Duration is an explicit 3–10s integer param; cost scales linearly per second.\n\n## Input conditioning (Slates handles this — know it exists)\n\nPhone footage stores rotation as a metadata flag; models ignore it and edit the raw sideways pixels. Clips must be rotation-normalized (and oversized sources downscaled) before upload — receipt 2026-07-09: a portrait Pixel clip came back sideways until conditioned. If an edit output comes back rotated, the source wasn't normalized.\n\n## Content notes\n\n- Google applies its own safety filters to input images/clips and output. Uploads containing recognizable real people are restricted by Google's policy — though own-footage editing of the uploader passed on our route 2026-07-09. See slates-content-policy.\n- Output carries an invisible SynthID watermark (Google-side, programmatic detection only).\n",
|
|
19
|
+
"slates-prompting-seedance": "---\nname: slates-prompting-seedance\ndescription: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance structures multi-beat prompts as a \"Shot 1 / Shot 2 / Shot 3\" storyboard against an 8-slot advanced formula — never per-second time stamps. Its syntax differs from Kling, Veo and the image models; don't cross-pollinate (in particular, no lens / aperture / film-stock vocabulary).\n---\n\n# Seedance 2.0 — prompting\n\nByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K — 4K video is Pro-only, default 1080p), 4–15s, first+last frame, and up to 9 reference images / 3 videos / 3 audio clips.\n\n> **How to read this file.**\n> **[official :NNNN]** — ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/seedance-2-modelark-docs.md`). Receipt-grade; treat as law.\n> **[community]** — third-party guides and our own field notes. Useful, but an `[official]` block always wins.\n> **[slates]** — how the Slates app composes or bills this; not ByteDance doctrine.\n>\n> The split is load-bearing. A community-sourced \"narrative timing beats\" doctrine shipped in this file for months teaching the **exact inverse** of ByteDance's published guidance. Never merge the two registers again.\n\n---\n\n# Part 1 — Official ByteDance doctrine\n\n## What Seedance actually is `[official :1450-1452]`\n\nSeedance 2.0 is a multimodal AI director. It reads text, images, video and audio **simultaneously** and internally decomposes them into two dimensions:\n\n- the **spatial layer** — what is in the frame\n- the **temporal layer** — how things change over time\n\nSo a good prompt is **not \"copywriting-style description\" but an \"engineering-style instruction\"**: who, in what scene, doing what action, how the camera moves, and in what chronological order events occur — delivered respectively to the spatial layer and the temporal layer.\n\n## The advanced formula — 8 slots `[official :1455]`\n\n```text\nprecise subject + action details + scene/environment + lighting & color tone\n+ camera movement + visual style + image quality + constraints\n```\n\n⚠️ There is **no official \"6-step formula.\"** `Subject + Action + Environment + Camera + Style + Constraints` is community branding with no ByteDance source, and it silently drops the **lighting & color tone** and **image quality** slots. Use the 8 slots above.\n\n## Task-type sentence patterns `[official :1389-1425]`\n\nSeedance classifies your request from the phrasing. Use the pattern that matches the task:\n\n| Task | Pattern |\n|---|---|\n| **Image reference** | ``Reference `<Subject_N>` in `<Image_N>` to generate…`` |\n| **Video reference** | ``Reference `<Action / Camera_movement / Style / Sound_effect>` in `<Video_N>` to generate…`` |\n| **Audio reference** | ``Reference the timbre in `<Audio_N>` to generate…`` |\n| **Video edit — modify** | ``Strictly edit `<Video_N>`, and modify `<Original_Characteristic>` in it to `<New_Characteristic>``` |\n| **Video edit — add** | ``<Element_Features>` + `<Timing>` + `<Location>`` |\n| **Video edit — delete** | Name what to delete; for anything that must stay, say so explicitly |\n| **Video extend** | ``Extend `<Video_N>` forward/backward to generate…`` |\n| **Combined** | ``Reference `[Dimension]` of `<Image/Video_N>`, strictly edit `<Video_X>`, `[Specific_Edits]``` |\n\n### ⚠️ Edit / extend phrasing landmine `[official :1431]`\n\n> *\"For edit / extend video tasks, directly use `<Video_N>` to refer to the video. **Do not use \"reference `<Video_N>`\"**, to avoid being incorrectly identified as a reference task.\"*\n\nThis is easy to trip: Slates has an edit lane<!-- slates-only --> (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)<!-- /slates-only -->. Writing *\"reference video 1 and change the jacket to red\"* gets classified as a **reference** task — the model generates a brand-new clip inspired by the source instead of editing it. Write *\"Strictly edit video 1, and modify the blue jacket to red.\"*\n\n## Shot structure — \"Shot 1 / Shot 2 / Shot 3\" `[official :1563-1598]`\n\n> *\"Use shot order, write a simple 'Shot 1 / Shot 2 / Shot 3' storyboard for each segment of the video, and then merge them into a complete prompt.\"*\n\n**❌ Never second-stamp.** No `0:00–0:03`, no \"At 4 seconds\", no per-segment durations.\n\n> *\"Do not impose strict limits on the duration of each segment; prioritize allowing the model to naturally generate the pacing based on the plot.\"* `[:1580]`\n>\n> *\"The model's support for precise timing (such as 0–3 seconds) is **unstable**, and forcibly limiting duration may lead to **abnormal generation results**.\"* `[:1586]`\n\nOrder shots by when events occur — primary first, secondary later. Let the plot set the pacing.\n\n**Per-shot internal order** `[official :1590-1598]` — organize each shot in exactly this sequence:\n\n1. **Camera movement or shot transition** — \"slowly push in from a wide shot\", \"fixed camera position\", \"cut to…\"\n2. **Subject actions and expressions** — the key actions and expression changes of the core character/object\n3. **Position or spatial change** — where the subject is, and how that relationship shifts\n4. **Audio** — sound effects, voices, background music for that shot\n\n**One primary camera move per shot** — see Camera below. `[official :1648]`\n\n## Subject binding — names + image indexes `[official :1488-1556]`\n\nEvery time a subject appears, it must be **explicitly referred to**. Two supported forms:\n\n- **Undefined subjects** — bind inline every mention: `<Subject_N>@<Image_N>`. Official example: **`Zhang San@Image 1`**. `[:1540]`\n- **Pre-declared subjects** — define once, then reuse the same label verbatim: *\"Define the tall man in **Video 1** as **police officer**, and define the other short man as **thief**\"*, then say \"police officer\" every time after. `[:1514]`\n\n**One subject spread across several assets** — bind them together: *\"Define `[…]` in **Image 1** and `[…]` in **Image 2** as `<Subject N>`.\"* `[:1514]`\n\n⚠️ **An Asset ID must never substitute for `<Image/Video_N>`.** `[:1546]` *\"the model cannot directly associate the Asset ID with the reference content.\"* Always cite by index.\n\nAlso official: keep descriptions concise, avoid redundancy, avoid semantic conflicts (contradictory traits for one subject), and prefer expressing spatial relationships through reference images rather than dense text. `[:1550-1556]`\n\n**`[slates]`** — the app composes this for you. `composeReferences()` cites each canonical character or environment reference inline as `Name (image N)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*\"Zhang San (corresponding to image 1)\"* `[:1976]`). You never hand-write role labels or index numbers.\n\n## Action description `[official :1602-1621]`\n\n- **Body-part specificity + quantified degree.** Name hands, legs, head, shoulders, back — and supplement **range, speed, and force**. *\"slowly raise a hand\", \"quickly turn the head\", \"push hard off the ground\", \"slightly lower the head.\"*\n- **Prioritize slow, gentle, continuous small movements.** Avoid high-burst, large-dynamic actions — sprinting, big jumps, violent rolls. *\"walk slowly\", \"gently raise a hand\", \"sit down naturally with the motion.\"* **This is the official basis for the folk rule that \"fast\" degrades quality** — it is not a banned token, it is a class of motion the model handles badly.\n- **Supplement transitions between actions.** Specify inertia and continuity between consecutive beats so movement reads coherent: *\"use the inertia of turning around to naturally raise a hand\", \"naturally transition from a pause into raising a hand.\"*\n\n## Externalize emotion `[official :1623-1636]`\n\nReplace abstract emotion words (\"very sad\", \"extremely angry\") with **specific physical detail**. This is the highest-leverage single habit in the official guide:\n\n| Abstract | Externalized as actions and details |\n|---|---|\n| **Sadness** | head lowering, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the corner of clothing, tears welling but not falling |\n| **Joy** | corners of the mouth rising uncontrollably, brows and eyes relaxing, steps becoming light, unconsciously humming a tune |\n| **Nervousness / anxiety** | frequently checking the watch, fingers constantly tapping the tabletop, rapid breathing, eyes darting away |\n| **Anger** | both fists clenched, jawline tense, chest heaving, eyes sharp, squeezing words out through gritted teeth |\n| **Relief** | letting out a long breath, tense shoulders completely relaxing, a faint smile appearing, looking up toward the distance |\n\n## Camera `[official :1643-1648]`\n\n> *\"The model has a **strong understanding of camera movement terms**, so you can **directly use standard camera movement terminology**, such as 'medium shot, close-up, wide shot, slow push-in, smooth lateral tracking, fixed shot.'\"*\n\nThis is an **open vocabulary, not a fixed list** — and it explicitly includes **shot size** (close-up / medium / wide / long shot), which is as much a camera instruction as the move itself.\n\n> ⚠️ *\"Try to specify only 1 type of camera movement in a single shot. Do not require push, pull, pan, and move at the same time, as this will increase image instability.\"* `[:1648]`\n\n## Image quality, style, and constraints `[official :1656-1679]`\n\nThese three slots \"define creative boundaries for the model, unify image quality and artistic tone, and avoid visual flaws and random deviations.\"\n\n**1. Image quality** — define clarity, texture detail, and lighting quality. Official vocabulary: `HD` · `rich details` · `cinematic texture` · `natural colors` · `soft lighting`.\n\n> ⚠️ This is a **real slot with real vocabulary** — do not confuse it with Stable-Diffusion-era quality incantations. `8K` / `masterpiece` / `trending on artstation` remain banned slop tokens (see Part 3); *\"cinematic texture, rich details, natural colors\"* is the officially sanctioned way to ask for the same thing.\n\n**2. Style** — the overall art style and visual tone: `cyberpunk cool blue-purple tone` · `retro film` · `fresh Japanese style`.\n\n**3. Constraint words** — *\"Constraint words are very important. They can effectively avoid visual flaws, deformities, breakdowns, and unreasonable elements.\"* Official templates, verbatim:\n\n- **No subtitles** — \"keep it subtitle-free\" / \"avoid generating any text or subtitles\"\n- **No logo** — \"do not generate a logo\"\n- **No watermark** — \"do not generate a watermark\"\n\nSeedance has **no `negativePrompt` field** — constraints go inline in this slot. See Part 3 for the wider inline-negative kit.\n\n## 🔴 Duplicated characters — the twin problem `[official :1948-1994]`\n\n**Symptom:** in frames with **many characters**, where **three-view / multi-view character images** are supplied as references, two identical characters appear in the same generated frame.\n\n**Root causes** `[:1954-1959]`:\n1. Character subjects are not clearly defined in the prompt, so the model cannot distinguish roles.\n2. *\"When character **three-view / multi-view images** are used as reference assets, it is easy to confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"*\n\n**Official fixes, in their order** `[:1971-1994]` — ByteDance is explicit that these *reduce probability*, not eliminate it:\n\n1. **Bind each character to its image explicitly**, in a consistent format. Official example: *\"Zhang San (corresponding to image 1) throws the green passbook toward Li Si (corresponding to image 2), who is standing.\"*\n2. **Append the global constraint verbatim** at the end of the prompt `[:1982]`:\n > *\"Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect. Keep only a single corresponding character in the same frame, and do not reproduce repeated copies of characters.\"*\n3. **Optimize reference assets** `[:1988]` — *\"For character reference images, prioritize independent single-person photos. Three-view or multi-view assets are not recommended.\"*\n4. **Simplify the prompt** — do not paste a whole script; redundant copy confuses the model.\n\n**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on identity sheets. Practical rule for Slates:\n\n- **Multi-character Seedance shot** → bind every character to its image, append the anti-twin constraint, and prefer single-person / dominant-portrait references over multi-view sheets.\n- **Single-character shot** → the standard character-sheet flow is fine.\n\n**Too many reference people** `[official :2048-2052]` — past **4 reference people**, output stability drops (wrong headcount, duplicates). Official workaround: group the cast into images of ≤4 people each, generate those stills first, then drive the video from them.\n\n## Worked examples `[official :1689-1745]`\n\nThese are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere.\n\n**Example 1 — dormitory emotional short drama (dialogue-focused).** Assets: `@Image 1` half-body photo of the female lead · `@Image 2` dormitory scene reference · `@Video 1` camera-movement reference · `@Audio 1` indoor ambience.\n\n> Use the girl in @Image 1 as the main character, use @Image 2 as the dormitory scene style reference, and refer to the camera movement in @Video 1.\n>\n> **Shot 1**: At dusk, **girl @Image 1** walks briskly to the **dormitory entrance @Image 2**. The camera follows steadily in a medium shot. Warm yellow sunlight spills into the hallway from the window. She pauses at the doorway, takes a deep breath, and looks slightly nervous.\n>\n> **Shot 2**: **Girl @Image 1** pushes the door open and enters the dormitory. The camera cuts to an indoor medium shot. Her roommates look up at her while organizing their books. One of them smiles and asks {How did the exam go? Did you pass?}. The camera slowly cuts between half-body close-ups of several people.\n>\n> **Shot 3**: **Girl @Image 1** first lowers her head with a dejected expression. The camera gives her a close-up. Then she raises her head, unable to hold back a smile, laughs out loud, and says {I was kidding}. Her roommates start chasing and play-fighting with her. The camera slowly pulls back and freezes on a wide shot of the dormitory filled with laughter.\n>\n> The entire video should have a high-definition cinematic documentary style, with warm tones and soft lighting. The character's face remains stable without deformation; movements are natural and smooth, with no stutter or flicker. The ambient sound blends naturally with @Audio 1.\n\n**Example 2 — ancient-style cliff confrontation (action/atmosphere-focused).** Assets: `@Image 1` female lead in red · `@Image 2` assassin in black · `@Image 3` cliff and bamboo forest · `@Video 1` martial-arts camera reference · `@Audio 1` drum beats.\n\n> Use the woman in red from @Image 1 as the female lead, use the woman in black from @Image 2 as the opponent, use the cliff and bamboo forest environment in @Image 3 as the scene reference, refer to the overall camera movement and action rhythm in @Video 1, and synchronize the background sound effects with @Audio 1.\n>\n> **Shot 1**: At dusk, the camera slowly pushes in from a side medium shot of **woman in red @Image 1**. She stands at the edge of the cliff and lifts a wine flask to drink. Her sleeves and robe hem sway gently in the mountain wind. The camera circles halfway around her, moving from the front to her back. In the distance, a figure in black is faintly visible in the bamboo forest.\n>\n> **Shot 2**: The camera zooms and fades into a long shot. From a drone perspective, it overlooks the entire cliff and bamboo forest. The two characters stand at opposite ends of the cliff. The mountain wind lifts their robe hems and dust, and the rhythm slightly accelerates with the drum beats.\n>\n> **Shot 3**: The camera cuts back to a ground-level close shot. The two slowly draw their swords and face off. **Woman in red @Image 1** shifts from a careless expression to a cold gaze. **Woman in black @Image 2** looks determined, and the sword tip trembles slightly. The camera steadily follows the two as they circle each other, finally freezing on a close-up of the instant before the two swords meet.\n>\n> The overall visual style should feel like a cinematic wuxia world in misty rain, with cool tones, low saturation, a film-grain texture, and rich light-and-shadow layers. The characters' faces and body proportions remain stable without deformation. Movements are continuous and natural, not stiff, with no clipping or stutter.\n\n## Other official notes\n\n- **On-screen text** `[official :1758]` — Seedance can render common text (ad slogans, subtitles, speech bubbles) and will auto-match style/colour from context, or take an explicit colour / style / timing / position. Prefer **common characters**; avoid rare glyphs and special symbols. (For *guaranteed* legible text, the start-frame route in Part 3 is still safer.)\n- **Extension degrades quality** `[official :2004-2024]` — using a generated video as the input for extension compounds degradation, with mottled colour blocks in face regions. Limit repeated continuations; prefer HD assets as input.\n- **Special effects that miss** `[official :2031-2044]` — when a described effect comes out wrong (a countdown that scrolls randomly), define it with a **reference video** instead of words: *\"the way the number '2999' appears should reference video 1.\"*\n\n---\n\n# Part 2 — Slates-specific `[slates]`\n\n## Reference media — caps and transport\n\nReference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.\n\n**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `\"first/last frame content cannot be mixed with reference media content.\"` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*\n\n### Motion transfer & lip-sync recipes (reference video / audio)\n\nThese aren't separate Seedance features — they're prompting strategies over reference media.<!-- slates-only --> The Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:<!-- /slates-only -->\n\n- **Motion transfer:** subject image as a reference + the driving clip<!-- slates-only --> via `videoReferenceAssetId`<!-- /slates-only --> (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`\n- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: \"…\"` — with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`\n- **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script → their likeness AND voice deliver the new line.\n<!-- slates-only -->\n- **Billing:** a reference VIDEO switches the cost key to `seedance-2*-vref-{res}-{T}s` where T = clip seconds + output seconds — quote before confirming. Audio references are free (audio is included on every route).\n<!-- /slates-only -->\n\n<!-- slates-only -->\n## Faces — set `seedanceFace` for AI-character faces\n\nSeedance routes through **three tiers** depending on the face in the reference, exposed as the \"Face in Reference\" toggle plus the real-face params on `slates_generate_video`:\n\n- **Faceless / object / environment refs → default route (cheapest).** Leave `seedanceFace` off.\n- **An AI-character's FACE in a reference → `seedanceFace: true`.** The default route's baseline moderation rejects or degrades faces, so this reroutes to the face-capable provider. It costs **~45% more** — the cost key becomes `seedance-2-face-{res}-{N}s`, so the pre-flight quote already reflects it. Announce the face-route price, not the faceless one.\n- **A REAL person's photo (the user themselves, an actor) → the consent-gated premium route.** If a `seedanceFace` gen fails with `[REAL_FACE_DETECTED]`, the provider classified the reference as a real person: confirm with the user that (a) they hold the rights/consent to the likeness and (b) they accept the higher price (cost key `seedance-2-realface-{res}-{N}s`, roughly 2× the AI-face rate — quote via `slates_estimate_generation_cost`), then retry with `seedanceRealFace: true` + `realFaceConsent: true`. Never set `realFaceConsent` without the user's explicit confirmation.\n\nRules:\n- **The real-vs-AI call is the PROVIDER'S, not yours.** ByteDance's classifier is probabilistic — some real photos pass the standard face route (billed at the cheap rate; fine), others get rejected with `[REAL_FACE_DETECTED]` (auto-refunded). Don't preemptively route to the real-face tier just because a photo looks real; try `seedanceFace: true` first and escalate only on the marked rejection. Public figures / celebrities fail on every route.\n- It's about the **reference, not the output.** If your character identity or generated portrait shows a face, turn it on. A product shot with no person stays off.\n- Don't toggle it on \"just in case\" — a faceless gen on the face route burns ~45% extra for nothing.\n<!-- /slates-only -->\n\n## Reference rules (the verified ones)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Seedance specifically\n\n- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say \"still / scene / from a movie / from the image.\" The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one<!-- slates-only --> — see slates-cost-discipline<!-- /slates-only -->).\n- **Seedance's own idiom for rule 2 is `Reference <Subject_N> in <Image_N>`** `[official :1389]` — `Image_N` indexes the order the refs are attached, so the name plus the index carries the role. The full binding grammar is in Part 1 (Subject binding).\n- **Rule 3 has an official ceiling here.** The trend is MORE references (video and audio into Seedance), all addressed by name — but for **multi-character frames** see the twin-problem section above: bind every character to its image, append the anti-twin constraint, and prefer single-person references. Past 4 reference people, stability drops `[official :2048-2052]`.\n- **Rule 8 holds even though Seedance can render common text natively** `[official :1758]`. A baked NB2 start frame is still the reliable route for text that must be legible.\n- **Rule 5 pairs with the first/last-frame exclusion** — frames and reference images are mutually exclusive on this model (see Reference media above), so an environment you must match exactly costs you the frame lane.\n\n<!-- slates-only -->\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** — if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Beach Sunset`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.\n\n- ✅ \"I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame — the camera move is going to be a slow dolly forward through the gap.\"\n- ❌ \"I'm using the first beach image and the last one...\" (which? They have four.)\n<!-- /slates-only -->\n\n---\n\n# Part 3 — Community field notes `[community]`\n\nThird-party guides and Slates field experience. Useful heuristics — but if one of these ever appears to contradict Part 1, **Part 1 wins**.\n\n## Length\n\n**Sweet spot 60-150 words** for a single shot (not 150-300 — that's the upper bound). Multi-shot storyboards run longer; official Example 1 above is ~230 words across three shots.\n\n## Pin the subject in the first 20-30 words\n\nThe opening sentence is the **identity anchor**. If the subject isn't locked early, the model hallucinates new subjects mid-clip. (Compatible with Part 1: the binding preamble comes before `Shot 1`.)\n\n```\nA matte black earbud case sits on a polished obsidian surface...\n```\n\n## Lighting is a top quality lever\n\nLighting has an outsized impact on output quality — which is why it has its own slot in the official 8-slot formula. Describe it before or alongside the subject.\n\n```\nA cool-white diagonal beam from upper left, dust particles drifting through.\nSoft golden hour lighting from low west angle.\nDramatic rim light against dark background.\n```\n\n## Camera and subject motion — separate sentences\n\nMixing them is a common cause of glitchy / shaky output.\n\n❌ \"The camera speed ramps as the earbud rises.\"\n✅ \"The earbud rises smoothly. The camera tracks upward.\"\n\n## Slow-motion works; \"fast\" is a known bad token\n\nSpeed ramps and slow-motion are supported in natural language, and `fast` is widely reported as a quality-degrading keyword. **The official version of this rule is stronger and better founded** — prioritize slow, gentle, continuous small movements and avoid high-burst action (Part 1, Action description `[:1611-1615]`). Prompt the motion class, not the adjective.\n\n```\nthe lid opens in slow-motion · the blade whips through the air\n```\n\n**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`. These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).\n\n## Style block at the end\n\nOne primary anchor + 2-3 supporting details, as the trailing paragraph (both official examples do exactly this). End with `Single continuous take` if you want one shot with no cuts. **Never** write `no cut` or `seamless transition` — not in the training vocabulary.\n\n## ⚠️ Don't cross-pollinate image-model syntax\n\nNamed **lenses, apertures, film stocks, and camera bodies** — `85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`, `shot on Sony A7S3` — are an **image-model lever** (correct and encouraged in `slates-prompting-nano-banana-2`) and a **Seedance anti-pattern**. ByteDance's guide uses shot sizes, camera moves, pacing words, and the image-quality/style vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.\n\nIf you are carrying a look over from an NB2 start frame, translate it: `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`.\n\n## Negative prompting — inline only\n\nSeedance has **no `negativePrompt` field**. Put negatives in the constraints slot, led by the three official templates (Part 1):\n\n```\nkeep it subtitle-free · do not generate a logo · do not generate a watermark\navoid jitter and bent limbs\navoid temporal flicker\navoid identity drift\nno distortion, no stretching\n```\n\nAlso fine: positive reframing (\"empty street\" not \"no cars\").\n\n## Image-to-video / first-frame guidance\n\n**Describe motion, not image.** The model already sees the visual; tokens spent re-describing appearance are wasted.\n\nStability phrases that help:\n- `preserve composition and colors`\n- `maintain exact appearance from reference image`\n- `consistent character throughout, no deformation or drift`\n\n**Cap I2V prompts under 60 words** when possible. Over 100 words frequently triggers silent generation failure.\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Hallucinated subject mid-clip | First 20-30 words = identity anchor |\n| Bent limbs / extra fingers | `avoid jitter and bent limbs` in Constraints |\n| Identity drift across multi-shot | Re-name the bound subject in **every** `Shot N` block `[official :1537]` |\n| Two identical characters in one frame | The twin fix in Part 1 — bind each character to its image + append the global anti-twin constraint |\n| Silent generation failure on I2V | Cut prompt under 100 words, single primary camera move |\n| Speech / motion conflict | Limit dialogue to one line per action shot |\n| Erratic/random pacing | You second-stamped. Remove all time markers and use `Shot N` `[official :1586]` |\n\n## Sources\n\n**Official (authoritative):**\n- BytePlus ModelArk — Seedance 2.0 prompting guide, archived at `research/seedance-2-modelark-docs.md` (all `:NNNN` refs above)\n\n**Community (secondary):**\n- [fal.ai — How to Use Seedance 2.0](https://fal.ai/learn/tools/how-to-use-seedance-2-0)\n- [apiyi.com — Seedance 2.0 Prompt Guide](https://help.apiyi.com/en/seedance-2-0-prompt-guide-video-generation-camera-style-tips-en.html)\n- [atlabs.ai — Ultimate Seedance 2.0 Prompting Guide](https://www.atlabs.ai/blog/the-ultimate-seedance-2.0-prompting-guide-47-prompts-2026)\n",
|
|
20
20
|
"slates-prompting-seedream-5-lite": "---\nname: slates-prompting-seedream-5-lite\ndescription: How to prompt Seedream 5 Lite (ByteDance image model — the cheap volume option in Slates). Read before calling slates_generate_image with model seedream-5-lite, or slates_edit_image with editModel seedream-5-lite. Seedream front-loads attention, likes 30-100 focused words, and takes quoted strings for in-image text.\n---\n\n# Seedream 5 Lite — prompting\n\nByteDance's Seedream image model, Lite tier, routed via fal.ai. In Slates: `slates_generate_image` with `model: seedream-5-lite` (REQUIRES projectId — no headless path). **Flat-priced regardless of resolution** — the cheapest image model in Slates, which makes it the right default for high-volume drafting, storyboard exploration, and variant grids. Call `slates_estimate_generation_cost` for the current number; never quote prices from memory. Less censored than Nano Banana 2.\n\n**When to pick it:** lots of frames cheap (storyboard passes, 3-4 variant exploration), posters/layouts with text, quick look-dev. Step up to NB2 or FLUX.2 Max for the locked hero shot.\n\n## Core structure — five components, most important first\n\n```\nSubject + Style + Composition + Lighting/Atmosphere + Technical parameters\n```\n\nSeedream weights concepts mentioned **earlier in the prompt** more heavily. Lead with the subject; close with camera/technical details.\n\n**Length sweet spot: 30-100 words.** Unlike models that reward verbosity, Seedream gets confused by very long prompts. Focused beats exhaustive.\n\n## Style, composition, lighting vocabulary it responds to\n\n- **Style:** portrait photography, macro photography, cinematic, photorealistic, minimalist, oil painting, watercolor, digital art\n- **Composition:** symmetrical composition, rule of thirds, foreground detail with blurred background, wide-angle view, overhead perspective, medium shot, close-up\n- **Lighting:** golden hour lighting, dramatic side lighting, soft diffused light, moody low-key lighting, bright high-key lighting\n- **Technical:** shot on 85mm lens, shallow depth of field, high resolution\n\n## Worked examples\n\n**Portrait:**\n> \"Professional headshot of a female CEO with short blonde hair, confident expression, wearing a navy blue suit, neutral office background, studio lighting, shallow depth of field, high-end corporate photography style\"\n\n**Product:**\n> \"Modern smartphone floating in space, dark background with subtle blue gradient, product photography, studio lighting highlighting the glossy screen, ultra-detailed, commercial quality, photorealistic rendering\"\n\n## In-image text: double-quote it\n\nPut the exact string in double quotation marks — Seedream treats quoted text as render-this-verbatim:\n\n```\nA minimalist poster with the headline \"SUMMER SALE\" in bold sans-serif, centered\n```\n\nSeedream is one of the stronger models for layout-heavy work (posters, mockups, diagrams): call out the layout explicitly — \"centered headline, subtitle beneath, clean margins.\"\n\n## Edits: change one thing, lock the rest\n\nVia `slates_edit_image` with `editModel: seedream-5-lite`. Seedream edits respond well to instructions that name the change AND the preserved elements:\n\n```\nChange the bag to brown leather. Keep the person's face, pose, and the room unchanged.\n```\n\nNote: Seedream edits in Slates ignore extra `referenceAssetIds` — that path is Nano Banana 2 only.\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Subject inconsistent / mutates | Put the subject description first; break complex subjects into clear components |\n| Style drift | Reinforce the aesthetic with 2-3 related terms (\"cinematic, photorealistic, shallow depth of field\") |\n| Compositional confusion | Use photography terms (\"medium shot,\" \"overhead view\"); simplify the scene |\n| Garbled text | Double-quote the exact string; keep it short; state placement |\n| Mushy long-prompt output | Cut to under 100 words — Seedream rewards focus, not volume |\n\n## Iterate cheap, lock expensive\n\nFlat pricing makes Seedream the iterate-fast model: run the 3-strike loop here (draft → evaluate inline → one specific delta → regenerate), and only re-render the winning composition on a pricier model if the project's hero shot demands it. Cost rules live in `slates-cost-discipline` — the batch-authorization pattern applies when generating variant grids.\n\n## Sources\n\n- [fal.ai — Seedream Prompt Guide](https://fal.ai/learn/devs/seedream-v4-5-prompt-guide)\n- [BytePlus ModelArk — Seedream Prompt Guide](https://docs.byteplus.com/en/docs/ModelArk/1829186)\n",
|
|
21
|
-
"slates-prompting-veo-3": "---\nname: slates-prompting-veo-3\ndescription: How to prompt Veo 3.1 (Google). Read before calling slates_generate_video with veo-3.1-fast or veo-3.1-standard. Veo is a NICHE pick, never the default (route per slates-model-selection — Kling is the general default, Seedance the premium tier) — reach for it only when native synchronized audio must generate WITH the video in one gen. 16:9 only. Different cinematography formula than Seedance/Kling. (no subtitles) is mandatory after every dialogue line.\n---\n\n# Veo 3.1 — prompting\n\nGoogle DeepMind's video model. Two tiers: `veo-3.1-fast` (cheaper, quick) and `veo-3.1-standard` (higher quality). 4k variants exist for both (4K video requires Slates Pro).\n\n**Native single-shot duration: 4, 6, or 8 seconds.** Longer durations require chaining clips via Extend / last-frame reuse — quality degrades if naively requested past 8s in a single generation. Aspect ratio: **16:9 only** — `slates_generate_video` locks Veo to 16:9; anything else is ignored or fails. For 9:16 vertical, use Kling or Seedance instead.\n\nNative synchronized audio at 48kHz: dialogue, SFX, ambient — generated WITH video, not added after.\n\n## Official Google formula\n\n```\n[Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]\n```\n\nSweet spot length: 50-150 words. Cloud's official benchmark is ~50 words.\n\nVerbatim official benchmark:\n> \"Medium shot, a tired corporate worker, rubbing his temples in exhaustion, in front of a bulky 1980s computer in a cluttered office late at night. The scene is lit by the harsh fluorescent overhead lights and the green glow of the monochrome monitor. Retro aesthetic, shot as if on 1980s color film, slightly grainy.\"\n\n## Cinematography vocabulary (Vertex AI docs)\n\n**Lenses:** wide-angle, telephoto, fisheye, anamorphic, 35mm, 85mm, shallow/deep depth of field\n\n**Lighting:** Rembrandt lighting, volumetric lighting, backlighting, golden hour glow, lens flare, rack focus, **vertigo effect** (dolly zoom)\n\n**Camera moves:** dolly (in/out), truck (left/right), pan, tilt, crane, aerial/drone, handheld, whip pan, arc shot, zoom\n\n## Texture-realism phrases (counter the AI-plastic look)\n\n```\nfine skin pores · visible fabric weave · subtle contrast, no gloss or sharpening\n```\n\nSpecify materials concretely: `charcoal cotton hoodie`, `matte concrete`, `silk lapel`. Generic \"smooth, beautiful\" rendering is the failure mode you're avoiding.\n\n## Dialogue — `(no subtitles)` is mandatory\n\nEvery dialogue line you don't want burned in as text overlay needs `(no subtitles)`. Verbatim from the founder talking-head benchmark:\n\n```\nThe founder says, \"This update cuts setup time in half, helping teams get started faster.\" (no subtitles).\n```\n\nWithout this, Veo will overlay subtitle text on top of your generation.\n\n## Voice direction — keep it terse\n\nVeo is less responsive to long voice-direction blocks than Kling. Use brief modifiers:\n\n```\nsays in a weary voice\nwhispers\nshouts\nmutters\n```\n\nMulti-character: handles 2-3 speakers natively. Past 3, sync degrades — use first-frame/last-frame chaining for 4+.\n\n## SFX with cause\n\n```\n✅ SFX: thunder cracks in the distance\n❌ SFX: thunder\n```\n\nAlways specify direction or distance.\n\n## Ambient is mandatory\n\nAlways include an ambience line per scene. Without it, the audio mix feels dead.\n\n```\nSoft office ambience.\nWind on the open ridge.\nDistant city hum.\n```\n\n## First-frame + last-frame workflow (Veo's strength)\n\n1. Generate start frame (Gemini 2.5 Flash Image is the recommended pair — Slates' Nano Banana 2 works)\n2. Generate end frame\n3. Animate with both frames as anchors\n\n**Motion-Lock hack:** Keep ~60% of the same background pixels between start and end frames. Prevents latent drift across the clip.\n\nVerbatim arc-shot example:\n> \"The camera performs a smooth 180-degree arc shot, starting with the front-facing view of the singer and circling around her to seamlessly end on the POV shot from behind her on stage. The singer sings 'when you look me in the eyes, I can see a million stars.'\"\n\n## Ingredients-to-Video (multiple references)\n\nVerbatim example:\n> \"Using the provided images for the detective, the woman, and the office setting, create a medium shot of the detective behind his desk. He looks up at the woman and says in a weary voice, 'Of all the offices in this town, you had to walk into mine.'\"\n\n## Reference discipline (character / environment refs)\n\n- **2-4 strong refs per role**, named inline (Slates cites each as \"image N\" from your `@mentions`) and reused across every shot. The model doesn't infer a ref's role from order — the name does it.\n- **Flat-lit identity refs.** A studio-lit / scene-lit character sheet bleeds its lighting into the clip. Prep refs flat and plain.\n- **Attach both character sheets, named as one entity** — the turnaround (body/proportion/outfit) and the close-up expression sheet (face detail), cited under the same name. The shared name keeps the varied expressions from averaging the face; don't write a role essay or \"render neutral\" instruction — the user's prompt owns the expression, wardrobe, and lighting.\n- **Environment: describe it, don't feed a multi-panel grid.** Reserve an environment ref for a hard exact-match, then use ONE clean establishing image.\n\n## Negative prompting — nouns, not instructions\n\nVeo has a `negativePrompt` field. **Verbatim Vertex AI rule:**\n> \"Describe unwanted elements as nouns rather than instructions. Use 'wall, frame' instead of 'no walls' or 'don't show walls.'\"\n\nInline: positive reframing in the body too.\n- ✅ `\"a desolate landscape with no buildings or roads\"`\n- ❌ `\"no man-made structures\"`\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Subject identity shifts mid-clip | Front-load identity at prompt start; use material cues (`charcoal canvas`, `cotton`, `silk`) to stabilize |\n| Floaty / weightless motion | Weight verbs (`trudges`, `drops heavily`), ground contact (`boots crunch on gravel`) |\n| AI-plastic look | `fine skin pores`, `visible fabric weave`, `subtle contrast` |\n| Subtitles baked into video | `(no subtitles)` after every dialogue line |\n| Rushed dialogue | Lines fit one natural breath in 8s |\n| Mismatched ambience | Always include an ambience line |\n| Warped geometry | `photorealistic stability` |\n\n## Timestamp shot syntax (for chained / multi-beat scenes)\n\nVeo accepts `[00:00-00:02]` brackets for timed sequences within an 8s clip. **Do NOT cross syntaxes** — Veo timestamps in a Seedance prompt cause subject drift; Seedance \"single continuous take\" in a Veo prompt suppresses cuts.\n\nVerbatim multi-beat:\n> \"[00:00-00:02] Medium shot from behind a young female explorer with a leather satchel and messy brown hair in a ponytail, as she pushes aside a large jungle vine to reveal a hidden path.\n> [00:02-00:04] Reverse shot of the explorer's freckled face, her expression filled with awe as she gazes upon ancient, moss-covered ruins. SFX: The rustle of dense leaves, distant exotic bird calls.\n> [00:04-00:06] Tracking shot following the explorer as she steps into the clearing and runs her hand over the intricate carvings on a crumbling stone wall.\n> [00:06-00:08] Wide, high-angle crane shot, revealing the lone explorer standing small in the center of the vast, forgotten temple complex, half-swallowed by the jungle. SFX: A swelling, gentle orchestral score begins to play.\"\n\n## Benchmark prompt — founder talking head (full)\n\n> \"Camera locked at eye level, medium close-up on a 35mm lens: a startup founder in his late 30s with short black hair and light stubble, wearing a charcoal cotton hoodie, speaking directly to camera, leaning slightly forward as he speaks, lifting one hand to emphasize a point, then relaxing back to neutral, in a quiet office during late afternoon, with blurred monitors glowing faintly in the background, lit by soft daylight from a side window with gentle fill on the opposite side and natural falloff across his face. Style: fine skin pores, visible fabric weave, subtle contrast, no gloss or sharpening. Audio: The founder says, 'This update cuts setup time in half, helping teams get started faster.' (no subtitles). Soft office ambience.\"\n\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with `firstFrameAssetId` / `lastFrameAssetId` / `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. Veo's strongest move is first-frame + last-frame; the pre-flight is where you confirm the two frames actually anchor the motion you wrote. Revise the prompt before `confirm=true` if needed.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Founder Headshot`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.\n\n- ✅ \"I'm anchoring on **IMG-A12** as the open shot and **IMG-A18** as the close — the 180° arc lands on her looking offscreen left.\"\n- ❌ \"I'm using two of the founder shots...\" (which two? They have six.)\n\n## Sources\n\n- [Google Cloud — Ultimate Prompting Guide for Veo 3.1](https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1)\n- [Google DeepMind — Veo Prompt Guide](https://deepmind.google/models/veo/prompt-guide/)\n- [Google Cloud Docs — Vertex AI Video Generation Prompt Guide](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide)\n- [Atlas Cloud — Veo 3.1 Master Guide](https://www.atlascloud.ai/blog/guides/google-veo-3-1-guide-master-image-to-video-ai-with-native-sound-and-4k-realism)\n- [Invideo — Veo 3.1 Prompt Guide](https://invideo.io/blog/google-veo-prompt-guide/)\n",
|
|
22
|
-
"slates-storyboard-from-script": "---\nname: slates-storyboard-from-script\ndescription: Turn a script or treatment into a Slates storyboard with scenes and frames. Use when the user has a script, treatment, shot list, or scene-by-scene description and wants to materialize it as a Slates storyboard, optionally generating frame images per shot.\n---\n\n# Storyboard from script — Slates workflow\n\nThe user has a script, treatment, or shot list. You're turning it into a Slates storyboard with scene → frame structure, optionally generating images for each frame.\n\n## Workflow\n\n### 1. Parse the script\nRead the user's script. Decide:\n\n- **Scene count** — usually 1 scene per location/setting change. Don't fragment into one-frame scenes.\n- **Frames per scene** — match the shot list. Default is 3-6 frames per scene unless the script specifies more.\n- **Shot labels** — pull them from the script (e.g., \"Wide\", \"Close-up\", \"Over-the-shoulder\").\n\nIf the user hasn't named the storyboard, suggest one based on the project tone.\n\n### 2. Materialize the structure first (no generation yet)\n- `slates_create_storyboard` with the chosen name.\n- For each scene: `slates_add_scene` with a descriptive name and order.\n- For each frame: write a *visual-only* prompt (image, not action), the shot label, and any director notes. **Don't generate yet.**\n\nSurface the planned structure back to the user as a tight summary:\n> Storyboard \"X\" • 4 scenes • 12 frames total\n> Scene 1: Forest opening (3 frames)\n> Scene 2: Confrontation (4 frames)\n> ...\n\nAsk: **\"Generate frame images now? (y/N)\"**\n\n### 3. Generate frames if requested\nFor each frame:\n- Estimate cost (`slates_estimate_generation_cost`, `count = total frames`). Confirm with user if total > ~17 credits.\n- Generate sequentially, with character/environment/style references attached when present in the project (`slates_list_characters`, `slates_list_environments`).\n- Each result returns inline. Evaluate. If wrong, refine prompt + regenerate (charge once, not multiple).\n- Bind to the frame via `slates_add_frame`.\n\n### 4. Hand back\n- Total frames generated, total credits spent, storyboard id.\n- Suggest next steps: review via `slates_get_storyboard_with_frames`, or take the frames to motion — `slates_generate_video` per frame (`firstFrameAssetId`, `background: true`, poll `slates_get_generation_status`), then `slates_add_clip_to_timeline` in story order and `slates_export_video`. The full frames-to-film pipeline (batch cost authorization, model mixing) is `slates-one-prompt-film`.\n\n## Anti-patterns\n\n- **Don't** auto-generate without asking. Generation is the expensive step. Always confirm first.\n- **Don't** invent shot details the script doesn't mention. If the script says \"they argue,\" ask what the shot looks like, don't fabricate \"she clenches her fists in a wide shot.\"\n- **Don't** mix scene structure and frame generation in one pass — building the skeleton first lets the user catch errors before spending credits.\n",
|
|
21
|
+
"slates-prompting-veo-3": "---\nname: slates-prompting-veo-3\ndescription: How to prompt Veo 3.1 (Google). Read before calling slates_generate_video with veo-3.1-fast or veo-3.1-standard. Veo is a NICHE pick, never the default (route per slates-model-selection — Kling is the general default, Seedance the premium tier) — reach for it only when native synchronized audio must generate WITH the video in one gen. 16:9 only. Different cinematography formula than Seedance/Kling. (no subtitles) is mandatory after every dialogue line.\n---\n\n# Veo 3.1 — prompting\n\nGoogle DeepMind's video model. Two tiers: `veo-3.1-fast` (cheaper, quick) and `veo-3.1-standard` (higher quality). 4k variants exist for both (4K video requires Slates Pro).\n\n**Native single-shot duration: 4, 6, or 8 seconds.** Longer durations require chaining clips via Extend / last-frame reuse — quality degrades if naively requested past 8s in a single generation. Aspect ratio: **16:9 only** — `slates_generate_video` locks Veo to 16:9; anything else is ignored or fails. For 9:16 vertical, use Kling or Seedance instead.\n\nNative synchronized audio at 48kHz: dialogue, SFX, ambient — generated WITH video, not added after.\n\n## Official Google formula\n\n```\n[Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]\n```\n\nSweet spot length: 50-150 words. Cloud's official benchmark is ~50 words.\n\nVerbatim official benchmark:\n> \"Medium shot, a tired corporate worker, rubbing his temples in exhaustion, in front of a bulky 1980s computer in a cluttered office late at night. The scene is lit by the harsh fluorescent overhead lights and the green glow of the monochrome monitor. Retro aesthetic, shot as if on 1980s color film, slightly grainy.\"\n\n## Cinematography vocabulary (Vertex AI docs)\n\n**Lenses:** wide-angle, telephoto, fisheye, anamorphic, 35mm, 85mm, shallow/deep depth of field\n\n**Lighting:** Rembrandt lighting, volumetric lighting, backlighting, golden hour glow, lens flare, rack focus, **vertigo effect** (dolly zoom)\n\n**Camera moves:** dolly (in/out), truck (left/right), pan, tilt, crane, aerial/drone, handheld, whip pan, arc shot, zoom\n\n## Texture-realism phrases (counter the AI-plastic look)\n\n```\nfine skin pores · visible fabric weave · subtle contrast, no gloss or sharpening\n```\n\nSpecify materials concretely: `charcoal cotton hoodie`, `matte concrete`, `silk lapel`. Generic \"smooth, beautiful\" rendering is the failure mode you're avoiding.\n\n## Dialogue — `(no subtitles)` is mandatory\n\nEvery dialogue line you don't want burned in as text overlay needs `(no subtitles)`. Verbatim from the founder talking-head benchmark:\n\n```\nThe founder says, \"This update cuts setup time in half, helping teams get started faster.\" (no subtitles).\n```\n\nWithout this, Veo will overlay subtitle text on top of your generation.\n\n## Voice direction — keep it terse\n\nVeo is less responsive to long voice-direction blocks than Kling. Use brief modifiers:\n\n```\nsays in a weary voice\nwhispers\nshouts\nmutters\n```\n\nMulti-character: handles 2-3 speakers natively. Past 3, sync degrades — use first-frame/last-frame chaining for 4+.\n\n## SFX with cause\n\n```\n✅ SFX: thunder cracks in the distance\n❌ SFX: thunder\n```\n\nAlways specify direction or distance.\n\n## Ambient is mandatory\n\nAlways include an ambience line per scene. Without it, the audio mix feels dead.\n\n```\nSoft office ambience.\nWind on the open ridge.\nDistant city hum.\n```\n\n## First-frame + last-frame workflow (Veo's strength)\n\n1. Generate start frame (Gemini 2.5 Flash Image is the recommended pair — Slates' Nano Banana 2 works)\n2. Generate end frame\n3. Animate with both frames as anchors\n\n**Motion-Lock hack:** Keep ~60% of the same background pixels between start and end frames. Prevents latent drift across the clip.\n\nVerbatim arc-shot example:\n> \"The camera performs a smooth 180-degree arc shot, starting with the front-facing view of the singer and circling around her to seamlessly end on the POV shot from behind her on stage. The singer sings 'when you look me in the eyes, I can see a million stars.'\"\n\n## Ingredients-to-Video (multiple references)\n\nVerbatim example:\n> \"Using the provided images for the detective, the woman, and the office setting, create a medium shot of the detective behind his desk. He looks up at the woman and says in a weary voice, 'Of all the offices in this town, you had to walk into mine.'\"\n\n## Reference discipline (character / environment refs)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime → real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Veo specifically\n\n- **Veo's idiom for rule 2 is plain-English role naming in the sentence itself** — *\"Using the provided images for the detective, the woman, and the office setting, create a medium shot of…\"* (see Ingredients-to-Video above). The role rides in the noun phrase, not in a separate label block.\n- **Rule 8 has a second reason to matter here:** Veo bakes subtitle text into the frame unless every dialogue line carries `(no subtitles)`. Text you did not ask for is the failure mode, not just text you did.\n\n## Negative prompting — nouns, not instructions\n\nVeo has a `negativePrompt` field. **Verbatim Vertex AI rule:**\n> \"Describe unwanted elements as nouns rather than instructions. Use 'wall, frame' instead of 'no walls' or 'don't show walls.'\"\n\nInline: positive reframing in the body too.\n- ✅ `\"a desolate landscape with no buildings or roads\"`\n- ❌ `\"no man-made structures\"`\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Subject identity shifts mid-clip | Front-load identity at prompt start; use material cues (`charcoal canvas`, `cotton`, `silk`) to stabilize |\n| Floaty / weightless motion | Weight verbs (`trudges`, `drops heavily`), ground contact (`boots crunch on gravel`) |\n| AI-plastic look | `fine skin pores`, `visible fabric weave`, `subtle contrast` |\n| Subtitles baked into video | `(no subtitles)` after every dialogue line |\n| Rushed dialogue | Lines fit one natural breath in 8s |\n| Mismatched ambience | Always include an ambience line |\n| Warped geometry | `photorealistic stability` |\n\n## Timestamp shot syntax (for chained / multi-beat scenes)\n\nVeo accepts `[00:00-00:02]` brackets for timed sequences within an 8s clip. **Do NOT cross syntaxes** — Veo timestamps in a Seedance prompt cause subject drift; Seedance \"single continuous take\" in a Veo prompt suppresses cuts.\n\nVerbatim multi-beat:\n> \"[00:00-00:02] Medium shot from behind a young female explorer with a leather satchel and messy brown hair in a ponytail, as she pushes aside a large jungle vine to reveal a hidden path.\n> [00:02-00:04] Reverse shot of the explorer's freckled face, her expression filled with awe as she gazes upon ancient, moss-covered ruins. SFX: The rustle of dense leaves, distant exotic bird calls.\n> [00:04-00:06] Tracking shot following the explorer as she steps into the clearing and runs her hand over the intricate carvings on a crumbling stone wall.\n> [00:06-00:08] Wide, high-angle crane shot, revealing the lone explorer standing small in the center of the vast, forgotten temple complex, half-swallowed by the jungle. SFX: A swelling, gentle orchestral score begins to play.\"\n\n## Benchmark prompt — founder talking head (full)\n\n> \"Camera locked at eye level, medium close-up on a 35mm lens: a startup founder in his late 30s with short black hair and light stubble, wearing a charcoal cotton hoodie, speaking directly to camera, leaning slightly forward as he speaks, lifting one hand to emphasize a point, then relaxing back to neutral, in a quiet office during late afternoon, with blurred monitors glowing faintly in the background, lit by soft daylight from a side window with gentle fill on the opposite side and natural falloff across his face. Style: fine skin pores, visible fabric weave, subtle contrast, no gloss or sharpening. Audio: The founder says, 'This update cuts setup time in half, helping teams get started faster.' (no subtitles). Soft office ambience.\"\n\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with `firstFrameAssetId` / `lastFrameAssetId` / `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. Veo's strongest move is first-frame + last-frame; the pre-flight is where you confirm the two frames actually anchor the motion you wrote. Revise the prompt before `confirm=true` if needed.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Founder Headshot`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.\n\n- ✅ \"I'm anchoring on **IMG-A12** as the open shot and **IMG-A18** as the close — the 180° arc lands on her looking offscreen left.\"\n- ❌ \"I'm using two of the founder shots...\" (which two? They have six.)\n\n## Sources\n\n- [Google Cloud — Ultimate Prompting Guide for Veo 3.1](https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1)\n- [Google DeepMind — Veo Prompt Guide](https://deepmind.google/models/veo/prompt-guide/)\n- [Google Cloud Docs — Vertex AI Video Generation Prompt Guide](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide)\n- [Atlas Cloud — Veo 3.1 Master Guide](https://www.atlascloud.ai/blog/guides/google-veo-3-1-guide-master-image-to-video-ai-with-native-sound-and-4k-realism)\n- [Invideo — Veo 3.1 Prompt Guide](https://invideo.io/blog/google-veo-prompt-guide/)\n",
|
|
22
|
+
"slates-storyboard-from-script": "---\nname: slates-storyboard-from-script\ndescription: Turn a script or treatment into a Slates storyboard with scenes and frames. Use when the user has a script, treatment, shot list, or scene-by-scene description and wants to materialize it as a Slates storyboard, optionally generating frame images per shot.\n---\n\n# Storyboard from script — Slates workflow\n\nThe user has a script, treatment, or shot list. You're turning it into a Slates storyboard with scene → frame structure, optionally generating images for each frame.\n\n## Workflow\n\n### 1. Parse the script\nRead the user's script. Decide:\n\n- **Scene count** — usually 1 scene per location/setting change. Don't fragment into one-frame scenes.\n- **Frames per scene** — match the shot list. Default is 3-6 frames per scene unless the script specifies more.\n- **Shot labels** — pull them from the script (e.g., \"Wide\", \"Close-up\", \"Over-the-shoulder\").\n\nIf the user hasn't named the storyboard, suggest one based on the project tone.\n\n### 2. Materialize the structure first (no generation yet)\n- `slates_create_storyboard` with the chosen name.\n- For each scene: `slates_add_scene` with a descriptive name and order.\n- For each frame: write a *visual-only* prompt (image, not action), the shot label, and any director notes. **Don't generate yet.**\n\nSurface the planned structure back to the user as a tight summary:\n> Storyboard \"X\" • 4 scenes • 12 frames total\n> Scene 1: Forest opening (3 frames)\n> Scene 2: Confrontation (4 frames)\n> ...\n\n**Surface a decision log alongside that summary.**\n\n<!-- @inject:decision-log -->\nWhen you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:\n\n```\nsource phrase or declared default → what you wrote → what it resolves\n\"in a diner\" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable\n(no time of day) → late afternoon, low warm key → default; say the word and it changes\n(no camera) → slow push-in, single move → one move per shot; stacking increases instability\n```\n\n**Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the \"why did you add that?\" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.\n\n> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.\n<!-- @end:decision-log -->\n\nTurning a script into *visual* frame prompts means resolving things the script left open — what the room looks like, where the light comes from, how the shot is framed. Those are your decisions, not the writer's; name them.\n\nAsk: **\"Generate frame images now? (y/N)\"**\n\n### 3. Generate frames if requested\nFor each frame:\n- Estimate cost (`slates_estimate_generation_cost`, `count = total frames`). Confirm with user if total > ~17 credits.\n- Generate sequentially, with character/environment/style references attached when present in the project (`slates_list_characters`, `slates_list_environments`).\n- Each result returns inline. Evaluate. If wrong, refine prompt + regenerate (charge once, not multiple).\n- Bind to the frame via `slates_add_frame`.\n\n### 4. Hand back\n- Total frames generated, total credits spent, storyboard id.\n- Suggest next steps: review via `slates_get_storyboard_with_frames`, or take the frames to motion — `slates_generate_video` per frame (`firstFrameAssetId`, `background: true`, poll `slates_get_generation_status`), then `slates_add_clip_to_timeline` in story order and `slates_export_video`. The full frames-to-film pipeline (batch cost authorization, model mixing) is `slates-one-prompt-film`.\n\n## Anti-patterns\n\n- **Don't** auto-generate without asking. Generation is the expensive step. Always confirm first.\n- **Don't** invent shot details the script doesn't mention. If the script says \"they argue,\" ask what the shot looks like, don't fabricate \"she clenches her fists in a wide shot.\"\n- **Don't** mix scene structure and frame generation in one pass — building the skeleton first lets the user catch errors before spending credits.\n",
|
|
23
23
|
"slates-style-prompting": "---\nname: slates-style-prompting\ndescription: Per-style prompting depth — how photoreal, anime, painterly, and 3d-render are prompted DIFFERENTLY on Seedance 2.0 vs Kling V3 vs Nano Banana 2, and the style-routing recipe (reference-first, styled start-frame → i2v). Load when the user asks for any visual style (\"make it anime\", \"painterly look\", \"like a Pixar film\") or when style consistency across shots matters.\n---\n\n# Per-style prompting (photoreal · anime · painterly · 3d-render)\n\nThe style library (`slates_create_style` / the app's style ids) defines what each style IS. This guide is how to PROMPT each style per model. Derived from `research/style-prompting-research.md` (second-brain) — claims marked *(hypothesis)* are untested; don't present them to users as fact.\n\n## The four ground rules (all styles)\n\n1. **Reference beats adjectives.** A style reference image outperforms prose style instructions. If the user has an on-style image — or you can cheaply generate one — attach it and let the default `inherit` behavior match it. Prose styling is the fallback.\n2. **Style lives in ONE slot per model:**\n - **Nano Banana 2** — narrative prose; the style is the opening framing of the sentence (\"A hand-drawn 2D anime cel illustration of…\"), never a comma tag.\n - **Seedance 2.0** — the 8-part formula reserves \"visual style\" (slot 6) and \"image quality\" (slot 7). One clause each. Don't scatter style words through the action text.\n - **Kling V3** — prose scene direction; style rides the lighting/style tail of Scene → Subject → Action → Camera → Lighting/Style. Tag soup underperforms badly.\n3. **Multi-shot consistency = the SAME style clause, byte-identical, in every shot's prompt** (plus shared references). Paraphrasing the style clause between shots invites drift.\n4. **The styled start-frame is the cheapest reliable style lever for video.** Compose the styled frame in NB2 (cheap), hand it to Seedance/Kling image-to-video, and describe only what CHANGES (motion). Never re-describe the style in the i2v prompt — the frame already encodes it.\n\nNever stack style buzzwords (\"ARRI ALEXA, 35mm, film grain, depth-of-field mastery…\"). One or two register tokens maximum — piles of specs dull the image.\n\n## Photoreal\n\n- **NB2:** never the literal word \"photorealistic\". Describe *a real photograph*: natural skin texture and imperfection, motivated lighting, one lens/film register (\"shot on a 50mm, soft window light\"). Photographic composition terms: wide-angle / macro / low-angle.\n- **Seedance:** put \"sharp focus, natural color, high detail\" in the image-quality slot and always include a lighting clause. Keep motion slow and coherent — fast/burst action is the #1 quality killer and reads most fake in photoreal.\n- **Kling:** the photoreal-PEOPLE lane — convincing acting, dialogue, lip-sync. It breaks on close-up hands, fine fluids, and crowds beyond ~5 faces: route those beats to Seedance or reframe.\n- **Faces on Seedance:** photoreal humans trigger the face-tier routing (AI face vs consented real face — see slates-prompting-seedance §Faces). Set the face flags honestly; never skip them to save credits.\n\n## Anime\n\n- **NB2:** open with the medium — \"A hand-drawn 2D anime cel illustration of…\" — then normal narrative Subject/Setting/Action. Clean line art, flat-shaded color, expressive eyes. NB2 has no negative prompt: phrase exclusions positively (\"flat cel shading with uniform focus\", not \"no depth of field\").\n- **Seedance:** visual-style slot = \"2D anime style, clean line art, flat cel shading\". The slow/coherent-motion preference still applies — burst sakuga actions are the same instability trap as in photoreal.\n- **Kling:** weakest anime lane (its strength is live-action-like acting); expect style drift on long prose-only shots. Prefer ground rule 4: NB2 anime start-frame → i2v with a motion-only prompt. *(hypothesis: refs hold Kling's anime better than prose — verify before promising.)*\n- Anime faces drift under multiple references faster than photoreal — the named-entity two-sheet doctrine applies unchanged.\n\n## Painterly\n\n- **NB2:** medium + technique in the style framing: \"digital concept-art painting, visible brushwork, painted edges\". At most ONE school/era register (\"classic gouache illustration\") — a register, not an artist-name pile.\n- **Video:** the least-supported style lane. Use ground rule 4 (painterly NB2 frame → i2v, motion-only prompt) and expect some cleanup of painterliness over the clip *(hypothesis — set user expectations, don't promise a perfectly painterly clip)*.\n- Camera language still applies — painterly ≠ static; \"slow push-in\" works the same.\n\n## 3D render\n\n- **NB2:** name the lineage register in the style framing: \"stylized 3D render, soft global illumination, subsurface skin\". Lighting vocabulary (GI, rim light) is unusually load-bearing for the 3D read.\n- **Seedance:** the physics/effects lane flatters 3D content — visual-style slot \"stylized 3D animation\", image-quality slot \"clean render, high detail\".\n- **Kling:** same start-frame preference as anime.\n- *(hypothesis)* An engine token (\"Unreal Engine 5 render\") may help NB2; if used, ONE token, style slot only — never on Seedance where spec-stuffing hurts.\n\n## Routing recipe (what to actually do)\n\n1. Style reference available → attach it, rely on inherit. Done.\n2. No reference, image request → styled NB2 prose per the section above.\n3. No reference, video request → NB2 styled start-frame first, then i2v with motion-only prompt. Direct styled text-to-video is the fallback when a start frame doesn't fit (e.g. dialogue-first Kling shots).\n4. Multi-shot run → byte-identical style clause per shot + shared references.\n",
|
|
24
|
-
"slates-vision-feedback-loop": "---\nname: slates-vision-feedback-loop\ndescription: Lower-level utility skill for any Slates workflow that needs to \"generate, look at the result, refine, regenerate.\" Defines the standard inline-vision pattern. Other Slates skills compose this. Use when generating images and you need to confirm they match the brief before moving on, or when the user asks to \"iterate\" on an image.\n---\n\n# Vision feedback loop — Slates utility skill\n\nSlates returns generated images inline as base64. You see the actual pixels. Use that — don't trust prompt-following blindly.\n\n## Asset codes are your shared vocabulary with the user\n\nEvery asset in Slates has a short stable code (e.g. `IMG-A12`, `VID-V3`, `AUD-S1`) and a label derived from its prompt (e.g. `Beach Sunset`). These are visible in the gallery as a corner badge on each thumbnail. **Always refer to assets by their code in chat** so the user can match what you're saying to a specific card in their gallery.\n\n- ✅ \"I'm using **IMG-A12 — Beach Sunset** as the first frame. The second-frame candidate **IMG-A15** has the right composition but warmer light — want me to use that one instead?\"\n- ❌ \"I'm using the beach sunset image...\" (user has four beach sunset variants — which one?)\n- ❌ \"I'm using asset `7a3f9e4b-...`\" (UUIDs aren't readable; user can't match to a badge)\n\nThe code is the FORMAL reference. The label is human texture. Use both: `IMG-A12 — Beach Sunset`.\n\n## Vision tools at your disposal\n\n- `slates_get_asset_image` — pull one image into context. Returns its code+label.\n- `slates_get_assets_batch` — pull up to 8 images in one call. Use when picking from a candidate set; cheaper than N individual fetches.\n- `slates_get_asset_video_frames` — extract N keyframes (default 3) from a video and inline them as JPEGs. You can't see video natively; this is how you \"look at\" a clip before refining its motion prompt.\n\n## Pre-flight is automatic on the gen tools\n\n`slates_generate_video`, `slates_generate_motion_transfer`, and `slates_generate_lip_sync` now show you their reference assets **inline** on the confirm response. You don't need to fetch them yourself — but you DO need to look at what comes back, revise the prompt if the references suggest a different motion/framing, and only then re-call with `confirm=true`.\n\n## The pattern\n\n1. **Generate.** Call `slates_generate_image` with a prompt. The result is in your context as an image content block.\n2. **Evaluate
|
|
24
|
+
"slates-vision-feedback-loop": "---\nname: slates-vision-feedback-loop\ndescription: Lower-level utility skill for any Slates workflow that needs to \"generate, look at the result, refine, regenerate.\" Defines the standard inline-vision pattern. Other Slates skills compose this. Use when generating images and you need to confirm they match the brief before moving on, or when the user asks to \"iterate\" on an image.\n---\n\n# Vision feedback loop — Slates utility skill\n\nSlates returns generated images inline as base64. You see the actual pixels. Use that — don't trust prompt-following blindly.\n\n## Asset codes are your shared vocabulary with the user\n\nEvery asset in Slates has a short stable code (e.g. `IMG-A12`, `VID-V3`, `AUD-S1`) and a label derived from its prompt (e.g. `Beach Sunset`). These are visible in the gallery as a corner badge on each thumbnail. **Always refer to assets by their code in chat** so the user can match what you're saying to a specific card in their gallery.\n\n- ✅ \"I'm using **IMG-A12 — Beach Sunset** as the first frame. The second-frame candidate **IMG-A15** has the right composition but warmer light — want me to use that one instead?\"\n- ❌ \"I'm using the beach sunset image...\" (user has four beach sunset variants — which one?)\n- ❌ \"I'm using asset `7a3f9e4b-...`\" (UUIDs aren't readable; user can't match to a badge)\n\nThe code is the FORMAL reference. The label is human texture. Use both: `IMG-A12 — Beach Sunset`.\n\n## Vision tools at your disposal\n\n- `slates_get_asset_image` — pull one image into context. Returns its code+label.\n- `slates_get_assets_batch` — pull up to 8 images in one call. Use when picking from a candidate set; cheaper than N individual fetches.\n- `slates_get_asset_video_frames` — extract N keyframes (default 3) from a video and inline them as JPEGs. You can't see video natively; this is how you \"look at\" a clip before refining its motion prompt.\n\n## Pre-flight is automatic on the gen tools\n\n`slates_generate_video`, `slates_generate_motion_transfer`, and `slates_generate_lip_sync` now show you their reference assets **inline** on the confirm response. You don't need to fetch them yourself — but you DO need to look at what comes back, revise the prompt if the references suggest a different motion/framing, and only then re-call with `confirm=true`.\n\n## 🔴 The still-gate — never animate a bad frame\n\n<!-- @inject:still-gate -->\n**A visible defect in the still is already a STOP.** Do not animate it. Fix the frame first, then move to motion — and go to motion only when the crop passes the still scan and you genuinely need movement to confirm an uncertain edge, reflection, or object.\n\nThis is a **cost** rule as much as a craft rule: a 1080p/10s premium video generation costs many multiples of an image re-roll, and video is where a defect stops being fixable. Anything wrong in the still gets worse in motion — soft geometry mushes, broken-but-plausible objects fall apart, oily textures start crawling. **Animating a known-bad frame is the single most expensive mistake in the pipeline.** Re-rolling the image is the cheap move; re-rolling the video is not.\n<!-- @end:still-gate -->\n\n## The pattern\n\n1. **Generate.** Call `slates_generate_image` with a prompt. The result is in your context as an image content block.\n2. **Evaluate on TWO axes — they are different questions:**\n - **Brief-conformance** — what did the user actually want? Are the elements right? Composition? Lighting? Subject identity?\n - **Defects** — run the slop rubric below. *A frame can match the brief perfectly and still be slop that mushes the moment it moves.* Checking only the first axis is how a bad frame reaches an expensive video call.\n3. **One of three outcomes:**\n - **Right** → save it (bind to a frame, character slot, etc.) and move on.\n - **Close, but adjustable** → refine with a specific delta, regenerate **once**.\n - **Wrong direction** → ask the user before regenerating. Don't burn credits on prompt-thrashing.\n\n## The defect rubric — four slop tells\n\n| Tell | What it looks like | Why it matters downstream |\n|---|---|---|\n| **Light with no transitions** | Flat-black pits instead of a shadow ramp; light that stops rather than falls off | Transfers onto every character or object added into that plate later |\n| **Broken-but-plausible objects** | Crates, railings, hardware, mechanisms you can *almost* read but that don't resolve | Turn to mush in motion, and the model multiplies them |\n| **Local logic breaks** | An effect present in only part of the frame — rain scratching one corner, wet ground under one figure | The video model's physical logic breaks along with it |\n| **Oily textures** | Soapy, licked-smooth surfaces that have lost their material identity | Reflections crawl in motion; the plate can't hold continuity |\n\n### Per-model accents — check the one you actually used\n\n- **Nano Banana Pro** (`nano-banana-pro`) — ruler-straight symmetry, everything parallel and square, flat even light, pretty but staged/stock, textures reading as 3D render rather than photograph. **It hyperbolizes every edit**: ask for graffiti on one wall and the whole location gets tagged.\n- **GPT Image 2** (`gpt-image-2`) — microcontrast to the ceiling, hard halos on every edge, no depth or bokeh, white balance pulled warm until the frame yellows, plastic licked-smooth materials. Worst tell: **one sickly texture pattern laid over the entire frame**.\n\n> ⚠️ These are accents for **`nano-banana-pro`** and **`gpt-image-2`** specifically. `nano-banana-2` is a **different model** (Gemini 3.1 Flash Image vs NB Pro's Gemini 3 Pro Image) and we have **no evidence** about its accent. Do not inherit one — say nothing rather than warn about a failure mode you can't substantiate.\n\n## Where the fault lives — triage before you change anything\n\nWe say \"one specific delta per regeneration\" but that only helps once you know *which* variable to move. Diagnose first:\n\n| Visible pattern | Diagnosis | Fix |\n|---|---|---|\n| The defect exists in the source asset, or stays tied to the same feature when the direction changes | **Source asset** | Fix the sheet / plate, not the prompt |\n| Source is clean, and the defect changes when only the suspect motion clause changes | **Motion direction** | Fix the prompt |\n| Controls conflict, or the failure follows neither variable | **Inconclusive** | Narrow the test — change less, not more |\n\n**Review routes; it is not pass/fail.** Geography melts → fix the location. Identity drifts → fix the character sheet. Assets are sound but the action is wrong → fix the video direction. Wrong idea entirely → reopen the brief with the user.\n\n**Correct the earliest broken handoff.** Polishing a downstream symptom hides the source and guarantees it resurfaces in the next shot built from the same asset.\n\n## Baseline hygiene — isolate the variable you're testing\n\nWhen the **character** is the question, keep the location out of it: test on a plate that already holds its own geometry, depth, materials, and light. **A broken plate gives every character failure a second plausible cause**, and you will spend re-rolls deciding which one you're looking at. The same applies in reverse — test a plate empty before you populate it.\n\n## Refinement rules\n\n- **One specific delta per regeneration.** Don't change five things at once — you won't know what helped.\n- **Rewrite the FULL prompt on every iteration — never a diff, never a fragment.** Change one decision, then re-emit the whole prompt so every slot still agrees with every other slot. This composes with the rule above rather than replacing it: *one delta* governs **what changes**, *full rewrite* governs **how you re-emit it**. A patched fragment leaves the old slots stale and silently contradicting the new one.\n - On **Seedance**, a re-emit must keep the `Shot N` structure intact — see `slates-prompting-seedance`.\n - **Exception — Omni Flash Edit.** Long prompts documentedly destroy its fidelity. There the rule inverts: one short instruction plus *\"Keep everything else the same.\"*\n- **Anchor with references.** If the result drifted from the user's intent, attach the *previous best* generation as a reference image alongside the original brief.\n- **Use `slates_get_asset_image`** to pull a previously-generated image back into context if you need to compare against a fresh generation.\n- **Use `slates_edit_image`** for surgical tweaks instead of full regeneration when ~90% of the image is right — `sourceAssetId` = the asset, `prompt` = the change only. Edits preserve composition and identity; full regen rolls the dice. Recipe: `slates-edit-and-iterate`.\n\n## Cost discipline\n\n- Track total credits spent across the loop. Surface to the user every 3 iterations.\n- Stop after 3 failed iterations on the same prompt — escalate to the user with what you tried and what's not working. The slot machine never converges.\n- For high-cost generations (above ~17 credits), confirm before *every* attempt, not just the first.\n\n## When to break the loop\n\n- The user said \"good enough\" or \"ship it.\" Stop iterating.\n- You've burned >5 generations on one frame. Hand back and ask.\n- The user changes brief mid-loop. Treat it as a new brief, not a continuation.\n\n## Voice when narrating to the user\n\nTight, observational, no editorializing.\n- ✅ \"Frame 2 has the wrong lighting direction — back-lit instead of side. Regenerating with side light.\"\n- ❌ \"I notice that the lighting in frame 2 isn't quite what we were going for. I'll go ahead and try again with a different approach.\"\n",
|
|
25
25
|
};
|
|
26
26
|
//# sourceMappingURL=content.js.map
|