@slatesvideo/shared 0.6.6 β†’ 0.6.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -153,7 +153,7 @@ export const MODEL_CAPABILITIES = {
153
153
  // than a number of ours: `image_urls` carries `maxItems: 16` on both 2.5
154
154
  // endpoints AND on both gpt-image-2 endpoints (schema, read 2026-09-09,
155
155
  // ripped verbatim to second-brain/business/projects/slates/research/
156
- // gpt-image-2-5-fal-api-docs.md).
156
+ // fal-gpt-image-2-5-openapi-schemas.md).
157
157
  //
158
158
  // 🚨 IT WAS 10 UNTIL 2026-09-09, AND 10 WAS NEVER ANYBODY'S LIMIT. The
159
159
  // comment here used to call it "the 10-reference ceiling ... unchanged by the
@@ -25,8 +25,8 @@ export const SKILLS = {
25
25
  "slates-prompting-nano-banana-2": "---\nname: slates-prompting-nano-banana-2\ndescription: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3.1 Flash Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work β€” the rules differ.\n---\n\n# Nano Banana 2 β€” cinematic & photorealistic prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Nano Banana 2 (Gemini 3.1 Flash Image).** Brief it like a creative director, not a tag list. Structure: `Film still from [director] [genre]. Shot on [camera] with [lens]. [Subject and action]. [3-5 specific visual details]. [Lighting β€” direction + quality]. [Color palette]. [Film stock]. [1-2 word tone].`\n\n**The five levers**\n1. **Named lens + aperture** beats \"shallow depth of field\" β€” `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin), `Panavision anamorphic`, `400mm telephoto`.\n2. **Light by direction and quality**, never \"good lighting\" β€” `hard sidelight from a single window, deep falloff`, `overcast north light`, `practical tungsten spill`.\n3. **A named film stock or sensor** carries a whole palette β€” `Kodak Portra 400`, `Cinestill 800T`, `ARRI Alexa 65`.\n4. **Composition as a shot** β€” `low angle`, `aerial view`, `rule of thirds with the subject camera-left`, `foreground occlusion`.\n5. **Positive framing only.** Describe what is there. \"Empty street\", never \"no cars\"; \"unstaged documentary photography\", never \"not anime\".\n\n**Examples**\n- `Film still from a Denis Villeneuve thriller. Shot on ARRI Alexa 65, 85mm f/1.4. A woman in a charcoal wool coat stands at a rain-slick bus stop, breath visible. Hard sodium light from a single overhead lamp, deep falloff into blue night. Kodak Vision3 500T. Isolated.`\n- `Editorial still life on seamless bone paper. 100mm macro, f/8. A cracked ceramic bowl holding three figs. Soft north light from camera-left, one gentle shadow. Muted earth palette. Portra 400 grain. Quiet.`\n\n**Hard constraint:** there is no `negativePrompt` field. Suppress by reframing positively, or inline `without` / `free of`. Knowledge cutoff January 2025 β€” anything later needs reference images.\n<!-- @card:end -->\n\nNano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the default model behind `slates_generate_image` β€” the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels β€” brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.\n\nKnowledge cutoff: January 2025. Anything after needs explicit reference images.\n\n## Google's 4 official rules (verbatim)\n\n1. **Be specific.** Provide concrete details on subject, lighting, and composition.\n2. **Use positive framing.** Describe what you want, not what you don't want.\n3. **Control the camera.** Use photographic and cinematic terms like \"low angle\" and \"aerial view.\"\n4. **Iterate.** Refine images with follow-up prompts in a conversational manner.\n\n## Official prompt formula\n\n```\n[Subject] + [Action] + [Location/context] + [Composition] + [Style]\n```\n\nFor the cinematic / photoreal use case, expand to:\n\n```\nFilm still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and action]. [3-5 specific visual details]. [LIGHTING β€” direction + quality]. [COLOR PALETTE]. [FILM STOCK or sensor language]. [1-2 word emotional tone].\n```\n\n## Photorealism positives β€” what consistently works\n\n> ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern β€” do not carry it across.**\n> Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.\n> The leak happens in one specific way β€” you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` β†’ `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `slates-prompting-seedance` (Part 3, \"Don't cross-pollinate image-model syntax\").\n\n**Named lenses + apertures** beat generic \"shallow depth of field\":\n- `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`\n- `Panavision anamorphic` for horizontal flares + cinematic width\n- `400mm telephoto` for compression + isolation\n- `24mm` for environmental interiors\n\n**Named cameras / sensors:**\n- `ARRI Alexa 65`, `Hasselblad X2D`, `Canon EOS R5`, `Sony A7III`, `Fujifilm X-T5`\n- \"Specific gear\" beats \"DSLR\"\n\n**Named film stocks** (one per prompt β€” never mix):\n- `Kodak Portra 400` β€” natural skin, warm\n- `Fuji Velvia 50` β€” saturated, landscape\n- `Ilford HP5 Plus` β€” black and white, gritty grain\n- `CineStill 800T` β€” tungsten night, halation\n\n**Physics-based lighting** (direction + quality):\n- `Single key light at 45 degrees from upper left`\n- `Late afternoon sun at 15 degrees above horizon`\n- `Color temperature 4500K` beats `slightly warm`\n- `Practicals only β€” no fill` for Deakins-style realism\n\n**Imperfection vocabulary** (forces away from AI-clean):\n- `visible pores`, `natural skin grain`, `peach fuzz`, `slight hyperpigmentation`\n- `unretouched raw photography`, `ISO noise`, `sweat beading`\n- `crisp catchlights in the eyes`, `skin micro-detail`\n\n**Director references** (use when locking style):\n| Director | Tone | Visual signature |\n|---|---|---|\n| Denis Villeneuve | Cold, vast, existential | Desaturated, overwhelming scale |\n| Roger Deakins | Precise motivated light | Single source, deep shadows, practicals |\n| Emmanuel Lubezki | Natural, spiritual | Available light, golden hour |\n| Bradford Young | Warm darkness | Underexposed, rich shadows, skin tones |\n\n**Genre cues that move the model:**\n- `unstaged documentary photography style`\n- `fashion magazine editorial, shot on medium-format analog film, pronounced grain`\n- `Film still from [Director] [genre]`\n\n## The anti-list β€” phrases that DEGRADE realism\n\nThese are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is extracted\n by src/prompts/banned-tokens.ts, inlined verbatim into the slates_generate_image\n op description (always in context on both surfaces), and matched against every\n submitted prompt. Editing this list changes what the agent is told AND what it\n is warned about β€” keep every entry backticked, and keep prose outside the\n backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- `8k`, `4k` (as a quality token)\n- `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone\n- `masterpiece`, `best quality`, `highly detailed`, `ultra-detailed`\n- `trending on ArtStation`, `award-winning`\n- `perfect skin`, `flawless`, `airbrushed`, `smooth skin`\n- `cinematic` standing alone β€” always specify *which cinema* (director, lens, era, stock)\n- `not anime, not cartoon, not 3D` β€” negation tag soup, replace with a positive style cue\n<!-- @banned:end -->\n\n## Negative prompting β€” there is no field\n\nNano Banana 2 has **no `negativePrompt` parameter**. Three patterns to suppress unwanted content:\n\n1. **Positive reframing (preferred):** \"empty street\" not \"no cars\". \"Unstaged documentary photography\" not \"not anime.\"\n2. **Inline `without` / `free of`:** \"without any people, vehicles, or man-made structures\", \"free of text overlays, logos, or watermarks.\"\n3. **Constraint clauses for anatomy/quality:** \"accurate anatomy with five fingers per hand, symmetrical features, natural proportions\"; \"sharp, well-exposed, free of blur or JPEG artifacts.\"\n\nDefault to #1. Reach for #2 only when positive framing can't suppress the unwanted element.\n\n## Reference images\n\n- **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade β€” you can't use 14 object slots even if no characters are referenced.\n- **Name each reference inline β€” Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as \"image N\" β€” e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **\"assign a distinct name to each character/object\"**. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render the scene's expression\") β€” that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n\n### Reference rules (the verified ones)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it β€” lighting, medium, texture, symmetry, competing identities β€” is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** β€” because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** β€” because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** β€” the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** β€” mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* β€” the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere β€” fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs β€” each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** β€” identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` β€” you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") β€” that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation β€” the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light β€” never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference β€” the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text β†’ bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media β€” describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime β†’ real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Nano Banana 2 specifically\n\n- **NB2's own consistency lever is \"assign a distinct name to each character/object.\"** That is Google's phrasing for rule 3 β€” cite each canonical identity inline by name.\n- **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model β€” when a downstream video shot needs legible text, render it here and animate from this frame.\n- **Character consistency is officially \"not 100% perfect\"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.\n- **Injection is stochastic β€” budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.\n\n## Common failure modes + fixes\n\n**Hands:** Append `accurate anatomy with five fingers per hand, symmetrical features, natural proportions, relaxed open palm`. Avoid heavy jewelry, props intersecting fingers, motion blur in references.\n\n**Text in images:** Quote-wrap target text. Specify font (`Century Gothic, 12pt`). Long phrases work; small text degrades. Two-step works best β€” generate text concepts conversationally first, then ask for the image.\n\n**Left/right confusion:** Default is **viewer's perspective**, not subject's. Append `left and right are from the character's perspective, NOT the camera's` when scene-blocking matters.\n\n**Surreal / absurd prompts trip uncanny valley:** The model drags toward realism. If you want surrealism, lean hard into stylization keywords (`painted`, `illustrated`, `stop-motion`).\n\n**Soft faces / dead eyes:** Add `crisp catchlights in the eyes`, `skin micro-detail`, `peach fuzz visible`. Don't stack quality enhancers β€” single clean prompt beats multiple re-interpretations.\n\n**Post-cutoff content (anything after Jan 2025):** Use reference images. The model has no knowledge of recent franchises, products, events.\n\n## Resolution tactics\n\n- Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change β€” check current numbers<!-- slates-only --> by calling `slates_estimate_generation_cost`<!-- /slates-only -->. Pick the cheapest resolution that serves the use case.\n- **At 2K and above, the model allocates more tokens to surface detail** β€” explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.\n- 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.\n- 2K generations vary 20-60s+. Don't time-budget tightly.\n\n## Boring vs cinema β€” examples\n\n❌ **Boring:** \"Wide shot of a man on a dock looking at the forest.\"\n\nβœ… **Cinema:** \"Direct overhead drone shot on weathered dock surface. Single figure standing center frame, climbing up from frame bottom. Boot prints leading away from him toward shore. Pale winter light. Anamorphic lens flare from low sun. Desaturated blue and slate grey palette. Kodak Portra 400 grain. The path already walked by someone else. Map of threat.\"\n\n❌ **Boring:** \"Close up of a woman looking scared.\"\n\nβœ… **Cinema:** \"Extreme close on subject's mouth and nose, 135mm f/2.8, shallow depth of field. Breath pluming out, catching cold light from upper-left key. Lips slightly parted, peach fuzz visible. The breath holds. CineStill 800T halation around catchlights. Waiting.\"\n\n## The 3-strike rule\n\nIf three iterations on the same prompt haven't produced what the user wants, stop. Hand back to the user with what you tried and what isn't working. The slot machine doesn't converge β€” the prompt structure is wrong, not the seed.\n\n## Family variants β€” Lite and Pro\n\nEverything in this skill applies to the whole Nano Banana family; two variants trade speed/ceiling around NB2 full:\n\n- **nano-banana-2-lite** β€” ~half the price, ~2.7Γ— faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.\n- **nano-banana-pro** β€” the hero-frame/typography ceiling (~2Γ— NB2, 4K native). NB2 β‰ˆ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs β€” it takes a full subject library in one call.\n\n<!-- slates-only -->\nRouting between them (and vs GPT Image 2.5 / FLUX / Seedream): `slates-model-selection`.\n<!-- /slates-only -->\n",
26
26
  "slates-prompting-omni-flash": "---\nname: slates-prompting-omni-flash\ndescription: How to prompt Gemini Omni Flash (Google, via fal). Read before calling slates_generate_video with omni-flash or slates_edit_video with omni-flash-edit. Cheap 720p tier with native synced audio included β€” 3-10s, 16:9/9:16 only; t2v, single-start-frame i2v, or reference-to-video with up to 7 reference images. The edit variant is the EDIT-FIDELITY WINNER for footage-synced VFX (receipt 2026-07-09) β€” but ONLY with short prompts: one change + \"Keep everything else the same.\" Long descriptive prompts destroy fidelity.\n---\n\n# Gemini Omni Flash β€” prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Gemini Omni Flash.** Two different jobs with OPPOSITE prompt rules, and getting them the wrong way round is the whole failure mode.\n\n**The five levers**\n1. **Editing: short prompt, ONE change, nothing else.** Google's own doc says so and a 2026-07-09 receipt confirms it β€” a long \"keep every frame identical\" preamble produced WORSE drift than two sentences.\n2. **Editing: always end with `Keep everything else the same.`** β€” the one documented preservation lever.\n\n3. **Editing: describe the EFFECT, never a real object as a metaphor.** \"Candle-like flame\" rendered a literal candle in the subject's hand.\n4. **Editing: no conditional timing cues.** \"…when he calls it, as he walks…\" hard-fails with `invalid_request`. Collapse to one continuous action; the model syncs to the footage's own motion.\n5. **Generation: the opposite β€” describe fully.** Subject, action, setting, `camera tracking alongside`, `overcast flat light`, tone. Audio is prompt-driven with no parameters: dialogue in quotes, sound in plain language β€” `rain patters on the tin roof`, `spray from tyres`, `a horn somewhere behind`.\n\n**Examples**\n- Edit: `Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.`\n- Generate: `A courier in a yellow shell jacket weaves between stalled cars on a wet arterial road, camera tracking alongside at shoulder height. Overcast flat light, spray from tyres. Rain patters on car roofs, a horn somewhere behind.`\n\n**Hard constraint:** it is a CHEAP DRAFT seat for generation and the EDIT-fidelity winner for footage-synced VFX β€” never a hero generation shot. Expect a possible jitter or doubled speech beat in the last half second of an edit: trim the tail rather than burning a re-roll.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use in an EDIT prompt** (each one has a receipt above):\n- a long preservation preamble β€” it produces WORSE drift than `Keep everything else the same.`\n- a real object as a metaphor: `candle-like`, `flame-like`, `laser-like`\n- a conditional timing cue: `when he`, `as she`, `once they` β€” these hard-fail, they do not merely drift\n- harm-to-person framing: `ignite`, `catch fire`, `on fire` applied to a person trips the safety filter\n<!-- @banned:end -->\n\nGoogle's fast video generation + editing model (\"Nano Banana Pro for video\" in creator slang β€” a nickname; it is NOT the NB Pro image model). Carried on fal (`google/gemini-omni-flash*`). 720p only, 24fps, 3–10 second clips, 16:9 or 9:16. **Audio is native and included** β€” dialogue, SFX, and ambient generate WITH the video at no extra cost.\n\n## Where it routes\n\n- **Video editing (`omni-flash-edit`) β€” its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) β€” Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.\n- **Cheap drafts and iteration volume** β€” lowest-cost audio-native video seat (~6.4 cr/s at 720p).\n- **NOT hero GENERATION shots** β€” Kling 3.0 stays the general gen default, Seedance 2.0 the premium tier; Omni Flash's *generation* quality seat is still unproven.\n\n## Editing (`slates_edit_video`, model `omni-flash-edit`) β€” THE RULES (receipts, not theory)\n\n1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *\"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes.\"* Live receipt 2026-07-09: a long \"keep every frame/word/movement identical…\" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *\"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.\"*\n2. **Always end with \"Keep everything else the same.\"** β€” the one documented preservation lever.\n3. **Never name a real-world object as a metaphor.** \"Candle-like flame\" rendered a literal candle in his hand. Describe the effect itself (\"small magical flames on his fingertips\").\n3b. **No conditional timing cues β€” they HARD-FAIL, not drift.** Receipt 2026-07-09: \"a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…\" β†’ deterministic `invalid_request` (2Γ—, \"could not generate with the given inputs\"); collapsing to one continuous action β€” \"A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke.\" β€” succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video.\n4. **Safety filter (Google's, strict about harm-to-person):** \"fingertips ignite / catch fire\" β†’ `content_policy_violation`. Frame effects as magical/harmless VFX: \"small magical flames appear on his fingertips\" passed. See slates-content-policy Β§Gemini for the substitution patterns.\n5. **Expect a possible tail artifact** β€” jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.\n6. **Prompt + source clip ONLY.** No element/style reference images β€” identity swaps that need refs go to `kling-v3.0-omni-edit`.\n7. Source clip 3–10s (trim longer clips first). Output length follows the source; billing per output second, rounded up. Voice editing unsupported β€” never ask it to change dialogue.\n8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage β€” this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).\n9. Chain edits one change at a time β€” each edit saves as a new asset linked to its parent.\n\n## Generation (`slates_generate_video`, model `omni-flash`)\n\n- **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params β€” they merge into one reference list). No last frame, no video/audio references β€” the op rejects them.\n- Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.\n- **Name references inline** the standard Slates way (\"Marcus (image 1) walks…\"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) β€” useful when a specific image must bind to a specific role.\n- **Audio is prompt-driven** β€” no audio parameters. Dialogue in quotes; direct sound in plain language (\"rain patters on the tin roof\"). Negative direction as plain instructions (\"Do not show text\").\n- Duration is an explicit 3–10s integer param; cost scales linearly per second.\n\n## Input conditioning (Slates handles this β€” know it exists)\n\nPhone footage stores rotation as a metadata flag; models ignore it and edit the raw sideways pixels. Clips must be rotation-normalized (and oversized sources downscaled) before upload β€” receipt 2026-07-09: a portrait Pixel clip came back sideways until conditioned. If an edit output comes back rotated, the source wasn't normalized.\n\n## Content notes\n\n- Google applies its own safety filters to input images/clips and output. Uploads containing recognizable real people are restricted by Google's policy β€” though own-footage editing of the uploader passed on our route 2026-07-09. See slates-content-policy.\n- Output carries an invisible SynthID watermark (Google-side, programmatic detection only).\n",
27
27
  "slates-prompting-seed-audio": "---\nname: slates-prompting-seed-audio\ndescription: How to prompt Seed Audio 1.0 (ByteDance, via fal). Read before calling slates_generate_audio with model seed-audio. The one-pass audio SCENE model β€” dialogue, SFX and ambience together from ONE plain sentence. CRITICAL - it has NO duration parameter, so length must be named IN THE PROMPT TEXT and Slates bills the duration you request. Covers the one-sentence doctrine, the crowd-size rule, why Kling \"SFX:\" syntax hurts here, and the audio-refs-XOR-image input rule.\n---\n\n# Seed Audio 1.0 β€” prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Seed Audio 1.0.** One plain sentence describing a SCENE, and it returns dialogue, effects and ambience together in one pass. For ONE named voice saying ONE line, `inworld-tts-2` is the seat instead; this one renders the whole room.\n\n**The five levers**\n1. **Write one sentence in plain language.** Describe the room and what is happening in it; the model casts and performs it.\n2. **Name the crowd size, the room size and the distance.** This is the highest-leverage single edit on any bed β€” unqualified nouns default BIG. \"Tiny applause of two or three people at an open mic\", not \"applause\".\n3. **Put the length in the prompt** and make it deliberate.\n4. **Dialogue is performed inside the scene** β€” write the line as spoken in the room, then re-roll until a take is right and lip-sync against it.\n5. **Describe sounds directly** β€” \"one coffee machine hissing, cutlery somewhere behind the counter\" β€” with the object, the action and where it is.\n\n**Examples**\n- `a diner at 2am, one coffee machine hissing, cutlery somewhere behind the counter, one man says quietly \"you're late again\". 15 seconds`\n- `nature soundscape, a wide open field, cicadas near, birds mid-distance, one loon far off across water. 20 seconds`\n\n**Hard constraint:** it has NO duration parameter β€” the length you request is written into the prompt AND is what you are billed, whatever comes back, so choose it deliberately. Kling's `SFX:` / `Ambient noise:` / `Background music:` labels have no parser here and measurably worsen the output. Shot language, camera moves and lighting are video-prompt words the model must ignore. It is not a music model and it cannot produce pixels.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** β€” Kling video syntax has no parser here and measurably worsens the output:\n- `SFX`, `Ambient noise`, `Background music` as labels β€” describe the sounds directly\n- shot language: `wide shot`, `slow push in`, `warm tungsten` β€” camera and lighting words are video-prompt words the model has to ignore\n<!-- @banned:end -->\n\nByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence. 1–120 seconds. It is the default audio model in Slates and the workhorse for continuity beds.\n\n## Where it routes\n\n- **Scene audio, room tone, ambience beds, crowd/nature soundscapes** β€” anything where several sounds share a space. One generation, not three layered ones.\n- **Dialogue and scratch VO inside a scene.** The line is performed in the room, by a voice the scene casts. When WHO is speaking matters β€” a character's own voice, a clean narrator β€” that is `inworld-tts-2` instead. Lock the read by re-rolling until a take is right, then lip-sync against it with `slates_generate_lip_sync`.\n- **NOT** a single effect that must land on a known frame β€” that is `eleven-sfx`, which takes an exact duration.\n- **NOT** music. Slates has no music model; import a track and drop it on an audio track.\n- **AUDIO-ONLY.** It cannot produce images or video.\n\n## THE RULES\n\n### 1. 🚨 There is no duration parameter β€” the words set the length\n\nThis is the single most important fact about this model. Output length is driven by the prompt text (\"… 15 seconds\"), capped at 120s.\n\n**Slates handles this for you:** the `durationSeconds` param appends the duration to the prompt and **bills that number of seconds**. So:\n\n- Set `durationSeconds` to what you actually want.\n- **Do not also write a different length into your sentence.** Two numbers fight, and you pay for the one you selected, not the one you got.\n- If the returned clip is shorter than requested you still paid for the request β€” that is the deal that keeps the displayed price equal to the charge. Ask for what you need.\n\n<!-- slates-only -->\nThe server re-derives the billed key from `durationSeconds` (a client cannot under-bill), probes the returned `audio.duration` after completion, and logs `SEED AUDIO BILLING DRIFT` if the model overshot. No auto-charge, no refund β€” the request is the contract.\n<!-- /slates-only -->\n\n### 2. One plain sentence. No production jargon.\n\nField-proven (Higgsfield sprint, 2026-07-27/28). Working prompts look like this:\n\n```\ntiny applause of 2 or 3 people at an open mic. 15 seconds\nnature soundscape, wide open field cicadas and birds and a loon.\na diner at 2am, one coffee machine hissing, cutlery somewhere behind the counter\n```\n\nNot this:\n\n```\nβœ— AMBIENCE: interior diner, night. SFX: espresso machine (hiss, 2s), cutlery.\nβœ— Wide shot of a diner. Slow push in. Warm tungsten. Ambient noise: ...\n```\n\nShot language, camera moves and lighting belong to video prompts. Here they are just words the model has to ignore.\n\n### 3. Never bring Kling's audio syntax to this model\n\n`SFX:` and `Ambient noise:` prefixes and `Background music:` labels are **Kling 3.0 video** syntax. Seed Audio has no parser for them β€” it reads them as text in the scene and the output gets measurably worse. Describe the sounds directly instead.\n\n### 4. Name the crowd size, the room size, the distance\n\nThe highest-leverage single edit on any bed. Unqualified nouns default big:\n\n| Vague | What it returns | Fixed |\n|---|---|---|\n| `applause` | a full auditorium | `tiny applause of 2 or 3 people` |\n| `traffic` | a highway | `one car passing on a wet residential street` |\n| `crowd` | a stadium | `four people talking at the next table` |\n\nDistance words (`far off`, `muffled through a wall`, `right next to the mic`) work the same way and are how you build depth in one sentence.\n\n### 5. Beds must outlast the cut\n\nAsk for a few seconds more than the clip needs so the edit has handles to fade through. A bed that ends exactly on the cut always sounds clipped. This is a product requirement, not a preference β€” it is why the duration control exists at all.\n\n### 6. Dialogue goes in quotes, inside the same sentence as the room\n\n```\na tired bartender says, \"we closed twenty minutes ago\", glasses clinking behind him\n```\n\nPick a preset voice when a specific speaker matters. Leave `voice` unset and the scene casts itself β€” which is usually right for crowd and background dialogue.\n\nPreset voices (20): `vivi_mixed_en_zh_ja_es_id`, `mindy_en_es_id_pt_zh`, `kian_en_zh`, `cedric_en_zh`, `sophie_en_zh`, `jean_en_zh`, `magnus_en_zh`, `mabel_en_zh`, `nadia_en_zh`, `opal_en_zh`, `pearl_en_zh`, `quentin_en_zh`, `corinne_mixed_en_zh`, `esther_mixed_en_zh`, `lyla_mixed_en_zh`, `tracy_es_zh`, `sandy_es_mixed_en_zh`, `felix_zh`, `celeste_zh`, `monkey_king_zh`.\n\nSet `multilingual: true` for non-English or mixed-language lines.\n\n### 7. Inputs: up to 3 audio clips **XOR** one image. Never both.\n\n- **Audio references** β€” up to 3 clips, each ≀30s and ≀10MB (wav/mp3/pcm/ogg_opus). Refer to them in the prompt as `@Audio1`, `@Audio2`, `@Audio3`: *\"match the room tone of @Audio1\"*.\n- **Image reference** β€” one image (jpeg/png/webp ≀10MB). The model scores what it sees.\n- Sending both is rejected by the API. Pick the one that carries the intent.\n\n### 8. The knobs, and when to touch them\n\n| Param | Range | Reach for it when |\n|---|---|---|\n| `speed` | 0.5–2.0 | Dialogue is racing or dragging against picture. |\n| `volume` | 0.5–2.0 | Rarely β€” normalize on the timeline instead. |\n| `pitch` | βˆ’12…+12 semitones | Ageing or shifting a voice. Small moves only; Β±3 is already a lot. |\n| `multilingual` | bool | Non-English or code-switched lines. |\n| `sampleRate` | 8k–48k | Leave at 24000 unless you are matching an existing stem. |\n| `outputFormat` | mp3 / wav / pcm / ogg_opus | wav when this is going into a mix; mp3 otherwise. |\n\n## Iterating\n\n- A bed that came back wrong is almost always a **scale** problem (crowd/room too big) or a **jargon** problem (the sentence reads like a spec). Fix those two before touching `speed`/`pitch`.\n- Three failed takes on the same sentence means the sentence is wrong, not the seed. Rewrite it the way you would say it out loud.\n- Generations are cheap enough at short durations that auditioning two phrasings beats agonizing over one.\n\n## Content notes\n\nProvider-side moderation applies to voices and to recognizable real people. See slates-content-policy.\n",
28
- "slates-prompting-seedance-2-5": "---\nname: slates-prompting-seedance-2-5\ndescription: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is a SECOND SEAT next to 2.0, not an upgrade β€” it buys 30-second takes, 30 image references, audio-only references and INTEGER-SECOND TIMESTAMPS, and it gives up native 4K and costs more than 2.0 at every resolution they share. Timestamps are the one grammar difference that matters: 2.0 ignores them and answers only to shot numbers, 2.5 acts on them. Otherwise it shares 2.0's grammar (read slates-prompting-seedance for subject binding, camera and constraint vocabulary); this file covers what is different, plus the two hazards unique to 2.5 β€” the prompt-intent task classifier and the cost trap that comes with 30-second takes.\n---\n\n# Seedance 2.5 β€” prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Seedance 2.5.** Shares 2.0's grammar exactly (subject binding, camera vocabulary, externalised emotion, inline constraints β€” read `slates-prompting-seedance` for those). Two things are different, and both matter.\n\n**The five levers**\n1. **Timestamps work here** β€” integer seconds, and the model acts on them: `[0-4] she reads the letter. [4-9] she folds it and looks up.` 2.0 ignores exactly this syntax.\n2. **Length is the reason to be here** β€” takes up to 30 seconds, where 2.0 stops at 15. Write the beats as `[0-6]`, `[6-12]`, `[12-18]`; do not hope for them.\n3. **Up to 30 image references**, and a multi-view image can serve as ONE subject reference (up to 5 subjects). 2.0 cannot do either.\n4. **Audio-only references are accepted** without an image or video alongside β€” the only Seedance seat that takes one.\n5. **Keep the 2.0 discipline**: one camera move per beat (`slow track right`, `handheld follow`), physical action instead of stated emotion, and quality asked for in the image-quality slot vocabulary β€” `rich details`, `natural colors`, `cinematic texture`, `soft lighting`.\n\n**Examples**\n- `[0-6] Wide shot, <Subject_1>@<Image_1> crosses an empty car park toward a idling van, slow track right. [6-12] Medium, she stops as the driver's window comes down. [12-18] Close-up, she looks off past the lens and does not answer. Rich details, natural colors. Keep it subtitle-free.`\n- `[0-10] A single continuous handheld follow behind a courier climbing a fire escape, rain. [10-20] She reaches the landing, turns, and the city opens behind her. Cinematic texture, soft lighting.`\n\n**Hard constraint:** it is the EXPENSIVE seat and it has NO 4K β€” 480p/720p/1080p only, and dearer than 2.0 at every resolution they share. It is a second seat, never an upgrade. Long takes multiply cost linearly: quote a 30-second take before you fire it.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** (2.5 reclassifies the task and fails a fresh generation on these):\n- `edit`, `extend`, `continue the video`, `same video but` β€” they make the provider read a fresh generation as an edit\n- `85mm`, `f/1.4`, `Portra 400` and any other lens, aperture, film-stock or camera-body token β€” image-model vocabulary, a Seedance anti-pattern on both seats\n<!-- @banned:end -->\n\n**Read `slates-prompting-seedance` first.** The prompt GRAMMAR is the same model family: the\n8-slot advanced formula, subject binding by `<Subject_N>@<Image_N>`, camera vocabulary,\nexternalised emotion, inline constraint words, the anti-twin fix. None of it is restated here.\nThis file is only what 2.5 changes β€” and the biggest change is that **2.5 acts on timestamps\nwhere 2.0 ignores them.**\n\n---\n\n## The one fact that decides whether you use it at all\n\n**Seedance 2.5 is the EXPENSIVE seat, not the cheap one β€” and it has no 4K.**\n\nIt runs at 480p, 720p or 1080p (1080p landed on all three routes on 2026-08-24), and at every\nresolution the two seats share it costs MORE than 2.0 β€” 720p $0.231/s against $0.15/s, **54% more**.\nSo 2.5 does not replace 2.0; it sits beside it, and you pay for what it buys:\n\n| | Seedance 2.0 | Seedance 2.5 |\n|---|---|---|\n| Resolution | 480p / 720p / 1080p / **native 4K** | 480p / 720p / 1080p β€” **no 4K** |\n| Price at 720p (faceless) | **$0.15/s** | $0.231/s |\n| Length | 4–15s | **4–30s in one take** |\n| Reference budget | 15 (9 image + 3 video + 3 audio) | **50 (30 image + 10 video + 10 audio)** |\n| Combined reference video/audio | ≀15s | **≀30s** |\n| Audio-only reference | βœ— (needs an image or video alongside) | **βœ“** |\n| **Timestamps in the prompt** | **βœ— β€” ignored; shot numbers only** | **βœ“ β€” integer seconds, acted on** |\n| Multi-view image as ONE subject reference | βœ— (not recommended) | **βœ“ (up to 5 subjects)** |\n| Video edit as its own task type | βœ— | **βœ“ (`seedance-2.5-edit`)** |\n| Default video model | **yes** | no |\n\n**Route to 2.5 when the shot needs LENGTH, MANY REFERENCES, or an audio-only reference.\nRoute to 2.0 when resolution matters at all** β€” which, for anything a client will see full-screen,\nis most of the time.\n\n---\n\n## 🚨 Hazard 1 β€” the prompt-intent task classifier\n\nThis is the one that costs money and time, and it has no equivalent on 2.0.\n\n**Seedance 2.5 sorts every request into one of five task types** β€” text-to-video,\nreference-to-video, first/last-frame, **video edit**, **video extend** β€” from the reference roles\nattached **plus the intent of your sentence**. Each type then has its own parameter constraints,\nand a violation comes back **asynchronously**: the task queues, credits are reserved, and only then\ndoes it fail.\n\nThe trigger words are ordinary English:\n\n| Reclassified as | Words that do it (ByteDance's own list) |\n|---|---|\n| **video edit** | `edit video` Β· `add` Β· `insert` Β· `remove` Β· `delete` Β· `modify` Β· `replace` Β· `change to` |\n| **video extend** | `extend forward` Β· `extend backward` Β· `continue` Β· `continue from` Β· `extend the story` |\n\nSo a perfectly legitimate reference-to-video prompt β€” *\"a wide shot of the workshop, **remove** the\ntripod from frame\"* β€” gets classified as an edit and fails on constraints it never set.\n\n**What to do:**\n\n1. **If you mean to edit an existing clip, say so with the MODEL, not the sentence.** Call\n `slates_edit_video` with `model: 'seedance-2.5-edit'`. That routes to a dedicated\n task-typed endpoint and the classifier never has to guess.\n2. **If you mean a fresh shot, describe the finished frame rather than an instruction to change\n one.** Not *\"remove the tripod\"* β†’ *\"the workshop bench, clear and uncluttered\"*. Not\n *\"add rain\"* β†’ *\"heavy rain falling through the streetlight\"*. This is better prompting anyway:\n the model renders what you describe, it does not take edits to an imagined draft.\n3. The trigger only fires when **references are attached**. A plain text-to-video prompt is safe\n however it is worded.\n\n**Slates will warn you, and it will never rewrite your prompt.** When a 2.5 reference generation's\nprompt contains one of these words, the composer shows a warning that NAMES the words and the agent\nroute returns the same string. Silently editing the user's sentence to dodge a provider classifier\nis forbidden β€” the words that reach the model are always the words the user can see.\n\n---\n\n## 🚨 Hazard 2 β€” resolution is not the price dial here. LENGTH is.\n\nEvery other model in Slates trains the habit that lower resolution means lower cost. 2.5 breaks it,\nbecause the thing that moves the bill is **length**, and 2.5's length ceiling is double 2.0's.\n\nWorked, at the shipped rates:\n\n| Generation | Credits |\n|---|---|\n| 2.5 Β· 480p Β· 5s Β· faceless | 26 |\n| 2.5 Β· 720p Β· 5s Β· faceless | 58 |\n| 2.5 Β· 1080p Β· 5s Β· faceless | 103 |\n| 2.5 Β· 720p Β· 30s Β· faceless | 347 |\n| 2.5 Β· 720p Β· 30s Β· AI-face route | **489** |\n| 2.5 Β· 720p Β· 30s Β· consented real-face route | **710** |\n| 2.5 Β· 1080p Β· 30s Β· faceless | **614** |\n| 2.5 Β· 1080p Β· 30s Β· consented real-face route | **1,749** |\n| *(for scale)* 2.0 Β· 1080p Β· 15s Β· AI-face route | 411 |\n\n**A 30-second 720p clip can cost more than a 15-second 1080p one** β€” and a base licence starts\nwith 1,000 credits. Someone who reads \"720p\" as \"cheap\" and asks for a 30-second take on the\nreal-face route has spent 71% of their welcome grant on one clip; **on the real-face route a single\n30-second 1080p take is more than the whole grant.**\n\n**Discipline:**\n\n- **Always quote with `slates_estimate_generation_cost` before a take over ~10 seconds,** and say\n the number out loud before generating.\n- **Find the shot at short LENGTH, not at low resolution.** Length is what moves the price, so cut\n seconds while you are still exploring β€” 4–8s β€” and stay at the resolution you actually want.\n **A 480p pass does not de-risk a 720p or 1080p render.** Generation is stochastic: the higher-\n resolution run is a different take, not the same shot rendered better. So a 480p draft that looks\n right buys you no guarantee, and one that looks wrong may have been fine at 720p β€” you paid 26\n credits to learn nothing, when 58 would have bought a real candidate.\n- **Length is a creative decision, not a default.** 30 seconds is available; it is rarely the right\n answer for a single shot. Multi-shot storyboards inside one 30s generation are what the length is\n actually for.\n- Read `slates-cost-discipline` β€” all of it applies, more sharply here.\n\n---\n\n## Timestamps β€” the one grammar change\n\n<!-- @inject:seedance-25-timestamps -->\n**2.0 does not respond to timestamps and answers only to shot numbers. 2.5 responds to\ninteger-second timestamps.** That is ByteDance's own first line under \"Differences from Seedance\n2.0\", and it is why a 30-second take is usable at all: the length is only worth buying if you can\nsay *when* things happen inside it.\n\nBoth formats are valid on 2.5, and you can mix them β€” `Shot N` blocks for a storyboard whose\npacing you are happy to leave to the model, timestamps when a beat has to land at a moment.\n\n**Three ways to control time, all first-party:**\n\n| Form | Write it like |\n|---|---|\n| **Interval** | `0-3 seconds… 3-7 seconds… 7-15 seconds` or `[1s-4s]… [4s-8s]… [8s-12s]` |\n| **Time point** | *\"Quick left sideways transition at the 5-second mark.\"* |\n| **Relative** | *\"After 3 seconds, everyone around him shakes their head.\"* Β· *\"The frame freezes for 1 second after he presses the shutter.\"* |\n\n**The rules that come with them:**\n\n- **One second is the smallest unit.** Integers only β€” no `2.5s`, no frames.\n- **No gaps in the timeline.** `0-3s… 5-6s…` leaves 3-5s unspecified and the model fills it however\n it likes. Intervals must abut: `0-3s`, `3-7s`, `7-15s`.\n- **Budget the plot to the seconds.** Too little content in a range and the model improvises to\n fill it; too much and you get extra cuts or dropped beats. This is the actual craft of a 30s take.\n- **Never time-code a high-frequency action.** *\"Shake your head three times per second\"* is\n explicitly called out as a misuse β€” timestamps schedule beats, they don't choreograph frames.\n- **Transitions want both halves:** the moment AND the method β€” *\"At the 5-second mark, the camera\n transitions leftward with a left wipe into a natural dissolve.\"*\n- **Timestamps work on an EDIT too**, and that is where they earn the most: they scope a change in\n time as well as in content β€” *\"Change the man's action from drinking coffee to mopping the floor\n from 4-6 seconds in Video 1, and leave the rest of the content unchanged.\"* Without a range, a\n whole-clip instruction is applied to the whole clip.\n\nDo **not** carry this back to 2.0, and do not carry Veo's `[00:00-00:02]` bracket syntax into\neither β€” 2.0 ignores time entirely, and the cross-model syntax swap is its own known failure.\n<!-- @end:seedance-25-timestamps -->\n\n---\n\n## What the extra reference budget is actually for\n\n30 image references (up from 9) does **not** mean \"attach 30 images\". Every rule in\n`slates-prompting-seedance` about references still holds β€” 2–4 strong references beat both\nextremes, and one reference per role.\n\n**Where 2.5 moves the ceiling, per ByteDance's own input recommendations:**\n\n| | Stable | Works, but expect re-rolls |\n|---|---|---|\n| Subjects bound by IMAGE reference | 1–8 | 9–12 |\n| Subjects bound by VIDEO or AUDIO reference | 1–5 | 6–10 |\n| Reference clip length, per subject | 5–10s | longer drops stability |\n\n**Multi-view images of one subject are supported on 2.5** β€” a turnaround sheet can be a single\nreference image, where 2.0 wanted one authoritative rendering per subject. Past **5 subjects**,\ngo back to single-view images, one per view, rather than one image carrying several viewpoints.\n\nThe larger budget earns its keep in exactly two places:\n\n- **A long multi-shot take** where different shots need different subjects and locations bound β€”\n the budget is spread across the storyboard, not stacked on one frame.\n- **Video and audio references alongside images**, which is where 2.5's 10 + 10 matters far more\n than the image count.\n\n### Audio-only references β€” the genuinely new input\n\n2.0 required an image or video alongside any audio reference. **2.5 accepts audio on its own.**\nThat makes one recipe possible that was not before: drive a scene's timing, voice or ambience from\na recording with no visual anchor at all β€” a voice line, a music bed, a room tone β€” and let the\nmodel build the picture to it. Cite it the same way as any other reference\n(`Reference the timbre in <Audio_N> to generate…`), and remember that audio references carry **no\nbilling dimension** on any Seedance route: audio is included.\n\n### Video references\n\nUp to 10 clips, ≀30s combined (2.0: 3 clips, ≀15s). A reference VIDEO switches the cost key to\n`seedance-2.5*-vref-{res}-{T}s`, where **T = Ξ£ input seconds + output seconds** β€” the sum is across\n**every** clip attached, not just the longest. Three 6-second references on a 12-second output bills\n30 seconds, not 12 and not 18. Quote before confirming.\n\n**The cap is a refusal, not a trim.** Attach an eleventh clip, or push past 30 combined seconds, and\nthe composition is rejected before anything uploads. That asymmetry is deliberate: reference images\nwarn-and-trim because dropping one doesn't change the price, and a dropped reference VIDEO would be\none you were quoted for and the model never saw.\n\n### Mixing all three in one call\n\n50 files total (30 image + 10 video + 10 audio) is a shared budget. Everything is cited positionally\nby type β€” `image 1`, `video 2`, `audio 1` β€” in attachment order, so reordering the attachments\nrenumbers the citations. Write the prompt against those numbers:\n\n```\nMarcus (image 1) performs the motion from video 1 in the workshop from image 2,\nusing the voice timbre from audio 1. Preserve his identity, appearance and outfit.\n```\n\n🚨 **SAY WHAT AN AUDIO REFERENCE IS FOR.** It can mean music, dialogue, voice, tone or timbre β€” five roles on one attachment β€” so an unroled clip falls back to **dialogue**: the model re-transcribes it and speaks ITS words. A real take came back as *\"a map called Slates\"* for *\"an app called Slates\"*. Name it as the voice timbre and the clip carries the voice while the prompt carries the words. ByteDance's own sentence: *\"Image 1 depicts the protagonist John and uses the voice timbre from Audio 1.\"* Bind each speaker in a sentence, never by attachment order β€” position carries nothing.\n\nFrames and reference media stay mutually exclusive, in every combination β€” the reference endpoint\nhas no first/last-frame parameters at all, so this is a shape mismatch rather than a preference.\n\n---\n\n## Seedance 2.5 Edit (`slates_edit_video`, `model: 'seedance-2.5-edit'`)\n\nIts own picker row and its own op call, deliberately: the task type is **the model you chose**,\nnever something inferred from your sentence.\n\n**Why route here at all:** it is the **only edit engine in Slates that accepts a clip longer than 15\nseconds** (4–30s, versus Kling O3 Edit's 3–15s and Omni Flash Edit's 3–10s). For a clip inside the\nothers' range, choose on fidelity instead β€” Omni Flash Edit won the prompt-only head-to-head, and\nKling O3 Edit is the one that takes element and style reference images.\n\n**How it behaves:**\n\n- **Output length follows the SOURCE clip**, and the bill is the ceiled source length. The provider\n requires an automatic duration on this task type, so there is no length knob β€” the clip you attach\n is the quote. The returned clip can differ from the source by up to ~0.3s, which only compresses\n transition frames; a clip that 2.5 itself generated comes back at exactly its input length.\n- **Source clips under 20 seconds edit more reliably.** 4–30s is what the task type accepts;\n ByteDance's own recommendation is to stay inside 20 for quality. A 28-second source is legal and\n will need more attempts.\n- **The aspect ratio follows the source clip too.** No ratio control; the frame is the clip's frame.\n- **480p, 720p or 1080p output**, native audio β€” and an edit bills the video-reference tier Γ—2,\n so 1080p on this row is the most expensive second in the app. Quote it.\n- **Prompt and source clip only** on this op. The MODEL takes reference images on an edit\n (ByteDance recommends 1–5 β€” *\"replace the man in dark clothing in @Video 1 with @Image 2\"*);\n **Slates has not wired that path**, so today an edit that must lock an identity from a photo\n goes to Kling O3 Edit. Constraint of our build, not of the model β€” worth revisiting.\n- **An edit bills roughly DOUBLE a plain 2.5 generation of the same length**, because every provider\n charges an edit on input + output seconds. Read the confirm gate's number; do not reason from the\n generation rate.\n- **Set `seedanceFace: true` when a character's face is visible in the clip.** The faceless provider\n blocks faces outright β€” this is not a price optimisation, it is whether the job runs at all.\n- **There is no consented-real-face route for editing.** Real-person footage that the AI-face route\n rejects has to go to Kling O3 Edit.\n\n**Prompting an edit** β€” the same discipline as every other edit engine: **describe only what\nchanges.** The source already carries its composition, motion, timing and performance; re-describing\nthem fights the model. Use Seedance's own edit grammar from `slates-prompting-seedance`\n(*\"Strictly edit `<Video_1>`, and modify `<Original_Characteristic>` to `<New_Characteristic>`\"*) and\n**never** write *\"reference video 1\"* in an edit β€” the official guide is explicit that this phrasing\ngets the request reclassified as a reference task, which is the same landmine as Hazard 1.\n\nTwo things sharpen an edit prompt, both first-party:\n\n- **Say it as A β†’ B, not as an outcome.** *\"Change the man's action from drinking coffee to mopping\n the floor\"* beats *\"the man mops the floor\"* β€” naming what it currently is tells the model what\n to overwrite.\n- **Timestamp a partial edit** β€” the edit task type reads the same integer-second timestamps the\n generation path does. Rules and forms are in Β§ Timestamps above; this is the single most useful\n thing they buy.\n\n**Audio is editable too, and it is the least obvious use of this row.** The same op rewrites what\nis heard while the picture stays put: change a spoken line, change the accent, translate the\ndialogue and re-fit the lip movement, strip or replace the BGM or a sound effect. *\"Only edit the\nman's dialogue in Video 1: change it to 'Don't come over here,' in an American accent\"* is an edit,\nnot a lip-sync job. Bill it like any other edit β€” on the source clip's length.\n\n---\n\n## Faces, and what does NOT change\n\nThe three-tier face routing is identical to 2.0 β€” faceless β†’ default route, an AI character's face β†’\n`seedanceFace: true` (the relaxed provider, a real cost premium), a real person's photo β†’ the\nconsent-gated premium route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`\nset **only** after the user explicitly confirms they hold the rights to the likeness. The full rules,\nincluding why the real-vs-AI call is the provider's and not yours, are in\n`slates-prompting-seedance`.\n\nAlso unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:\n\n- **One primary camera move per shot.**\n- **No lens / aperture / film-stock vocabulary.** That is image-model syntax and a Seedance\n anti-pattern.\n- **No `negativePrompt` field** β€” constraints go inline, and 2.5 acts on negative phrasing in\n exactly two dimensions: subtitles (*\"no subtitles\"*) and audio (*\"no BGM; environmental and\n action sounds only\"*, *\"no audio\"*). Everywhere else, describe what you want, not what you don't.\n- **Legible in-shot text still belongs in a baked start frame**, not in the video prompt.\n",
29
- "slates-prompting-seedance": "---\nname: slates-prompting-seedance\ndescription: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance 2.0 structures multi-beat prompts as a \"Shot 1 / Shot 2 / Shot 3\" storyboard against an 8-slot advanced formula β€” never per-second time stamps, which 2.0 does not respond to (Seedance 2.5 does; see slates-prompting-seedance-2-5). Its syntax differs from Kling, Veo and the image models; don't cross-pollinate (in particular, no lens / aperture / film-stock vocabulary).\n---\n\n# Seedance 2.0 β€” prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Seedance 2.0.** Not copywriting β€” an ENGINEERING instruction to a spatial layer and a temporal layer: who, in what scene, doing what, how the camera moves, and in what order. Multi-beat work is a `Shot 1 / Shot 2 / Shot 3` storyboard.\n\n**The five levers**\n1. **Bind every subject to its reference** β€” `<Subject_1>@<Image_1>` β€” and keep the descriptions identical across shots. Unbound subjects are where twins come from.\n2. **Shot sizes and camera MOVES, not lens data** β€” `medium close-up`, `slow push in`, `handheld follow`, `whip pan`. One primary move per shot.\n3. **Externalise emotion as physical action.** Not \"she is nervous\": `she turns the ring on her finger twice, then stops`.\n4. **Use the image-quality slot vocabulary** for quality β€” `HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`. That is the officially sanctioned way to ask.\n5. **Constraints go INLINE**, led by the official templates β€” `keep it subtitle-free`, `do not generate a watermark`, `avoid jitter and bent limbs`, `avoid temporal flicker`.\n\n**Examples**\n- `Shot 1: medium shot, <Subject_1>@<Image_1> steps out of the freight lift into a wet loading bay, slow push in. Shot 2: close-up, she turns the ring on her finger twice and stops, handheld. Rich details, cinematic texture, natural colors. Keep it subtitle-free.`\n- `Single continuous take. Wide shot of a fishing skiff crossing a grey swell, camera tracks from the starboard rail. Spray hits the lens once. Soft lighting, natural colors, film-grain texture. Avoid jitter and bent limbs.`\n\n**Hard constraint:** NO timestamps β€” 2.0 ignores them and answers only to shot numbers (2.5 acts on them). No lens, aperture, film stock or camera body: that is image-model vocabulary and a Seedance anti-pattern. There is no negativePrompt field.\n<!-- @card:end -->\n\nByteDance's video model β€” first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K β€” 4K video is Pro-only, default 1080p), 4–15s, first+last frame, and up to 9 reference images / 3 videos / 3 audio clips.\n\n> **How to read this file.**\n> **[official :NNNN]** β€” ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/seedance-2-modelark-docs.md`). Receipt-grade; treat as law.\n> **[community]** β€” third-party guides and our own field notes. Useful, but an `[official]` block always wins.\n> **[slates]** β€” how the Slates app composes or bills this; not ByteDance doctrine.\n>\n> The split is load-bearing. A community-sourced \"narrative timing beats\" doctrine shipped in this file for months teaching the **exact inverse** of ByteDance's published guidance. Never merge the two registers again.\n\n---\n\n# Part 1 β€” Official ByteDance doctrine\n\n## What Seedance actually is `[official :1450-1452]`\n\nSeedance 2.0 is a multimodal AI director. It reads text, images, video and audio **simultaneously** and internally decomposes them into two dimensions:\n\n- the **spatial layer** β€” what is in the frame\n- the **temporal layer** β€” how things change over time\n\nSo a good prompt is **not \"copywriting-style description\" but an \"engineering-style instruction\"**: who, in what scene, doing what action, how the camera moves, and in what chronological order events occur β€” delivered respectively to the spatial layer and the temporal layer.\n\n## The advanced formula β€” 8 slots `[official :1455]`\n\n```text\nprecise subject + action details + scene/environment + lighting & color tone\n+ camera movement + visual style + image quality + constraints\n```\n\n⚠️ There is **no official \"6-step formula.\"** `Subject + Action + Environment + Camera + Style + Constraints` is community branding with no ByteDance source, and it silently drops the **lighting & color tone** and **image quality** slots. Use the 8 slots above.\n\n## Task-type sentence patterns `[official :1389-1425]`\n\nSeedance classifies your request from the phrasing. Use the pattern that matches the task:\n\n| Task | Pattern |\n|---|---|\n| **Image reference** | ``Reference `<Subject_N>` in `<Image_N>` to generate…`` |\n| **Video reference** | ``Reference `<Action / Camera_movement / Style / Sound_effect>` in `<Video_N>` to generate…`` |\n| **Audio reference** | ``Reference the timbre in `<Audio_N>` to generate…`` |\n| **Video edit β€” modify** | ``Strictly edit `<Video_N>`, and modify `<Original_Characteristic>` in it to `<New_Characteristic>``` |\n| **Video edit β€” add** | ``<Element_Features>` + `<Timing>` + `<Location>`` |\n| **Video edit β€” delete** | Name what to delete; for anything that must stay, say so explicitly |\n| **Video extend** | ``Extend `<Video_N>` forward/backward to generate…`` |\n| **Combined** | ``Reference `[Dimension]` of `<Image/Video_N>`, strictly edit `<Video_X>`, `[Specific_Edits]``` |\n\n### ⚠️ Edit / extend phrasing landmine `[official :1431]`\n\n> *\"For edit / extend video tasks, directly use `<Video_N>` to refer to the video. **Do not use \"reference `<Video_N>`\"**, to avoid being incorrectly identified as a reference task.\"*\n\nThis is easy to trip: Slates has an edit lane<!-- slates-only --> (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)<!-- /slates-only -->. Writing *\"reference video 1 and change the jacket to red\"* gets classified as a **reference** task β€” the model generates a brand-new clip inspired by the source instead of editing it. Write *\"Strictly edit video 1, and modify the blue jacket to red.\"*\n\n## Shot structure β€” \"Shot 1 / Shot 2 / Shot 3\" `[official :1563-1598]`\n\n> *\"Use shot order, write a simple 'Shot 1 / Shot 2 / Shot 3' storyboard for each segment of the video, and then merge them into a complete prompt.\"*\n\n**❌ Never second-stamp.** No `0:00–0:03`, no \"At 4 seconds\", no per-segment durations.\n\n> *\"Do not impose strict limits on the duration of each segment; prioritize allowing the model to naturally generate the pacing based on the plot.\"* `[:1580]`\n>\n> *\"The model's support for precise timing (such as 0–3 seconds) is **unstable**, and forcibly limiting duration may lead to **abnormal generation results**.\"* `[:1586]`\n\nOrder shots by when events occur β€” primary first, secondary later. Let the plot set the pacing.\n\n**Per-shot internal order** `[official :1590-1598]` β€” organize each shot in exactly this sequence:\n\n1. **Camera movement or shot transition** β€” \"slowly push in from a wide shot\", \"fixed camera position\", \"cut to…\"\n2. **Subject actions and expressions** β€” the key actions and expression changes of the core character/object\n3. **Position or spatial change** β€” where the subject is, and how that relationship shifts\n4. **Audio** β€” sound effects, voices, background music for that shot\n\n**One primary camera move per shot** β€” see Camera below. `[official :1648]`\n\n## Subject binding β€” names + image indexes `[official :1488-1556]`\n\nEvery time a subject appears, it must be **explicitly referred to**. Two supported forms:\n\n- **Undefined subjects** β€” bind inline every mention: `<Subject_N>@<Image_N>`. Official example: **`Zhang San@Image 1`**. `[:1540]`\n- **Pre-declared subjects** β€” define once, then reuse the same label verbatim: *\"Define the tall man in **Video 1** as **police officer**, and define the other short man as **thief**\"*, then say \"police officer\" every time after. `[:1514]`\n\n**One subject spread across several assets** β€” bind them together: *\"Define `[…]` in **Image 1** and `[…]` in **Image 2** as `<Subject N>`.\"* `[:1514]`\n\n⚠️ **An Asset ID must never substitute for `<Image/Video_N>`.** `[:1546]` *\"the model cannot directly associate the Asset ID with the reference content.\"* Always cite by index.\n\nAlso official: keep descriptions concise, avoid redundancy, avoid semantic conflicts (contradictory traits for one subject), and prefer expressing spatial relationships through reference images rather than dense text. `[:1550-1556]`\n\n**`[slates]`** β€” the app composes this for you. `composeReferences()` cites each canonical character or environment reference inline as `Name (image N)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*\"Zhang San (corresponding to image 1)\"* `[:1976]`). You never hand-write role labels or index numbers.\n\n## Action description `[official :1602-1621]`\n\n- **Body-part specificity + quantified degree.** Name hands, legs, head, shoulders, back β€” and supplement **range, speed, and force**. *\"slowly raise a hand\", \"quickly turn the head\", \"push hard off the ground\", \"slightly lower the head.\"*\n- **Prioritize slow, gentle, continuous small movements.** Avoid high-burst, large-dynamic actions β€” sprinting, big jumps, violent rolls. *\"walk slowly\", \"gently raise a hand\", \"sit down naturally with the motion.\"* **This is the official basis for the folk rule that \"fast\" degrades quality** β€” it is not a banned token, it is a class of motion the model handles badly.\n- **Supplement transitions between actions.** Specify inertia and continuity between consecutive beats so movement reads coherent: *\"use the inertia of turning around to naturally raise a hand\", \"naturally transition from a pause into raising a hand.\"*\n\n## Externalize emotion `[official :1623-1636]`\n\nReplace abstract emotion words (\"very sad\", \"extremely angry\") with **specific physical detail**. This is the highest-leverage single habit in the official guide:\n\n| Abstract | Externalized as actions and details |\n|---|---|\n| **Sadness** | head lowering, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the corner of clothing, tears welling but not falling |\n| **Joy** | corners of the mouth rising uncontrollably, brows and eyes relaxing, steps becoming light, unconsciously humming a tune |\n| **Nervousness / anxiety** | frequently checking the watch, fingers constantly tapping the tabletop, rapid breathing, eyes darting away |\n| **Anger** | both fists clenched, jawline tense, chest heaving, eyes sharp, squeezing words out through gritted teeth |\n| **Relief** | letting out a long breath, tense shoulders completely relaxing, a faint smile appearing, looking up toward the distance |\n\n## Camera `[official :1643-1648]`\n\n> *\"The model has a **strong understanding of camera movement terms**, so you can **directly use standard camera movement terminology**, such as 'medium shot, close-up, wide shot, slow push-in, smooth lateral tracking, fixed shot.'\"*\n\nThis is an **open vocabulary, not a fixed list** β€” and it explicitly includes **shot size** (close-up / medium / wide / long shot), which is as much a camera instruction as the move itself.\n\n> ⚠️ *\"Try to specify only 1 type of camera movement in a single shot. Do not require push, pull, pan, and move at the same time, as this will increase image instability.\"* `[:1648]`\n\n## Image quality, style, and constraints `[official :1656-1679]`\n\nThese three slots \"define creative boundaries for the model, unify image quality and artistic tone, and avoid visual flaws and random deviations.\"\n\n**1. Image quality** β€” define clarity, texture detail, and lighting quality. Official vocabulary: `HD` Β· `rich details` Β· `cinematic texture` Β· `natural colors` Β· `soft lighting`.\n\n> ⚠️ This is a **real slot with real vocabulary** β€” do not confuse it with Stable-Diffusion-era quality incantations. `8K` / `masterpiece` / `trending on artstation` remain banned slop tokens (see Part 3); *\"cinematic texture, rich details, natural colors\"* is the officially sanctioned way to ask for the same thing.\n\n**2. Style** β€” the overall art style and visual tone: `cyberpunk cool blue-purple tone` Β· `retro film` Β· `fresh Japanese style`.\n\n**3. Constraint words** β€” *\"Constraint words are very important. They can effectively avoid visual flaws, deformities, breakdowns, and unreasonable elements.\"* Official templates, verbatim:\n\n- **No subtitles** β€” \"keep it subtitle-free\" / \"avoid generating any text or subtitles\"\n- **No logo** β€” \"do not generate a logo\"\n- **No watermark** β€” \"do not generate a watermark\"\n\nSeedance has **no `negativePrompt` field** β€” constraints go inline in this slot. See Part 3 for the wider inline-negative kit.\n\n## πŸ”΄ Duplicated characters β€” the twin problem `[official :1948-1994]`\n\n**Symptom:** in frames with **many characters**, where **three-view / multi-view character images** are supplied as references, two identical characters appear in the same generated frame.\n\n**Root causes** `[:1954-1959]`:\n1. Character subjects are not clearly defined in the prompt, so the model cannot distinguish roles.\n2. *\"When character **three-view / multi-view images** are used as reference assets, it is easy to confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"*\n\n**Official fixes, in their order** `[:1971-1994]` β€” ByteDance is explicit that these *reduce probability*, not eliminate it:\n\n1. **Bind each character to its image explicitly**, in a consistent format. Official example: *\"Zhang San (corresponding to image 1) throws the green passbook toward Li Si (corresponding to image 2), who is standing.\"*\n2. **Append the global constraint verbatim** at the end of the prompt `[:1982]`:\n > *\"Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect. Keep only a single corresponding character in the same frame, and do not reproduce repeated copies of characters.\"*\n3. **Optimize reference assets** `[:1988]` β€” *\"For character reference images, prioritize independent single-person photos. Three-view or multi-view assets are not recommended.\"*\n4. **Simplify the prompt** β€” do not paste a whole script; redundant copy confuses the model.\n\n**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on identity sheets. Practical rule for Slates:\n\n- **Multi-character Seedance shot** β†’ bind every character to its image, append the anti-twin constraint, and prefer single-person / dominant-portrait references over multi-view sheets.\n- **Single-character shot** β†’ the standard character-sheet flow is fine.\n\n**Too many reference people** `[official :2048-2052]` β€” past **4 reference people**, output stability drops (wrong headcount, duplicates). Official workaround: group the cast into images of ≀4 people each, generate those stills first, then drive the video from them.\n\n## Worked examples `[official :1689-1745]`\n\nThese are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere β€” 2.0 does not respond to them at all, which is version-scoped and reverses on 2.5.\n\n**Example 1 β€” dormitory emotional short drama (dialogue-focused).** Assets: `@Image 1` half-body photo of the female lead Β· `@Image 2` dormitory scene reference Β· `@Video 1` camera-movement reference Β· `@Audio 1` indoor ambience.\n\n> Use the girl in @Image 1 as the main character, use @Image 2 as the dormitory scene style reference, and refer to the camera movement in @Video 1.\n>\n> **Shot 1**: At dusk, **girl @Image 1** walks briskly to the **dormitory entrance @Image 2**. The camera follows steadily in a medium shot. Warm yellow sunlight spills into the hallway from the window. She pauses at the doorway, takes a deep breath, and looks slightly nervous.\n>\n> **Shot 2**: **Girl @Image 1** pushes the door open and enters the dormitory. The camera cuts to an indoor medium shot. Her roommates look up at her while organizing their books. One of them smiles and asks {How did the exam go? Did you pass?}. The camera slowly cuts between half-body close-ups of several people.\n>\n> **Shot 3**: **Girl @Image 1** first lowers her head with a dejected expression. The camera gives her a close-up. Then she raises her head, unable to hold back a smile, laughs out loud, and says {I was kidding}. Her roommates start chasing and play-fighting with her. The camera slowly pulls back and freezes on a wide shot of the dormitory filled with laughter.\n>\n> The entire video should have a high-definition cinematic documentary style, with warm tones and soft lighting. The character's face remains stable without deformation; movements are natural and smooth, with no stutter or flicker. The ambient sound blends naturally with @Audio 1.\n\n**Example 2 β€” ancient-style cliff confrontation (action/atmosphere-focused).** Assets: `@Image 1` female lead in red Β· `@Image 2` assassin in black Β· `@Image 3` cliff and bamboo forest Β· `@Video 1` martial-arts camera reference Β· `@Audio 1` drum beats.\n\n> Use the woman in red from @Image 1 as the female lead, use the woman in black from @Image 2 as the opponent, use the cliff and bamboo forest environment in @Image 3 as the scene reference, refer to the overall camera movement and action rhythm in @Video 1, and synchronize the background sound effects with @Audio 1.\n>\n> **Shot 1**: At dusk, the camera slowly pushes in from a side medium shot of **woman in red @Image 1**. She stands at the edge of the cliff and lifts a wine flask to drink. Her sleeves and robe hem sway gently in the mountain wind. The camera circles halfway around her, moving from the front to her back. In the distance, a figure in black is faintly visible in the bamboo forest.\n>\n> **Shot 2**: The camera zooms and fades into a long shot. From a drone perspective, it overlooks the entire cliff and bamboo forest. The two characters stand at opposite ends of the cliff. The mountain wind lifts their robe hems and dust, and the rhythm slightly accelerates with the drum beats.\n>\n> **Shot 3**: The camera cuts back to a ground-level close shot. The two slowly draw their swords and face off. **Woman in red @Image 1** shifts from a careless expression to a cold gaze. **Woman in black @Image 2** looks determined, and the sword tip trembles slightly. The camera steadily follows the two as they circle each other, finally freezing on a close-up of the instant before the two swords meet.\n>\n> The overall visual style should feel like a cinematic wuxia world in misty rain, with cool tones, low saturation, a film-grain texture, and rich light-and-shadow layers. The characters' faces and body proportions remain stable without deformation. Movements are continuous and natural, not stiff, with no clipping or stutter.\n\n## Other official notes\n\n- **On-screen text** `[official :1758]` β€” Seedance can render common text (ad slogans, subtitles, speech bubbles) and will auto-match style/colour from context, or take an explicit colour / style / timing / position. Prefer **common characters**; avoid rare glyphs and special symbols. (For *guaranteed* legible text, the start-frame route in Part 3 is still safer.)\n- **Extension degrades quality** `[official :2004-2024]` β€” using a generated video as the input for extension compounds degradation, with mottled colour blocks in face regions. Limit repeated continuations; prefer HD assets as input.\n- **Special effects that miss** `[official :2031-2044]` β€” when a described effect comes out wrong (a countdown that scrolls randomly), define it with a **reference video** instead of words: *\"the way the number '2999' appears should reference video 1.\"*\n\n---\n\n# Part 2 β€” Slates-specific `[slates]`\n\n## Reference media β€” caps and transport\n\nReference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.\n\n**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `\"first/last frame content cannot be mixed with reference media content.\"` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.\n\n### All three modalities go in ONE call\n\nThe caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.\n\nCite each by type and index, in the order they were attached β€” `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.\n\n```\nMarcus (image 1) performs the motion from video 1, in the workshop from image 2,\nusing the voice timbre from audio 1. Preserve his identity, appearance and outfit.\n```\n\n🚨 **SAY WHAT AN AUDIO REFERENCE IS FOR.** It can mean music, dialogue, voice, tone or timbre β€” five roles on one attachment β€” so an unroled clip falls back to **dialogue**: the model re-transcribes it and speaks ITS words. A real take came back as *\"a map called Slates\"* for *\"an app called Slates\"*. Name it as the voice timbre and the clip carries the voice while the prompt carries the words. ByteDance's own sentence: *\"Image 1 depicts the protagonist John and uses the voice timbre from Audio 1.\"* Bind each speaker in a sentence, never by attachment order β€” position carries nothing.\n<!-- slates-only -->\n**Attaching a clip is NOT the same as editing it.** \"Add as reference\" puts it in the composer alongside everything else and wipes nothing; \"Edit with AI\" makes the clip the canvas and clears the tray for a fresh instruction. Two different jobs, two different menu entries β€” never infer one from the other.\n\n**Over the cap is REFUSED, never trimmed.** A reference video is priced into the quote before it is sent, so a clip silently dropped after the quote would be a clip you paid for and the model never saw. Remove one and retry.\n<!-- /slates-only -->\n\n### Motion transfer & lip-sync recipes (reference video / audio)\n\nThese aren't separate Seedance features β€” they're prompting strategies over reference media.<!-- slates-only --> The Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:<!-- /slates-only -->\n\n- **Motion transfer:** subject image as a reference + the driving clip<!-- slates-only --> via `videoReferenceAssetId`<!-- /slates-only --> (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`\n- **Lip-sync / dialogue:** write the line in the prompt β€” `The person in video 1 says: \"…\"` β€” with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≀15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`\n- **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script β†’ their likeness AND voice deliver the new line.\n<!-- slates-only -->\n- **Billing:** a reference VIDEO switches the cost key to `seedance-2*-vref-{res}-{T}s` where T = clip seconds + output seconds β€” quote before confirming. Audio references are free (audio is included on every route).\n<!-- /slates-only -->\n\n<!-- slates-only -->\n## Faces β€” set `seedanceFace` for AI-character faces\n\nSeedance routes through **three tiers** depending on the face in the reference, exposed as the \"Face in Reference\" toggle plus the real-face params on `slates_generate_video`:\n\n- **Faceless / object / environment refs β†’ default route (cheapest).** Leave `seedanceFace` off.\n- **An AI-character's FACE in a reference β†’ `seedanceFace: true`.** The default route's baseline moderation rejects or degrades faces, so this reroutes to the face-capable provider. It costs **~45% more** β€” the cost key becomes `seedance-2-face-{res}-{N}s`, so the pre-flight quote already reflects it. Announce the face-route price, not the faceless one.\n- **A REAL person's photo (the user themselves, an actor) β†’ the consent-gated premium route.** If a `seedanceFace` gen fails with `[REAL_FACE_DETECTED]`, the provider classified the reference as a real person: confirm with the user that (a) they hold the rights/consent to the likeness and (b) they accept the higher price (cost key `seedance-2-realface-{res}-{N}s`, roughly 2Γ— the AI-face rate β€” quote via `slates_estimate_generation_cost`), then retry with `seedanceRealFace: true` + `realFaceConsent: true`. Never set `realFaceConsent` without the user's explicit confirmation.\n\nRules:\n- **The real-vs-AI call is the PROVIDER'S, not yours.** ByteDance's classifier is probabilistic β€” some real photos pass the standard face route (billed at the cheap rate; fine), others get rejected with `[REAL_FACE_DETECTED]` (auto-refunded). Don't preemptively route to the real-face tier just because a photo looks real; try `seedanceFace: true` first and escalate only on the marked rejection. Public figures / celebrities fail on every route.\n- It's about the **reference, not the output.** If your character identity or generated portrait shows a face, turn it on. A product shot with no person stays off.\n- Don't toggle it on \"just in case\" β€” a faceless gen on the face route burns ~45% extra for nothing.\n<!-- /slates-only -->\n\n## Reference rules (the verified ones)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it β€” lighting, medium, texture, symmetry, competing identities β€” is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** β€” because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** β€” because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** β€” the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** β€” mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* β€” the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere β€” fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs β€” each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** β€” identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` β€” you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") β€” that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation β€” the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light β€” never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference β€” the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text β†’ bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media β€” describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime β†’ real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Seedance specifically\n\n- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* β€” motion, change, camera. Never re-describe what's in the reference, and never say \"still / scene / from a movie / from the image.\" The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic β€” if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one<!-- slates-only --> β€” see slates-cost-discipline<!-- /slates-only -->).\n- **Seedance's own idiom for rule 2 is `Reference <Subject_N> in <Image_N>`** `[official :1389]` β€” `Image_N` indexes the order the refs are attached, so the name plus the index carries the role. The full binding grammar is in Part 1 (Subject binding).\n- **Rule 3 has an official ceiling here.** The trend is MORE references (video and audio into Seedance), all addressed by name β€” but for **multi-character frames** see the twin-problem section above: bind every character to its image, append the anti-twin constraint, and prefer single-person references. Past 4 reference people, stability drops `[official :2048-2052]`.\n- **Rule 8 holds even though Seedance can render common text natively** `[official :1758]`. A baked NB2 start frame is still the reliable route for text that must be legible.\n- **Rule 5 pairs with the first/last-frame exclusion** β€” frames and reference images are mutually exclusive on this model (see Reference media above), so an environment you must match exactly costs you the frame lane.\n\n<!-- slates-only -->\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** β€” if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 β€” Beach Sunset`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.\n\n- βœ… \"I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame β€” the camera move is going to be a slow dolly forward through the gap.\"\n- ❌ \"I'm using the first beach image and the last one...\" (which? They have four.)\n<!-- /slates-only -->\n\n---\n\n# Part 3 β€” Community field notes `[community]`\n\nThird-party guides and Slates field experience. Useful heuristics β€” but if one of these ever appears to contradict Part 1, **Part 1 wins**.\n\n## Length\n\n**Sweet spot 60-150 words** for a single shot (not 150-300 β€” that's the upper bound). Multi-shot storyboards run longer; official Example 1 above is ~230 words across three shots.\n\n## Pin the subject in the first 20-30 words\n\nThe opening sentence is the **identity anchor**. If the subject isn't locked early, the model hallucinates new subjects mid-clip. (Compatible with Part 1: the binding preamble comes before `Shot 1`.)\n\n```\nA matte black earbud case sits on a polished obsidian surface...\n```\n\n## Lighting is a top quality lever\n\nLighting has an outsized impact on output quality β€” which is why it has its own slot in the official 8-slot formula. Describe it before or alongside the subject.\n\n```\nA cool-white diagonal beam from upper left, dust particles drifting through.\nSoft golden hour lighting from low west angle.\nDramatic rim light against dark background.\n```\n\n## Camera and subject motion β€” separate sentences\n\nMixing them is a common cause of glitchy / shaky output.\n\n❌ \"The camera speed ramps as the earbud rises.\"\nβœ… \"The earbud rises smoothly. The camera tracks upward.\"\n\n## Slow-motion works; \"fast\" is a known bad token\n\nSpeed ramps and slow-motion are supported in natural language, and `fast` is widely reported as a quality-degrading keyword. **The official version of this rule is stronger and better founded** β€” prioritize slow, gentle, continuous small movements and avoid high-burst action (Part 1, Action description `[:1611-1615]`). Prompt the motion class, not the adjective.\n\n```\nthe lid opens in slow-motion Β· the blade whips through the air\n```\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ β€” same contract as the anti-list in slates-prompting-nano-banana-2.\n Extracted by src/prompts/banned-tokens.ts into the slates_generate_video op\n description and matched against submitted prompts. The RECOMMENDED vocabulary\n below sits OUTSIDE the markers on purpose β€” it is backticked too. -->\n<!-- /slates-only -->\n**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.\n<!-- @banned:end -->\nThese are quality *incantations* β€” the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).\n\n## Style block at the end\n\nOne primary anchor + 2-3 supporting details, as the trailing paragraph (both official examples do exactly this). End with `Single continuous take` if you want one shot with no cuts. **Never** write `no cut` or `seamless transition` β€” not in the training vocabulary.\n\n## ⚠️ Don't cross-pollinate image-model syntax\n\nNamed **lenses, apertures, film stocks, and camera bodies** β€” `85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`, `shot on Sony A7S3` β€” are an **image-model lever** (correct and encouraged in `slates-prompting-nano-banana-2`) and a **Seedance anti-pattern**. ByteDance's guide uses shot sizes, camera moves, pacing words, and the image-quality/style vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.\n\nIf you are carrying a look over from an NB2 start frame, translate it: `85mm f/1.4, Portra 400` β†’ `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`.\n\n## Negative prompting β€” inline only\n\nSeedance has **no `negativePrompt` field**. Put negatives in the constraints slot, led by the three official templates (Part 1):\n\n```\nkeep it subtitle-free Β· do not generate a logo Β· do not generate a watermark\navoid jitter and bent limbs\navoid temporal flicker\navoid identity drift\nno distortion, no stretching\n```\n\nAlso fine: positive reframing (\"empty street\" not \"no cars\").\n\n## Image-to-video / first-frame guidance\n\n**Describe motion, not image.** The model already sees the visual; tokens spent re-describing appearance are wasted.\n\nStability phrases that help:\n- `preserve composition and colors`\n- `maintain exact appearance from reference image`\n- `consistent character throughout, no deformation or drift`\n\n**Cap I2V prompts under 60 words** when possible. Over 100 words frequently triggers silent generation failure.\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Hallucinated subject mid-clip | First 20-30 words = identity anchor |\n| Bent limbs / extra fingers | `avoid jitter and bent limbs` in Constraints |\n| Identity drift across multi-shot | Re-name the bound subject in **every** `Shot N` block `[official :1537]` |\n| Two identical characters in one frame | The twin fix in Part 1 β€” bind each character to its image + append the global anti-twin constraint |\n| Silent generation failure on I2V | Cut prompt under 100 words, single primary camera move |\n| Speech / motion conflict | Limit dialogue to one line per action shot |\n| Erratic/random pacing | You second-stamped. Remove all time markers and use `Shot N` `[official :1586]` |\n\n## Sources\n\n**Official (authoritative):**\n- BytePlus ModelArk β€” Seedance 2.0 prompting guide, archived at `research/seedance-2-modelark-docs.md` (all `:NNNN` refs above)\n\n**Community (secondary):**\n- [fal.ai β€” How to Use Seedance 2.0](https://fal.ai/learn/tools/how-to-use-seedance-2-0)\n- [apiyi.com β€” Seedance 2.0 Prompt Guide](https://help.apiyi.com/en/seedance-2-0-prompt-guide-video-generation-camera-style-tips-en.html)\n- [atlabs.ai β€” Ultimate Seedance 2.0 Prompting Guide](https://www.atlabs.ai/blog/the-ultimate-seedance-2.0-prompting-guide-47-prompts-2026)\n",
28
+ "slates-prompting-seedance-2-5": "---\nname: slates-prompting-seedance-2-5\ndescription: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is a SECOND SEAT next to 2.0, not an upgrade β€” it buys 30-second takes, 30 image references, audio-only references and INTEGER-SECOND TIMESTAMPS, and it gives up native 4K and costs more than 2.0 at every resolution they share. Timestamps are the one grammar difference that matters: 2.0 ignores them and answers only to shot numbers, 2.5 acts on them. Otherwise it shares 2.0's grammar (read slates-prompting-seedance for subject binding, camera and constraint vocabulary); this file covers what is different, plus the two hazards unique to 2.5 β€” the prompt-intent task classifier and the cost trap that comes with 30-second takes.\n---\n\n# Seedance 2.5 β€” prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Seedance 2.5.** Shares 2.0's grammar exactly (subject binding, camera vocabulary, externalised emotion, inline constraints β€” read `slates-prompting-seedance` for those). Two things are different, and both matter.\n\n**The five levers**\n1. **Timestamps work here** β€” integer seconds, and the model acts on them: `[0-4] she reads the letter. [4-9] she folds it and looks up.` 2.0 ignores exactly this syntax.\n2. **Length is the reason to be here** β€” takes up to 30 seconds, where 2.0 stops at 15. Write the beats as `[0-6]`, `[6-12]`, `[12-18]`; do not hope for them.\n3. **Up to 30 image references**, and a multi-view image can serve as ONE subject reference (up to 5 subjects). 2.0 cannot do either.\n4. **Audio-only references are accepted** without an image or video alongside β€” the only Seedance seat that takes one.\n5. **Keep the 2.0 discipline**: one camera move per beat (`slow track right`, `handheld follow`), physical action instead of stated emotion, and quality asked for in the image-quality slot vocabulary β€” `rich details`, `natural colors`, `cinematic texture`, `soft lighting`.\n\n**Examples**\n- `[0-6] Wide shot, <Subject_1>@<Image_1> crosses an empty car park toward a idling van, slow track right. [6-12] Medium, she stops as the driver's window comes down. [12-18] Close-up, she looks off past the lens and does not answer. Rich details, natural colors. Keep it subtitle-free.`\n- `[0-10] A single continuous handheld follow behind a courier climbing a fire escape, rain. [10-20] She reaches the landing, turns, and the city opens behind her. Cinematic texture, soft lighting.`\n\n**Hard constraint:** it is the EXPENSIVE seat and it has NO 4K β€” 480p/720p/1080p only, and dearer than 2.0 at every resolution they share. It is a second seat, never an upgrade. Long takes multiply cost linearly: quote a 30-second take before you fire it.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** (2.5 reclassifies the task and fails a fresh generation on these):\n- `edit`, `extend`, `continue the video`, `same video but` β€” they make the provider read a fresh generation as an edit\n- `85mm`, `f/1.4`, `Portra 400` and any other lens, aperture, film-stock or camera-body token β€” image-model vocabulary, a Seedance anti-pattern on both seats\n<!-- @banned:end -->\n\n**Read `slates-prompting-seedance` first.** The prompt GRAMMAR is the same model family: the\n8-slot advanced formula, subject binding by `<Subject_N>@<Image_N>`, camera vocabulary,\nexternalised emotion, inline constraint words, the anti-twin fix. None of it is restated here.\nThis file is only what 2.5 changes β€” and the biggest change is that **2.5 acts on timestamps\nwhere 2.0 ignores them.**\n\n---\n\n## The one fact that decides whether you use it at all\n\n**Seedance 2.5 is the EXPENSIVE seat, not the cheap one β€” and it has no 4K.**\n\nIt runs at 480p, 720p or 1080p (1080p landed on all three routes on 2026-08-24), and at every\nresolution the two seats share it costs MORE than 2.0 β€” 720p $0.231/s against $0.15/s, **54% more**.\nSo 2.5 does not replace 2.0; it sits beside it, and you pay for what it buys:\n\n| | Seedance 2.0 | Seedance 2.5 |\n|---|---|---|\n| Resolution | 480p / 720p / 1080p / **native 4K** | 480p / 720p / 1080p β€” **no 4K** |\n| Price at 720p (faceless) | **$0.15/s** | $0.231/s |\n| Length | 4–15s | **4–30s in one take** |\n| Reference budget | 15 (9 image + 3 video + 3 audio) | **50 (30 image + 10 video + 10 audio)** |\n| Combined reference video/audio | ≀15s | **≀30s** |\n| Audio-only reference | βœ— (needs an image or video alongside) | **βœ“** |\n| **Timestamps in the prompt** | **βœ— β€” ignored; shot numbers only** | **βœ“ β€” integer seconds, acted on** |\n| Multi-view image as ONE subject reference | βœ— (not recommended) | **βœ“ (up to 5 subjects)** |\n| Video edit as its own task type | βœ— | **βœ“ (`seedance-2.5-edit`)** |\n| Default video model | **yes** | no |\n\n**Route to 2.5 when the shot needs LENGTH, MANY REFERENCES, or an audio-only reference.\nRoute to 2.0 when resolution matters at all** β€” which, for anything a client will see full-screen,\nis most of the time.\n\n---\n\n## 🚨 Hazard 1 β€” the prompt-intent task classifier\n\nThis is the one that costs money and time, and it has no equivalent on 2.0.\n\n**Seedance 2.5 sorts every request into one of five task types** β€” text-to-video,\nreference-to-video, first/last-frame, **video edit**, **video extend** β€” from the reference roles\nattached **plus the intent of your sentence**. Each type then has its own parameter constraints,\nand a violation comes back **asynchronously**: the task queues, credits are reserved, and only then\ndoes it fail.\n\nThe trigger words are ordinary English:\n\n| Reclassified as | Words that do it (ByteDance's own list) |\n|---|---|\n| **video edit** | `edit video` Β· `add` Β· `insert` Β· `remove` Β· `delete` Β· `modify` Β· `replace` Β· `change to` |\n| **video extend** | `extend forward` Β· `extend backward` Β· `continue` Β· `continue from` Β· `extend the story` |\n\nSo a perfectly legitimate reference-to-video prompt β€” *\"a wide shot of the workshop, **remove** the\ntripod from frame\"* β€” gets classified as an edit and fails on constraints it never set.\n\n**What to do:**\n\n1. **If you mean to edit an existing clip, say so with the MODEL, not the sentence.** Call\n `slates_edit_video` with `model: 'seedance-2.5-edit'`. That routes to a dedicated\n task-typed endpoint and the classifier never has to guess.\n2. **If you mean a fresh shot, describe the finished frame rather than an instruction to change\n one.** Not *\"remove the tripod\"* β†’ *\"the workshop bench, clear and uncluttered\"*. Not\n *\"add rain\"* β†’ *\"heavy rain falling through the streetlight\"*. This is better prompting anyway:\n the model renders what you describe, it does not take edits to an imagined draft.\n3. The trigger only fires when **references are attached**. A plain text-to-video prompt is safe\n however it is worded.\n\n**Slates will warn you, and it will never rewrite your prompt.** When a 2.5 reference generation's\nprompt contains one of these words, the composer shows a warning that NAMES the words and the agent\nroute returns the same string. Silently editing the user's sentence to dodge a provider classifier\nis forbidden β€” the words that reach the model are always the words the user can see.\n\n---\n\n## 🚨 Hazard 2 β€” resolution is not the price dial here. LENGTH is.\n\nEvery other model in Slates trains the habit that lower resolution means lower cost. 2.5 breaks it,\nbecause the thing that moves the bill is **length**, and 2.5's length ceiling is double 2.0's.\n\nWorked, at the shipped rates:\n\n| Generation | Credits |\n|---|---|\n| 2.5 Β· 480p Β· 5s Β· faceless | 26 |\n| 2.5 Β· 720p Β· 5s Β· faceless | 58 |\n| 2.5 Β· 1080p Β· 5s Β· faceless | 103 |\n| 2.5 Β· 720p Β· 30s Β· faceless | 347 |\n| 2.5 Β· 720p Β· 30s Β· AI-face route | **489** |\n| 2.5 Β· 720p Β· 30s Β· consented real-face route | **710** |\n| 2.5 Β· 1080p Β· 30s Β· faceless | **614** |\n| 2.5 Β· 1080p Β· 30s Β· consented real-face route | **1,749** |\n| *(for scale)* 2.0 Β· 1080p Β· 15s Β· AI-face route | 411 |\n\n**A 30-second 720p clip can cost more than a 15-second 1080p one** β€” and a base licence starts\nwith 1,000 credits. Someone who reads \"720p\" as \"cheap\" and asks for a 30-second take on the\nreal-face route has spent 71% of their welcome grant on one clip; **on the real-face route a single\n30-second 1080p take is more than the whole grant.**\n\n**Discipline:**\n\n- **Always quote with `slates_estimate_generation_cost` before a take over ~10 seconds,** and say\n the number out loud before generating.\n- **Find the shot at short LENGTH, not at low resolution.** Length is what moves the price, so cut\n seconds while you are still exploring β€” 4–8s β€” and stay at the resolution you actually want.\n **A 480p pass does not de-risk a 720p or 1080p render.** Generation is stochastic: the higher-\n resolution run is a different take, not the same shot rendered better. So a 480p draft that looks\n right buys you no guarantee, and one that looks wrong may have been fine at 720p β€” you paid 26\n credits to learn nothing, when 58 would have bought a real candidate.\n- **Length is a creative decision, not a default.** 30 seconds is available; it is rarely the right\n answer for a single shot. Multi-shot storyboards inside one 30s generation are what the length is\n actually for.\n- Read `slates-cost-discipline` β€” all of it applies, more sharply here.\n\n---\n\n## Timestamps β€” the one grammar change\n\n<!-- @inject:seedance-25-timestamps -->\n**2.0 does not respond to timestamps and answers only to shot numbers. 2.5 responds to\ninteger-second timestamps.** That is ByteDance's own first line under \"Differences from Seedance\n2.0\", and it is why a 30-second take is usable at all: the length is only worth buying if you can\nsay *when* things happen inside it.\n\nBoth formats are valid on 2.5, and you can mix them β€” `Shot N` blocks for a storyboard whose\npacing you are happy to leave to the model, timestamps when a beat has to land at a moment.\n\n**Three ways to control time, all first-party:**\n\n| Form | Write it like |\n|---|---|\n| **Interval** | `0-3 seconds… 3-7 seconds… 7-15 seconds` or `[1s-4s]… [4s-8s]… [8s-12s]` |\n| **Time point** | *\"Quick left sideways transition at the 5-second mark.\"* |\n| **Relative** | *\"After 3 seconds, everyone around him shakes their head.\"* Β· *\"The frame freezes for 1 second after he presses the shutter.\"* |\n\n**The rules that come with them:**\n\n- **One second is the smallest unit.** Integers only β€” no `2.5s`, no frames.\n- **No gaps in the timeline.** `0-3s… 5-6s…` leaves 3-5s unspecified and the model fills it however\n it likes. Intervals must abut: `0-3s`, `3-7s`, `7-15s`.\n- **Budget the plot to the seconds.** Too little content in a range and the model improvises to\n fill it; too much and you get extra cuts or dropped beats. This is the actual craft of a 30s take.\n- **Never time-code a high-frequency action.** *\"Shake your head three times per second\"* is\n explicitly called out as a misuse β€” timestamps schedule beats, they don't choreograph frames.\n- **Transitions want both halves:** the moment AND the method β€” *\"At the 5-second mark, the camera\n transitions leftward with a left wipe into a natural dissolve.\"*\n- **Timestamps work on an EDIT too**, and that is where they earn the most: they scope a change in\n time as well as in content β€” *\"Change the man's action from drinking coffee to mopping the floor\n from 4-6 seconds in Video 1, and leave the rest of the content unchanged.\"* Without a range, a\n whole-clip instruction is applied to the whole clip.\n\nDo **not** carry this back to 2.0, and do not carry Veo's `[00:00-00:02]` bracket syntax into\neither β€” 2.0 ignores time entirely, and the cross-model syntax swap is its own known failure.\n<!-- @end:seedance-25-timestamps -->\n\n---\n\n## What the extra reference budget is actually for\n\n30 image references (up from 9) does **not** mean \"attach 30 images\". Every rule in\n`slates-prompting-seedance` about references still holds β€” 2–4 strong references beat both\nextremes, and one reference per role.\n\n**Where 2.5 moves the ceiling, per ByteDance's own input recommendations:**\n\n| | Stable | Works, but expect re-rolls |\n|---|---|---|\n| Subjects bound by IMAGE reference | 1–8 | 9–12 |\n| Subjects bound by VIDEO or AUDIO reference | 1–5 | 6–10 |\n| Reference clip length, per subject | 5–10s | longer drops stability |\n\n**Multi-view images of one subject are supported on 2.5** β€” a turnaround sheet can be a single\nreference image, where 2.0 wanted one authoritative rendering per subject. Past **5 subjects**,\ngo back to single-view images, one per view, rather than one image carrying several viewpoints.\n\nThe larger budget earns its keep in exactly two places:\n\n- **A long multi-shot take** where different shots need different subjects and locations bound β€”\n the budget is spread across the storyboard, not stacked on one frame.\n- **Video and audio references alongside images**, which is where 2.5's 10 + 10 matters far more\n than the image count.\n\n### Audio-only references β€” the genuinely new input\n\n2.0 required an image or video alongside any audio reference. **2.5 accepts audio on its own.**\nThat makes one recipe possible that was not before: drive a scene's timing, voice or ambience from\na recording with no visual anchor at all β€” a voice line, a music bed, a room tone β€” and let the\nmodel build the picture to it. Cite it the same way as any other reference\n(`Reference the timbre in <Audio_N> to generate…`), and remember that audio references carry **no\nbilling dimension** on any Seedance route: audio is included.\n\n### Video references\n\nUp to 10 clips, ≀30s combined (2.0: 3 clips, ≀15s). A reference VIDEO switches the cost key to\n`seedance-2.5*-vref-{res}-{T}s`, where **T = Ξ£ input seconds + output seconds** β€” the sum is across\n**every** clip attached, not just the longest. Three 6-second references on a 12-second output bills\n30 seconds, not 12 and not 18. Quote before confirming.\n\n**The cap is a refusal, not a trim.** Attach an eleventh clip, or push past 30 combined seconds, and\nthe composition is rejected before anything uploads. That asymmetry is deliberate: reference images\nwarn-and-trim because dropping one doesn't change the price, and a dropped reference VIDEO would be\none you were quoted for and the model never saw.\n\n### Mixing all three in one call\n\n50 files total (30 image + 10 video + 10 audio) is a shared budget. Everything is cited positionally\nby type β€” `image 1`, `video 2`, `audio 1` β€” in attachment order, so reordering the attachments\nrenumbers the citations. Write the prompt against those numbers:\n\n```\nMarcus (image 1) performs the motion from video 1 in the workshop from image 2,\nusing the voice timbre from audio 1. Preserve his identity, appearance and outfit.\n```\n\n🚨 **SAY WHAT AN AUDIO REFERENCE IS FOR.** It can mean music, dialogue, voice, tone or timbre β€” five roles on one attachment β€” so an unroled clip falls back to **dialogue**: the model re-transcribes it and speaks ITS words. A real take came back as *\"a map called Slates\"* for *\"an app called Slates\"*. Name it as the voice timbre and the clip carries the voice while the prompt carries the words. ByteDance's own sentence: *\"Image 1 depicts the protagonist John and uses the voice timbre from Audio 1.\"* Bind each speaker in a sentence, never by attachment order β€” position carries nothing.\n\nFrames and reference media stay mutually exclusive, in every combination β€” the reference endpoint\nhas no first/last-frame parameters at all, so this is a shape mismatch rather than a preference.\n\n---\n\n## Sound: four bracket types, and they are the vendor's syntax\n\nByteDance's 2.5 API tutorial states this as a **prompt rule**, not a suggestion β€” verbatim: *\"Use\nspecial characters to distinguish sounds: `()` for music, `<>` for sound effects, `{}` for dialogue,\nand `【】` for subtitles. For non-Chinese dialogue, it is recommended to specify the language before\nthe dialogue.\"*\n\n```\nShe sets the cup down {English: \"We open in ten minutes.\"} <ceramic clink on wood>\n(low piano, unhurried)\n```\n\n- `()` **music** Β· `<>` **sound effects** Β· `{}` **dialogue** Β· `【】` **on-screen subtitles**\n- **Name the language before non-Chinese dialogue.** `{English: \"...\"}`.\n- Unbracketed sound description still works β€” this is a disambiguator, not a required wrapper. Reach\n for it when one sentence carries more than one kind of sound and you need the model to tell them\n apart, which is exactly where an unmarked prompt puts a line of dialogue into the score.\n\n⚠️ **These four are SEEDANCE 2.5's.** MiniMax H3 has its own three-layer scheme (body / soundscape /\nscore) and its angle brackets are documentation notation that must never be typed. Do not carry\neither grammar onto the other model.\n\n## Say what a reference is NOT for\n\nThe same rule adds a half nobody uses: *\"Specify what each asset provides, such as appearance,\naction, or timbre, **and what should not be referenced**.\"* Negative scoping is a first-class part of\nthe citation, not a fallback β€” *\"use her face and wardrobe from image 1, not its lighting or\nbackground\"* is a stronger instruction than naming the positive alone, because an unscoped reference\nbrings its whole frame with it.\n\n🚨 **The vendor writes `@Image 1`; Slates writes `image 1`, and that difference is deliberate.**\nBytePlus's API tutorial says *\"Use `@Image 1`, `@Video 1`, and `@Audio 1`\"*, while its own 2.5 prompt\nguide states the bare form (`Image 1 / Video 1 / Audio 1`) in the one normative sentence it has.\n**Two first-party docs, two forms** β€” the disagreement is recorded, not resolved, in\n`second-brain/business/projects/slates/research/model-prompting-research.md`. What settles it FOR US\nis neither: **`@` is a reference-token sigil in the Slates prompt composer, and an unresolved one is\nsilently deleted from the prompt before it is sent.** Typing `@Image 1` here does not produce\n`@Image 1`, it produces nothing. The bare form is confirmed working on both models. Never hand-type\nthe sigil.\n\n## Seedance 2.5 Edit (`slates_edit_video`, `model: 'seedance-2.5-edit'`)\n\nIts own picker row and its own op call, deliberately: the task type is **the model you chose**,\nnever something inferred from your sentence.\n\n**Why route here at all:** it is the **only edit engine in Slates that accepts a clip longer than 15\nseconds** (4–30s, versus Kling O3 Edit's 3–15s and Omni Flash Edit's 3–10s). For a clip inside the\nothers' range, choose on fidelity instead β€” Omni Flash Edit won the prompt-only head-to-head, and\nKling O3 Edit is the one that takes element and style reference images.\n\n**How it behaves:**\n\n- **Output length follows the SOURCE clip**, and the bill is the ceiled source length. The provider\n requires an automatic duration on this task type, so there is no length knob β€” the clip you attach\n is the quote. The returned clip can differ from the source by up to ~0.3s, which only compresses\n transition frames; a clip that 2.5 itself generated comes back at exactly its input length.\n- **Source clips under 20 seconds edit more reliably.** 4–30s is what the task type accepts;\n ByteDance's own recommendation is to stay inside 20 for quality. A 28-second source is legal and\n will need more attempts.\n- **The aspect ratio follows the source clip too.** No ratio control; the frame is the clip's frame.\n- **480p, 720p or 1080p output**, native audio β€” and an edit bills the video-reference tier Γ—2,\n so 1080p on this row is the most expensive second in the app. Quote it.\n- **Prompt and source clip only** on this op. The MODEL takes reference images on an edit\n (ByteDance recommends 1–5 β€” *\"replace the man in dark clothing in @Video 1 with @Image 2\"*);\n **Slates has not wired that path**, so today an edit that must lock an identity from a photo\n goes to Kling O3 Edit. Constraint of our build, not of the model β€” worth revisiting.\n- **An edit bills roughly DOUBLE a plain 2.5 generation of the same length**, because every provider\n charges an edit on input + output seconds. Read the confirm gate's number; do not reason from the\n generation rate.\n- **Set `seedanceFace: true` when a character's face is visible in the clip.** The faceless provider\n blocks faces outright β€” this is not a price optimisation, it is whether the job runs at all.\n- **There is no consented-real-face route for editing.** Real-person footage that the AI-face route\n rejects has to go to Kling O3 Edit.\n\n**Prompting an edit** β€” the same discipline as every other edit engine: **describe only what\nchanges.** The source already carries its composition, motion, timing and performance; re-describing\nthem fights the model. Use Seedance's own edit grammar from `slates-prompting-seedance`\n(*\"Strictly edit `<Video_1>`, and modify `<Original_Characteristic>` to `<New_Characteristic>`\"*) and\n**never** write *\"reference video 1\"* in an edit β€” the official guide is explicit that this phrasing\ngets the request reclassified as a reference task, which is the same landmine as Hazard 1.\n\nTwo things sharpen an edit prompt, both first-party:\n\n- **Say it as A β†’ B, not as an outcome.** *\"Change the man's action from drinking coffee to mopping\n the floor\"* beats *\"the man mops the floor\"* β€” naming what it currently is tells the model what\n to overwrite.\n- **Timestamp a partial edit** β€” the edit task type reads the same integer-second timestamps the\n generation path does. Rules and forms are in Β§ Timestamps above; this is the single most useful\n thing they buy.\n\n**Audio is editable too, and it is the least obvious use of this row.** The same op rewrites what\nis heard while the picture stays put: change a spoken line, change the accent, translate the\ndialogue and re-fit the lip movement, strip or replace the BGM or a sound effect. *\"Only edit the\nman's dialogue in Video 1: change it to 'Don't come over here,' in an American accent\"* is an edit,\nnot a lip-sync job. Bill it like any other edit β€” on the source clip's length.\n\n---\n\n## Faces, and what does NOT change\n\nThe three-tier face routing is identical to 2.0 β€” faceless β†’ default route, an AI character's face β†’\n`seedanceFace: true` (the relaxed provider, a real cost premium), a real person's photo β†’ the\nconsent-gated premium route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`\nset **only** after the user explicitly confirms they hold the rights to the likeness. The full rules,\nincluding why the real-vs-AI call is the provider's and not yours, are in\n`slates-prompting-seedance`.\n\nAlso unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:\n\n- **One primary camera move per shot.**\n- **No lens / aperture / film-stock vocabulary.** That is image-model syntax and a Seedance\n anti-pattern.\n- **No `negativePrompt` field** β€” constraints go inline, and 2.5 acts on negative phrasing in\n exactly two dimensions: subtitles (*\"no subtitles\"*) and audio (*\"no BGM; environmental and\n action sounds only\"*, *\"no audio\"*). Everywhere else, describe what you want, not what you don't.\n- **Legible in-shot text still belongs in a baked start frame**, not in the video prompt.\n",
29
+ "slates-prompting-seedance": "---\nname: slates-prompting-seedance\ndescription: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance 2.0 structures multi-beat prompts as a \"Shot 1 / Shot 2 / Shot 3\" storyboard against an 8-slot advanced formula β€” never per-second time stamps, which 2.0 does not respond to (Seedance 2.5 does; see slates-prompting-seedance-2-5). Its syntax differs from Kling, Veo and the image models; don't cross-pollinate (in particular, no lens / aperture / film-stock vocabulary).\n---\n\n# Seedance 2.0 β€” prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Seedance 2.0.** Not copywriting β€” an ENGINEERING instruction to a spatial layer and a temporal layer: who, in what scene, doing what, how the camera moves, and in what order. Multi-beat work is a `Shot 1 / Shot 2 / Shot 3` storyboard.\n\n**The five levers**\n1. **Bind every subject to its reference** β€” `<Subject_1>@<Image_1>` β€” and keep the descriptions identical across shots. Unbound subjects are where twins come from.\n2. **Shot sizes and camera MOVES, not lens data** β€” `medium close-up`, `slow push in`, `handheld follow`, `whip pan`. One primary move per shot.\n3. **Externalise emotion as physical action.** Not \"she is nervous\": `she turns the ring on her finger twice, then stops`.\n4. **Use the image-quality slot vocabulary** for quality β€” `HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`. That is the officially sanctioned way to ask.\n5. **Constraints go INLINE**, led by the official templates β€” `keep it subtitle-free`, `do not generate a watermark`, `avoid jitter and bent limbs`, `avoid temporal flicker`.\n\n**Examples**\n- `Shot 1: medium shot, <Subject_1>@<Image_1> steps out of the freight lift into a wet loading bay, slow push in. Shot 2: close-up, she turns the ring on her finger twice and stops, handheld. Rich details, cinematic texture, natural colors. Keep it subtitle-free.`\n- `Single continuous take. Wide shot of a fishing skiff crossing a grey swell, camera tracks from the starboard rail. Spray hits the lens once. Soft lighting, natural colors, film-grain texture. Avoid jitter and bent limbs.`\n\n**Hard constraint:** NO timestamps β€” 2.0 ignores them and answers only to shot numbers (2.5 acts on them). No lens, aperture, film stock or camera body: that is image-model vocabulary and a Seedance anti-pattern. There is no negativePrompt field.\n<!-- @card:end -->\n\nByteDance's video model β€” first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K β€” 4K video is Pro-only, default 1080p), 4–15s, first+last frame, and up to 9 reference images / 3 videos / 3 audio clips.\n\n> **How to read this file.**\n> **[official :NNNN]** β€” ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/byteplus-seedance-2-0-api-docs.md`). Receipt-grade; treat as law.\n> **[community]** β€” third-party guides and our own field notes. Useful, but an `[official]` block always wins.\n> **[slates]** β€” how the Slates app composes or bills this; not ByteDance doctrine.\n>\n> The split is load-bearing. A community-sourced \"narrative timing beats\" doctrine shipped in this file for months teaching the **exact inverse** of ByteDance's published guidance. Never merge the two registers again.\n\n---\n\n# Part 1 β€” Official ByteDance doctrine\n\n## What Seedance actually is `[official :1450-1452]`\n\nSeedance 2.0 is a multimodal AI director. It reads text, images, video and audio **simultaneously** and internally decomposes them into two dimensions:\n\n- the **spatial layer** β€” what is in the frame\n- the **temporal layer** β€” how things change over time\n\nSo a good prompt is **not \"copywriting-style description\" but an \"engineering-style instruction\"**: who, in what scene, doing what action, how the camera moves, and in what chronological order events occur β€” delivered respectively to the spatial layer and the temporal layer.\n\n## The advanced formula β€” 8 slots `[official :1455]`\n\n```text\nprecise subject + action details + scene/environment + lighting & color tone\n+ camera movement + visual style + image quality + constraints\n```\n\n⚠️ There is **no official \"6-step formula.\"** `Subject + Action + Environment + Camera + Style + Constraints` is community branding with no ByteDance source, and it silently drops the **lighting & color tone** and **image quality** slots. Use the 8 slots above.\n\n## Task-type sentence patterns `[official :1389-1425]`\n\nSeedance classifies your request from the phrasing. Use the pattern that matches the task:\n\n| Task | Pattern |\n|---|---|\n| **Image reference** | ``Reference `<Subject_N>` in `<Image_N>` to generate…`` |\n| **Video reference** | ``Reference `<Action / Camera_movement / Style / Sound_effect>` in `<Video_N>` to generate…`` |\n| **Audio reference** | ``Reference the timbre in `<Audio_N>` to generate…`` |\n| **Video edit β€” modify** | ``Strictly edit `<Video_N>`, and modify `<Original_Characteristic>` in it to `<New_Characteristic>``` |\n| **Video edit β€” add** | ``<Element_Features>` + `<Timing>` + `<Location>`` |\n| **Video edit β€” delete** | Name what to delete; for anything that must stay, say so explicitly |\n| **Video extend** | ``Extend `<Video_N>` forward/backward to generate…`` |\n| **Combined** | ``Reference `[Dimension]` of `<Image/Video_N>`, strictly edit `<Video_X>`, `[Specific_Edits]``` |\n\n### ⚠️ Edit / extend phrasing landmine `[official :1431]`\n\n> *\"For edit / extend video tasks, directly use `<Video_N>` to refer to the video. **Do not use \"reference `<Video_N>`\"**, to avoid being incorrectly identified as a reference task.\"*\n\nThis is easy to trip: Slates has an edit lane<!-- slates-only --> (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)<!-- /slates-only -->. Writing *\"reference video 1 and change the jacket to red\"* gets classified as a **reference** task β€” the model generates a brand-new clip inspired by the source instead of editing it. Write *\"Strictly edit video 1, and modify the blue jacket to red.\"*\n\n## Shot structure β€” \"Shot 1 / Shot 2 / Shot 3\" `[official :1563-1598]`\n\n> *\"Use shot order, write a simple 'Shot 1 / Shot 2 / Shot 3' storyboard for each segment of the video, and then merge them into a complete prompt.\"*\n\n**❌ Never second-stamp.** No `0:00–0:03`, no \"At 4 seconds\", no per-segment durations.\n\n> *\"Do not impose strict limits on the duration of each segment; prioritize allowing the model to naturally generate the pacing based on the plot.\"* `[:1580]`\n>\n> *\"The model's support for precise timing (such as 0–3 seconds) is **unstable**, and forcibly limiting duration may lead to **abnormal generation results**.\"* `[:1586]`\n\nOrder shots by when events occur β€” primary first, secondary later. Let the plot set the pacing.\n\n**Per-shot internal order** `[official :1590-1598]` β€” organize each shot in exactly this sequence:\n\n1. **Camera movement or shot transition** β€” \"slowly push in from a wide shot\", \"fixed camera position\", \"cut to…\"\n2. **Subject actions and expressions** β€” the key actions and expression changes of the core character/object\n3. **Position or spatial change** β€” where the subject is, and how that relationship shifts\n4. **Audio** β€” sound effects, voices, background music for that shot\n\n**One primary camera move per shot** β€” see Camera below. `[official :1648]`\n\n## Subject binding β€” names + image indexes `[official :1488-1556]`\n\nEvery time a subject appears, it must be **explicitly referred to**. Two supported forms:\n\n- **Undefined subjects** β€” bind inline every mention: `<Subject_N>@<Image_N>`. Official example: **`Zhang San@Image 1`**. `[:1540]`\n- **Pre-declared subjects** β€” define once, then reuse the same label verbatim: *\"Define the tall man in **Video 1** as **police officer**, and define the other short man as **thief**\"*, then say \"police officer\" every time after. `[:1514]`\n\n**One subject spread across several assets** β€” bind them together: *\"Define `[…]` in **Image 1** and `[…]` in **Image 2** as `<Subject N>`.\"* `[:1514]`\n\n⚠️ **An Asset ID must never substitute for `<Image/Video_N>`.** `[:1546]` *\"the model cannot directly associate the Asset ID with the reference content.\"* Always cite by index.\n\nAlso official: keep descriptions concise, avoid redundancy, avoid semantic conflicts (contradictory traits for one subject), and prefer expressing spatial relationships through reference images rather than dense text. `[:1550-1556]`\n\n**`[slates]`** β€” the app composes this for you. `composeReferences()` cites each canonical character or environment reference inline as `Name (image N)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*\"Zhang San (corresponding to image 1)\"* `[:1976]`). You never hand-write role labels or index numbers.\n\n## Action description `[official :1602-1621]`\n\n- **Body-part specificity + quantified degree.** Name hands, legs, head, shoulders, back β€” and supplement **range, speed, and force**. *\"slowly raise a hand\", \"quickly turn the head\", \"push hard off the ground\", \"slightly lower the head.\"*\n- **Prioritize slow, gentle, continuous small movements.** Avoid high-burst, large-dynamic actions β€” sprinting, big jumps, violent rolls. *\"walk slowly\", \"gently raise a hand\", \"sit down naturally with the motion.\"* **This is the official basis for the folk rule that \"fast\" degrades quality** β€” it is not a banned token, it is a class of motion the model handles badly.\n- **Supplement transitions between actions.** Specify inertia and continuity between consecutive beats so movement reads coherent: *\"use the inertia of turning around to naturally raise a hand\", \"naturally transition from a pause into raising a hand.\"*\n\n## Externalize emotion `[official :1623-1636]`\n\nReplace abstract emotion words (\"very sad\", \"extremely angry\") with **specific physical detail**. This is the highest-leverage single habit in the official guide:\n\n| Abstract | Externalized as actions and details |\n|---|---|\n| **Sadness** | head lowering, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the corner of clothing, tears welling but not falling |\n| **Joy** | corners of the mouth rising uncontrollably, brows and eyes relaxing, steps becoming light, unconsciously humming a tune |\n| **Nervousness / anxiety** | frequently checking the watch, fingers constantly tapping the tabletop, rapid breathing, eyes darting away |\n| **Anger** | both fists clenched, jawline tense, chest heaving, eyes sharp, squeezing words out through gritted teeth |\n| **Relief** | letting out a long breath, tense shoulders completely relaxing, a faint smile appearing, looking up toward the distance |\n\n## Camera `[official :1643-1648]`\n\n> *\"The model has a **strong understanding of camera movement terms**, so you can **directly use standard camera movement terminology**, such as 'medium shot, close-up, wide shot, slow push-in, smooth lateral tracking, fixed shot.'\"*\n\nThis is an **open vocabulary, not a fixed list** β€” and it explicitly includes **shot size** (close-up / medium / wide / long shot), which is as much a camera instruction as the move itself.\n\n> ⚠️ *\"Try to specify only 1 type of camera movement in a single shot. Do not require push, pull, pan, and move at the same time, as this will increase image instability.\"* `[:1648]`\n\n## Image quality, style, and constraints `[official :1656-1679]`\n\nThese three slots \"define creative boundaries for the model, unify image quality and artistic tone, and avoid visual flaws and random deviations.\"\n\n**1. Image quality** β€” define clarity, texture detail, and lighting quality. Official vocabulary: `HD` Β· `rich details` Β· `cinematic texture` Β· `natural colors` Β· `soft lighting`.\n\n> ⚠️ This is a **real slot with real vocabulary** β€” do not confuse it with Stable-Diffusion-era quality incantations. `8K` / `masterpiece` / `trending on artstation` remain banned slop tokens (see Part 3); *\"cinematic texture, rich details, natural colors\"* is the officially sanctioned way to ask for the same thing.\n\n**2. Style** β€” the overall art style and visual tone: `cyberpunk cool blue-purple tone` Β· `retro film` Β· `fresh Japanese style`.\n\n**3. Constraint words** β€” *\"Constraint words are very important. They can effectively avoid visual flaws, deformities, breakdowns, and unreasonable elements.\"* Official templates, verbatim:\n\n- **No subtitles** β€” \"keep it subtitle-free\" / \"avoid generating any text or subtitles\"\n- **No logo** β€” \"do not generate a logo\"\n- **No watermark** β€” \"do not generate a watermark\"\n\nSeedance has **no `negativePrompt` field** β€” constraints go inline in this slot. See Part 3 for the wider inline-negative kit.\n\n## πŸ”΄ Duplicated characters β€” the twin problem `[official :1948-1994]`\n\n**Symptom:** in frames with **many characters**, where **three-view / multi-view character images** are supplied as references, two identical characters appear in the same generated frame.\n\n**Root causes** `[:1954-1959]`:\n1. Character subjects are not clearly defined in the prompt, so the model cannot distinguish roles.\n2. *\"When character **three-view / multi-view images** are used as reference assets, it is easy to confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"*\n\n**Official fixes, in their order** `[:1971-1994]` β€” ByteDance is explicit that these *reduce probability*, not eliminate it:\n\n1. **Bind each character to its image explicitly**, in a consistent format. Official example: *\"Zhang San (corresponding to image 1) throws the green passbook toward Li Si (corresponding to image 2), who is standing.\"*\n2. **Append the global constraint verbatim** at the end of the prompt `[:1982]`:\n > *\"Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect. Keep only a single corresponding character in the same frame, and do not reproduce repeated copies of characters.\"*\n3. **Optimize reference assets** `[:1988]` β€” *\"For character reference images, prioritize independent single-person photos. Three-view or multi-view assets are not recommended.\"*\n4. **Simplify the prompt** β€” do not paste a whole script; redundant copy confuses the model.\n\n**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on identity sheets. Practical rule for Slates:\n\n- **Multi-character Seedance shot** β†’ bind every character to its image, append the anti-twin constraint, and prefer single-person / dominant-portrait references over multi-view sheets.\n- **Single-character shot** β†’ the standard character-sheet flow is fine.\n\n**Too many reference people** `[official :2048-2052]` β€” past **4 reference people**, output stability drops (wrong headcount, duplicates). Official workaround: group the cast into images of ≀4 people each, generate those stills first, then drive the video from them.\n\n## Worked examples `[official :1689-1745]`\n\nThese are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere β€” 2.0 does not respond to them at all, which is version-scoped and reverses on 2.5.\n\n**Example 1 β€” dormitory emotional short drama (dialogue-focused).** Assets: `@Image 1` half-body photo of the female lead Β· `@Image 2` dormitory scene reference Β· `@Video 1` camera-movement reference Β· `@Audio 1` indoor ambience.\n\n> Use the girl in @Image 1 as the main character, use @Image 2 as the dormitory scene style reference, and refer to the camera movement in @Video 1.\n>\n> **Shot 1**: At dusk, **girl @Image 1** walks briskly to the **dormitory entrance @Image 2**. The camera follows steadily in a medium shot. Warm yellow sunlight spills into the hallway from the window. She pauses at the doorway, takes a deep breath, and looks slightly nervous.\n>\n> **Shot 2**: **Girl @Image 1** pushes the door open and enters the dormitory. The camera cuts to an indoor medium shot. Her roommates look up at her while organizing their books. One of them smiles and asks {How did the exam go? Did you pass?}. The camera slowly cuts between half-body close-ups of several people.\n>\n> **Shot 3**: **Girl @Image 1** first lowers her head with a dejected expression. The camera gives her a close-up. Then she raises her head, unable to hold back a smile, laughs out loud, and says {I was kidding}. Her roommates start chasing and play-fighting with her. The camera slowly pulls back and freezes on a wide shot of the dormitory filled with laughter.\n>\n> The entire video should have a high-definition cinematic documentary style, with warm tones and soft lighting. The character's face remains stable without deformation; movements are natural and smooth, with no stutter or flicker. The ambient sound blends naturally with @Audio 1.\n\n**Example 2 β€” ancient-style cliff confrontation (action/atmosphere-focused).** Assets: `@Image 1` female lead in red Β· `@Image 2` assassin in black Β· `@Image 3` cliff and bamboo forest Β· `@Video 1` martial-arts camera reference Β· `@Audio 1` drum beats.\n\n> Use the woman in red from @Image 1 as the female lead, use the woman in black from @Image 2 as the opponent, use the cliff and bamboo forest environment in @Image 3 as the scene reference, refer to the overall camera movement and action rhythm in @Video 1, and synchronize the background sound effects with @Audio 1.\n>\n> **Shot 1**: At dusk, the camera slowly pushes in from a side medium shot of **woman in red @Image 1**. She stands at the edge of the cliff and lifts a wine flask to drink. Her sleeves and robe hem sway gently in the mountain wind. The camera circles halfway around her, moving from the front to her back. In the distance, a figure in black is faintly visible in the bamboo forest.\n>\n> **Shot 2**: The camera zooms and fades into a long shot. From a drone perspective, it overlooks the entire cliff and bamboo forest. The two characters stand at opposite ends of the cliff. The mountain wind lifts their robe hems and dust, and the rhythm slightly accelerates with the drum beats.\n>\n> **Shot 3**: The camera cuts back to a ground-level close shot. The two slowly draw their swords and face off. **Woman in red @Image 1** shifts from a careless expression to a cold gaze. **Woman in black @Image 2** looks determined, and the sword tip trembles slightly. The camera steadily follows the two as they circle each other, finally freezing on a close-up of the instant before the two swords meet.\n>\n> The overall visual style should feel like a cinematic wuxia world in misty rain, with cool tones, low saturation, a film-grain texture, and rich light-and-shadow layers. The characters' faces and body proportions remain stable without deformation. Movements are continuous and natural, not stiff, with no clipping or stutter.\n\n## Other official notes\n\n- **On-screen text** `[official :1758]` β€” Seedance can render common text (ad slogans, subtitles, speech bubbles) and will auto-match style/colour from context, or take an explicit colour / style / timing / position. Prefer **common characters**; avoid rare glyphs and special symbols. (For *guaranteed* legible text, the start-frame route in Part 3 is still safer.)\n- **Extension degrades quality** `[official :2004-2024]` β€” using a generated video as the input for extension compounds degradation, with mottled colour blocks in face regions. Limit repeated continuations; prefer HD assets as input.\n- **Special effects that miss** `[official :2031-2044]` β€” when a described effect comes out wrong (a countdown that scrolls randomly), define it with a **reference video** instead of words: *\"the way the number '2999' appears should reference video 1.\"*\n\n---\n\n# Part 2 β€” Slates-specific `[slates]`\n\n## Reference media β€” caps and transport\n\nReference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.\n\n**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `\"first/last frame content cannot be mixed with reference media content.\"` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.\n\n### All three modalities go in ONE call\n\nThe caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.\n\nCite each by type and index, in the order they were attached β€” `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.\n\n```\nMarcus (image 1) performs the motion from video 1, in the workshop from image 2,\nusing the voice timbre from audio 1. Preserve his identity, appearance and outfit.\n```\n\n🚨 **SAY WHAT AN AUDIO REFERENCE IS FOR.** It can mean music, dialogue, voice, tone or timbre β€” five roles on one attachment β€” so an unroled clip falls back to **dialogue**: the model re-transcribes it and speaks ITS words. A real take came back as *\"a map called Slates\"* for *\"an app called Slates\"*. Name it as the voice timbre and the clip carries the voice while the prompt carries the words. ByteDance's own sentence: *\"Image 1 depicts the protagonist John and uses the voice timbre from Audio 1.\"* Bind each speaker in a sentence, never by attachment order β€” position carries nothing.\n<!-- slates-only -->\n**Attaching a clip is NOT the same as editing it.** \"Add as reference\" puts it in the composer alongside everything else and wipes nothing; \"Edit with AI\" makes the clip the canvas and clears the tray for a fresh instruction. Two different jobs, two different menu entries β€” never infer one from the other.\n\n**Over the cap is REFUSED, never trimmed.** A reference video is priced into the quote before it is sent, so a clip silently dropped after the quote would be a clip you paid for and the model never saw. Remove one and retry.\n<!-- /slates-only -->\n\n### Motion transfer & lip-sync recipes (reference video / audio)\n\nThese aren't separate Seedance features β€” they're prompting strategies over reference media.<!-- slates-only --> The Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:<!-- /slates-only -->\n\n- **Motion transfer:** subject image as a reference + the driving clip<!-- slates-only --> via `videoReferenceAssetId`<!-- /slates-only --> (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`\n- **Lip-sync / dialogue:** write the line in the prompt β€” `The person in video 1 says: \"…\"` β€” with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≀15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`\n- **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script β†’ their likeness AND voice deliver the new line.\n<!-- slates-only -->\n- **Billing:** a reference VIDEO switches the cost key to `seedance-2*-vref-{res}-{T}s` where T = clip seconds + output seconds β€” quote before confirming. Audio references are free (audio is included on every route).\n<!-- /slates-only -->\n\n<!-- slates-only -->\n## Faces β€” set `seedanceFace` for AI-character faces\n\nSeedance routes through **three tiers** depending on the face in the reference, exposed as the \"Face in Reference\" toggle plus the real-face params on `slates_generate_video`:\n\n- **Faceless / object / environment refs β†’ default route (cheapest).** Leave `seedanceFace` off.\n- **An AI-character's FACE in a reference β†’ `seedanceFace: true`.** The default route's baseline moderation rejects or degrades faces, so this reroutes to the face-capable provider. It costs **~45% more** β€” the cost key becomes `seedance-2-face-{res}-{N}s`, so the pre-flight quote already reflects it. Announce the face-route price, not the faceless one.\n- **A REAL person's photo (the user themselves, an actor) β†’ the consent-gated premium route.** If a `seedanceFace` gen fails with `[REAL_FACE_DETECTED]`, the provider classified the reference as a real person: confirm with the user that (a) they hold the rights/consent to the likeness and (b) they accept the higher price (cost key `seedance-2-realface-{res}-{N}s`, roughly 2Γ— the AI-face rate β€” quote via `slates_estimate_generation_cost`), then retry with `seedanceRealFace: true` + `realFaceConsent: true`. Never set `realFaceConsent` without the user's explicit confirmation.\n\nRules:\n- **The real-vs-AI call is the PROVIDER'S, not yours.** ByteDance's classifier is probabilistic β€” some real photos pass the standard face route (billed at the cheap rate; fine), others get rejected with `[REAL_FACE_DETECTED]` (auto-refunded). Don't preemptively route to the real-face tier just because a photo looks real; try `seedanceFace: true` first and escalate only on the marked rejection. Public figures / celebrities fail on every route.\n- It's about the **reference, not the output.** If your character identity or generated portrait shows a face, turn it on. A product shot with no person stays off.\n- Don't toggle it on \"just in case\" β€” a faceless gen on the face route burns ~45% extra for nothing.\n<!-- /slates-only -->\n\n## Reference rules (the verified ones)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it β€” lighting, medium, texture, symmetry, competing identities β€” is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** β€” because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** β€” because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** β€” the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** β€” mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* β€” the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere β€” fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs β€” each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** β€” identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` β€” you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") β€” that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation β€” the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light β€” never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference β€” the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text β†’ bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media β€” describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime β†’ real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Seedance specifically\n\n- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* β€” motion, change, camera. Never re-describe what's in the reference, and never say \"still / scene / from a movie / from the image.\" The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic β€” if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one<!-- slates-only --> β€” see slates-cost-discipline<!-- /slates-only -->).\n- **Seedance's own idiom for rule 2 is `Reference <Subject_N> in <Image_N>`** `[official :1389]` β€” `Image_N` indexes the order the refs are attached, so the name plus the index carries the role. The full binding grammar is in Part 1 (Subject binding).\n- **Rule 3 has an official ceiling here.** The trend is MORE references (video and audio into Seedance), all addressed by name β€” but for **multi-character frames** see the twin-problem section above: bind every character to its image, append the anti-twin constraint, and prefer single-person references. Past 4 reference people, stability drops `[official :2048-2052]`.\n- **Rule 8 holds even though Seedance can render common text natively** `[official :1758]`. A baked NB2 start frame is still the reliable route for text that must be legible.\n- **Rule 5 pairs with the first/last-frame exclusion** β€” frames and reference images are mutually exclusive on this model (see Reference media above), so an environment you must match exactly costs you the frame lane.\n\n<!-- slates-only -->\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** β€” if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 β€” Beach Sunset`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.\n\n- βœ… \"I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame β€” the camera move is going to be a slow dolly forward through the gap.\"\n- ❌ \"I'm using the first beach image and the last one...\" (which? They have four.)\n<!-- /slates-only -->\n\n---\n\n# Part 3 β€” Community field notes `[community]`\n\nThird-party guides and Slates field experience. Useful heuristics β€” but if one of these ever appears to contradict Part 1, **Part 1 wins**.\n\n## Length\n\n**Sweet spot 60-150 words** for a single shot (not 150-300 β€” that's the upper bound). Multi-shot storyboards run longer; official Example 1 above is ~230 words across three shots.\n\n## Pin the subject in the first 20-30 words\n\nThe opening sentence is the **identity anchor**. If the subject isn't locked early, the model hallucinates new subjects mid-clip. (Compatible with Part 1: the binding preamble comes before `Shot 1`.)\n\n```\nA matte black earbud case sits on a polished obsidian surface...\n```\n\n## Lighting is a top quality lever\n\nLighting has an outsized impact on output quality β€” which is why it has its own slot in the official 8-slot formula. Describe it before or alongside the subject.\n\n```\nA cool-white diagonal beam from upper left, dust particles drifting through.\nSoft golden hour lighting from low west angle.\nDramatic rim light against dark background.\n```\n\n## Camera and subject motion β€” separate sentences\n\nMixing them is a common cause of glitchy / shaky output.\n\n❌ \"The camera speed ramps as the earbud rises.\"\nβœ… \"The earbud rises smoothly. The camera tracks upward.\"\n\n## Slow-motion works; \"fast\" is a known bad token\n\nSpeed ramps and slow-motion are supported in natural language, and `fast` is widely reported as a quality-degrading keyword. **The official version of this rule is stronger and better founded** β€” prioritize slow, gentle, continuous small movements and avoid high-burst action (Part 1, Action description `[:1611-1615]`). Prompt the motion class, not the adjective.\n\n```\nthe lid opens in slow-motion Β· the blade whips through the air\n```\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ β€” same contract as the anti-list in slates-prompting-nano-banana-2.\n Extracted by src/prompts/banned-tokens.ts into the slates_generate_video op\n description and matched against submitted prompts. The RECOMMENDED vocabulary\n below sits OUTSIDE the markers on purpose β€” it is backticked too. -->\n<!-- /slates-only -->\n**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.\n<!-- @banned:end -->\nThese are quality *incantations* β€” the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).\n\n## Style block at the end\n\nOne primary anchor + 2-3 supporting details, as the trailing paragraph (both official examples do exactly this). End with `Single continuous take` if you want one shot with no cuts. **Never** write `no cut` or `seamless transition` β€” not in the training vocabulary.\n\n## ⚠️ Don't cross-pollinate image-model syntax\n\nNamed **lenses, apertures, film stocks, and camera bodies** β€” `85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`, `shot on Sony A7S3` β€” are an **image-model lever** (correct and encouraged in `slates-prompting-nano-banana-2`) and a **Seedance anti-pattern**. ByteDance's guide uses shot sizes, camera moves, pacing words, and the image-quality/style vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.\n\nIf you are carrying a look over from an NB2 start frame, translate it: `85mm f/1.4, Portra 400` β†’ `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`.\n\n## Negative prompting β€” inline only\n\nSeedance has **no `negativePrompt` field**. Put negatives in the constraints slot, led by the three official templates (Part 1):\n\n```\nkeep it subtitle-free Β· do not generate a logo Β· do not generate a watermark\navoid jitter and bent limbs\navoid temporal flicker\navoid identity drift\nno distortion, no stretching\n```\n\nAlso fine: positive reframing (\"empty street\" not \"no cars\").\n\n## Image-to-video / first-frame guidance\n\n**Describe motion, not image.** The model already sees the visual; tokens spent re-describing appearance are wasted.\n\nStability phrases that help:\n- `preserve composition and colors`\n- `maintain exact appearance from reference image`\n- `consistent character throughout, no deformation or drift`\n\n**Cap I2V prompts under 60 words** when possible. Over 100 words frequently triggers silent generation failure.\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Hallucinated subject mid-clip | First 20-30 words = identity anchor |\n| Bent limbs / extra fingers | `avoid jitter and bent limbs` in Constraints |\n| Identity drift across multi-shot | Re-name the bound subject in **every** `Shot N` block `[official :1537]` |\n| Two identical characters in one frame | The twin fix in Part 1 β€” bind each character to its image + append the global anti-twin constraint |\n| Silent generation failure on I2V | Cut prompt under 100 words, single primary camera move |\n| Speech / motion conflict | Limit dialogue to one line per action shot |\n| Erratic/random pacing | You second-stamped. Remove all time markers and use `Shot N` `[official :1586]` |\n\n## Sources\n\n**Official (authoritative):**\n- BytePlus ModelArk β€” Seedance 2.0 prompting guide, archived at `research/byteplus-seedance-2-0-api-docs.md` (all `:NNNN` refs above)\n\n**Community (secondary):**\n- [fal.ai β€” How to Use Seedance 2.0](https://fal.ai/learn/tools/how-to-use-seedance-2-0)\n- [apiyi.com β€” Seedance 2.0 Prompt Guide](https://help.apiyi.com/en/seedance-2-0-prompt-guide-video-generation-camera-style-tips-en.html)\n- [atlabs.ai β€” Ultimate Seedance 2.0 Prompting Guide](https://www.atlabs.ai/blog/the-ultimate-seedance-2.0-prompting-guide-47-prompts-2026)\n",
30
30
  "slates-prompting-seedream-5-lite": "---\nname: slates-prompting-seedream-5-lite\ndescription: How to prompt Seedream 5 Lite (ByteDance image model β€” the cheap volume option in Slates). Read before calling slates_generate_image with model seedream-5-lite, or slates_edit_image with editModel seedream-5-lite. Seedream front-loads attention, likes 30-100 focused words, and takes quoted strings for in-image text.\n---\n\n# Seedream 5 Lite β€” prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Seedream 5 Lite.** The cheap volume seat: flat-priced at every resolution, which makes it the right default for storyboard passes, variant grids and look-dev. Structure, most important first: `Subject + Style + Composition + Lighting/Atmosphere + Technical`.\n\n**The five levers**\n1. **Lead with the subject.** Earlier words weigh more; close with the camera and technical detail.\n2. **Keep it 30-100 words.** Unlike models that reward verbosity, this one gets confused by very long prompts. Focused beats exhaustive.\n3. **Name the composition** β€” `symmetrical composition`, `rule of thirds`, `foreground detail with blurred background`, `overhead perspective`.\n4. **Name the light as a named condition** β€” `golden hour`, `dramatic side lighting`, `soft diffused light`, `moody low-key`, `bright high-key`.\n5. **Quote in-image text.** It takes quoted strings for posters and layouts, which is half of why it is the drafting seat.\n\n**Examples**\n- `Professional headshot of a female CEO, short blonde hair, confident expression, navy suit, neutral office background. Studio lighting, shallow depth of field, high-end corporate photography, shot on 85mm.`\n- `A rain-soaked night market stall, cinematic, rule of thirds with the vendor camera-right, foreground steam blurred, moody low-key lighting with practical neon, shot on 35mm.`\n\n**Hard constraint:** it is the DRAFTING seat, not the hero seat. Explore here, then re-run the winner on Nano Banana 2 or FLUX.2 Max for the locked shot.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use:**\n- `masterpiece`, `best quality`, `highly detailed`, `8k`, `award-winning` β€” quality incantations, not description\n- a prompt past about 100 words: this model gets confused by very long prompts, and focused beats exhaustive\n<!-- @banned:end -->\n\nByteDance's Seedream image model, Lite tier, routed via fal.ai. In Slates: `slates_generate_image` with `model: seedream-5-lite` (REQUIRES projectId β€” no headless path). **Flat-priced regardless of resolution** β€” the cheapest image model in Slates, which makes it the right default for high-volume drafting, storyboard exploration, and variant grids. Call `slates_estimate_generation_cost` for the current number; never quote prices from memory. Less censored than Nano Banana 2.\n\n**When to pick it:** lots of frames cheap (storyboard passes, 3-4 variant exploration), posters/layouts with text, quick look-dev. Step up to NB2 or FLUX.2 Max for the locked hero shot.\n\n## Core structure β€” five components, most important first\n\n```\nSubject + Style + Composition + Lighting/Atmosphere + Technical parameters\n```\n\nSeedream weights concepts mentioned **earlier in the prompt** more heavily. Lead with the subject; close with camera/technical details.\n\n**Length sweet spot: 30-100 words.** Unlike models that reward verbosity, Seedream gets confused by very long prompts. Focused beats exhaustive.\n\n## Style, composition, lighting vocabulary it responds to\n\n- **Style:** portrait photography, macro photography, cinematic, photorealistic, minimalist, oil painting, watercolor, digital art\n- **Composition:** symmetrical composition, rule of thirds, foreground detail with blurred background, wide-angle view, overhead perspective, medium shot, close-up\n- **Lighting:** golden hour lighting, dramatic side lighting, soft diffused light, moody low-key lighting, bright high-key lighting\n- **Technical:** shot on 85mm lens, shallow depth of field, high resolution\n\n## Worked examples\n\n**Portrait:**\n> \"Professional headshot of a female CEO with short blonde hair, confident expression, wearing a navy blue suit, neutral office background, studio lighting, shallow depth of field, high-end corporate photography style\"\n\n**Product:**\n> \"Modern smartphone floating in space, dark background with subtle blue gradient, product photography, studio lighting highlighting the glossy screen, ultra-detailed, commercial quality, photorealistic rendering\"\n\n## In-image text: double-quote it\n\nPut the exact string in double quotation marks β€” Seedream treats quoted text as render-this-verbatim:\n\n```\nA minimalist poster with the headline \"SUMMER SALE\" in bold sans-serif, centered\n```\n\nSeedream is one of the stronger models for layout-heavy work (posters, mockups, diagrams): call out the layout explicitly β€” \"centered headline, subtitle beneath, clean margins.\"\n\n## Edits: change one thing, lock the rest\n\nVia `slates_edit_image` with `editModel: seedream-5-lite`. Seedream edits respond well to instructions that name the change AND the preserved elements:\n\n```\nChange the bag to brown leather. Keep the person's face, pose, and the room unchanged.\n```\n\nNote: Seedream edits in Slates ignore extra `referenceAssetIds` β€” that path is Nano Banana 2 only.\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Subject inconsistent / mutates | Put the subject description first; break complex subjects into clear components |\n| Style drift | Reinforce the aesthetic with 2-3 related terms (\"cinematic, photorealistic, shallow depth of field\") |\n| Compositional confusion | Use photography terms (\"medium shot,\" \"overhead view\"); simplify the scene |\n| Garbled text | Double-quote the exact string; keep it short; state placement |\n| Mushy long-prompt output | Cut to under 100 words β€” Seedream rewards focus, not volume |\n\n## Iterate cheap, lock expensive\n\nFlat pricing makes Seedream the iterate-fast model: run the 3-strike loop here (draft β†’ evaluate inline β†’ one specific delta β†’ regenerate), and only re-render the winning composition on a pricier model if the project's hero shot demands it. Cost rules live in `slates-cost-discipline` β€” the batch-authorization pattern applies when generating variant grids.\n\n## Sources\n\n- [fal.ai β€” Seedream Prompt Guide](https://fal.ai/learn/devs/seedream-v4-5-prompt-guide)\n- [BytePlus ModelArk β€” Seedream Prompt Guide](https://docs.byteplus.com/en/docs/ModelArk/1829186)\n",
31
31
  "slates-prompting-veo-3": "---\nname: slates-prompting-veo-3\ndescription: How to prompt Veo 3.1 (Google). Read before calling slates_generate_video with veo-3.1-fast or veo-3.1-standard. Veo is a NICHE pick, never the default (route per slates-model-selection β€” Kling is the general default, Seedance the premium tier) β€” reach for it only when native synchronized audio must generate WITH the video in one gen. 16:9 or 9:16, 4/6/8s. Different cinematography formula than Seedance/Kling. (no subtitles) is mandatory after every dialogue line.\n---\n\n# Veo 3.1 β€” prompting\n\n<!-- @card:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Everything between the @card markers is extracted by\n src/prompts/craft-cards.ts and returned on every cost estimate for this\n model, so it is the ONE piece of positive craft guidance the agent cannot\n skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved\n compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.\n Keep it under 2,400 characters (the build fails above that) and keep the\n rationale, the receipts and the worked examples in the body below. -->\n<!-- /slates-only -->\n**Card β€” Veo 3.1.** Google's formula, in order: `[Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]`. Sweet spot 50-150 words; the official benchmark is about 50.\n\n**The five levers**\n1. **Open with the cinematography** β€” `Medium shot`, `low angle`, `aerial`, `dolly in`, `rack focus`, `vertigo effect`. Veo reads the first clause as the camera.\n2. **Texture words counter the AI-plastic look** β€” `fine skin pores`, `visible fabric weave`, `subtle contrast, no gloss or sharpening`. Name materials concretely: `charcoal cotton hoodie`, `matte concrete`, `silk lapel`.\n3. **Weight verbs stop floaty motion** β€” `trudges`, `drops heavily`, and a ground contact: `boots crunch on gravel`.\n4. **Terse voice direction only** β€” `says in a weary voice`, `whispers`, `mutters`. Veo is far less responsive to long voice blocks than Kling.\n5. **Always include an ambience line.** Without one the mix feels dead. `Soft office ambience.` `Wind on the open ridge.` And SFX always carries a cause: `SFX: thunder cracks in the distance`, never `SFX: thunder`.\n\n**Examples**\n- `Medium shot, a tired founder rubbing her temples in front of a bulky monitor in a cluttered office late at night. Harsh fluorescent overheads and the green glow of the screen. Fine skin pores, visible fabric weave on a charcoal cotton hoodie. Soft office ambience, a fan hum. Retro, slightly grainy.`\n- `Low angle, a farrier trudges across a wet yard carrying a shoeing box, boots crunching on gravel. Overcast north light, matte concrete and wet steel. Wind and distant livestock. He says in a weary voice, \"One more and we're done.\" (no subtitles).`\n\n**Hard constraint:** `(no subtitles)` after EVERY dialogue line you do not want burned in as text. Negatives are NOUNS, not instructions β€” `wall, frame`, never `no walls`. And do not cross syntaxes: Seedance's `single continuous take` suppresses Veo's cuts, and Veo timestamps in a Seedance prompt cause drift.\n<!-- @card:end -->\n\n<!-- @banned:start -->\n<!-- slates-only -->\n<!-- MACHINE-READ. Every `backticked` token between the @banned markers is\n extracted by src/prompts/banned-tokens.ts and returned on this model's cost\n estimate, and every submitted prompt is matched against it. Keep entries\n backticked and prose outside the backticks. -->\n<!-- /slates-only -->\n**Never use** (each one is spelled out above):\n- `single continuous take` β€” Seedance's phrase; in a Veo prompt it suppresses the cuts you asked for\n- `no subtitles` is REQUIRED after dialogue, but instruction-shaped negatives are not: `no walls`, `no man-made structures`, `don't show` β€” negatives are NOUNS here\n- `SFX: thunder` and any label-only effect β€” every effect carries a cause and a distance\n<!-- @banned:end -->\n\nGoogle DeepMind's video model. Two tiers: `veo-3.1-fast` (cheaper, quick) and `veo-3.1-standard` (higher quality). 4k variants exist for both (4K video requires Slates Pro).\n\n**Native single-shot duration: 4, 6, or 8 seconds** β€” and **8s only** at 1080p or 4K, or whenever you attach reference images (that endpoint is 8s-fixed). 4s and 6s exist at 720p, text-to-video or single-start-frame only. Longer durations require chaining clips via Extend / last-frame reuse β€” quality degrades if naively requested past 8s in a single generation. Aspect ratio: **16:9 or 9:16** on the route Slates uses. `slates_generate_video` REFUSES anything outside these before submit and names the legal set β€” nothing is silently ignored or downgraded.\n\nNative synchronized audio at 48kHz: dialogue, SFX, ambient β€” generated WITH video, not added after.\n\n## Official Google formula\n\n```\n[Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]\n```\n\nSweet spot length: 50-150 words. Cloud's official benchmark is ~50 words.\n\nVerbatim official benchmark:\n> \"Medium shot, a tired corporate worker, rubbing his temples in exhaustion, in front of a bulky 1980s computer in a cluttered office late at night. The scene is lit by the harsh fluorescent overhead lights and the green glow of the monochrome monitor. Retro aesthetic, shot as if on 1980s color film, slightly grainy.\"\n\n## Cinematography vocabulary (Vertex AI docs)\n\n**Lenses:** wide-angle, telephoto, fisheye, anamorphic, 35mm, 85mm, shallow/deep depth of field\n\n**Lighting:** Rembrandt lighting, volumetric lighting, backlighting, golden hour glow, lens flare, rack focus, **vertigo effect** (dolly zoom)\n\n**Camera moves:** dolly (in/out), truck (left/right), pan, tilt, crane, aerial/drone, handheld, whip pan, arc shot, zoom\n\n## Texture-realism phrases (counter the AI-plastic look)\n\n```\nfine skin pores Β· visible fabric weave Β· subtle contrast, no gloss or sharpening\n```\n\nSpecify materials concretely: `charcoal cotton hoodie`, `matte concrete`, `silk lapel`. Generic \"smooth, beautiful\" rendering is the failure mode you're avoiding.\n\n## Dialogue β€” `(no subtitles)` is mandatory\n\nEvery dialogue line you don't want burned in as text overlay needs `(no subtitles)`. Verbatim from the founder talking-head benchmark:\n\n```\nThe founder says, \"This update cuts setup time in half, helping teams get started faster.\" (no subtitles).\n```\n\nWithout this, Veo will overlay subtitle text on top of your generation.\n\n## Voice direction β€” keep it terse\n\nVeo is less responsive to long voice-direction blocks than Kling. Use brief modifiers:\n\n```\nsays in a weary voice\nwhispers\nshouts\nmutters\n```\n\nMulti-character: handles 2-3 speakers natively. Past 3, sync degrades β€” use first-frame/last-frame chaining for 4+.\n\n## SFX with cause\n\n```\nβœ… SFX: thunder cracks in the distance\n❌ SFX: thunder\n```\n\nAlways specify direction or distance.\n\n## Ambient is mandatory\n\nAlways include an ambience line per scene. Without it, the audio mix feels dead.\n\n```\nSoft office ambience.\nWind on the open ridge.\nDistant city hum.\n```\n\n## First-frame + last-frame workflow (Veo's strength)\n\n1. Generate start frame (Gemini 2.5 Flash Image is the recommended pair β€” Slates' Nano Banana 2 works)\n2. Generate end frame\n3. Animate with both frames as anchors\n\n**Motion-Lock hack:** Keep ~60% of the same background pixels between start and end frames. Prevents latent drift across the clip.\n\nVerbatim arc-shot example:\n> \"The camera performs a smooth 180-degree arc shot, starting with the front-facing view of the singer and circling around her to seamlessly end on the POV shot from behind her on stage. The singer sings 'when you look me in the eyes, I can see a million stars.'\"\n\n## Ingredients-to-Video (multiple references)\n\nVerbatim example:\n> \"Using the provided images for the detective, the woman, and the office setting, create a medium shot of the detective behind his desk. He looks up at the woman and says in a weary voice, 'Of all the offices in this town, you had to walk into mine.'\"\n\n## Reference discipline (character / environment refs)\n\n<!-- @inject:references-read-literally -->\n> **The general law: the model reads a reference literally.**\n> A reference image is not a suggestion. Whatever is baked into it β€” lighting, medium, texture, symmetry, competing identities β€” is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.\n\nEvery reference rule below is a corollary of that one sentence, which is why \"prep the reference\" beats \"prompt around the reference\" every time:\n\n- **Flat, plain identity refs** β€” because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).\n- **One authoritative rendering per subject** β€” because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets \"confuse the model's character recognition, causing it to generate duplicate characters of the same appearance.\"\n- **No 3D-game-render look in a reference** β€” the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.\n- **Break perfect symmetry** β€” mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.\n\n**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* β€” the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere β€” fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.\n<!-- @end:references-read-literally -->\n\n<!-- @inject:reference-rules-core -->\nIdentity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.\n\n1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs β€” each one adds context AND another variable to balance.\n2. **One reference per ROLE, named in the prompt** β€” identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two \"identity\" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` β€” you never hand-write role labels.\n3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a \"Reference Image Instructions\" block or role essays** (\"use for identity, ignore the outfit, render a neutral expression\") β€” that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.\n4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation β€” the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.\n5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light β€” never a multi-panel grid fed whole.\n6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference β€” the cells share a split detail budget and were generated jointly, so their flaws propagate.\n7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.\n8. **Legible in-shot text β†’ bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.\n9. **Working from existing media β€” describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on \"video one\"; marker-object insertion; video-as-reference for a series.)\n10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction (\"anime β†’ real person\"). There are no preset pickers, and there is no style slider.\n<!-- @end:reference-rules-core -->\n\n### For Veo specifically\n\n- **Veo's idiom for rule 2 is plain-English role naming in the sentence itself** β€” *\"Using the provided images for the detective, the woman, and the office setting, create a medium shot of…\"* (see Ingredients-to-Video above). The role rides in the noun phrase, not in a separate label block.\n- **Rule 8 has a second reason to matter here:** Veo bakes subtitle text into the frame unless every dialogue line carries `(no subtitles)`. Text you did not ask for is the failure mode, not just text you did.\n\n## Negative prompting β€” nouns, not instructions\n\nVeo has a `negativePrompt` field. **Verbatim Vertex AI rule:**\n> \"Describe unwanted elements as nouns rather than instructions. Use 'wall, frame' instead of 'no walls' or 'don't show walls.'\"\n\nInline: positive reframing in the body too.\n- βœ… `\"a desolate landscape with no buildings or roads\"`\n- ❌ `\"no man-made structures\"`\n\n## Common failure modes + fixes\n\n| Failure | Fix |\n|---|---|\n| Subject identity shifts mid-clip | Front-load identity at prompt start; use material cues (`charcoal canvas`, `cotton`, `silk`) to stabilize |\n| Floaty / weightless motion | Weight verbs (`trudges`, `drops heavily`), ground contact (`boots crunch on gravel`) |\n| AI-plastic look | `fine skin pores`, `visible fabric weave`, `subtle contrast` |\n| Subtitles baked into video | `(no subtitles)` after every dialogue line |\n| Rushed dialogue | Lines fit one natural breath in 8s |\n| Mismatched ambience | Always include an ambience line |\n| Warped geometry | `photorealistic stability` |\n\n## Timestamp shot syntax (for chained / multi-beat scenes)\n\nVeo accepts `[00:00-00:02]` brackets for timed sequences within an 8s clip. **Do NOT cross syntaxes** β€” Veo timestamps in a Seedance prompt cause subject drift; Seedance \"single continuous take\" in a Veo prompt suppresses cuts.\n\nVerbatim multi-beat:\n> \"[00:00-00:02] Medium shot from behind a young female explorer with a leather satchel and messy brown hair in a ponytail, as she pushes aside a large jungle vine to reveal a hidden path.\n> [00:02-00:04] Reverse shot of the explorer's freckled face, her expression filled with awe as she gazes upon ancient, moss-covered ruins. SFX: The rustle of dense leaves, distant exotic bird calls.\n> [00:04-00:06] Tracking shot following the explorer as she steps into the clearing and runs her hand over the intricate carvings on a crumbling stone wall.\n> [00:06-00:08] Wide, high-angle crane shot, revealing the lone explorer standing small in the center of the vast, forgotten temple complex, half-swallowed by the jungle. SFX: A swelling, gentle orchestral score begins to play.\"\n\n## Benchmark prompt β€” founder talking head (full)\n\n> \"Camera locked at eye level, medium close-up on a 35mm lens: a startup founder in his late 30s with short black hair and light stubble, wearing a charcoal cotton hoodie, speaking directly to camera, leaning slightly forward as he speaks, lifting one hand to emphasize a point, then relaxing back to neutral, in a quiet office during late afternoon, with blurred monitors glowing faintly in the background, lit by soft daylight from a side window with gentle fill on the opposite side and natural falloff across his face. Style: fine skin pores, visible fabric weave, subtle contrast, no gloss or sharpening. Audio: The founder says, 'This update cuts setup time in half, helping teams get started faster.' (no subtitles). Soft office ambience.\"\n\n## Pre-flight: references arrive inline, refer by code\n\nWhen you call `slates_generate_video` with `firstFrameAssetId` / `lastFrameAssetId` / `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. Veo's strongest move is first-frame + last-frame; the pre-flight is where you confirm the two frames actually anchor the motion you wrote. Revise the prompt before `confirm=true` if needed.\n\nWhen talking to the user about the gen, refer to each reference by its short code: `IMG-A12 β€” Founder Headshot`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.\n\n- βœ… \"I'm anchoring on **IMG-A12** as the open shot and **IMG-A18** as the close β€” the 180Β° arc lands on her looking offscreen left.\"\n- ❌ \"I'm using two of the founder shots...\" (which two? They have six.)\n\n## Sources\n\n- [Google Cloud β€” Ultimate Prompting Guide for Veo 3.1](https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1)\n- [Google DeepMind β€” Veo Prompt Guide](https://deepmind.google/models/veo/prompt-guide/)\n- [Google Cloud Docs β€” Vertex AI Video Generation Prompt Guide](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide)\n- [Atlas Cloud β€” Veo 3.1 Master Guide](https://www.atlascloud.ai/blog/guides/google-veo-3-1-guide-master-image-to-video-ai-with-native-sound-and-4k-realism)\n- [Invideo β€” Veo 3.1 Prompt Guide](https://invideo.io/blog/google-veo-prompt-guide/)\n",
32
32
  "slates-restyle-from-blocking": "---\nname: slates-restyle-from-blocking\ndescription: Render one blocking pass as several different visual worlds β€” live action, 2.5D painted, 2D ink, toybox β€” matching cut for cut. Use when a client needs style options, when someone wants to see the same edit in another look, or when an approved edit needs a new treatment without re-blocking.\n---\n\n# Restyle β€” one edit, many worlds\n\nThe commercial payoff of the whole previs workflow, and the reason a blocking file is an asset rather than a step.\n\n## The idea\n\nEvery prompt has two halves:\n\n- **Structure** β€” cuts, camera, timing, who is where. Lives in the blocking clip. **Never changes.**\n- **Style** β€” what any of it looks like. Lives in the references and the prompt text. **Changes freely.**\n\nHold the structure, swap the style, and the same edit comes back as live action, painted 2.5D, ink on paper or a toybox β€” **matching frame for frame across all of them.** Cuts land on the same frames, the car drifts at the same moment, the same head turns at the same beat.\n\nFor anyone pitching work: three visual worlds in a day, off one edit the client has already approved. The foundation is not up for renegotiation, so the conversation is only about look.\n\n## Before you restyle\n\nYou need a blocking clip whose structure you are happy with, and a finished prompt for at least one style (per `slates-blocking-to-prompt`). The first style is the expensive one; every later style is an edit of its text.\n\n## What stays fixed\n\nCopy these across every style **verbatim**. Changing them is what desynchronises the outputs:\n\n- The blocking reference's own contract β€” that it is the master for all movement, the placement-only clause, the tie-break clause, the disambiguation clause\n- The shot count and every timestamp\n- Every shot's camera position, angle, framing and cut point\n- Screen direction and seating\n- The `HOLD FOR THE FULL TIMELINE` block\n- `videoReferenceAssetIds` and `videoReferenceSecondsEach`\n\nLead each style's prompt with a lock so the style layer cannot leak into the structure:\n\n> VIDEO LOCK β€” the dominant rule of this prompt: the reference defines 100% of the motion, editing and object choreography. The text below defines only look, materials, locations and effects layered onto that motion. Wherever the text and the video could be read differently about motion, the video decides.\n\n## What changes\n\n| Layer | What you swap |\n|---|---|\n| Rendering style | photoreal Β· painted 2.5D Β· 2D ink Β· miniature/toybox |\n| Characters | different sheets entirely β€” a couple, grandparents, a robot and a cat |\n| Locations | the same four beats set in a different world |\n| Time of day / weather | night after rain Β· golden hour Β· hard noon |\n| Lighting and colour | per style |\n| Audio | SFX-only, or scored, or lip-synced dialogue |\n\nCharacters can change species and still land, because the blocking only supplies where a body is and how it moves.\n\n## Dummy mapping β€” the mechanism that makes it work\n\nEach style needs its own explicit mapping from grey proxy to real object. The proxy is a slot; the style fills it:\n\n> DUMMY MAPPING: the front-LEFT sphere-head dummy (with its grey arm at the shifter and grey leg at the pedals) is THE GRANDPA; the front-RIGHT sphere-head dummy is THE GRANDMA; a front-seat dummy together with its loose blocks is that ONE whole person. Blocks on the rear bench are the luggage. The low-poly flying model in SHOT 18 is THE HELICOPTER. The two vehicles behind the hero car in SHOT 19 are THE POLICE CARS.\n\nSame clause per style, different right-hand side. And restate the placement-only rule in style terms:\n\n> The source defines only placement and motion, never appearance: every placeholder becomes the real object its position implies β€” spheres are always people, cabin blocks are always cases and bags, fully drawn.\n\n## Location continuity\n\nIf the piece travels, name the places and pin each shot to one. Reusing labels across styles keeps the four prompts diffable:\n\n```\nLOCATION CONTINUITY β€” one journey through four fixed places; each looks\nidentical in every shot where it appears:\n LOC-A <opening> LOC-B <middle> LOC-C <turn> LOC-D <finale>\n```\n\nThen tag every beat: `SHOT 9 β€” 7.79-9.33s β€” LOCKED, LOC-B: <description>`.\n\n**On a piece that visits many places, make the map absolute and countable** β€” otherwise the model reuses a room it liked and you get the same interior three times:\n\n> The location map is absolute β€” SEVEN locations, each appearing EXACTLY ONCE, in this exact order: 1) yard 00:00-00:03.3 … 7) rooftop 00:20-00:30. No location ever appears twice, and the three interiors are three COMPLETELY DIFFERENT rooms β€” different walls, furniture, people and light β€” never the same room repeated.\n\n## Style references\n\nA style reference is **not a keyframe**, and saying so prevents the model reproducing its composition as a shot:\n\n> STYLE MASTER β€” defines the painting and rendering style only: hand-painted look with visible brushstrokes, sculpted painterly volumes, textured matte surfaces, dramatic coloured rim light, deep moody shadows. NOT a keyframe, NOT a location to reproduce, NOT a frame that ever appears in the film. Its own subject, framing and composition are never seen in any shot.\n\nA style can also be **text-only** β€” no reference image at all. Ink and toybox looks usually specify better in words than they match from a still.\n\n## Keep performance inside the existing shots\n\nStyle changes tempt the model to earn new coverage. Refuse it:\n\n> ACTING β€” inside the existing shots only: performance is visible only at the size and distance the reference already gives it, only where the source already shows a face; everywhere else it reads through posture and hands alone. The performance NEVER earns a new shot, a new angle or a closer framing.\n\n## Text-free worlds\n\nStylised worlds are where invented signage and garbled lettering appear. One clause kills it:\n\n> TEXT-FREE WORLD: every sign is a blank painted shape, every gauge face carries tick marks only, every licence plate is a blank plate.\n\n## Running it\n\nGenerate each style as its own `slates_generate_video` call against the **same** `videoReferenceAssetIds`. Keep them in one project so they sit side by side; name assets by style so the comparison reads at a glance.\n\nQuote the whole set before firing β€” `slates_estimate_generation_cost` per style β€” and confirm. Four styles is four generations, not one.\n\n🚨 Never fire a batch of style variants without showing the user the prompts and the total cost first.\n\n## Checklist per style\n\n- [ ] Same blocking asset, same `videoReferenceSecondsEach`\n- [ ] VIDEO LOCK leads the prompt\n- [ ] Every timestamp and shot count identical to style 1\n- [ ] Dummy mapping written for this style's cast\n- [ ] Style reference declared as style-only, or none used\n- [ ] Location labels reused; on a travelling piece the map is absolute and countable\n- [ ] Acting-inside-existing-shots clause present\n- [ ] HOLD block copied verbatim\n- [ ] Cost quoted and confirmed\n\n## Related\n\n`slates-previs-blocking` Β· `slates-blocking-to-prompt` Β· `slates-style-prompting` (style vocabulary per model) Β· `slates-cost-discipline` (batch quoting)\n",
@@ -24,7 +24,7 @@
24
24
  ByteDance's video model β€” first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K β€” 4K video is Pro-only, default 1080p), 4–15s, first+last frame, and up to 9 reference images / 3 videos / 3 audio clips.
25
25
 
26
26
  > **How to read this file.**
27
- > **[official :NNNN]** β€” ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/seedance-2-modelark-docs.md`). Receipt-grade; treat as law.
27
+ > **[official :NNNN]** β€” ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/byteplus-seedance-2-0-api-docs.md`). Receipt-grade; treat as law.
28
28
  > **[community]** β€” third-party guides and our own field notes. Useful, but an `[official]` block always wins.
29
29
  > **[slates]** β€” how the Slates app composes or bills this; not ByteDance doctrine.
30
30
  >
@@ -376,7 +376,7 @@ Stability phrases that help:
376
376
  ## Sources
377
377
 
378
378
  **Official (authoritative):**
379
- - BytePlus ModelArk β€” Seedance 2.0 prompting guide, archived at `research/seedance-2-modelark-docs.md` (all `:NNNN` refs above)
379
+ - BytePlus ModelArk β€” Seedance 2.0 prompting guide, archived at `research/byteplus-seedance-2-0-api-docs.md` (all `:NNNN` refs above)
380
380
 
381
381
  **Community (secondary):**
382
382
  - [fal.ai β€” How to Use Seedance 2.0](https://fal.ai/learn/tools/how-to-use-seedance-2-0)
@@ -12,7 +12,7 @@
12
12
  },
13
13
  {
14
14
  "path": "skills/slates-prompting-seedance.md",
15
- "sha256": "e3210015ed5079fdd1cd21e3e5fea186ad12126015af5f0fd53d65e833004382"
15
+ "sha256": "3f2d1f02096169a34da80a1f8a4ef70fb1b22d3dfc8f95cfbb40be908fc7e031"
16
16
  },
17
17
  {
18
18
  "path": "skills/slates-prompting-kling-v3.md",
@@ -44,8 +44,8 @@
44
44
  },
45
45
  {
46
46
  "path": "reference-seedance.md",
47
- "bytes": 35031,
48
- "sha256": "19f26396f7fc9f688d0084cbae8570594a2eab8035113ecf853dd16e6d02d520"
47
+ "bytes": 35043,
48
+ "sha256": "de7e38a0e54b5c244c2c6be62b74fff0298d1ebf08a77727cc14b093a7487b34"
49
49
  },
50
50
  {
51
51
  "path": "reference-kling.md",
@@ -65,8 +65,8 @@
65
65
  ],
66
66
  "archive": {
67
67
  "path": "slates-prompt-builder.skill",
68
- "bytes": 40494,
69
- "sha256": "4e412da5d81d93ec4f93f88f8478855021df4c8f64292ec84018855078d9a184",
68
+ "bytes": 40497,
69
+ "sha256": "6289a7827816f38fa6cf6c9afe8c6a0d9efa865861a4e732170a40b584edc8a7",
70
70
  "entries": [
71
71
  "SKILL.md",
72
72
  "reference-character.md",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@slatesvideo/shared",
3
- "version": "0.6.6",
3
+ "version": "0.6.7",
4
4
  "description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -263,6 +263,46 @@ has no first/last-frame parameters at all, so this is a shape mismatch rather th
263
263
 
264
264
  ---
265
265
 
266
+ ## Sound: four bracket types, and they are the vendor's syntax
267
+
268
+ ByteDance's 2.5 API tutorial states this as a **prompt rule**, not a suggestion β€” verbatim: *"Use
269
+ special characters to distinguish sounds: `()` for music, `<>` for sound effects, `{}` for dialogue,
270
+ and `【】` for subtitles. For non-Chinese dialogue, it is recommended to specify the language before
271
+ the dialogue."*
272
+
273
+ ```
274
+ She sets the cup down {English: "We open in ten minutes."} <ceramic clink on wood>
275
+ (low piano, unhurried)
276
+ ```
277
+
278
+ - `()` **music** Β· `<>` **sound effects** Β· `{}` **dialogue** Β· `【】` **on-screen subtitles**
279
+ - **Name the language before non-Chinese dialogue.** `{English: "..."}`.
280
+ - Unbracketed sound description still works β€” this is a disambiguator, not a required wrapper. Reach
281
+ for it when one sentence carries more than one kind of sound and you need the model to tell them
282
+ apart, which is exactly where an unmarked prompt puts a line of dialogue into the score.
283
+
284
+ ⚠️ **These four are SEEDANCE 2.5's.** MiniMax H3 has its own three-layer scheme (body / soundscape /
285
+ score) and its angle brackets are documentation notation that must never be typed. Do not carry
286
+ either grammar onto the other model.
287
+
288
+ ## Say what a reference is NOT for
289
+
290
+ The same rule adds a half nobody uses: *"Specify what each asset provides, such as appearance,
291
+ action, or timbre, **and what should not be referenced**."* Negative scoping is a first-class part of
292
+ the citation, not a fallback β€” *"use her face and wardrobe from image 1, not its lighting or
293
+ background"* is a stronger instruction than naming the positive alone, because an unscoped reference
294
+ brings its whole frame with it.
295
+
296
+ 🚨 **The vendor writes `@Image 1`; Slates writes `image 1`, and that difference is deliberate.**
297
+ BytePlus's API tutorial says *"Use `@Image 1`, `@Video 1`, and `@Audio 1`"*, while its own 2.5 prompt
298
+ guide states the bare form (`Image 1 / Video 1 / Audio 1`) in the one normative sentence it has.
299
+ **Two first-party docs, two forms** β€” the disagreement is recorded, not resolved, in
300
+ `second-brain/business/projects/slates/research/model-prompting-research.md`. What settles it FOR US
301
+ is neither: **`@` is a reference-token sigil in the Slates prompt composer, and an unresolved one is
302
+ silently deleted from the prompt before it is sent.** Typing `@Image 1` here does not produce
303
+ `@Image 1`, it produces nothing. The bare form is confirmed working on both models. Never hand-type
304
+ the sigil.
305
+
266
306
  ## Seedance 2.5 Edit (`slates_edit_video`, `model: 'seedance-2.5-edit'`)
267
307
 
268
308
  Its own picker row and its own op call, deliberately: the task type is **the model you chose**,
@@ -34,7 +34,7 @@ description: How to prompt Seedance 2.0 (ByteDance video model). Read before cal
34
34
  ByteDance's video model β€” first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K β€” 4K video is Pro-only, default 1080p), 4–15s, first+last frame, and up to 9 reference images / 3 videos / 3 audio clips.
35
35
 
36
36
  > **How to read this file.**
37
- > **[official :NNNN]** β€” ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/seedance-2-modelark-docs.md`). Receipt-grade; treat as law.
37
+ > **[official :NNNN]** β€” ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/byteplus-seedance-2-0-api-docs.md`). Receipt-grade; treat as law.
38
38
  > **[community]** β€” third-party guides and our own field notes. Useful, but an `[official]` block always wins.
39
39
  > **[slates]** β€” how the Slates app composes or bills this; not ByteDance doctrine.
40
40
  >
@@ -428,7 +428,7 @@ Stability phrases that help:
428
428
  ## Sources
429
429
 
430
430
  **Official (authoritative):**
431
- - BytePlus ModelArk β€” Seedance 2.0 prompting guide, archived at `research/seedance-2-modelark-docs.md` (all `:NNNN` refs above)
431
+ - BytePlus ModelArk β€” Seedance 2.0 prompting guide, archived at `research/byteplus-seedance-2-0-api-docs.md` (all `:NNNN` refs above)
432
432
 
433
433
  **Community (secondary):**
434
434
  - [fal.ai β€” How to Use Seedance 2.0](https://fal.ai/learn/tools/how-to-use-seedance-2-0)