@slatesvideo/shared 0.5.3 → 0.5.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (32) hide show
  1. package/README.md +18 -0
  2. package/dist/operations/index.d.ts +1 -0
  3. package/dist/operations/index.js +7 -2
  4. package/dist/prompts/character-sheet.d.ts +24 -12
  5. package/dist/prompts/character-sheet.js +78 -28
  6. package/dist/prompts/environment-sheet.d.ts +9 -1
  7. package/dist/prompts/environment-sheet.js +17 -3
  8. package/dist/prompts/model-facts.js +4 -1
  9. package/dist/prompts/partials.generated.d.ts +2 -0
  10. package/dist/prompts/partials.generated.js +14 -0
  11. package/dist/prompts/prompting-tips.js +49 -18
  12. package/dist/prompts/reference-rules.d.ts +43 -14
  13. package/dist/prompts/reference-rules.js +50 -26
  14. package/dist/skills/content.js +12 -12
  15. package/package.json +4 -3
  16. package/skills/_partials/decision-log.md +12 -0
  17. package/skills/_partials/reference-rules-core.md +12 -0
  18. package/skills/_partials/reference-tips-short.md +2 -0
  19. package/skills/_partials/references-read-literally.md +11 -0
  20. package/skills/_partials/still-gate.md +3 -0
  21. package/skills/slates-character-turnaround.md +64 -29
  22. package/skills/slates-cost-discipline.md +10 -0
  23. package/skills/slates-edit-and-iterate.md +16 -1
  24. package/skills/slates-model-selection.md +24 -1
  25. package/skills/slates-one-prompt-film.md +19 -0
  26. package/skills/slates-prompting-flux-2-max.md +36 -5
  27. package/skills/slates-prompting-kling-v3.md +33 -4
  28. package/skills/slates-prompting-nano-banana-2.md +40 -10
  29. package/skills/slates-prompting-seedance.md +284 -85
  30. package/skills/slates-prompting-veo-3.md +33 -4
  31. package/skills/slates-storyboard-from-script.md +19 -0
  32. package/skills/slates-vision-feedback-loop.md +49 -2
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@slatesvideo/shared",
3
- "version": "0.5.3",
3
+ "version": "0.5.4",
4
4
  "description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -27,8 +27,9 @@
27
27
  "README.md"
28
28
  ],
29
29
  "scripts": {
30
- "build": "node scripts/embed-skills.mjs && tsc",
31
- "typecheck": "node scripts/embed-skills.mjs && tsc --noEmit",
30
+ "sync-partials": "node scripts/sync-partials.mjs",
31
+ "build": "node scripts/sync-partials.mjs --check && node scripts/embed-skills.mjs && tsc",
32
+ "typecheck": "node scripts/sync-partials.mjs --check && node scripts/embed-skills.mjs && tsc --noEmit",
32
33
  "prepublishOnly": "npm run build"
33
34
  },
34
35
  "repository": {
@@ -0,0 +1,12 @@
1
+ When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:
2
+
3
+ ```
4
+ source phrase or declared default → what you wrote → what it resolves
5
+ "in a diner" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable
6
+ (no time of day) → late afternoon, low warm key → default; say the word and it changes
7
+ (no camera) → slow push-in, single move → one move per shot; stacking increases instability
8
+ ```
9
+
10
+ **Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the "why did you add that?" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.
11
+
12
+ > ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
@@ -0,0 +1,12 @@
1
+ Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
2
+
3
+ 1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
4
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
5
+ 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
6
+ 4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
7
+ 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
8
+ 6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
9
+ 7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
10
+ 8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
11
+ 9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
12
+ 10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
@@ -0,0 +1,2 @@
1
+ <!-- consumer:ts -->
2
+ Name each reference inline; never write role essays. Slates does this for you: `@mention` a subject or environment and it composes `Marcus (images 1 and 2) in the cafe (image 3)`, citing them in the exact order it sends them. Citing both of a character's sheets under the SAME name is what tells the model they are one person — a "Reference Image Instructions" block does the opposite and drags the sheet's studio lighting into your scene. Start with 2-3 focused refs.
@@ -0,0 +1,11 @@
1
+ > **The general law: the model reads a reference literally.**
2
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
3
+
4
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
5
+
6
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
7
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
8
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
9
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
10
+
11
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
@@ -0,0 +1,3 @@
1
+ **A visible defect in the still is already a STOP.** Do not animate it. Fix the frame first, then move to motion — and go to motion only when the crop passes the still scan and you genuinely need movement to confirm an uncertain edge, reflection, or object.
2
+
3
+ This is a **cost** rule as much as a craft rule: a 1080p/10s premium video generation costs many multiples of an image re-roll, and video is where a defect stops being fixable. Anything wrong in the still gets worse in motion — soft geometry mushes, broken-but-plausible objects fall apart, oily textures start crawling. **Animating a known-bad frame is the single most expensive mistake in the pipeline.** Re-rolling the image is the cheap move; re-rolling the video is not.
@@ -1,11 +1,43 @@
1
1
  ---
2
2
  name: slates-character-turnaround
3
- description: Build a Slates character from a reference image — generate the turnaround sheet and expression sheet, bind both to the character's slots so the user sees the character card update live. Use when the user wants to "create a character", "build a character from this image", "generate a turnaround for X", or starts any storyboard-flow that needs consistent character references.
3
+ description: Build a Slates character from a reference image — generate its identity reference sheet and bind it to the character so the card updates live. Use when the user wants to "create a character", "build a character from this image", "generate a turnaround for X", or starts any storyboard flow that needs consistent character references.
4
4
  ---
5
5
 
6
- # Character turnaround — Slates workflow
6
+ # Character identity sheet — Slates workflow
7
7
 
8
- Slates stores characters with two image slots: turnaround (full-body multi-angle) and expression sheet (face close-ups). At scene time Slates attaches BOTH via `@character` mentions — the turnaround for body/proportion/outfit, the expression-sheet close-ups for high-res facial detail. Building them well = consistent character across every frame.
8
+ A character's identity sheet is attached to **every** downstream generation that mentions it, so a flaw in the sheet becomes a flaw in every shot made from it. Building it well is the highest-leverage thing you can do for a project.
9
+
10
+ <!-- @inject:references-read-literally -->
11
+ > **The general law: the model reads a reference literally.**
12
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
13
+
14
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
15
+
16
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
17
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
18
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
19
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
20
+
21
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
22
+ <!-- @end:references-read-literally -->
23
+
24
+ ## The shape: ONE sheet, three panels
25
+
26
+ Slates generates **one identity sheet per character**, bound to the character's turnaround slot:
27
+
28
+ | Panel | What it carries |
29
+ |---|---|
30
+ | **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
31
+ | **Full-body front, relaxed A-pose** | Build, proportion, wardrobe |
32
+ | **Full-body back** | Hair fall and the back of the outfit — the only panel where either reads |
33
+
34
+ On a deep neutral-grey plate (`#3a3a3c`), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look.
35
+
36
+ **The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
37
+
38
+ **Why one sheet and not two.** Every `@character` mention pushes *all* of that character's bound sheets into one reference group, so a two-sheet character costs **two reference slots on every generation**. Against real caps that is brutal — Kling 3.0 takes 4 ingredients (2 characters, zero room for an environment), NB2 has 4 character slots, Seedance 9. One sheet each **doubles the cast you can stage on every model.** It also takes competing facial renderings from six down to two, which is what stops a face from averaging (see the general law above), and it halves the per-character sheet spend.
39
+
40
+ **The expression slot still exists** and is still read — characters built before this change have one bound and keep working. Generate one only when a character genuinely needs a dedicated expression range, and tell the user it costs a reference slot on every shot from then on.
9
41
 
10
42
  ## Workflow
11
43
 
@@ -15,7 +47,7 @@ The user has either:
15
47
  - Described the character in text only.
16
48
 
17
49
  If image: upload it as a reference (`slates_upload_reference_image`).
18
- If text only: skip step 2's reference and generate the turnaround from prompt-only (less consistent — warn the user).
50
+ If text only: generate from prompt-only — less consistent, so warn the user.
19
51
 
20
52
  ### 2. Create the character record
21
53
  `slates_create_character` with:
@@ -23,33 +55,36 @@ If text only: skip step 2's reference and generate the turnaround from prompt-on
23
55
  - `description` — 1-2 sentences, *visual* only ("tall, dark hair, scar over left eye"), not personality.
24
56
  - `style` — leave as the source's own medium by default. Only name a transform if the user wants one (e.g. anime → realistic).
25
57
 
26
- ### 3. Generate the turnaround sheet — the identity anchor
27
- - Use `slates_generate_image` with a prompt like:
28
- > "Character model reference sheet with 4 full body views of the same character: front view, back view, left side profile, right side profile. Neutral pose, neutral expression, consistent appearance across all views. Preserve the artistic medium and visual style of the reference (photo / anime / illustration / 3D / painterly). Render on a plain neutral-grey background with flat, even, shadowless lighting so the sheet captures identity, not scene lighting. No text, no labels, no captions. {character description}"
29
- - **Flat light + plain background is the whole point.** A studio-lit or scene-lit sheet bleeds its lighting into every later generation (the "green-screen-pasted-in-front-of-mountains" failure). Reference prep beats prompting here.
30
- - To transform the medium (e.g. "make her a real person"), append a plain-language instruction; otherwise the source style is preserved.
31
- - Pass the reference image as a reference. Resolution: **2k** (4k wastes credits at sheet scale — no identity gain).
32
- - Estimate cost first.
33
- - When the result returns inline: evaluate against the reference. Same character? Right angle count? If off, refine the prompt with the explicit angle list and regenerate once.
34
- - When right: bind via `slates_set_character_turnaround_asset` (the user sees the card update).
35
-
36
- ### 4. Generate the expression sheet — the close-up face reference
37
- - Prompt:
38
- > "Character expression reference sheet with 3 head-and-shoulder portraits side by side: neutral on left, genuine smile showing teeth in center, serious frown on right. Same character, same flat even shadowless lighting and plain neutral-grey background as the turnaround. {character description}"
39
- - Pass BOTH the original reference AND the just-generated turnaround as references for max consistency.
40
- - Same model, **2k**.
41
- - On success bind via `slates_set_character_expression_asset`.
42
-
43
- ### 5. Hand back
44
- > "Character {name} ready. Turnaround + expressions bound. Use `@{name}` in any prompt — Slates attaches both sheets and names them as one person so the face stays consistent."
45
-
46
- ## Why both sheets (don't gate them)
47
- At scene time Slates attaches the turnaround AND the expression sheet and cites BOTH inline under the same name — `{name} (images N and M)`. That shared NAME — not a withheld expression sheet, and not an injected "use for identity / render neutral" essay — is what tells the model they're ONE person and stops the multiple expressions from averaging the face to a midpoint. (Naming-as-one-entity is each model's own official lever: NB2 "assign a distinct name", Seedance "Reference Subject_N in Image_N".) The close-ups carry far more facial signal (eyes, teeth, skin) than the postage-stamp faces in a full-body turnaround, so attaching both is a fidelity win. Critically, the app injects NO wardrobe/expression/lighting directive — the user's scene prompt owns all of that, so `@{name}` in a movie-still injection keeps the still's own clothing and lighting. The trend is MORE references (video/audio/3D into newer models), all addressed by name — lean into attaching rich refs and let the naming do the work.
58
+ ### 3. Generate the sheet
59
+ `slates_generate_character_sheets` with `characterId`, `projectId`, and `baseAssetId` (the source portrait).
60
+
61
+ **Do not hand-write the sheet prompt.** Slates builds it from the canonical template in `@slatesvideo/shared/prompts` (`buildCharacterTurnaroundPrompt`) — panels, plate, lighting and craft clauses included — and appends your `userNotes`. Use `userNotes` for what the template can't know: *"use the woman on the left"*, *"keep the scar on the right cheek"*. A hand-written prompt is a fork of the template and will drift from it.
62
+
63
+ - Estimate cost first with `slates_estimate_generation_cost` and announce in **credits** — never quote a price from memory. Default is Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted credits.
64
+ - When the result returns inline, **evaluate it before binding**:
65
+ - Is the portrait clearly the largest panel, and is it off-frontal?
66
+ - Is it the same person across all three panels?
67
+ - Catchlights present, irises readable rather than black holes?
68
+ - Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
69
+ - Plate a flat deep grey, not white and not black?
70
+ - If off: one focused refinement, then regenerate. The sheet is upstream of everything — it is worth a re-roll that a scene frame is not.
71
+ - The op binds the result to the turnaround slot automatically.
72
+
73
+ ### 4. Hand back
74
+ > "Character {name} ready — identity sheet bound. Use `@{name}` in any prompt and Slates attaches it and names it inline, so the face stays consistent."
75
+
76
+ ## How the reference gets used at scene time
77
+
78
+ Slates cites the sheet inline under the character's name — `{name} (image N)` — in the exact order it sends references. That **name** is the anti-averaging lever, and it is each model's own official mechanism (NB2: "assign a distinct name"; Seedance: `Reference <Subject_N> in <Image_N>`; Kling: reuse a fixed label verbatim). If a character has both slots bound, both are cited under the *same* name so the model reads them as one person.
79
+
80
+ Critically, the app injects **no** wardrobe, expression, or lighting directive. The user's scene prompt owns all of that — which is why `@{name}` dropped into a movie-still injection keeps the still's own clothing and lighting instead of dragging the sheet's.
48
81
 
49
82
  ## Anti-patterns
50
83
 
51
- - **Don't** studio-light or white-background the sheets. Flat, even, shadowless light on a plain neutral-grey background — or the lighting bleeds into every scene.
52
- - **Don't** generate turnaround and expressions in one prompt. Slates expects them as two separate assets in two separate slots.
84
+ - **Don't** studio-light, white-background, or black-background the sheet. White bleeds into the video and washes out the location; black eats edge detail. Flat, even, shadowless light on a deep neutral grey.
85
+ - **Don't** hand-write the sheet prompt when the op will build it — that is how the template and the shipped prompt fork.
86
+ - **Don't** generate an expression sheet by reflex. It is opt-in now, and it costs a reference slot on every downstream shot.
53
87
  - **Don't** skip binding. The slots are what the storyboard pipeline reads — an unbound asset doesn't help downstream.
54
88
  - **Don't** invent character details. Stick to what's in the reference image and the user's description.
55
- - **Don't** use 4k unless asked — wastes credits, no quality gain at sheet scale.
89
+ - **Don't** use 4K — wastes credits, no quality gain at sheet scale.
90
+ - **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `slates-prompting-seedance`.
@@ -103,6 +103,16 @@ Video gens take minutes (Seedance 4K can run far longer). A client/CLI timeout o
103
103
  - **Poll, don't re-roll.** Use `background: true` on `slates_generate_video`, then poll `slates_get_generation_status` (free, read-only) until it reports `completed` or `failed`. In-flight jobs survive app restarts and are recovered.
104
104
  - A gen has only failed when the status comes back `failed` — and a provider *rejection* **refunds** the credits, so failed isolation tests are ~free. Until you see a terminal status, the job is in flight. Wait.
105
105
 
106
+ ## 🔴 The still-gate — the most expensive mistake in the pipeline
107
+
108
+ <!-- @inject:still-gate -->
109
+ **A visible defect in the still is already a STOP.** Do not animate it. Fix the frame first, then move to motion — and go to motion only when the crop passes the still scan and you genuinely need movement to confirm an uncertain edge, reflection, or object.
110
+
111
+ This is a **cost** rule as much as a craft rule: a 1080p/10s premium video generation costs many multiples of an image re-roll, and video is where a defect stops being fixable. Anything wrong in the still gets worse in motion — soft geometry mushes, broken-but-plausible objects fall apart, oily textures start crawling. **Animating a known-bad frame is the single most expensive mistake in the pipeline.** Re-rolling the image is the cheap move; re-rolling the video is not.
112
+ <!-- @end:still-gate -->
113
+
114
+ The check itself lives in `slates-vision-feedback-loop` (the four slop tells and the per-model accents). The **stop** is a cost rule and belongs here: before every image→video call, confirm the source frame passed the still scan. If it didn't, spending video credits on it is not iteration — it is buying a more expensive copy of a defect you already found.
115
+
106
116
  ## The 3-strike rule
107
117
 
108
118
  Stop after 3 iterations on the same prompt. Hand back to the user with what you tried and what's not working. The slot machine doesn't converge — if it's not landing, the prompt structure is wrong, not the seed.
@@ -7,6 +7,20 @@ description: Iterate on an existing Slates asset — re-evaluate, refine prompt,
7
7
 
8
8
  The user already has a generated image in Slates and wants to refine it. The vision-feedback-loop skill defines the general pattern; this skill is the specific recipe for "I have asset X, here's what's wrong with it."
9
9
 
10
+ ## 🔴 The master rule — an edit is a LEAF, not a node
11
+
12
+ **Never re-edit an edit. Always go back and re-edit the master.**
13
+
14
+ Every edit model silently re-renders the **whole frame**, not just the region you named. So the parts you didn't ask to change come back slightly different every pass — softer texture, drifted colour, mushier fine detail. It is barely visible after one edit and obvious by the second. Chaining edits compounds the damage and there is no way to undo it, because each generation *is* the new source.
15
+
16
+ The fix is structural, not a matter of care:
17
+
18
+ - **Want two changes?** Make them in ONE edit off the master, or make them as two separate edits **both taken from the master**, then keep whichever you prefer.
19
+ - **An edit came back wrong?** Do NOT edit the result to fix it. Discard it and re-edit the master with a better instruction.
20
+ - **Only the changed region is worth keeping?** That is a compositing job — the edit supplies the new region, the untouched master supplies everything else.
21
+
22
+ Slates records this: an edit result carries `sourceAssetIds` pointing at the asset it was made from, so **you can tell whether the thing you are about to edit is itself an edit.** Check before you edit — `[Edit]`-prefixed prompts and a populated source lineage both say "this is a leaf; go back to its parent."
23
+
10
24
  ## Workflow
11
25
 
12
26
  ### 1. Pull the current asset back into context
@@ -45,4 +59,5 @@ The user's request is one of:
45
59
  - **Don't** delete the original asset until the user confirms the new one. Slates keeps both; the user picks.
46
60
  - **Don't** mix surgical and wholesale changes in one regeneration. The user said "make it warmer" — don't also reframe the shot.
47
61
  - **Don't** re-generate when `slates_edit_image` would work. Edits preserve composition and identity; full regen rolls the dice.
48
- - **Don't** chain >3 iterations without checking in. If three tries didn't land, the brief is wrong, not the model.
62
+ - **Don't** edit an edit — ever. Not once, not "just a small one." Go back to the master (see the master rule above). Every attempt re-renders the full frame and the degradation is cumulative and permanent.
63
+ - **Don't** keep re-rolling the same failed edit. If three tries off the master didn't land, the brief is wrong, not the model — check in with the user.
@@ -7,6 +7,18 @@ description: Which model to pick for a given job — the routing doctrine. Read
7
7
 
8
8
  Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.
9
9
 
10
+ ## 🔑 The meta-rule — above the table
11
+
12
+ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:
13
+
14
+ > **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.
15
+ >
16
+ > **Inspect the output at its intended crop.** A frame that holds up as a thumbnail can fall apart at the size it will actually be watched. For a location, look at atmosphere, material texture, and anchor objects; for a character, identity, skin, pose, and gradients.
17
+ >
18
+ > **Choose the model that PROVES that requirement** and leaves only failures you can afford to rerun or mask.
19
+ >
20
+ > **When the roster changes, repeat the evidence test.** Do not carry today's ranking forward on reputation.
21
+
10
22
  ## Video routing
11
23
 
12
24
  | Job | Model | Why |
@@ -16,6 +28,17 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
16
28
  | Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
17
29
  | **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
18
30
  | The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
31
+
32
+ ### Named Seedance escalation triggers
33
+
34
+ "Physics matter" is an abstract category and it under-fires. These are the beats Seedance is **observably** good at — if the shot contains one, escalate without deliberating:
35
+
36
+ - **Real-time → slow-motion contrast.** The signature beat; nearly every strong clip rides it.
37
+ - **The camera moving while debris, meteors, sparks or particles crash around the subject.** Distinctly a feature of this model, not just a thing it survives.
38
+ - **Massive scale that has to read as genuinely huge** — not "a big thing", a thing whose size is the point of the shot.
39
+ - **One continuous unbroken take.**
40
+
41
+ Concrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).
19
42
  | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | The only job Veo wins. |
20
43
 
21
44
  ## Video EDIT routing (changing an existing clip)
@@ -29,7 +52,7 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
29
52
  | AI-edit the user's OWN footage | Omni Flash Edit (3–10s) or Kling O3 Edit (3–15s, 720–3840px) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
30
53
 
31
54
  - **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
32
- - **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the pro demos (e.g. Higgsfield's split-screen short) actually work, plus gesture-only beats with voiceover laid over in post.
55
+ - **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.
33
56
  - **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
34
57
  - Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
35
58
 
@@ -12,6 +12,25 @@ The user gives an idea. You hand back an MP4 on disk. Everything in between is y
12
12
  ### 1. Script the beats
13
13
  Turn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.
14
14
 
15
+ **Surface a decision log with the plan.**
16
+
17
+ <!-- @inject:decision-log -->
18
+ When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:
19
+
20
+ ```
21
+ source phrase or declared default → what you wrote → what it resolves
22
+ "in a diner" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable
23
+ (no time of day) → late afternoon, low warm key → default; say the word and it changes
24
+ (no camera) → slow push-in, single move → one move per shot; stacking increases instability
25
+ ```
26
+
27
+ **Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the "why did you add that?" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.
28
+
29
+ > ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
30
+ <!-- @end:decision-log -->
31
+
32
+ A 4-10 shot script is where you invent the most on the user's behalf — time of day, wardrobe, weather, lens feel, camera moves the brief never mentioned. The log is what makes those visible while they are still free to change.
33
+
15
34
  ### 2. Set up the project
16
35
  - `slates_create_project` named for the piece.
17
36
  - Recurring character? Build it properly — `slates_create_character` + the `slates-character-turnaround` recipe — so every frame references the same turnaround.
@@ -82,11 +82,42 @@ Use natural language for exploration, JSON when the layout is locked and you're
82
82
 
83
83
  In Slates, pass `referenceAssetIds` on `slates_generate_image` — FLUX routes them through its edit endpoint. Slates names each reference inline in the prompt ("the subject (image 1), the style (image 2)") in the order it sends them, so you don't hand-write role labels; the name carries the role and unnamed-by-position blending is avoided. For surgical changes to one existing image use `slates_edit_image` with `editModel: flux-2-max` (note: FLUX edits ignore extra referenceAssetIds — that's NB2-only).
84
84
 
85
- Reference discipline (FLUX caps refs lower than NB2's 14, so be deliberate):
86
- - **2-4 strong refs**, one per role, named — not 1 (warps), not many (blends).
87
- - **Flat-lit identity refs** — a studio-lit / scene-lit character sheet bleeds its lighting into the output.
88
- - **Attach both character sheets, named as one entity** — turnaround (body/proportion/outfit) + close-up expression sheet (face detail), cited under the same name; the shared name keeps the expressions from averaging the face. Don't write a role essay or "render neutral" instruction — the user's prompt owns the expression, wardrobe, and lighting.
89
- - **Environment: describe it, don't feed a multi-panel grid** — reserve a ref for a hard exact-match, then use ONE clean establishing image.
85
+ ### Reference rules (the verified ones)
86
+
87
+ <!-- @inject:references-read-literally -->
88
+ > **The general law: the model reads a reference literally.**
89
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
90
+
91
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
92
+
93
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
94
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
95
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
96
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
97
+
98
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
99
+ <!-- @end:references-read-literally -->
100
+
101
+ <!-- @inject:reference-rules-core -->
102
+ Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
103
+
104
+ 1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
105
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
106
+ 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
107
+ 4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
108
+ 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
109
+ 6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
110
+ 7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
111
+ 8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
112
+ 9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
113
+ 10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
114
+ <!-- @end:reference-rules-core -->
115
+
116
+ ### For FLUX.2 Max specifically
117
+
118
+ - **FLUX caps references well below NB2's 14, so rule 1's "2-4" is a ceiling here, not a starting point.** Be deliberate about which roles earn a slot.
119
+ - **Rule 9 has a hard edge on this model:** `slates_edit_image` with `editModel: flux-2-max` ignores extra `referenceAssetIds` — that is NB2-only. A FLUX edit sees the source image and the prompt, nothing else.
120
+ - **FLUX has no memory between generations, so rule 7 is enforced by repetition.** Define the character exhaustively once and repeat those exact descriptors verbatim in every subsequent prompt — see Character consistency across a series below.
90
121
 
91
122
  ## Character consistency across a series
92
123
 
@@ -106,10 +106,39 @@ Upload 2-4 multi-angle reference photos per character/object. Tag inline:
106
106
 
107
107
  ## Reference discipline (character / environment refs)
108
108
 
109
- - **2-4 strong refs per role**, named (the same fixed label reused verbatim) and reused across every shot — swapping mid-sequence drifts. Kling's consistency lever is **"lock the subject with a fixed label reused verbatim"** (pronoun/synonym drift breaks it), so reusing the exact name on every mention is the whole game. Slates composes this for you from `@mentions`.
110
- - **Flat-lit identity refs.** A studio-lit / scene-lit character sheet bleeds its lighting into the clip. Prep refs flat and plain.
111
- - **Attach both character sheets, named as one entity** — the turnaround (body/proportion/outfit) and the close-up expression sheet (face detail), cited under the same name. The shared name keeps the varied expressions from averaging the face; don't write a role essay or tell it to "render neutral" — the user's prompt owns the expression, wardrobe, and lighting.
112
- - **Environment: describe it, don't feed a multi-panel grid.** Reserve an environment ref for a hard exact-match, then use ONE clean establishing image.
109
+ <!-- @inject:references-read-literally -->
110
+ > **The general law: the model reads a reference literally.**
111
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
112
+
113
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
114
+
115
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
116
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
117
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
118
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
119
+
120
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
121
+ <!-- @end:references-read-literally -->
122
+
123
+ <!-- @inject:reference-rules-core -->
124
+ Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
125
+
126
+ 1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
127
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
128
+ 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
129
+ 4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
130
+ 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
131
+ 6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
132
+ 7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
133
+ 8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
134
+ 9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
135
+ 10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
136
+ <!-- @end:reference-rules-core -->
137
+
138
+ ### For Kling specifically
139
+
140
+ - **Kling's consistency lever is "lock the subject with a fixed label reused verbatim."** That is Kling's phrasing for rules 2 and 3, and it is stricter than the others: **pronoun and synonym drift breaks it**, so the exact same label must appear on every single mention — not "he", not "the detective" after you named him. Reusing the label verbatim is the whole game. Slates composes this for you from `@mentions`.
141
+ - **Element references are the transport for rule 1** — 2-4 multi-angle photos per character/object, tagged `@element1` / `@element2` (see Element references above). The cap is 4 combined refs on the edit path.
113
142
 
114
143
  ## Negative prompting — has a real field
115
144
 
@@ -1,11 +1,11 @@
1
1
  ---
2
2
  name: slates-prompting-nano-banana-2
3
- description: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3 Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.
3
+ description: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3.1 Flash Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.
4
4
  ---
5
5
 
6
6
  # Nano Banana 2 — cinematic & photorealistic prompting
7
7
 
8
- The **default** model behind `slates_generate_image` is **Gemini 3 Image** (Nano Banana 2 / Flash) — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill. NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
8
+ The **default** model behind `slates_generate_image` is **Gemini 3.1 Flash Image** (Nano Banana 2) — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill. It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat. Verified against the runtime slug map in `slate/src/main/api/google.ts`. NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
9
9
 
10
10
  Knowledge cutoff: January 2025. Anything after needs explicit reference images.
11
11
 
@@ -30,6 +30,10 @@ Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and a
30
30
 
31
31
  ## Photorealism positives — what consistently works
32
32
 
33
+ > ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern — do not carry it across.**
34
+ > Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.
35
+ > The leak happens in one specific way — you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `slates-prompting-seedance` (Part 3, "Don't cross-pollinate image-model syntax").
36
+
33
37
  **Named lenses + apertures** beat generic "shallow depth of field":
34
38
  - `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`
35
39
  - `Panavision anamorphic` for horizontal flares + cinematic width
@@ -99,14 +103,40 @@ Default to #1. Reach for #2 only when positive framing can't suppress the unwant
99
103
  - **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style (or pass `referenceAssetIds`), Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (images 1 and 2) sits across from the woman (images 3 and 4) in the cafe (image 5)`, with a trailing `Render in the visual style of image 6.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**, so citing both of a subject's sheets under the SAME name ("Marcus") is what tells the model they are ONE person and stops the face averaging. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
100
104
 
101
105
  ### Reference rules (the verified ones)
102
- 1. **2-4 strong refs beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each adds context AND variables to balance.
103
- 2. **One reference per ROLE, named** (identity / style-grade / environment). Same-role competitors drift. The model doesn't infer roles from order — the inline name does it.
104
- 3. **Identity refs: attach both sheets, named as one entity — don't gate them.** A character's turnaround (body/proportion/outfit) AND its close-up expression sheet (high-res face: eyes, skin, teeth) both go in, cited under the SAME name ("Marcus (images 1 and 2)"). That shared name — not a role essay — is what stops the varied expressions from averaging the face. An *unnamed* expression sheet hurts; named as one entity, the close-ups are a fidelity win.
105
- 4. **Flat-light identity refs.** Prep them with flat, even, shadowless lighting on a plain neutral background. Studio-lit / scene-lit sheets bleed their lighting into the generation ("green-screen pasted in front of mountains").
106
- 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words. Reserve an environment ref for a mandatory exact-match, and then use ONE clean establishing image — never a multi-panel grid fed whole.
107
- 6. **Grids: explore, don't input.** Use grids to explore compositions, then pick a cell. Never feed a grid back in as a reference — cells share a split detail budget, so flaws propagate.
108
- 7. **Reuse the same refs across all shots.** Swapping mid-sequence causes drift.
109
- 8. **Legible in-shot text → bake it into the NB2 start frame**, then animate from it. Never trust text-to-video to render clean text.
106
+
107
+ <!-- @inject:references-read-literally -->
108
+ > **The general law: the model reads a reference literally.**
109
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
110
+
111
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
112
+
113
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
114
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
115
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
116
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
117
+
118
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
119
+ <!-- @end:references-read-literally -->
120
+
121
+ <!-- @inject:reference-rules-core -->
122
+ Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
123
+
124
+ 1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
125
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
126
+ 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
127
+ 4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
128
+ 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
129
+ 6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
130
+ 7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
131
+ 8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
132
+ 9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
133
+ 10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
134
+ <!-- @end:reference-rules-core -->
135
+
136
+ ### For Nano Banana 2 specifically
137
+
138
+ - **NB2's own consistency lever is "assign a distinct name to each character/object."** That is Google's phrasing for rule 3 — citing both of a subject's sheets under the SAME name is the officially-sanctioned mechanism, not a workaround.
139
+ - **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.
110
140
  - **Character consistency is officially "not 100% perfect"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.
111
141
  - **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.
112
142