@slatesvideo/shared 0.5.3 → 0.5.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/README.md +31 -0
  2. package/dist/operations/index.d.ts +4 -8
  3. package/dist/operations/index.js +26 -38
  4. package/dist/prompts/character-sheet.d.ts +19 -18
  5. package/dist/prompts/character-sheet.js +84 -38
  6. package/dist/prompts/environment-sheet.d.ts +9 -1
  7. package/dist/prompts/environment-sheet.js +17 -3
  8. package/dist/prompts/model-facts.js +4 -1
  9. package/dist/prompts/partials.generated.d.ts +2 -0
  10. package/dist/prompts/partials.generated.js +14 -0
  11. package/dist/prompts/prompting-tips.js +50 -19
  12. package/dist/prompts/reference-composer.d.ts +1 -1
  13. package/dist/prompts/reference-composer.js +3 -4
  14. package/dist/prompts/reference-rules.d.ts +43 -14
  15. package/dist/prompts/reference-rules.js +51 -27
  16. package/dist/skills/content.js +14 -14
  17. package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
  18. package/exports/slates-prompt-builder/generated/reference-character.md +78 -0
  19. package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
  20. package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
  21. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
  22. package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
  23. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
  24. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  25. package/package.json +8 -3
  26. package/skills/_partials/decision-log.md +12 -0
  27. package/skills/_partials/reference-rules-core.md +12 -0
  28. package/skills/_partials/reference-tips-short.md +2 -0
  29. package/skills/_partials/references-read-literally.md +11 -0
  30. package/skills/_partials/still-gate.md +3 -0
  31. package/skills/slates-character-identity.md +100 -0
  32. package/skills/slates-cost-discipline.md +10 -0
  33. package/skills/slates-edit-and-iterate.md +17 -2
  34. package/skills/slates-model-selection.md +24 -1
  35. package/skills/slates-one-prompt-film.md +22 -3
  36. package/skills/slates-prompting-flux-2-max.md +36 -5
  37. package/skills/slates-prompting-gpt-image-2.md +1 -1
  38. package/skills/slates-prompting-kling-v3.md +40 -9
  39. package/skills/slates-prompting-nano-banana-2.md +44 -12
  40. package/skills/slates-prompting-omni-flash.md +1 -1
  41. package/skills/slates-prompting-seedance.md +295 -90
  42. package/skills/slates-prompting-veo-3.md +33 -4
  43. package/skills/slates-storyboard-from-script.md +19 -0
  44. package/skills/slates-vision-feedback-loop.md +49 -2
  45. package/skills/slates-character-turnaround.md +0 -55
@@ -0,0 +1,100 @@
1
+ ---
2
+ name: slates-character-identity
3
+ description: Build a Slates character from a reference image — generate one identity sheet and bind it to the character so the card updates live. Use when the user wants to create a character, build a character from an image, or starts a storyboard flow that needs consistent character references.
4
+ ---
5
+
6
+ # Character identity sheet — Slates workflow
7
+
8
+ A character's identity sheet is attached to **every** downstream generation that mentions it, so a flaw in the sheet becomes a flaw in every shot made from it. Building it well is the highest-leverage thing you can do for a project.
9
+
10
+ <!-- @inject:references-read-literally -->
11
+ > **The general law: the model reads a reference literally.**
12
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
13
+
14
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
15
+
16
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
17
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
18
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
19
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
20
+
21
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
22
+ <!-- @end:references-read-literally -->
23
+
24
+ ## The shape: ONE sheet, three panels
25
+
26
+ Slates generates **one identity sheet per character**, bound as the character's canonical reference:
27
+
28
+ | Panel | What it carries |
29
+ |---|---|
30
+ | **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
31
+ | **Full-body front, relaxed A-pose — framed from the collarbone down, head not shown** | Build, proportion, wardrobe. Headless on purpose: a front-facing body panel renders a ~40px face that can't match the portrait's, so the sheet would carry two competing identities and the model averages them. |
32
+ | **Full-body back, head and hair visible** | Hair fall and the back of the outfit — the only panel where either reads. Keeps its head because there's no face to compete with. |
33
+
34
+ The rule is **kill every competing rendering of the FACE, not every head** — which is why exactly one body panel is headless.
35
+
36
+ On a deep neutral-grey plate (`#3a3a3c`), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look. Quadrupeds and non-bipedal characters are carved out — natural standing stance, head shown on both body panels.
37
+
38
+ **The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
39
+
40
+ **Why one sheet.** Every `@character` mention attaches that character's canonical identity image, so each character costs one reference slot. It also reduces competing facial renderings to **one** — with the front panel headless and the back panel turned away, the portrait is the only face on the sheet, so there is nothing left to average.
41
+
42
+ ## Workflow
43
+
44
+ ### Get the reference
45
+ The user has either:
46
+ - Pasted/uploaded an image of the character (real person, drawing, AI render).
47
+ - Described the character in text only.
48
+
49
+ If image: upload it as a reference<!-- slates-only --> (`slates_upload_reference_image`)<!-- /slates-only -->.
50
+ If text only: generate from prompt-only — less consistent, so warn the user.
51
+
52
+ <!-- slates-only -->
53
+ ### Create the character record
54
+ `slates_create_character` with:
55
+ - `name` (ask if not given)
56
+ - `description` — 1-2 sentences, *visual* only ("tall, dark hair, scar over left eye"), not personality.
57
+ - `style` — leave as the source's own medium by default. Only name a transform if the user wants one (e.g. anime → realistic).
58
+ <!-- /slates-only -->
59
+
60
+ ### Generate the sheet
61
+ <!-- slates-only -->
62
+ `slates_generate_character_identity` with `characterId`, `projectId`, and `baseAssetId` (the source portrait).
63
+
64
+ **Do not hand-write the sheet prompt.** Slates builds it from the canonical template in `@slatesvideo/shared/prompts` (`buildCharacterIdentityPrompt`) — panels, plate, lighting and craft clauses included — and appends your `userNotes`. Use `userNotes` for what the template can't know: *"use the woman on the left"*, *"keep the scar on the right cheek"*. A hand-written prompt is a fork of the template and will drift from it.
65
+
66
+ - Estimate cost first with `slates_estimate_generation_cost` and announce in **credits** — never quote a price from memory.
67
+ <!-- /slates-only -->
68
+
69
+ - Default to Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted spend.
70
+ - When the result returns inline, **evaluate it before binding**:
71
+ - Is the portrait clearly the largest panel, and is it off-frontal?
72
+ - **Is the front body panel cleanly headless** — an empty collar with the garment holding its shape, no partial face, no floating jaw, no smeared neck stump? A botched crop is worse than no crop.
73
+ - Do the body panels read as the same build, wardrobe and hair as the portrait?
74
+ - Catchlights present, irises readable rather than black holes?
75
+ - Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
76
+ - Plate a flat deep grey, not white and not black?
77
+ - If off: one focused refinement, then regenerate. The sheet is upstream of everything — it is worth a re-roll that a scene frame is not.
78
+ <!-- slates-only -->
79
+ - The op binds the result as the canonical identity automatically.
80
+
81
+ ### Hand back
82
+ > "Character {name} ready — identity sheet bound. Use `@{name}` in any prompt and Slates attaches it and names it inline, so the face stays consistent."
83
+ <!-- /slates-only -->
84
+
85
+ ## How the reference gets used at scene time
86
+
87
+ Slates cites the sheet inline under the character's name — `{name} (image N)` — in the exact order it sends references. That **name** is the anti-averaging lever, and it is each model's own official mechanism (NB2: "assign a distinct name"; Seedance: `Reference <Subject_N> in <Image_N>`; Kling: reuse a fixed label verbatim).
88
+
89
+ Critically, the app injects **no** wardrobe, expression, or lighting directive. The user's scene prompt owns all of that — which is why `@{name}` dropped into a movie-still injection keeps the still's own clothing and lighting instead of dragging the sheet's.
90
+
91
+ ## Anti-patterns
92
+
93
+ - **Don't** studio-light, white-background, or black-background the sheet. White bleeds into the video and washes out the location; black eats edge detail. Flat, even, shadowless light on a deep neutral grey.
94
+ - **Don't** hand-write the sheet prompt when the op will build it — that is how the template and the shipped prompt fork.
95
+ - **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
96
+ - **Don't** skip binding. An unbound asset doesn't help downstream.
97
+ - **Don't** invent character details. Stick to what's in the reference image and the user's description.
98
+ - **Don't** describe the headless front panel as removal or decapitation — in `userNotes` or any hand-written variant. The template asks for it as *framing* — "cropped at the collarbone, head not shown, invisible-mannequin presentation" — which is a standard e-commerce genre with deep training data. Removal phrasing is untested and invites a refusal.
99
+ - **Don't** use 4K — wastes credits, no quality gain at sheet scale.
100
+ - **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `slates-prompting-seedance`.
@@ -103,6 +103,16 @@ Video gens take minutes (Seedance 4K can run far longer). A client/CLI timeout o
103
103
  - **Poll, don't re-roll.** Use `background: true` on `slates_generate_video`, then poll `slates_get_generation_status` (free, read-only) until it reports `completed` or `failed`. In-flight jobs survive app restarts and are recovered.
104
104
  - A gen has only failed when the status comes back `failed` — and a provider *rejection* **refunds** the credits, so failed isolation tests are ~free. Until you see a terminal status, the job is in flight. Wait.
105
105
 
106
+ ## 🔴 The still-gate — the most expensive mistake in the pipeline
107
+
108
+ <!-- @inject:still-gate -->
109
+ **A visible defect in the still is already a STOP.** Do not animate it. Fix the frame first, then move to motion — and go to motion only when the crop passes the still scan and you genuinely need movement to confirm an uncertain edge, reflection, or object.
110
+
111
+ This is a **cost** rule as much as a craft rule: a 1080p/10s premium video generation costs many multiples of an image re-roll, and video is where a defect stops being fixable. Anything wrong in the still gets worse in motion — soft geometry mushes, broken-but-plausible objects fall apart, oily textures start crawling. **Animating a known-bad frame is the single most expensive mistake in the pipeline.** Re-rolling the image is the cheap move; re-rolling the video is not.
112
+ <!-- @end:still-gate -->
113
+
114
+ The check itself lives in `slates-vision-feedback-loop` (the four slop tells and the per-model accents). The **stop** is a cost rule and belongs here: before every image→video call, confirm the source frame passed the still scan. If it didn't, spending video credits on it is not iteration — it is buying a more expensive copy of a defect you already found.
115
+
106
116
  ## The 3-strike rule
107
117
 
108
118
  Stop after 3 iterations on the same prompt. Hand back to the user with what you tried and what's not working. The slot machine doesn't converge — if it's not landing, the prompt structure is wrong, not the seed.
@@ -7,6 +7,20 @@ description: Iterate on an existing Slates asset — re-evaluate, refine prompt,
7
7
 
8
8
  The user already has a generated image in Slates and wants to refine it. The vision-feedback-loop skill defines the general pattern; this skill is the specific recipe for "I have asset X, here's what's wrong with it."
9
9
 
10
+ ## 🔴 The master rule — an edit is a LEAF, not a node
11
+
12
+ **Never re-edit an edit. Always go back and re-edit the master.**
13
+
14
+ Every edit model silently re-renders the **whole frame**, not just the region you named. So the parts you didn't ask to change come back slightly different every pass — softer texture, drifted colour, mushier fine detail. It is barely visible after one edit and obvious by the second. Chaining edits compounds the damage and there is no way to undo it, because each generation *is* the new source.
15
+
16
+ The fix is structural, not a matter of care:
17
+
18
+ - **Want two changes?** Make them in ONE edit off the master, or make them as two separate edits **both taken from the master**, then keep whichever you prefer.
19
+ - **An edit came back wrong?** Do NOT edit the result to fix it. Discard it and re-edit the master with a better instruction.
20
+ - **Only the changed region is worth keeping?** That is a compositing job — the edit supplies the new region, the untouched master supplies everything else.
21
+
22
+ Slates records this: an edit result carries `sourceAssetIds` pointing at the asset it was made from, so **you can tell whether the thing you are about to edit is itself an edit.** Check before you edit — `[Edit]`-prefixed prompts and a populated source lineage both say "this is a leaf; go back to its parent."
23
+
10
24
  ## Workflow
11
25
 
12
26
  ### 1. Pull the current asset back into context
@@ -33,7 +47,7 @@ The user's request is one of:
33
47
  ### 4. Generate, evaluate, decide
34
48
  - Estimate cost first.
35
49
  - After generation, the result is inline. Compare side-by-side with the original (`slates_get_asset_image` again).
36
- - If the delta is correct: bind to the same slot (frame, character turnaround, etc.) the original was bound to.
50
+ - If the delta is correct: bind to the same role (frame, character identity, etc.) the original was bound to.
37
51
  - If the delta missed: one focused refinement, then regenerate. Cap at 3 tries.
38
52
 
39
53
  ### 5. Hand back
@@ -45,4 +59,5 @@ The user's request is one of:
45
59
  - **Don't** delete the original asset until the user confirms the new one. Slates keeps both; the user picks.
46
60
  - **Don't** mix surgical and wholesale changes in one regeneration. The user said "make it warmer" — don't also reframe the shot.
47
61
  - **Don't** re-generate when `slates_edit_image` would work. Edits preserve composition and identity; full regen rolls the dice.
48
- - **Don't** chain >3 iterations without checking in. If three tries didn't land, the brief is wrong, not the model.
62
+ - **Don't** edit an edit — ever. Not once, not "just a small one." Go back to the master (see the master rule above). Every attempt re-renders the full frame and the degradation is cumulative and permanent.
63
+ - **Don't** keep re-rolling the same failed edit. If three tries off the master didn't land, the brief is wrong, not the model — check in with the user.
@@ -7,6 +7,18 @@ description: Which model to pick for a given job — the routing doctrine. Read
7
7
 
8
8
  Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.
9
9
 
10
+ ## 🔑 The meta-rule — above the table
11
+
12
+ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:
13
+
14
+ > **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.
15
+ >
16
+ > **Inspect the output at its intended crop.** A frame that holds up as a thumbnail can fall apart at the size it will actually be watched. For a location, look at atmosphere, material texture, and anchor objects; for a character, identity, skin, pose, and gradients.
17
+ >
18
+ > **Choose the model that PROVES that requirement** and leaves only failures you can afford to rerun or mask.
19
+ >
20
+ > **When the roster changes, repeat the evidence test.** Do not carry today's ranking forward on reputation.
21
+
10
22
  ## Video routing
11
23
 
12
24
  | Job | Model | Why |
@@ -16,6 +28,17 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
16
28
  | Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
17
29
  | **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
18
30
  | The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
31
+
32
+ ### Named Seedance escalation triggers
33
+
34
+ "Physics matter" is an abstract category and it under-fires. These are the beats Seedance is **observably** good at — if the shot contains one, escalate without deliberating:
35
+
36
+ - **Real-time → slow-motion contrast.** The signature beat; nearly every strong clip rides it.
37
+ - **The camera moving while debris, meteors, sparks or particles crash around the subject.** Distinctly a feature of this model, not just a thing it survives.
38
+ - **Massive scale that has to read as genuinely huge** — not "a big thing", a thing whose size is the point of the shot.
39
+ - **One continuous unbroken take.**
40
+
41
+ Concrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).
19
42
  | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | The only job Veo wins. |
20
43
 
21
44
  ## Video EDIT routing (changing an existing clip)
@@ -29,7 +52,7 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
29
52
  | AI-edit the user's OWN footage | Omni Flash Edit (3–10s) or Kling O3 Edit (3–15s, 720–3840px) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
30
53
 
31
54
  - **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
32
- - **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the pro demos (e.g. Higgsfield's split-screen short) actually work, plus gesture-only beats with voiceover laid over in post.
55
+ - **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.
33
56
  - **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
34
57
  - Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
35
58
 
@@ -12,9 +12,28 @@ The user gives an idea. You hand back an MP4 on disk. Everything in between is y
12
12
  ### 1. Script the beats
13
13
  Turn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.
14
14
 
15
+ **Surface a decision log with the plan.**
16
+
17
+ <!-- @inject:decision-log -->
18
+ When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:
19
+
20
+ ```
21
+ source phrase or declared default → what you wrote → what it resolves
22
+ "in a diner" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable
23
+ (no time of day) → late afternoon, low warm key → default; say the word and it changes
24
+ (no camera) → slow push-in, single move → one move per shot; stacking increases instability
25
+ ```
26
+
27
+ **Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the "why did you add that?" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.
28
+
29
+ > ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
30
+ <!-- @end:decision-log -->
31
+
32
+ A 4-10 shot script is where you invent the most on the user's behalf — time of day, wardrobe, weather, lens feel, camera moves the brief never mentioned. The log is what makes those visible while they are still free to change.
33
+
15
34
  ### 2. Set up the project
16
35
  - `slates_create_project` named for the piece.
17
- - Recurring character? Build it properly — `slates_create_character` + the `slates-character-turnaround` recipe — so every frame references the same turnaround.
36
+ - Recurring character? Build it properly — `slates_create_character` + the `slates-character-identity` recipe — so every frame references the same identity.
18
37
  - Recurring location? `slates_create_environment`.
19
38
  - One-off shots don't need character/environment records; skip the ceremony.
20
39
 
@@ -30,7 +49,7 @@ Price the whole batch before the first generation: frame images (count × model
30
49
  Per `slates-cost-discipline` 3b: that single OK authorizes `confirm=true` for **every enumerated call in the batch** — no per-call re-asking. Re-confirm only if a call's price overruns the plan >25% or new calls get added (extra retakes, new shots).
31
50
 
32
51
  ### 5. Generate frame images
33
- Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the character turnaround / environment / prior frames for consistency (Slates names each reference inline as "image N" — you don't hand-write role labels; reuse the same subject name across shots). Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`.
52
+ Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the character identity / environment / prior frames for consistency (Slates names each reference inline as "image N" — you don't hand-write role labels; reuse the same subject name across shots). Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`.
34
53
 
35
54
  **Multi-take where it matters:** for the hook shot and any shot the whole film hangs on, generate 2-4 variants (cheap model or 1k), pull them back with `slates_get_assets_batch`, pick the strongest on composition + identity, discard the rest. Don't multi-take filler shots.
36
55
 
@@ -64,4 +83,4 @@ Shots delivered, total spent vs. approved plan, the export path, and the single
64
83
  - **Skeleton before spend.** Project + storyboard structure are free; generation isn't.
65
84
  - **Look at everything.** Every image inline, every video via `slates_get_asset_video_frames` if a clip seems off. Never assemble a timeline from clips you haven't evaluated.
66
85
  - **3-strike rule per shot.** Three failed takes on one shot = stop, show the user what you tried, ask.
67
- - **Consistency comes from references, not luck.** Same turnaround asset on every character frame; same environment refs across a location's shots.
86
+ - **Consistency comes from references, not luck.** Same identity asset on every character frame; same environment refs across a location's shots.
@@ -82,11 +82,42 @@ Use natural language for exploration, JSON when the layout is locked and you're
82
82
 
83
83
  In Slates, pass `referenceAssetIds` on `slates_generate_image` — FLUX routes them through its edit endpoint. Slates names each reference inline in the prompt ("the subject (image 1), the style (image 2)") in the order it sends them, so you don't hand-write role labels; the name carries the role and unnamed-by-position blending is avoided. For surgical changes to one existing image use `slates_edit_image` with `editModel: flux-2-max` (note: FLUX edits ignore extra referenceAssetIds — that's NB2-only).
84
84
 
85
- Reference discipline (FLUX caps refs lower than NB2's 14, so be deliberate):
86
- - **2-4 strong refs**, one per role, named — not 1 (warps), not many (blends).
87
- - **Flat-lit identity refs** — a studio-lit / scene-lit character sheet bleeds its lighting into the output.
88
- - **Attach both character sheets, named as one entity** — turnaround (body/proportion/outfit) + close-up expression sheet (face detail), cited under the same name; the shared name keeps the expressions from averaging the face. Don't write a role essay or "render neutral" instruction — the user's prompt owns the expression, wardrobe, and lighting.
89
- - **Environment: describe it, don't feed a multi-panel grid** — reserve a ref for a hard exact-match, then use ONE clean establishing image.
85
+ ### Reference rules (the verified ones)
86
+
87
+ <!-- @inject:references-read-literally -->
88
+ > **The general law: the model reads a reference literally.**
89
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
90
+
91
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
92
+
93
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
94
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
95
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
96
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
97
+
98
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
99
+ <!-- @end:references-read-literally -->
100
+
101
+ <!-- @inject:reference-rules-core -->
102
+ Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
103
+
104
+ 1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
105
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
106
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
107
+ 4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
108
+ 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
109
+ 6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
110
+ 7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
111
+ 8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
112
+ 9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
113
+ 10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
114
+ <!-- @end:reference-rules-core -->
115
+
116
+ ### For FLUX.2 Max specifically
117
+
118
+ - **FLUX caps references well below NB2's 14, so rule 1's "2-4" is a ceiling here, not a starting point.** Be deliberate about which roles earn a slot.
119
+ - **Rule 9 has a hard edge on this model:** `slates_edit_image` with `editModel: flux-2-max` ignores extra `referenceAssetIds` — that is NB2-only. A FLUX edit sees the source image and the prompt, nothing else.
120
+ - **FLUX has no memory between generations, so rule 7 is enforced by repetition.** Define the character exhaustively once and repeat those exact descriptors verbatim in every subsequent prompt — see Character consistency across a series below.
90
121
 
91
122
  ## Character consistency across a series
92
123
 
@@ -29,7 +29,7 @@ Never rely on the provider default (it's high — the priciest tier). The Slates
29
29
 
30
30
  - State the grid explicitly and number the cells: "a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom".
31
31
  - Give each cell ONE content clause: "Panel 3: the character mid-jump, side view".
32
- - Character sheets: "character turnaround sheet: front, 3/4 left, profile, back — same character, same outfit, flat even lighting, plain background". GPT Image 2 holds the layout; the Banana line holds the *face* better — for identity-critical turnarounds prefer NB2/NB Pro and use GPT Image 2 when labels/annotations matter.
32
+ - Character identity sheets: GPT Image 2 holds structured panel layouts; the Banana line holds the *face* better. Prefer NB2/NB Pro for identity-critical sheets and GPT Image 2 when labels or annotations are the main requirement.
33
33
 
34
34
  ## References & editing
35
35
 
@@ -106,10 +106,39 @@ Upload 2-4 multi-angle reference photos per character/object. Tag inline:
106
106
 
107
107
  ## Reference discipline (character / environment refs)
108
108
 
109
- - **2-4 strong refs per role**, named (the same fixed label reused verbatim) and reused across every shot — swapping mid-sequence drifts. Kling's consistency lever is **"lock the subject with a fixed label reused verbatim"** (pronoun/synonym drift breaks it), so reusing the exact name on every mention is the whole game. Slates composes this for you from `@mentions`.
110
- - **Flat-lit identity refs.** A studio-lit / scene-lit character sheet bleeds its lighting into the clip. Prep refs flat and plain.
111
- - **Attach both character sheets, named as one entity** — the turnaround (body/proportion/outfit) and the close-up expression sheet (face detail), cited under the same name. The shared name keeps the varied expressions from averaging the face; don't write a role essay or tell it to "render neutral" — the user's prompt owns the expression, wardrobe, and lighting.
112
- - **Environment: describe it, don't feed a multi-panel grid.** Reserve an environment ref for a hard exact-match, then use ONE clean establishing image.
109
+ <!-- @inject:references-read-literally -->
110
+ > **The general law: the model reads a reference literally.**
111
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
112
+
113
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
114
+
115
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
116
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
117
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
118
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
119
+
120
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
121
+ <!-- @end:references-read-literally -->
122
+
123
+ <!-- @inject:reference-rules-core -->
124
+ Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
125
+
126
+ 1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
127
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
128
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
129
+ 4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
130
+ 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
131
+ 6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
132
+ 7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
133
+ 8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
134
+ 9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
135
+ 10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
136
+ <!-- @end:reference-rules-core -->
137
+
138
+ ### For Kling specifically
139
+
140
+ - **Kling's consistency lever is "lock the subject with a fixed label reused verbatim."** That is Kling's phrasing for rules 2 and 3, and it is stricter than the others: **pronoun and synonym drift breaks it**, so the exact same label must appear on every single mention — not "he", not "the detective" after you named him. Reusing the label verbatim is the whole game. Slates composes this for you from `@mentions`.
141
+ - **Element references are the transport for rule 1** — 2-4 multi-angle photos per character/object, tagged `@element1` / `@element2` (see Element references above). The cap is 4 combined refs on the edit path.
113
142
 
114
143
  ## Negative prompting — has a real field
115
144
 
@@ -135,7 +164,7 @@ Layer scene-specific suppressions on top.
135
164
  - **Pro**: higher visual quality, no audio
136
165
  - **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references
137
166
 
138
- Pick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models` for current numbers before choosing a tier.
167
+ Pick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.
139
168
 
140
169
  ## Benchmark prompt structure
141
170
 
@@ -150,6 +179,7 @@ Cinematic example (paraphrasing fal blog patterns):
150
179
  > Shot 2: Medium shot of a detective in a trench coat ducking under an awning, water dripping from his hat brim. [Detective: weary, raspy]: 'I knew she'd come back.' Ambient noise: distant traffic, rain on metal.
151
180
  > Shot 3: Close-up on his eyes, narrowing as headlights flash across his face."
152
181
 
182
+ <!-- slates-only -->
153
183
  ## Pre-flight: references arrive inline, refer by code
154
184
 
155
185
  When you call `slates_generate_video` with `firstFrameAssetId` or `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside cost + `requires_confirm: true`. Look at them, revise prompt if needed, then re-call with `confirm=true`. Kling Omni multi-character with several ingredient images especially benefits — confirm each character image lands cleanly before spending.
@@ -158,14 +188,15 @@ When talking to the user about the gen, refer to each reference by its short cod
158
188
 
159
189
  - ✅ "I'm anchoring on **IMG-A12** as the detective and **IMG-A18** as the alleyway environment — Omni will handle the line delivery in EN."
160
190
  - ❌ "I'm using the detective image and the alley one..." (which alley? Three exist.)
191
+ <!-- /slates-only -->
161
192
 
162
- ## Video-to-video EDIT (`slates_edit_video`) — @Video1 / @ElementN / @ImageN
193
+ ## Video-to-video EDIT<!-- slates-only --> (`slates_edit_video`)<!-- /slates-only --> — @Video1 / @ElementN / @ImageN
163
194
 
164
195
  Kling O3 edit takes an EXISTING 3-15s clip and changes only what the prompt names — character swap, environment change, style transfer — in one pass, no masking. Original motion, camera, and audio are preserved by default. Its notation is Kling's own, different from the "image N" naming used everywhere else:
165
196
 
166
197
  - **`@Video1`** — the source clip (always; the transport anchors the instruction to it).
167
- - **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically).
168
- - **`@Image1..`** — style/appearance references (pass as `styleAssetIds`).
198
+ - **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images<!-- slates-only --> (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically)<!-- /slates-only -->.
199
+ - **`@Image1..`** — style/appearance references<!-- slates-only --> (pass as `styleAssetIds`)<!-- /slates-only -->.
169
200
  - Max **4 combined** element + image refs per edit.
170
201
 
171
202
  **Prompt shape — the change, not the whole scene:**
@@ -183,7 +214,7 @@ Rules:
183
214
  - One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).
184
215
  - Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.
185
216
  - Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.
186
- - Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings — see `slates-model-selection`.
217
+ - Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.
187
218
 
188
219
  ## Sources
189
220
 
@@ -1,11 +1,11 @@
1
1
  ---
2
2
  name: slates-prompting-nano-banana-2
3
- description: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3 Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.
3
+ description: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3.1 Flash Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.
4
4
  ---
5
5
 
6
6
  # Nano Banana 2 — cinematic & photorealistic prompting
7
7
 
8
- The **default** model behind `slates_generate_image` is **Gemini 3 Image** (Nano Banana 2 / Flash) — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill. NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
8
+ Nano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the default model behind `slates_generate_image` — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
9
9
 
10
10
  Knowledge cutoff: January 2025. Anything after needs explicit reference images.
11
11
 
@@ -30,6 +30,10 @@ Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and a
30
30
 
31
31
  ## Photorealism positives — what consistently works
32
32
 
33
+ > ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern — do not carry it across.**
34
+ > Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.
35
+ > The leak happens in one specific way — you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `slates-prompting-seedance` (Part 3, "Don't cross-pollinate image-model syntax").
36
+
33
37
  **Named lenses + apertures** beat generic "shallow depth of field":
34
38
  - `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`
35
39
  - `Panavision anamorphic` for horizontal flares + cinematic width
@@ -96,17 +100,43 @@ Default to #1. Reach for #2 only when positive framing can't suppress the unwant
96
100
  ## Reference images
97
101
 
98
102
  - **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.
99
- - **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style (or pass `referenceAssetIds`), Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (images 1 and 2) sits across from the woman (images 3 and 4) in the cafe (image 5)`, with a trailing `Render in the visual style of image 6.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**, so citing both of a subject's sheets under the SAME name ("Marcus") is what tells the model they are ONE person and stops the face averaging. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
103
+ - **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
100
104
 
101
105
  ### Reference rules (the verified ones)
102
- 1. **2-4 strong refs beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each adds context AND variables to balance.
103
- 2. **One reference per ROLE, named** (identity / style-grade / environment). Same-role competitors drift. The model doesn't infer roles from order — the inline name does it.
104
- 3. **Identity refs: attach both sheets, named as one entity — don't gate them.** A character's turnaround (body/proportion/outfit) AND its close-up expression sheet (high-res face: eyes, skin, teeth) both go in, cited under the SAME name ("Marcus (images 1 and 2)"). That shared name — not a role essay — is what stops the varied expressions from averaging the face. An *unnamed* expression sheet hurts; named as one entity, the close-ups are a fidelity win.
105
- 4. **Flat-light identity refs.** Prep them with flat, even, shadowless lighting on a plain neutral background. Studio-lit / scene-lit sheets bleed their lighting into the generation ("green-screen pasted in front of mountains").
106
- 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words. Reserve an environment ref for a mandatory exact-match, and then use ONE clean establishing image — never a multi-panel grid fed whole.
107
- 6. **Grids: explore, don't input.** Use grids to explore compositions, then pick a cell. Never feed a grid back in as a reference — cells share a split detail budget, so flaws propagate.
108
- 7. **Reuse the same refs across all shots.** Swapping mid-sequence causes drift.
109
- 8. **Legible in-shot text → bake it into the NB2 start frame**, then animate from it. Never trust text-to-video to render clean text.
106
+
107
+ <!-- @inject:references-read-literally -->
108
+ > **The general law: the model reads a reference literally.**
109
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
110
+
111
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
112
+
113
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
114
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
115
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
116
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
117
+
118
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
119
+ <!-- @end:references-read-literally -->
120
+
121
+ <!-- @inject:reference-rules-core -->
122
+ Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
123
+
124
+ 1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
125
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
126
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
127
+ 4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
128
+ 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
129
+ 6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
130
+ 7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
131
+ 8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
132
+ 9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
133
+ 10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
134
+ <!-- @end:reference-rules-core -->
135
+
136
+ ### For Nano Banana 2 specifically
137
+
138
+ - **NB2's own consistency lever is "assign a distinct name to each character/object."** That is Google's phrasing for rule 3 — cite each canonical identity inline by name.
139
+ - **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.
110
140
  - **Character consistency is officially "not 100% perfect"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.
111
141
  - **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.
112
142
 
@@ -126,7 +156,7 @@ Default to #1. Reach for #2 only when positive framing can't suppress the unwant
126
156
 
127
157
  ## Resolution tactics
128
158
 
129
- - Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — call `slates_estimate_generation_cost` for current numbers. Pick the cheapest resolution that serves the use case.
159
+ - Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — check current numbers<!-- slates-only --> by calling `slates_estimate_generation_cost`<!-- /slates-only -->. Pick the cheapest resolution that serves the use case.
130
160
  - **At 2K and above, the model allocates more tokens to surface detail** — explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.
131
161
  - 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.
132
162
  - 2K generations vary 20-60s+. Don't time-budget tightly.
@@ -152,4 +182,6 @@ Everything in this skill applies to the whole Nano Banana family; two variants t
152
182
  - **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.
153
183
  - **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.
154
184
 
185
+ <!-- slates-only -->
155
186
  Routing between them (and vs GPT Image 2 / FLUX / Seedream): `slates-model-selection`.
187
+ <!-- /slates-only -->
@@ -30,7 +30,7 @@ Google's fast video generation + editing model ("Nano Banana Pro for video" in c
30
30
 
31
31
  - **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
32
32
  - Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.
33
- - **Name references inline** the standard Slates way ("Marcus (images 1 and 2) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
33
+ - **Name references inline** the standard Slates way ("Marcus (image 1) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
34
34
  - **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language ("rain patters on the tin roof"). Negative direction as plain instructions ("Do not show text").
35
35
  - Duration is an explicit 3–10s integer param; cost scales linearly per second.
36
36