@slatesvideo/shared 0.5.2 → 0.5.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +18 -0
- package/dist/index.d.ts +1 -0
- package/dist/index.js +3 -0
- package/dist/operations/index.d.ts +25 -0
- package/dist/operations/index.js +74 -5
- package/dist/prompts/character-sheet.d.ts +24 -12
- package/dist/prompts/character-sheet.js +78 -28
- package/dist/prompts/environment-sheet.d.ts +9 -1
- package/dist/prompts/environment-sheet.js +17 -3
- package/dist/prompts/index.d.ts +1 -0
- package/dist/prompts/index.js +4 -0
- package/dist/prompts/model-facts.js +4 -1
- package/dist/prompts/partials.generated.d.ts +2 -0
- package/dist/prompts/partials.generated.js +14 -0
- package/dist/prompts/prompting-tips.d.ts +23 -0
- package/dist/prompts/prompting-tips.js +376 -0
- package/dist/prompts/reference-rules.d.ts +43 -14
- package/dist/prompts/reference-rules.js +50 -26
- package/dist/skills/content.js +12 -12
- package/package.json +4 -3
- package/skills/_partials/decision-log.md +12 -0
- package/skills/_partials/reference-rules-core.md +12 -0
- package/skills/_partials/reference-tips-short.md +2 -0
- package/skills/_partials/references-read-literally.md +11 -0
- package/skills/_partials/still-gate.md +3 -0
- package/skills/slates-character-turnaround.md +64 -29
- package/skills/slates-cost-discipline.md +10 -0
- package/skills/slates-edit-and-iterate.md +16 -1
- package/skills/slates-model-selection.md +24 -1
- package/skills/slates-one-prompt-film.md +19 -0
- package/skills/slates-prompting-flux-2-max.md +36 -5
- package/skills/slates-prompting-kling-v3.md +33 -4
- package/skills/slates-prompting-nano-banana-2.md +40 -10
- package/skills/slates-prompting-seedance.md +284 -85
- package/skills/slates-prompting-veo-3.md +33 -4
- package/skills/slates-storyboard-from-script.md +19 -0
- package/skills/slates-vision-feedback-loop.md +49 -2
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@slatesvideo/shared",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.4",
|
|
4
4
|
"description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -27,8 +27,9 @@
|
|
|
27
27
|
"README.md"
|
|
28
28
|
],
|
|
29
29
|
"scripts": {
|
|
30
|
-
"
|
|
31
|
-
"
|
|
30
|
+
"sync-partials": "node scripts/sync-partials.mjs",
|
|
31
|
+
"build": "node scripts/sync-partials.mjs --check && node scripts/embed-skills.mjs && tsc",
|
|
32
|
+
"typecheck": "node scripts/sync-partials.mjs --check && node scripts/embed-skills.mjs && tsc --noEmit",
|
|
32
33
|
"prepublishOnly": "npm run build"
|
|
33
34
|
},
|
|
34
35
|
"repository": {
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:
|
|
2
|
+
|
|
3
|
+
```
|
|
4
|
+
source phrase or declared default → what you wrote → what it resolves
|
|
5
|
+
"in a diner" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable
|
|
6
|
+
(no time of day) → late afternoon, low warm key → default; say the word and it changes
|
|
7
|
+
(no camera) → slow push-in, single move → one move per shot; stacking increases instability
|
|
8
|
+
```
|
|
9
|
+
|
|
10
|
+
**Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the "why did you add that?" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.
|
|
11
|
+
|
|
12
|
+
> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
2
|
+
|
|
3
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
4
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
5
|
+
3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
6
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
7
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
8
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
9
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
10
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
11
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
12
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
@@ -0,0 +1,2 @@
|
|
|
1
|
+
<!-- consumer:ts -->
|
|
2
|
+
Name each reference inline; never write role essays. Slates does this for you: `@mention` a subject or environment and it composes `Marcus (images 1 and 2) in the cafe (image 3)`, citing them in the exact order it sends them. Citing both of a character's sheets under the SAME name is what tells the model they are one person — a "Reference Image Instructions" block does the opposite and drags the sheet's studio lighting into your scene. Start with 2-3 focused refs.
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
> **The general law: the model reads a reference literally.**
|
|
2
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
3
|
+
|
|
4
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
5
|
+
|
|
6
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
7
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
8
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
9
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
10
|
+
|
|
11
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
@@ -0,0 +1,3 @@
|
|
|
1
|
+
**A visible defect in the still is already a STOP.** Do not animate it. Fix the frame first, then move to motion — and go to motion only when the crop passes the still scan and you genuinely need movement to confirm an uncertain edge, reflection, or object.
|
|
2
|
+
|
|
3
|
+
This is a **cost** rule as much as a craft rule: a 1080p/10s premium video generation costs many multiples of an image re-roll, and video is where a defect stops being fixable. Anything wrong in the still gets worse in motion — soft geometry mushes, broken-but-plausible objects fall apart, oily textures start crawling. **Animating a known-bad frame is the single most expensive mistake in the pipeline.** Re-rolling the image is the cheap move; re-rolling the video is not.
|
|
@@ -1,11 +1,43 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-character-turnaround
|
|
3
|
-
description: Build a Slates character from a reference image — generate
|
|
3
|
+
description: Build a Slates character from a reference image — generate its identity reference sheet and bind it to the character so the card updates live. Use when the user wants to "create a character", "build a character from this image", "generate a turnaround for X", or starts any storyboard flow that needs consistent character references.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Character
|
|
6
|
+
# Character identity sheet — Slates workflow
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
A character's identity sheet is attached to **every** downstream generation that mentions it, so a flaw in the sheet becomes a flaw in every shot made from it. Building it well is the highest-leverage thing you can do for a project.
|
|
9
|
+
|
|
10
|
+
<!-- @inject:references-read-literally -->
|
|
11
|
+
> **The general law: the model reads a reference literally.**
|
|
12
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
13
|
+
|
|
14
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
15
|
+
|
|
16
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
17
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
18
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
19
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
20
|
+
|
|
21
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
22
|
+
<!-- @end:references-read-literally -->
|
|
23
|
+
|
|
24
|
+
## The shape: ONE sheet, three panels
|
|
25
|
+
|
|
26
|
+
Slates generates **one identity sheet per character**, bound to the character's turnaround slot:
|
|
27
|
+
|
|
28
|
+
| Panel | What it carries |
|
|
29
|
+
|---|---|
|
|
30
|
+
| **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
|
|
31
|
+
| **Full-body front, relaxed A-pose** | Build, proportion, wardrobe |
|
|
32
|
+
| **Full-body back** | Hair fall and the back of the outfit — the only panel where either reads |
|
|
33
|
+
|
|
34
|
+
On a deep neutral-grey plate (`#3a3a3c`), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look.
|
|
35
|
+
|
|
36
|
+
**The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
|
|
37
|
+
|
|
38
|
+
**Why one sheet and not two.** Every `@character` mention pushes *all* of that character's bound sheets into one reference group, so a two-sheet character costs **two reference slots on every generation**. Against real caps that is brutal — Kling 3.0 takes 4 ingredients (2 characters, zero room for an environment), NB2 has 4 character slots, Seedance 9. One sheet each **doubles the cast you can stage on every model.** It also takes competing facial renderings from six down to two, which is what stops a face from averaging (see the general law above), and it halves the per-character sheet spend.
|
|
39
|
+
|
|
40
|
+
**The expression slot still exists** and is still read — characters built before this change have one bound and keep working. Generate one only when a character genuinely needs a dedicated expression range, and tell the user it costs a reference slot on every shot from then on.
|
|
9
41
|
|
|
10
42
|
## Workflow
|
|
11
43
|
|
|
@@ -15,7 +47,7 @@ The user has either:
|
|
|
15
47
|
- Described the character in text only.
|
|
16
48
|
|
|
17
49
|
If image: upload it as a reference (`slates_upload_reference_image`).
|
|
18
|
-
If text only:
|
|
50
|
+
If text only: generate from prompt-only — less consistent, so warn the user.
|
|
19
51
|
|
|
20
52
|
### 2. Create the character record
|
|
21
53
|
`slates_create_character` with:
|
|
@@ -23,33 +55,36 @@ If text only: skip step 2's reference and generate the turnaround from prompt-on
|
|
|
23
55
|
- `description` — 1-2 sentences, *visual* only ("tall, dark hair, scar over left eye"), not personality.
|
|
24
56
|
- `style` — leave as the source's own medium by default. Only name a transform if the user wants one (e.g. anime → realistic).
|
|
25
57
|
|
|
26
|
-
### 3. Generate the
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
-
|
|
30
|
-
|
|
31
|
-
-
|
|
32
|
-
-
|
|
33
|
-
-
|
|
34
|
-
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
-
|
|
38
|
-
|
|
39
|
-
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
58
|
+
### 3. Generate the sheet
|
|
59
|
+
`slates_generate_character_sheets` with `characterId`, `projectId`, and `baseAssetId` (the source portrait).
|
|
60
|
+
|
|
61
|
+
**Do not hand-write the sheet prompt.** Slates builds it from the canonical template in `@slatesvideo/shared/prompts` (`buildCharacterTurnaroundPrompt`) — panels, plate, lighting and craft clauses included — and appends your `userNotes`. Use `userNotes` for what the template can't know: *"use the woman on the left"*, *"keep the scar on the right cheek"*. A hand-written prompt is a fork of the template and will drift from it.
|
|
62
|
+
|
|
63
|
+
- Estimate cost first with `slates_estimate_generation_cost` and announce in **credits** — never quote a price from memory. Default is Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted credits.
|
|
64
|
+
- When the result returns inline, **evaluate it before binding**:
|
|
65
|
+
- Is the portrait clearly the largest panel, and is it off-frontal?
|
|
66
|
+
- Is it the same person across all three panels?
|
|
67
|
+
- Catchlights present, irises readable rather than black holes?
|
|
68
|
+
- Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
|
|
69
|
+
- Plate a flat deep grey, not white and not black?
|
|
70
|
+
- If off: one focused refinement, then regenerate. The sheet is upstream of everything — it is worth a re-roll that a scene frame is not.
|
|
71
|
+
- The op binds the result to the turnaround slot automatically.
|
|
72
|
+
|
|
73
|
+
### 4. Hand back
|
|
74
|
+
> "Character {name} ready — identity sheet bound. Use `@{name}` in any prompt and Slates attaches it and names it inline, so the face stays consistent."
|
|
75
|
+
|
|
76
|
+
## How the reference gets used at scene time
|
|
77
|
+
|
|
78
|
+
Slates cites the sheet inline under the character's name — `{name} (image N)` — in the exact order it sends references. That **name** is the anti-averaging lever, and it is each model's own official mechanism (NB2: "assign a distinct name"; Seedance: `Reference <Subject_N> in <Image_N>`; Kling: reuse a fixed label verbatim). If a character has both slots bound, both are cited under the *same* name so the model reads them as one person.
|
|
79
|
+
|
|
80
|
+
Critically, the app injects **no** wardrobe, expression, or lighting directive. The user's scene prompt owns all of that — which is why `@{name}` dropped into a movie-still injection keeps the still's own clothing and lighting instead of dragging the sheet's.
|
|
48
81
|
|
|
49
82
|
## Anti-patterns
|
|
50
83
|
|
|
51
|
-
- **Don't** studio-light or
|
|
52
|
-
- **Don't**
|
|
84
|
+
- **Don't** studio-light, white-background, or black-background the sheet. White bleeds into the video and washes out the location; black eats edge detail. Flat, even, shadowless light on a deep neutral grey.
|
|
85
|
+
- **Don't** hand-write the sheet prompt when the op will build it — that is how the template and the shipped prompt fork.
|
|
86
|
+
- **Don't** generate an expression sheet by reflex. It is opt-in now, and it costs a reference slot on every downstream shot.
|
|
53
87
|
- **Don't** skip binding. The slots are what the storyboard pipeline reads — an unbound asset doesn't help downstream.
|
|
54
88
|
- **Don't** invent character details. Stick to what's in the reference image and the user's description.
|
|
55
|
-
- **Don't** use
|
|
89
|
+
- **Don't** use 4K — wastes credits, no quality gain at sheet scale.
|
|
90
|
+
- **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `slates-prompting-seedance`.
|
|
@@ -103,6 +103,16 @@ Video gens take minutes (Seedance 4K can run far longer). A client/CLI timeout o
|
|
|
103
103
|
- **Poll, don't re-roll.** Use `background: true` on `slates_generate_video`, then poll `slates_get_generation_status` (free, read-only) until it reports `completed` or `failed`. In-flight jobs survive app restarts and are recovered.
|
|
104
104
|
- A gen has only failed when the status comes back `failed` — and a provider *rejection* **refunds** the credits, so failed isolation tests are ~free. Until you see a terminal status, the job is in flight. Wait.
|
|
105
105
|
|
|
106
|
+
## 🔴 The still-gate — the most expensive mistake in the pipeline
|
|
107
|
+
|
|
108
|
+
<!-- @inject:still-gate -->
|
|
109
|
+
**A visible defect in the still is already a STOP.** Do not animate it. Fix the frame first, then move to motion — and go to motion only when the crop passes the still scan and you genuinely need movement to confirm an uncertain edge, reflection, or object.
|
|
110
|
+
|
|
111
|
+
This is a **cost** rule as much as a craft rule: a 1080p/10s premium video generation costs many multiples of an image re-roll, and video is where a defect stops being fixable. Anything wrong in the still gets worse in motion — soft geometry mushes, broken-but-plausible objects fall apart, oily textures start crawling. **Animating a known-bad frame is the single most expensive mistake in the pipeline.** Re-rolling the image is the cheap move; re-rolling the video is not.
|
|
112
|
+
<!-- @end:still-gate -->
|
|
113
|
+
|
|
114
|
+
The check itself lives in `slates-vision-feedback-loop` (the four slop tells and the per-model accents). The **stop** is a cost rule and belongs here: before every image→video call, confirm the source frame passed the still scan. If it didn't, spending video credits on it is not iteration — it is buying a more expensive copy of a defect you already found.
|
|
115
|
+
|
|
106
116
|
## The 3-strike rule
|
|
107
117
|
|
|
108
118
|
Stop after 3 iterations on the same prompt. Hand back to the user with what you tried and what's not working. The slot machine doesn't converge — if it's not landing, the prompt structure is wrong, not the seed.
|
|
@@ -7,6 +7,20 @@ description: Iterate on an existing Slates asset — re-evaluate, refine prompt,
|
|
|
7
7
|
|
|
8
8
|
The user already has a generated image in Slates and wants to refine it. The vision-feedback-loop skill defines the general pattern; this skill is the specific recipe for "I have asset X, here's what's wrong with it."
|
|
9
9
|
|
|
10
|
+
## 🔴 The master rule — an edit is a LEAF, not a node
|
|
11
|
+
|
|
12
|
+
**Never re-edit an edit. Always go back and re-edit the master.**
|
|
13
|
+
|
|
14
|
+
Every edit model silently re-renders the **whole frame**, not just the region you named. So the parts you didn't ask to change come back slightly different every pass — softer texture, drifted colour, mushier fine detail. It is barely visible after one edit and obvious by the second. Chaining edits compounds the damage and there is no way to undo it, because each generation *is* the new source.
|
|
15
|
+
|
|
16
|
+
The fix is structural, not a matter of care:
|
|
17
|
+
|
|
18
|
+
- **Want two changes?** Make them in ONE edit off the master, or make them as two separate edits **both taken from the master**, then keep whichever you prefer.
|
|
19
|
+
- **An edit came back wrong?** Do NOT edit the result to fix it. Discard it and re-edit the master with a better instruction.
|
|
20
|
+
- **Only the changed region is worth keeping?** That is a compositing job — the edit supplies the new region, the untouched master supplies everything else.
|
|
21
|
+
|
|
22
|
+
Slates records this: an edit result carries `sourceAssetIds` pointing at the asset it was made from, so **you can tell whether the thing you are about to edit is itself an edit.** Check before you edit — `[Edit]`-prefixed prompts and a populated source lineage both say "this is a leaf; go back to its parent."
|
|
23
|
+
|
|
10
24
|
## Workflow
|
|
11
25
|
|
|
12
26
|
### 1. Pull the current asset back into context
|
|
@@ -45,4 +59,5 @@ The user's request is one of:
|
|
|
45
59
|
- **Don't** delete the original asset until the user confirms the new one. Slates keeps both; the user picks.
|
|
46
60
|
- **Don't** mix surgical and wholesale changes in one regeneration. The user said "make it warmer" — don't also reframe the shot.
|
|
47
61
|
- **Don't** re-generate when `slates_edit_image` would work. Edits preserve composition and identity; full regen rolls the dice.
|
|
48
|
-
- **Don't**
|
|
62
|
+
- **Don't** edit an edit — ever. Not once, not "just a small one." Go back to the master (see the master rule above). Every attempt re-renders the full frame and the degradation is cumulative and permanent.
|
|
63
|
+
- **Don't** keep re-rolling the same failed edit. If three tries off the master didn't land, the brief is wrong, not the model — check in with the user.
|
|
@@ -7,6 +7,18 @@ description: Which model to pick for a given job — the routing doctrine. Read
|
|
|
7
7
|
|
|
8
8
|
Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.
|
|
9
9
|
|
|
10
|
+
## 🔑 The meta-rule — above the table
|
|
11
|
+
|
|
12
|
+
The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:
|
|
13
|
+
|
|
14
|
+
> **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.
|
|
15
|
+
>
|
|
16
|
+
> **Inspect the output at its intended crop.** A frame that holds up as a thumbnail can fall apart at the size it will actually be watched. For a location, look at atmosphere, material texture, and anchor objects; for a character, identity, skin, pose, and gradients.
|
|
17
|
+
>
|
|
18
|
+
> **Choose the model that PROVES that requirement** and leaves only failures you can afford to rerun or mask.
|
|
19
|
+
>
|
|
20
|
+
> **When the roster changes, repeat the evidence test.** Do not carry today's ranking forward on reputation.
|
|
21
|
+
|
|
10
22
|
## Video routing
|
|
11
23
|
|
|
12
24
|
| Job | Model | Why |
|
|
@@ -16,6 +28,17 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
|
|
|
16
28
|
| Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
|
|
17
29
|
| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
|
|
18
30
|
| The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
|
|
31
|
+
|
|
32
|
+
### Named Seedance escalation triggers
|
|
33
|
+
|
|
34
|
+
"Physics matter" is an abstract category and it under-fires. These are the beats Seedance is **observably** good at — if the shot contains one, escalate without deliberating:
|
|
35
|
+
|
|
36
|
+
- **Real-time → slow-motion contrast.** The signature beat; nearly every strong clip rides it.
|
|
37
|
+
- **The camera moving while debris, meteors, sparks or particles crash around the subject.** Distinctly a feature of this model, not just a thing it survives.
|
|
38
|
+
- **Massive scale that has to read as genuinely huge** — not "a big thing", a thing whose size is the point of the shot.
|
|
39
|
+
- **One continuous unbroken take.**
|
|
40
|
+
|
|
41
|
+
Concrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).
|
|
19
42
|
| Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | The only job Veo wins. |
|
|
20
43
|
|
|
21
44
|
## Video EDIT routing (changing an existing clip)
|
|
@@ -29,7 +52,7 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
|
|
|
29
52
|
| AI-edit the user's OWN footage | Omni Flash Edit (3–10s) or Kling O3 Edit (3–15s, 720–3840px) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
|
|
30
53
|
|
|
31
54
|
- **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
|
|
32
|
-
- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the
|
|
55
|
+
- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.
|
|
33
56
|
- **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
|
|
34
57
|
- Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
|
|
35
58
|
|
|
@@ -12,6 +12,25 @@ The user gives an idea. You hand back an MP4 on disk. Everything in between is y
|
|
|
12
12
|
### 1. Script the beats
|
|
13
13
|
Turn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.
|
|
14
14
|
|
|
15
|
+
**Surface a decision log with the plan.**
|
|
16
|
+
|
|
17
|
+
<!-- @inject:decision-log -->
|
|
18
|
+
When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:
|
|
19
|
+
|
|
20
|
+
```
|
|
21
|
+
source phrase or declared default → what you wrote → what it resolves
|
|
22
|
+
"in a diner" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable
|
|
23
|
+
(no time of day) → late afternoon, low warm key → default; say the word and it changes
|
|
24
|
+
(no camera) → slow push-in, single move → one move per shot; stacking increases instability
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
**Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the "why did you add that?" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.
|
|
28
|
+
|
|
29
|
+
> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
|
|
30
|
+
<!-- @end:decision-log -->
|
|
31
|
+
|
|
32
|
+
A 4-10 shot script is where you invent the most on the user's behalf — time of day, wardrobe, weather, lens feel, camera moves the brief never mentioned. The log is what makes those visible while they are still free to change.
|
|
33
|
+
|
|
15
34
|
### 2. Set up the project
|
|
16
35
|
- `slates_create_project` named for the piece.
|
|
17
36
|
- Recurring character? Build it properly — `slates_create_character` + the `slates-character-turnaround` recipe — so every frame references the same turnaround.
|
|
@@ -82,11 +82,42 @@ Use natural language for exploration, JSON when the layout is locked and you're
|
|
|
82
82
|
|
|
83
83
|
In Slates, pass `referenceAssetIds` on `slates_generate_image` — FLUX routes them through its edit endpoint. Slates names each reference inline in the prompt ("the subject (image 1), the style (image 2)") in the order it sends them, so you don't hand-write role labels; the name carries the role and unnamed-by-position blending is avoided. For surgical changes to one existing image use `slates_edit_image` with `editModel: flux-2-max` (note: FLUX edits ignore extra referenceAssetIds — that's NB2-only).
|
|
84
84
|
|
|
85
|
-
Reference
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
85
|
+
### Reference rules (the verified ones)
|
|
86
|
+
|
|
87
|
+
<!-- @inject:references-read-literally -->
|
|
88
|
+
> **The general law: the model reads a reference literally.**
|
|
89
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
90
|
+
|
|
91
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
92
|
+
|
|
93
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
94
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
95
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
96
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
97
|
+
|
|
98
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
99
|
+
<!-- @end:references-read-literally -->
|
|
100
|
+
|
|
101
|
+
<!-- @inject:reference-rules-core -->
|
|
102
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
103
|
+
|
|
104
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
105
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
106
|
+
3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
107
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
108
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
109
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
110
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
111
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
112
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
113
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
114
|
+
<!-- @end:reference-rules-core -->
|
|
115
|
+
|
|
116
|
+
### For FLUX.2 Max specifically
|
|
117
|
+
|
|
118
|
+
- **FLUX caps references well below NB2's 14, so rule 1's "2-4" is a ceiling here, not a starting point.** Be deliberate about which roles earn a slot.
|
|
119
|
+
- **Rule 9 has a hard edge on this model:** `slates_edit_image` with `editModel: flux-2-max` ignores extra `referenceAssetIds` — that is NB2-only. A FLUX edit sees the source image and the prompt, nothing else.
|
|
120
|
+
- **FLUX has no memory between generations, so rule 7 is enforced by repetition.** Define the character exhaustively once and repeat those exact descriptors verbatim in every subsequent prompt — see Character consistency across a series below.
|
|
90
121
|
|
|
91
122
|
## Character consistency across a series
|
|
92
123
|
|
|
@@ -106,10 +106,39 @@ Upload 2-4 multi-angle reference photos per character/object. Tag inline:
|
|
|
106
106
|
|
|
107
107
|
## Reference discipline (character / environment refs)
|
|
108
108
|
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
109
|
+
<!-- @inject:references-read-literally -->
|
|
110
|
+
> **The general law: the model reads a reference literally.**
|
|
111
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
112
|
+
|
|
113
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
114
|
+
|
|
115
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
116
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
117
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
118
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
119
|
+
|
|
120
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
121
|
+
<!-- @end:references-read-literally -->
|
|
122
|
+
|
|
123
|
+
<!-- @inject:reference-rules-core -->
|
|
124
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
125
|
+
|
|
126
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
127
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
128
|
+
3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
129
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
130
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
131
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
132
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
133
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
134
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
135
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
136
|
+
<!-- @end:reference-rules-core -->
|
|
137
|
+
|
|
138
|
+
### For Kling specifically
|
|
139
|
+
|
|
140
|
+
- **Kling's consistency lever is "lock the subject with a fixed label reused verbatim."** That is Kling's phrasing for rules 2 and 3, and it is stricter than the others: **pronoun and synonym drift breaks it**, so the exact same label must appear on every single mention — not "he", not "the detective" after you named him. Reusing the label verbatim is the whole game. Slates composes this for you from `@mentions`.
|
|
141
|
+
- **Element references are the transport for rule 1** — 2-4 multi-angle photos per character/object, tagged `@element1` / `@element2` (see Element references above). The cap is 4 combined refs on the edit path.
|
|
113
142
|
|
|
114
143
|
## Negative prompting — has a real field
|
|
115
144
|
|
|
@@ -1,11 +1,11 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-nano-banana-2
|
|
3
|
-
description: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3 Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.
|
|
3
|
+
description: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3.1 Flash Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Nano Banana 2 — cinematic & photorealistic prompting
|
|
7
7
|
|
|
8
|
-
The **default** model behind `slates_generate_image` is **Gemini 3 Image** (Nano Banana 2
|
|
8
|
+
The **default** model behind `slates_generate_image` is **Gemini 3.1 Flash Image** (Nano Banana 2) — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill. It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat. Verified against the runtime slug map in `slate/src/main/api/google.ts`. NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
|
|
9
9
|
|
|
10
10
|
Knowledge cutoff: January 2025. Anything after needs explicit reference images.
|
|
11
11
|
|
|
@@ -30,6 +30,10 @@ Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and a
|
|
|
30
30
|
|
|
31
31
|
## Photorealism positives — what consistently works
|
|
32
32
|
|
|
33
|
+
> ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern — do not carry it across.**
|
|
34
|
+
> Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.
|
|
35
|
+
> The leak happens in one specific way — you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `slates-prompting-seedance` (Part 3, "Don't cross-pollinate image-model syntax").
|
|
36
|
+
|
|
33
37
|
**Named lenses + apertures** beat generic "shallow depth of field":
|
|
34
38
|
- `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`
|
|
35
39
|
- `Panavision anamorphic` for horizontal flares + cinematic width
|
|
@@ -99,14 +103,40 @@ Default to #1. Reach for #2 only when positive framing can't suppress the unwant
|
|
|
99
103
|
- **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style (or pass `referenceAssetIds`), Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (images 1 and 2) sits across from the woman (images 3 and 4) in the cafe (image 5)`, with a trailing `Render in the visual style of image 6.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**, so citing both of a subject's sheets under the SAME name ("Marcus") is what tells the model they are ONE person and stops the face averaging. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
100
104
|
|
|
101
105
|
### Reference rules (the verified ones)
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
106
|
+
|
|
107
|
+
<!-- @inject:references-read-literally -->
|
|
108
|
+
> **The general law: the model reads a reference literally.**
|
|
109
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
110
|
+
|
|
111
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
112
|
+
|
|
113
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
114
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
115
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
116
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
117
|
+
|
|
118
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
119
|
+
<!-- @end:references-read-literally -->
|
|
120
|
+
|
|
121
|
+
<!-- @inject:reference-rules-core -->
|
|
122
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
123
|
+
|
|
124
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
125
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
126
|
+
3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
127
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
128
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
129
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
130
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
131
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
132
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
133
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
134
|
+
<!-- @end:reference-rules-core -->
|
|
135
|
+
|
|
136
|
+
### For Nano Banana 2 specifically
|
|
137
|
+
|
|
138
|
+
- **NB2's own consistency lever is "assign a distinct name to each character/object."** That is Google's phrasing for rule 3 — citing both of a subject's sheets under the SAME name is the officially-sanctioned mechanism, not a workaround.
|
|
139
|
+
- **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.
|
|
110
140
|
- **Character consistency is officially "not 100% perfect"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.
|
|
111
141
|
- **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.
|
|
112
142
|
|