@slatesvideo/shared 0.5.3 → 0.5.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +31 -0
- package/dist/operations/index.d.ts +4 -8
- package/dist/operations/index.js +26 -38
- package/dist/prompts/character-sheet.d.ts +19 -18
- package/dist/prompts/character-sheet.js +84 -38
- package/dist/prompts/environment-sheet.d.ts +9 -1
- package/dist/prompts/environment-sheet.js +17 -3
- package/dist/prompts/model-facts.js +4 -1
- package/dist/prompts/partials.generated.d.ts +2 -0
- package/dist/prompts/partials.generated.js +14 -0
- package/dist/prompts/prompting-tips.js +50 -19
- package/dist/prompts/reference-composer.d.ts +1 -1
- package/dist/prompts/reference-composer.js +3 -4
- package/dist/prompts/reference-rules.d.ts +43 -14
- package/dist/prompts/reference-rules.js +51 -27
- package/dist/skills/content.js +14 -14
- package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
- package/exports/slates-prompt-builder/generated/reference-character.md +78 -0
- package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
- package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
- package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +8 -3
- package/skills/_partials/decision-log.md +12 -0
- package/skills/_partials/reference-rules-core.md +12 -0
- package/skills/_partials/reference-tips-short.md +2 -0
- package/skills/_partials/references-read-literally.md +11 -0
- package/skills/_partials/still-gate.md +3 -0
- package/skills/slates-character-identity.md +100 -0
- package/skills/slates-cost-discipline.md +10 -0
- package/skills/slates-edit-and-iterate.md +17 -2
- package/skills/slates-model-selection.md +24 -1
- package/skills/slates-one-prompt-film.md +22 -3
- package/skills/slates-prompting-flux-2-max.md +36 -5
- package/skills/slates-prompting-gpt-image-2.md +1 -1
- package/skills/slates-prompting-kling-v3.md +40 -9
- package/skills/slates-prompting-nano-banana-2.md +44 -12
- package/skills/slates-prompting-omni-flash.md +1 -1
- package/skills/slates-prompting-seedance.md +295 -90
- package/skills/slates-prompting-veo-3.md +33 -4
- package/skills/slates-storyboard-from-script.md +19 -0
- package/skills/slates-vision-feedback-loop.md +49 -2
- package/skills/slates-character-turnaround.md +0 -55
|
@@ -1,97 +1,221 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-seedance
|
|
3
|
-
description: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance prompts
|
|
3
|
+
description: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance structures multi-beat prompts as a "Shot 1 / Shot 2 / Shot 3" storyboard against an 8-slot advanced formula — never per-second time stamps. Its syntax differs from Kling, Veo and the image models; don't cross-pollinate (in particular, no lens / aperture / film-stock vocabulary).
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Seedance 2.0 — prompting
|
|
7
7
|
|
|
8
|
-
ByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K — 4K video is Pro-only, default 1080p), 4–15s, first+last frame
|
|
8
|
+
ByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K — 4K video is Pro-only, default 1080p), 4–15s, first+last frame, and up to 9 reference images / 3 videos / 3 audio clips.
|
|
9
9
|
|
|
10
|
-
|
|
10
|
+
> **How to read this file.**
|
|
11
|
+
> **[official :NNNN]** — ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/seedance-2-modelark-docs.md`). Receipt-grade; treat as law.
|
|
12
|
+
> **[community]** — third-party guides and our own field notes. Useful, but an `[official]` block always wins.
|
|
13
|
+
> **[slates]** — how the Slates app composes or bills this; not ByteDance doctrine.
|
|
14
|
+
>
|
|
15
|
+
> The split is load-bearing. A community-sourced "narrative timing beats" doctrine shipped in this file for months teaching the **exact inverse** of ByteDance's published guidance. Never merge the two registers again.
|
|
11
16
|
|
|
12
|
-
|
|
13
|
-
Subject + Action + Environment + Camera + Style + Constraints
|
|
14
|
-
```
|
|
17
|
+
---
|
|
15
18
|
|
|
16
|
-
|
|
19
|
+
# Part 1 — Official ByteDance doctrine
|
|
17
20
|
|
|
18
|
-
##
|
|
21
|
+
## What Seedance actually is `[official :1450-1452]`
|
|
19
22
|
|
|
20
|
-
|
|
23
|
+
Seedance 2.0 is a multimodal AI director. It reads text, images, video and audio **simultaneously** and internally decomposes them into two dimensions:
|
|
21
24
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
```
|
|
25
|
+
- the **spatial layer** — what is in the frame
|
|
26
|
+
- the **temporal layer** — how things change over time
|
|
25
27
|
|
|
26
|
-
|
|
28
|
+
So a good prompt is **not "copywriting-style description" but an "engineering-style instruction"**: who, in what scene, doing what action, how the camera moves, and in what chronological order events occur — delivered respectively to the spatial layer and the temporal layer.
|
|
27
29
|
|
|
28
|
-
|
|
30
|
+
## The advanced formula — 8 slots `[official :1455]`
|
|
29
31
|
|
|
30
|
-
```
|
|
31
|
-
|
|
32
|
-
|
|
32
|
+
```text
|
|
33
|
+
precise subject + action details + scene/environment + lighting & color tone
|
|
34
|
+
+ camera movement + visual style + image quality + constraints
|
|
33
35
|
```
|
|
34
36
|
|
|
35
|
-
|
|
37
|
+
⚠️ There is **no official "6-step formula."** `Subject + Action + Environment + Camera + Style + Constraints` is community branding with no ByteDance source, and it silently drops the **lighting & color tone** and **image quality** slots. Use the 8 slots above.
|
|
36
38
|
|
|
37
|
-
##
|
|
39
|
+
## Task-type sentence patterns `[official :1389-1425]`
|
|
38
40
|
|
|
39
|
-
|
|
41
|
+
Seedance classifies your request from the phrasing. Use the pattern that matches the task:
|
|
40
42
|
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
43
|
+
| Task | Pattern |
|
|
44
|
+
|---|---|
|
|
45
|
+
| **Image reference** | ``Reference `<Subject_N>` in `<Image_N>` to generate…`` |
|
|
46
|
+
| **Video reference** | ``Reference `<Action / Camera_movement / Style / Sound_effect>` in `<Video_N>` to generate…`` |
|
|
47
|
+
| **Audio reference** | ``Reference the timbre in `<Audio_N>` to generate…`` |
|
|
48
|
+
| **Video edit — modify** | ``Strictly edit `<Video_N>`, and modify `<Original_Characteristic>` in it to `<New_Characteristic>``` |
|
|
49
|
+
| **Video edit — add** | ``<Element_Features>` + `<Timing>` + `<Location>`` |
|
|
50
|
+
| **Video edit — delete** | Name what to delete; for anything that must stay, say so explicitly |
|
|
51
|
+
| **Video extend** | ``Extend `<Video_N>` forward/backward to generate…`` |
|
|
52
|
+
| **Combined** | ``Reference `[Dimension]` of `<Image/Video_N>`, strictly edit `<Video_X>`, `[Specific_Edits]``` |
|
|
44
53
|
|
|
45
|
-
|
|
54
|
+
### ⚠️ Edit / extend phrasing landmine `[official :1431]`
|
|
46
55
|
|
|
47
|
-
|
|
56
|
+
> *"For edit / extend video tasks, directly use `<Video_N>` to refer to the video. **Do not use "reference `<Video_N>`"**, to avoid being incorrectly identified as a reference task."*
|
|
48
57
|
|
|
49
|
-
|
|
50
|
-
A cool-white diagonal beam from upper left, dust particles drifting through.
|
|
51
|
-
Soft golden hour lighting from low west angle.
|
|
52
|
-
Dramatic rim light against dark background.
|
|
53
|
-
```
|
|
58
|
+
This is easy to trip: Slates has an edit lane<!-- slates-only --> (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)<!-- /slates-only -->. Writing *"reference video 1 and change the jacket to red"* gets classified as a **reference** task — the model generates a brand-new clip inspired by the source instead of editing it. Write *"Strictly edit video 1, and modify the blue jacket to red."*
|
|
54
59
|
|
|
55
|
-
##
|
|
60
|
+
## Shot structure — "Shot 1 / Shot 2 / Shot 3" `[official :1563-1598]`
|
|
56
61
|
|
|
57
|
-
|
|
62
|
+
> *"Use shot order, write a simple 'Shot 1 / Shot 2 / Shot 3' storyboard for each segment of the video, and then merge them into a complete prompt."*
|
|
58
63
|
|
|
59
|
-
|
|
60
|
-
✅ "The earbud rises smoothly. The camera tracks upward."
|
|
64
|
+
**❌ Never second-stamp.** No `0:00–0:03`, no "At 4 seconds", no per-segment durations.
|
|
61
65
|
|
|
62
|
-
|
|
66
|
+
> *"Do not impose strict limits on the duration of each segment; prioritize allowing the model to naturally generate the pacing based on the plot."* `[:1580]`
|
|
67
|
+
>
|
|
68
|
+
> *"The model's support for precise timing (such as 0–3 seconds) is **unstable**, and forcibly limiting duration may lead to **abnormal generation results**."* `[:1586]`
|
|
63
69
|
|
|
64
|
-
|
|
70
|
+
Order shots by when events occur — primary first, secondary later. Let the plot set the pacing.
|
|
65
71
|
|
|
66
|
-
|
|
67
|
-
the lid opens in slow-motion · the blade whips through the air
|
|
68
|
-
```
|
|
72
|
+
**Per-shot internal order** `[official :1590-1598]` — organize each shot in exactly this sequence:
|
|
69
73
|
|
|
70
|
-
|
|
74
|
+
1. **Camera movement or shot transition** — "slowly push in from a wide shot", "fixed camera position", "cut to…"
|
|
75
|
+
2. **Subject actions and expressions** — the key actions and expression changes of the core character/object
|
|
76
|
+
3. **Position or spatial change** — where the subject is, and how that relationship shifts
|
|
77
|
+
4. **Audio** — sound effects, voices, background music for that shot
|
|
71
78
|
|
|
72
|
-
|
|
79
|
+
**One primary camera move per shot** — see Camera below. `[official :1648]`
|
|
73
80
|
|
|
74
|
-
|
|
81
|
+
## Subject binding — names + image indexes `[official :1488-1556]`
|
|
75
82
|
|
|
76
|
-
|
|
83
|
+
Every time a subject appears, it must be **explicitly referred to**. Two supported forms:
|
|
77
84
|
|
|
78
|
-
|
|
85
|
+
- **Undefined subjects** — bind inline every mention: `<Subject_N>@<Image_N>`. Official example: **`Zhang San@Image 1`**. `[:1540]`
|
|
86
|
+
- **Pre-declared subjects** — define once, then reuse the same label verbatim: *"Define the tall man in **Video 1** as **police officer**, and define the other short man as **thief**"*, then say "police officer" every time after. `[:1514]`
|
|
79
87
|
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
88
|
+
**One subject spread across several assets** — bind them together: *"Define `[…]` in **Image 1** and `[…]` in **Image 2** as `<Subject N>`."* `[:1514]`
|
|
89
|
+
|
|
90
|
+
⚠️ **An Asset ID must never substitute for `<Image/Video_N>`.** `[:1546]` *"the model cannot directly associate the Asset ID with the reference content."* Always cite by index.
|
|
91
|
+
|
|
92
|
+
Also official: keep descriptions concise, avoid redundancy, avoid semantic conflicts (contradictory traits for one subject), and prefer expressing spatial relationships through reference images rather than dense text. `[:1550-1556]`
|
|
93
|
+
|
|
94
|
+
**`[slates]`** — the app composes this for you. `composeReferences()` cites each canonical character or environment reference inline as `Name (image N)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*"Zhang San (corresponding to image 1)"* `[:1976]`). You never hand-write role labels or index numbers.
|
|
95
|
+
|
|
96
|
+
## Action description `[official :1602-1621]`
|
|
97
|
+
|
|
98
|
+
- **Body-part specificity + quantified degree.** Name hands, legs, head, shoulders, back — and supplement **range, speed, and force**. *"slowly raise a hand", "quickly turn the head", "push hard off the ground", "slightly lower the head."*
|
|
99
|
+
- **Prioritize slow, gentle, continuous small movements.** Avoid high-burst, large-dynamic actions — sprinting, big jumps, violent rolls. *"walk slowly", "gently raise a hand", "sit down naturally with the motion."* **This is the official basis for the folk rule that "fast" degrades quality** — it is not a banned token, it is a class of motion the model handles badly.
|
|
100
|
+
- **Supplement transitions between actions.** Specify inertia and continuity between consecutive beats so movement reads coherent: *"use the inertia of turning around to naturally raise a hand", "naturally transition from a pause into raising a hand."*
|
|
101
|
+
|
|
102
|
+
## Externalize emotion `[official :1623-1636]`
|
|
103
|
+
|
|
104
|
+
Replace abstract emotion words ("very sad", "extremely angry") with **specific physical detail**. This is the highest-leverage single habit in the official guide:
|
|
105
|
+
|
|
106
|
+
| Abstract | Externalized as actions and details |
|
|
107
|
+
|---|---|
|
|
108
|
+
| **Sadness** | head lowering, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the corner of clothing, tears welling but not falling |
|
|
109
|
+
| **Joy** | corners of the mouth rising uncontrollably, brows and eyes relaxing, steps becoming light, unconsciously humming a tune |
|
|
110
|
+
| **Nervousness / anxiety** | frequently checking the watch, fingers constantly tapping the tabletop, rapid breathing, eyes darting away |
|
|
111
|
+
| **Anger** | both fists clenched, jawline tense, chest heaving, eyes sharp, squeezing words out through gritted teeth |
|
|
112
|
+
| **Relief** | letting out a long breath, tense shoulders completely relaxing, a faint smile appearing, looking up toward the distance |
|
|
113
|
+
|
|
114
|
+
## Camera `[official :1643-1648]`
|
|
115
|
+
|
|
116
|
+
> *"The model has a **strong understanding of camera movement terms**, so you can **directly use standard camera movement terminology**, such as 'medium shot, close-up, wide shot, slow push-in, smooth lateral tracking, fixed shot.'"*
|
|
117
|
+
|
|
118
|
+
This is an **open vocabulary, not a fixed list** — and it explicitly includes **shot size** (close-up / medium / wide / long shot), which is as much a camera instruction as the move itself.
|
|
119
|
+
|
|
120
|
+
> ⚠️ *"Try to specify only 1 type of camera movement in a single shot. Do not require push, pull, pan, and move at the same time, as this will increase image instability."* `[:1648]`
|
|
121
|
+
|
|
122
|
+
## Image quality, style, and constraints `[official :1656-1679]`
|
|
83
123
|
|
|
84
|
-
|
|
124
|
+
These three slots "define creative boundaries for the model, unify image quality and artistic tone, and avoid visual flaws and random deviations."
|
|
125
|
+
|
|
126
|
+
**1. Image quality** — define clarity, texture detail, and lighting quality. Official vocabulary: `HD` · `rich details` · `cinematic texture` · `natural colors` · `soft lighting`.
|
|
127
|
+
|
|
128
|
+
> ⚠️ This is a **real slot with real vocabulary** — do not confuse it with Stable-Diffusion-era quality incantations. `8K` / `masterpiece` / `trending on artstation` remain banned slop tokens (see Part 3); *"cinematic texture, rich details, natural colors"* is the officially sanctioned way to ask for the same thing.
|
|
129
|
+
|
|
130
|
+
**2. Style** — the overall art style and visual tone: `cyberpunk cool blue-purple tone` · `retro film` · `fresh Japanese style`.
|
|
131
|
+
|
|
132
|
+
**3. Constraint words** — *"Constraint words are very important. They can effectively avoid visual flaws, deformities, breakdowns, and unreasonable elements."* Official templates, verbatim:
|
|
133
|
+
|
|
134
|
+
- **No subtitles** — "keep it subtitle-free" / "avoid generating any text or subtitles"
|
|
135
|
+
- **No logo** — "do not generate a logo"
|
|
136
|
+
- **No watermark** — "do not generate a watermark"
|
|
137
|
+
|
|
138
|
+
Seedance has **no `negativePrompt` field** — constraints go inline in this slot. See Part 3 for the wider inline-negative kit.
|
|
139
|
+
|
|
140
|
+
## 🔴 Duplicated characters — the twin problem `[official :1948-1994]`
|
|
141
|
+
|
|
142
|
+
**Symptom:** in frames with **many characters**, where **three-view / multi-view character images** are supplied as references, two identical characters appear in the same generated frame.
|
|
143
|
+
|
|
144
|
+
**Root causes** `[:1954-1959]`:
|
|
145
|
+
1. Character subjects are not clearly defined in the prompt, so the model cannot distinguish roles.
|
|
146
|
+
2. *"When character **three-view / multi-view images** are used as reference assets, it is easy to confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."*
|
|
147
|
+
|
|
148
|
+
**Official fixes, in their order** `[:1971-1994]` — ByteDance is explicit that these *reduce probability*, not eliminate it:
|
|
149
|
+
|
|
150
|
+
1. **Bind each character to its image explicitly**, in a consistent format. Official example: *"Zhang San (corresponding to image 1) throws the green passbook toward Li Si (corresponding to image 2), who is standing."*
|
|
151
|
+
2. **Append the global constraint verbatim** at the end of the prompt `[:1982]`:
|
|
152
|
+
> *"Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect. Keep only a single corresponding character in the same frame, and do not reproduce repeated copies of characters."*
|
|
153
|
+
3. **Optimize reference assets** `[:1988]` — *"For character reference images, prioritize independent single-person photos. Three-view or multi-view assets are not recommended."*
|
|
154
|
+
4. **Simplify the prompt** — do not paste a whole script; redundant copy confuses the model.
|
|
155
|
+
|
|
156
|
+
**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on identity sheets. Practical rule for Slates:
|
|
157
|
+
|
|
158
|
+
- **Multi-character Seedance shot** → bind every character to its image, append the anti-twin constraint, and prefer single-person / dominant-portrait references over multi-view sheets.
|
|
159
|
+
- **Single-character shot** → the standard character-sheet flow is fine.
|
|
160
|
+
|
|
161
|
+
**Too many reference people** `[official :2048-2052]` — past **4 reference people**, output stability drops (wrong headcount, duplicates). Official workaround: group the cast into images of ≤4 people each, generate those stills first, then drive the video from them.
|
|
162
|
+
|
|
163
|
+
## Worked examples `[official :1689-1745]`
|
|
164
|
+
|
|
165
|
+
These are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere.
|
|
166
|
+
|
|
167
|
+
**Example 1 — dormitory emotional short drama (dialogue-focused).** Assets: `@Image 1` half-body photo of the female lead · `@Image 2` dormitory scene reference · `@Video 1` camera-movement reference · `@Audio 1` indoor ambience.
|
|
168
|
+
|
|
169
|
+
> Use the girl in @Image 1 as the main character, use @Image 2 as the dormitory scene style reference, and refer to the camera movement in @Video 1.
|
|
170
|
+
>
|
|
171
|
+
> **Shot 1**: At dusk, **girl @Image 1** walks briskly to the **dormitory entrance @Image 2**. The camera follows steadily in a medium shot. Warm yellow sunlight spills into the hallway from the window. She pauses at the doorway, takes a deep breath, and looks slightly nervous.
|
|
172
|
+
>
|
|
173
|
+
> **Shot 2**: **Girl @Image 1** pushes the door open and enters the dormitory. The camera cuts to an indoor medium shot. Her roommates look up at her while organizing their books. One of them smiles and asks {How did the exam go? Did you pass?}. The camera slowly cuts between half-body close-ups of several people.
|
|
174
|
+
>
|
|
175
|
+
> **Shot 3**: **Girl @Image 1** first lowers her head with a dejected expression. The camera gives her a close-up. Then she raises her head, unable to hold back a smile, laughs out loud, and says {I was kidding}. Her roommates start chasing and play-fighting with her. The camera slowly pulls back and freezes on a wide shot of the dormitory filled with laughter.
|
|
176
|
+
>
|
|
177
|
+
> The entire video should have a high-definition cinematic documentary style, with warm tones and soft lighting. The character's face remains stable without deformation; movements are natural and smooth, with no stutter or flicker. The ambient sound blends naturally with @Audio 1.
|
|
178
|
+
|
|
179
|
+
**Example 2 — ancient-style cliff confrontation (action/atmosphere-focused).** Assets: `@Image 1` female lead in red · `@Image 2` assassin in black · `@Image 3` cliff and bamboo forest · `@Video 1` martial-arts camera reference · `@Audio 1` drum beats.
|
|
180
|
+
|
|
181
|
+
> Use the woman in red from @Image 1 as the female lead, use the woman in black from @Image 2 as the opponent, use the cliff and bamboo forest environment in @Image 3 as the scene reference, refer to the overall camera movement and action rhythm in @Video 1, and synchronize the background sound effects with @Audio 1.
|
|
182
|
+
>
|
|
183
|
+
> **Shot 1**: At dusk, the camera slowly pushes in from a side medium shot of **woman in red @Image 1**. She stands at the edge of the cliff and lifts a wine flask to drink. Her sleeves and robe hem sway gently in the mountain wind. The camera circles halfway around her, moving from the front to her back. In the distance, a figure in black is faintly visible in the bamboo forest.
|
|
184
|
+
>
|
|
185
|
+
> **Shot 2**: The camera zooms and fades into a long shot. From a drone perspective, it overlooks the entire cliff and bamboo forest. The two characters stand at opposite ends of the cliff. The mountain wind lifts their robe hems and dust, and the rhythm slightly accelerates with the drum beats.
|
|
186
|
+
>
|
|
187
|
+
> **Shot 3**: The camera cuts back to a ground-level close shot. The two slowly draw their swords and face off. **Woman in red @Image 1** shifts from a careless expression to a cold gaze. **Woman in black @Image 2** looks determined, and the sword tip trembles slightly. The camera steadily follows the two as they circle each other, finally freezing on a close-up of the instant before the two swords meet.
|
|
188
|
+
>
|
|
189
|
+
> The overall visual style should feel like a cinematic wuxia world in misty rain, with cool tones, low saturation, a film-grain texture, and rich light-and-shadow layers. The characters' faces and body proportions remain stable without deformation. Movements are continuous and natural, not stiff, with no clipping or stutter.
|
|
190
|
+
|
|
191
|
+
## Other official notes
|
|
192
|
+
|
|
193
|
+
- **On-screen text** `[official :1758]` — Seedance can render common text (ad slogans, subtitles, speech bubbles) and will auto-match style/colour from context, or take an explicit colour / style / timing / position. Prefer **common characters**; avoid rare glyphs and special symbols. (For *guaranteed* legible text, the start-frame route in Part 3 is still safer.)
|
|
194
|
+
- **Extension degrades quality** `[official :2004-2024]` — using a generated video as the input for extension compounds degradation, with mottled colour blocks in face regions. Limit repeated continuations; prefer HD assets as input.
|
|
195
|
+
- **Special effects that miss** `[official :2031-2044]` — when a described effect comes out wrong (a countdown that scrolls randomly), define it with a **reference video** instead of words: *"the way the number '2999' appears should reference video 1."*
|
|
196
|
+
|
|
197
|
+
---
|
|
198
|
+
|
|
199
|
+
# Part 2 — Slates-specific `[slates]`
|
|
200
|
+
|
|
201
|
+
## Reference media — caps and transport
|
|
202
|
+
|
|
203
|
+
Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
|
|
204
|
+
|
|
205
|
+
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
|
|
85
206
|
|
|
86
207
|
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
87
208
|
|
|
88
|
-
These aren't separate Seedance features — they're prompting strategies over reference media
|
|
209
|
+
These aren't separate Seedance features — they're prompting strategies over reference media.<!-- slates-only --> The Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:<!-- /slates-only -->
|
|
89
210
|
|
|
90
|
-
- **Motion transfer:** subject image as a reference + the driving clip via `videoReferenceAssetId
|
|
91
|
-
- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with
|
|
211
|
+
- **Motion transfer:** subject image as a reference + the driving clip<!-- slates-only --> via `videoReferenceAssetId`<!-- /slates-only --> (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`
|
|
212
|
+
- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`
|
|
92
213
|
- **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script → their likeness AND voice deliver the new line.
|
|
214
|
+
<!-- slates-only -->
|
|
93
215
|
- **Billing:** a reference VIDEO switches the cost key to `seedance-2*-vref-{res}-{T}s` where T = clip seconds + output seconds — quote before confirming. Audio references are free (audio is included on every route).
|
|
216
|
+
<!-- /slates-only -->
|
|
94
217
|
|
|
218
|
+
<!-- slates-only -->
|
|
95
219
|
## Faces — set `seedanceFace` for AI-character faces
|
|
96
220
|
|
|
97
221
|
Seedance routes through **three tiers** depending on the face in the reference, exposed as the "Face in Reference" toggle plus the real-face params on `slates_generate_video`:
|
|
@@ -102,37 +226,121 @@ Seedance routes through **three tiers** depending on the face in the reference,
|
|
|
102
226
|
|
|
103
227
|
Rules:
|
|
104
228
|
- **The real-vs-AI call is the PROVIDER'S, not yours.** ByteDance's classifier is probabilistic — some real photos pass the standard face route (billed at the cheap rate; fine), others get rejected with `[REAL_FACE_DETECTED]` (auto-refunded). Don't preemptively route to the real-face tier just because a photo looks real; try `seedanceFace: true` first and escalate only on the marked rejection. Public figures / celebrities fail on every route.
|
|
105
|
-
- It's about the **reference, not the output.** If your character
|
|
229
|
+
- It's about the **reference, not the output.** If your character identity or generated portrait shows a face, turn it on. A product shot with no person stays off.
|
|
106
230
|
- Don't toggle it on "just in case" — a faceless gen on the face route burns ~45% extra for nothing.
|
|
231
|
+
<!-- /slates-only -->
|
|
107
232
|
|
|
108
233
|
## Reference rules (the verified ones)
|
|
109
234
|
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
- **Character identity: attach the turnaround AND the close-up expression sheet, named as one entity** — cite both under the same name. The shared name — not a role essay — is what keeps the varied expressions from averaging the face; don't gate the expression sheet, and don't tell it to "render neutral / ignore the outfit" (the user's prompt owns expression, wardrobe, and lighting). The trend is MORE references (video/audio into Seedance), all addressed by name — lean into attaching rich refs and let the naming do the work.
|
|
114
|
-
- **Flat-lit identity refs.** A studio-lit / scene-lit character sheet bleeds its lighting into the clip ("green-screen pasted in front of mountains"). Prep refs flat and plain.
|
|
115
|
-
- **Environment: describe it, don't feed a grid.** Default to words and let the model build the space to fit; reserve an environment ref for a hard exact-match, and then use ONE clean establishing image — never a multi-panel grid.
|
|
116
|
-
- **Reuse the same refs across every shot** in a sequence — swapping mid-sequence drifts.
|
|
117
|
-
- **Legible on-screen text → bake it into an NB2 start frame** and animate from it; Seedance won't render clean text from scratch.
|
|
118
|
-
- **Grids are for EXPLORING compositions, not for inputting** — pick a cell, don't feed the grid back as a reference.
|
|
235
|
+
<!-- @inject:references-read-literally -->
|
|
236
|
+
> **The general law: the model reads a reference literally.**
|
|
237
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
119
238
|
|
|
120
|
-
|
|
239
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
121
240
|
|
|
122
|
-
**
|
|
241
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
242
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
243
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
244
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
123
245
|
|
|
124
|
-
|
|
125
|
-
-
|
|
126
|
-
- `maintain exact appearance from reference image`
|
|
127
|
-
- `consistent character throughout, no deformation or drift`
|
|
246
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
247
|
+
<!-- @end:references-read-literally -->
|
|
128
248
|
|
|
129
|
-
|
|
249
|
+
<!-- @inject:reference-rules-core -->
|
|
250
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
251
|
+
|
|
252
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
253
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
254
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
255
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
256
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
257
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
258
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
259
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
260
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
261
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
262
|
+
<!-- @end:reference-rules-core -->
|
|
263
|
+
|
|
264
|
+
### For Seedance specifically
|
|
265
|
+
|
|
266
|
+
- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say "still / scene / from a movie / from the image." The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one<!-- slates-only --> — see slates-cost-discipline<!-- /slates-only -->).
|
|
267
|
+
- **Seedance's own idiom for rule 2 is `Reference <Subject_N> in <Image_N>`** `[official :1389]` — `Image_N` indexes the order the refs are attached, so the name plus the index carries the role. The full binding grammar is in Part 1 (Subject binding).
|
|
268
|
+
- **Rule 3 has an official ceiling here.** The trend is MORE references (video and audio into Seedance), all addressed by name — but for **multi-character frames** see the twin-problem section above: bind every character to its image, append the anti-twin constraint, and prefer single-person references. Past 4 reference people, stability drops `[official :2048-2052]`.
|
|
269
|
+
- **Rule 8 holds even though Seedance can render common text natively** `[official :1758]`. A baked NB2 start frame is still the reliable route for text that must be legible.
|
|
270
|
+
- **Rule 5 pairs with the first/last-frame exclusion** — frames and reference images are mutually exclusive on this model (see Reference media above), so an environment you must match exactly costs you the frame lane.
|
|
271
|
+
|
|
272
|
+
<!-- slates-only -->
|
|
273
|
+
## Pre-flight: references arrive inline, refer by code
|
|
274
|
+
|
|
275
|
+
When you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** — if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.
|
|
276
|
+
|
|
277
|
+
When talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Beach Sunset`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.
|
|
278
|
+
|
|
279
|
+
- ✅ "I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame — the camera move is going to be a slow dolly forward through the gap."
|
|
280
|
+
- ❌ "I'm using the first beach image and the last one..." (which? They have four.)
|
|
281
|
+
<!-- /slates-only -->
|
|
282
|
+
|
|
283
|
+
---
|
|
284
|
+
|
|
285
|
+
# Part 3 — Community field notes `[community]`
|
|
286
|
+
|
|
287
|
+
Third-party guides and Slates field experience. Useful heuristics — but if one of these ever appears to contradict Part 1, **Part 1 wins**.
|
|
288
|
+
|
|
289
|
+
## Length
|
|
290
|
+
|
|
291
|
+
**Sweet spot 60-150 words** for a single shot (not 150-300 — that's the upper bound). Multi-shot storyboards run longer; official Example 1 above is ~230 words across three shots.
|
|
292
|
+
|
|
293
|
+
## Pin the subject in the first 20-30 words
|
|
294
|
+
|
|
295
|
+
The opening sentence is the **identity anchor**. If the subject isn't locked early, the model hallucinates new subjects mid-clip. (Compatible with Part 1: the binding preamble comes before `Shot 1`.)
|
|
296
|
+
|
|
297
|
+
```
|
|
298
|
+
A matte black earbud case sits on a polished obsidian surface...
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
## Lighting is a top quality lever
|
|
302
|
+
|
|
303
|
+
Lighting has an outsized impact on output quality — which is why it has its own slot in the official 8-slot formula. Describe it before or alongside the subject.
|
|
304
|
+
|
|
305
|
+
```
|
|
306
|
+
A cool-white diagonal beam from upper left, dust particles drifting through.
|
|
307
|
+
Soft golden hour lighting from low west angle.
|
|
308
|
+
Dramatic rim light against dark background.
|
|
309
|
+
```
|
|
310
|
+
|
|
311
|
+
## Camera and subject motion — separate sentences
|
|
312
|
+
|
|
313
|
+
Mixing them is a common cause of glitchy / shaky output.
|
|
314
|
+
|
|
315
|
+
❌ "The camera speed ramps as the earbud rises."
|
|
316
|
+
✅ "The earbud rises smoothly. The camera tracks upward."
|
|
317
|
+
|
|
318
|
+
## Slow-motion works; "fast" is a known bad token
|
|
319
|
+
|
|
320
|
+
Speed ramps and slow-motion are supported in natural language, and `fast` is widely reported as a quality-degrading keyword. **The official version of this rule is stronger and better founded** — prioritize slow, gentle, continuous small movements and avoid high-burst action (Part 1, Action description `[:1611-1615]`). Prompt the motion class, not the adjective.
|
|
321
|
+
|
|
322
|
+
```
|
|
323
|
+
the lid opens in slow-motion · the blade whips through the air
|
|
324
|
+
```
|
|
325
|
+
|
|
326
|
+
**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`. These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
|
|
327
|
+
|
|
328
|
+
## Style block at the end
|
|
329
|
+
|
|
330
|
+
One primary anchor + 2-3 supporting details, as the trailing paragraph (both official examples do exactly this). End with `Single continuous take` if you want one shot with no cuts. **Never** write `no cut` or `seamless transition` — not in the training vocabulary.
|
|
331
|
+
|
|
332
|
+
## ⚠️ Don't cross-pollinate image-model syntax
|
|
333
|
+
|
|
334
|
+
Named **lenses, apertures, film stocks, and camera bodies** — `85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`, `shot on Sony A7S3` — are an **image-model lever** (correct and encouraged in `slates-prompting-nano-banana-2`) and a **Seedance anti-pattern**. ByteDance's guide uses shot sizes, camera moves, pacing words, and the image-quality/style vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.
|
|
335
|
+
|
|
336
|
+
If you are carrying a look over from an NB2 start frame, translate it: `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`.
|
|
130
337
|
|
|
131
338
|
## Negative prompting — inline only
|
|
132
339
|
|
|
133
|
-
Seedance has **no `negativePrompt` field**.
|
|
340
|
+
Seedance has **no `negativePrompt` field**. Put negatives in the constraints slot, led by the three official templates (Part 1):
|
|
134
341
|
|
|
135
342
|
```
|
|
343
|
+
keep it subtitle-free · do not generate a logo · do not generate a watermark
|
|
136
344
|
avoid jitter and bent limbs
|
|
137
345
|
avoid temporal flicker
|
|
138
346
|
avoid identity drift
|
|
@@ -141,38 +349,35 @@ no distortion, no stretching
|
|
|
141
349
|
|
|
142
350
|
Also fine: positive reframing ("empty street" not "no cars").
|
|
143
351
|
|
|
352
|
+
## Image-to-video / first-frame guidance
|
|
353
|
+
|
|
354
|
+
**Describe motion, not image.** The model already sees the visual; tokens spent re-describing appearance are wasted.
|
|
355
|
+
|
|
356
|
+
Stability phrases that help:
|
|
357
|
+
- `preserve composition and colors`
|
|
358
|
+
- `maintain exact appearance from reference image`
|
|
359
|
+
- `consistent character throughout, no deformation or drift`
|
|
360
|
+
|
|
361
|
+
**Cap I2V prompts under 60 words** when possible. Over 100 words frequently triggers silent generation failure.
|
|
362
|
+
|
|
144
363
|
## Common failure modes + fixes
|
|
145
364
|
|
|
146
365
|
| Failure | Fix |
|
|
147
366
|
|---|---|
|
|
148
367
|
| Hallucinated subject mid-clip | First 20-30 words = identity anchor |
|
|
149
368
|
| Bent limbs / extra fingers | `avoid jitter and bent limbs` in Constraints |
|
|
150
|
-
| Identity drift across multi-shot |
|
|
369
|
+
| Identity drift across multi-shot | Re-name the bound subject in **every** `Shot N` block `[official :1537]` |
|
|
370
|
+
| Two identical characters in one frame | The twin fix in Part 1 — bind each character to its image + append the global anti-twin constraint |
|
|
151
371
|
| Silent generation failure on I2V | Cut prompt under 100 words, single primary camera move |
|
|
152
372
|
| Speech / motion conflict | Limit dialogue to one line per action shot |
|
|
153
|
-
|
|
154
|
-
## Benchmark prompts (verbatim from authoritative sources)
|
|
155
|
-
|
|
156
|
-
**Single-shot (fal.ai):**
|
|
157
|
-
> "A golden retriever runs across a sandy beach at sunset, kicking up wet sand with each stride, the camera tracking alongside at ground level. Waves crash softly in the background."
|
|
158
|
-
|
|
159
|
-
**Multi-shot commercial (fal.ai):**
|
|
160
|
-
> "Shot 1: extreme close-up of condensation dripping down a glass bottle, the sound of ice clinking. Shot 2: the bottle rises from a bed of crushed ice, camera tilting up slowly, bright backlight creating a halo effect. Shot 3: a hand grabs the bottle against a sunset rooftop backdrop, the city humming below."
|
|
161
|
-
|
|
162
|
-
**Cinematic anchor (atlabs):**
|
|
163
|
-
> "Modern Rural Aesthetics, Cinematic Commercial quality, shot with Sony A7S3/cinema camera, 4K/8K ultra-clear, Extreme Macro, natural transparent lighting, healing ASMR, no historical costume drama feel."
|
|
164
|
-
|
|
165
|
-
## Pre-flight: references arrive inline, refer by code
|
|
166
|
-
|
|
167
|
-
When you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** — if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.
|
|
168
|
-
|
|
169
|
-
When talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Beach Sunset`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.
|
|
170
|
-
|
|
171
|
-
- ✅ "I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame — the camera move is going to be a slow dolly forward through the gap."
|
|
172
|
-
- ❌ "I'm using the first beach image and the last one..." (which? They have four.)
|
|
373
|
+
| Erratic/random pacing | You second-stamped. Remove all time markers and use `Shot N` `[official :1586]` |
|
|
173
374
|
|
|
174
375
|
## Sources
|
|
175
376
|
|
|
377
|
+
**Official (authoritative):**
|
|
378
|
+
- BytePlus ModelArk — Seedance 2.0 prompting guide, archived at `research/seedance-2-modelark-docs.md` (all `:NNNN` refs above)
|
|
379
|
+
|
|
380
|
+
**Community (secondary):**
|
|
176
381
|
- [fal.ai — How to Use Seedance 2.0](https://fal.ai/learn/tools/how-to-use-seedance-2-0)
|
|
177
382
|
- [apiyi.com — Seedance 2.0 Prompt Guide](https://help.apiyi.com/en/seedance-2-0-prompt-guide-video-generation-camera-style-tips-en.html)
|
|
178
383
|
- [atlabs.ai — Ultimate Seedance 2.0 Prompting Guide](https://www.atlabs.ai/blog/the-ultimate-seedance-2.0-prompting-guide-47-prompts-2026)
|
|
@@ -98,10 +98,39 @@ Verbatim example:
|
|
|
98
98
|
|
|
99
99
|
## Reference discipline (character / environment refs)
|
|
100
100
|
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
101
|
+
<!-- @inject:references-read-literally -->
|
|
102
|
+
> **The general law: the model reads a reference literally.**
|
|
103
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
104
|
+
|
|
105
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
106
|
+
|
|
107
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
108
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
109
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
110
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
111
|
+
|
|
112
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
113
|
+
<!-- @end:references-read-literally -->
|
|
114
|
+
|
|
115
|
+
<!-- @inject:reference-rules-core -->
|
|
116
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
117
|
+
|
|
118
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
119
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
120
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
121
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
122
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
123
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
124
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
125
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
126
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
127
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
128
|
+
<!-- @end:reference-rules-core -->
|
|
129
|
+
|
|
130
|
+
### For Veo specifically
|
|
131
|
+
|
|
132
|
+
- **Veo's idiom for rule 2 is plain-English role naming in the sentence itself** — *"Using the provided images for the detective, the woman, and the office setting, create a medium shot of…"* (see Ingredients-to-Video above). The role rides in the noun phrase, not in a separate label block.
|
|
133
|
+
- **Rule 8 has a second reason to matter here:** Veo bakes subtitle text into the frame unless every dialogue line carries `(no subtitles)`. Text you did not ask for is the failure mode, not just text you did.
|
|
105
134
|
|
|
106
135
|
## Negative prompting — nouns, not instructions
|
|
107
136
|
|
|
@@ -29,6 +29,25 @@ Surface the planned structure back to the user as a tight summary:
|
|
|
29
29
|
> Scene 2: Confrontation (4 frames)
|
|
30
30
|
> ...
|
|
31
31
|
|
|
32
|
+
**Surface a decision log alongside that summary.**
|
|
33
|
+
|
|
34
|
+
<!-- @inject:decision-log -->
|
|
35
|
+
When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:
|
|
36
|
+
|
|
37
|
+
```
|
|
38
|
+
source phrase or declared default → what you wrote → what it resolves
|
|
39
|
+
"in a diner" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable
|
|
40
|
+
(no time of day) → late afternoon, low warm key → default; say the word and it changes
|
|
41
|
+
(no camera) → slow push-in, single move → one move per shot; stacking increases instability
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
**Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the "why did you add that?" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.
|
|
45
|
+
|
|
46
|
+
> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
|
|
47
|
+
<!-- @end:decision-log -->
|
|
48
|
+
|
|
49
|
+
Turning a script into *visual* frame prompts means resolving things the script left open — what the room looks like, where the light comes from, how the shot is framed. Those are your decisions, not the writer's; name them.
|
|
50
|
+
|
|
32
51
|
Ask: **"Generate frame images now? (y/N)"**
|
|
33
52
|
|
|
34
53
|
### 3. Generate frames if requested
|