@slatesvideo/shared 0.5.3 → 0.5.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +31 -0
- package/dist/operations/index.d.ts +4 -8
- package/dist/operations/index.js +26 -38
- package/dist/prompts/character-sheet.d.ts +19 -18
- package/dist/prompts/character-sheet.js +84 -38
- package/dist/prompts/environment-sheet.d.ts +9 -1
- package/dist/prompts/environment-sheet.js +17 -3
- package/dist/prompts/model-facts.js +4 -1
- package/dist/prompts/partials.generated.d.ts +2 -0
- package/dist/prompts/partials.generated.js +14 -0
- package/dist/prompts/prompting-tips.js +50 -19
- package/dist/prompts/reference-composer.d.ts +1 -1
- package/dist/prompts/reference-composer.js +3 -4
- package/dist/prompts/reference-rules.d.ts +43 -14
- package/dist/prompts/reference-rules.js +51 -27
- package/dist/skills/content.js +14 -14
- package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
- package/exports/slates-prompt-builder/generated/reference-character.md +78 -0
- package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
- package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
- package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +8 -3
- package/skills/_partials/decision-log.md +12 -0
- package/skills/_partials/reference-rules-core.md +12 -0
- package/skills/_partials/reference-tips-short.md +2 -0
- package/skills/_partials/references-read-literally.md +11 -0
- package/skills/_partials/still-gate.md +3 -0
- package/skills/slates-character-identity.md +100 -0
- package/skills/slates-cost-discipline.md +10 -0
- package/skills/slates-edit-and-iterate.md +17 -2
- package/skills/slates-model-selection.md +24 -1
- package/skills/slates-one-prompt-film.md +22 -3
- package/skills/slates-prompting-flux-2-max.md +36 -5
- package/skills/slates-prompting-gpt-image-2.md +1 -1
- package/skills/slates-prompting-kling-v3.md +40 -9
- package/skills/slates-prompting-nano-banana-2.md +44 -12
- package/skills/slates-prompting-omni-flash.md +1 -1
- package/skills/slates-prompting-seedance.md +295 -90
- package/skills/slates-prompting-veo-3.md +33 -4
- package/skills/slates-storyboard-from-script.md +19 -0
- package/skills/slates-vision-feedback-loop.md +49 -2
- package/skills/slates-character-turnaround.md +0 -55
|
@@ -0,0 +1,353 @@
|
|
|
1
|
+
<!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
|
|
2
|
+
|
|
3
|
+
> **This is the real thing.** Every rule below is the working doctrine Slates runs in production against this model — not a summary written for a handout. Slates automates it end to end; the doctrine works by hand too.
|
|
4
|
+
|
|
5
|
+
# Seedance 2.0 — prompting
|
|
6
|
+
|
|
7
|
+
ByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K — 4K video is Pro-only, default 1080p), 4–15s, first+last frame, and up to 9 reference images / 3 videos / 3 audio clips.
|
|
8
|
+
|
|
9
|
+
> **How to read this file.**
|
|
10
|
+
> **[official :NNNN]** — ByteDance's own BytePlus ModelArk prompting guide, line `NNNN` of the archived doc dump (`research/seedance-2-modelark-docs.md`). Receipt-grade; treat as law.
|
|
11
|
+
> **[community]** — third-party guides and our own field notes. Useful, but an `[official]` block always wins.
|
|
12
|
+
> **[slates]** — how the Slates app composes or bills this; not ByteDance doctrine.
|
|
13
|
+
>
|
|
14
|
+
> The split is load-bearing. A community-sourced "narrative timing beats" doctrine shipped in this file for months teaching the **exact inverse** of ByteDance's published guidance. Never merge the two registers again.
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
# Part 1 — Official ByteDance doctrine
|
|
19
|
+
|
|
20
|
+
## What Seedance actually is `[official :1450-1452]`
|
|
21
|
+
|
|
22
|
+
Seedance 2.0 is a multimodal AI director. It reads text, images, video and audio **simultaneously** and internally decomposes them into two dimensions:
|
|
23
|
+
|
|
24
|
+
- the **spatial layer** — what is in the frame
|
|
25
|
+
- the **temporal layer** — how things change over time
|
|
26
|
+
|
|
27
|
+
So a good prompt is **not "copywriting-style description" but an "engineering-style instruction"**: who, in what scene, doing what action, how the camera moves, and in what chronological order events occur — delivered respectively to the spatial layer and the temporal layer.
|
|
28
|
+
|
|
29
|
+
## The advanced formula — 8 slots `[official :1455]`
|
|
30
|
+
|
|
31
|
+
```text
|
|
32
|
+
precise subject + action details + scene/environment + lighting & color tone
|
|
33
|
+
+ camera movement + visual style + image quality + constraints
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
⚠️ There is **no official "6-step formula."** `Subject + Action + Environment + Camera + Style + Constraints` is community branding with no ByteDance source, and it silently drops the **lighting & color tone** and **image quality** slots. Use the 8 slots above.
|
|
37
|
+
|
|
38
|
+
## Task-type sentence patterns `[official :1389-1425]`
|
|
39
|
+
|
|
40
|
+
Seedance classifies your request from the phrasing. Use the pattern that matches the task:
|
|
41
|
+
|
|
42
|
+
| Task | Pattern |
|
|
43
|
+
|---|---|
|
|
44
|
+
| **Image reference** | ``Reference `<Subject_N>` in `<Image_N>` to generate…`` |
|
|
45
|
+
| **Video reference** | ``Reference `<Action / Camera_movement / Style / Sound_effect>` in `<Video_N>` to generate…`` |
|
|
46
|
+
| **Audio reference** | ``Reference the timbre in `<Audio_N>` to generate…`` |
|
|
47
|
+
| **Video edit — modify** | ``Strictly edit `<Video_N>`, and modify `<Original_Characteristic>` in it to `<New_Characteristic>``` |
|
|
48
|
+
| **Video edit — add** | ``<Element_Features>` + `<Timing>` + `<Location>`` |
|
|
49
|
+
| **Video edit — delete** | Name what to delete; for anything that must stay, say so explicitly |
|
|
50
|
+
| **Video extend** | ``Extend `<Video_N>` forward/backward to generate…`` |
|
|
51
|
+
| **Combined** | ``Reference `[Dimension]` of `<Image/Video_N>`, strictly edit `<Video_X>`, `[Specific_Edits]``` |
|
|
52
|
+
|
|
53
|
+
### ⚠️ Edit / extend phrasing landmine `[official :1431]`
|
|
54
|
+
|
|
55
|
+
> *"For edit / extend video tasks, directly use `<Video_N>` to refer to the video. **Do not use "reference `<Video_N>`"**, to avoid being incorrectly identified as a reference task."*
|
|
56
|
+
|
|
57
|
+
This is easy to trip: Slates has an edit lane. Writing *"reference video 1 and change the jacket to red"* gets classified as a **reference** task — the model generates a brand-new clip inspired by the source instead of editing it. Write *"Strictly edit video 1, and modify the blue jacket to red."*
|
|
58
|
+
|
|
59
|
+
## Shot structure — "Shot 1 / Shot 2 / Shot 3" `[official :1563-1598]`
|
|
60
|
+
|
|
61
|
+
> *"Use shot order, write a simple 'Shot 1 / Shot 2 / Shot 3' storyboard for each segment of the video, and then merge them into a complete prompt."*
|
|
62
|
+
|
|
63
|
+
**❌ Never second-stamp.** No `0:00–0:03`, no "At 4 seconds", no per-segment durations.
|
|
64
|
+
|
|
65
|
+
> *"Do not impose strict limits on the duration of each segment; prioritize allowing the model to naturally generate the pacing based on the plot."* `[:1580]`
|
|
66
|
+
>
|
|
67
|
+
> *"The model's support for precise timing (such as 0–3 seconds) is **unstable**, and forcibly limiting duration may lead to **abnormal generation results**."* `[:1586]`
|
|
68
|
+
|
|
69
|
+
Order shots by when events occur — primary first, secondary later. Let the plot set the pacing.
|
|
70
|
+
|
|
71
|
+
**Per-shot internal order** `[official :1590-1598]` — organize each shot in exactly this sequence:
|
|
72
|
+
|
|
73
|
+
1. **Camera movement or shot transition** — "slowly push in from a wide shot", "fixed camera position", "cut to…"
|
|
74
|
+
2. **Subject actions and expressions** — the key actions and expression changes of the core character/object
|
|
75
|
+
3. **Position or spatial change** — where the subject is, and how that relationship shifts
|
|
76
|
+
4. **Audio** — sound effects, voices, background music for that shot
|
|
77
|
+
|
|
78
|
+
**One primary camera move per shot** — see Camera below. `[official :1648]`
|
|
79
|
+
|
|
80
|
+
## Subject binding — names + image indexes `[official :1488-1556]`
|
|
81
|
+
|
|
82
|
+
Every time a subject appears, it must be **explicitly referred to**. Two supported forms:
|
|
83
|
+
|
|
84
|
+
- **Undefined subjects** — bind inline every mention: `<Subject_N>@<Image_N>`. Official example: **`Zhang San@Image 1`**. `[:1540]`
|
|
85
|
+
- **Pre-declared subjects** — define once, then reuse the same label verbatim: *"Define the tall man in **Video 1** as **police officer**, and define the other short man as **thief**"*, then say "police officer" every time after. `[:1514]`
|
|
86
|
+
|
|
87
|
+
**One subject spread across several assets** — bind them together: *"Define `[…]` in **Image 1** and `[…]` in **Image 2** as `<Subject N>`."* `[:1514]`
|
|
88
|
+
|
|
89
|
+
⚠️ **An Asset ID must never substitute for `<Image/Video_N>`.** `[:1546]` *"the model cannot directly associate the Asset ID with the reference content."* Always cite by index.
|
|
90
|
+
|
|
91
|
+
Also official: keep descriptions concise, avoid redundancy, avoid semantic conflicts (contradictory traits for one subject), and prefer expressing spatial relationships through reference images rather than dense text. `[:1550-1556]`
|
|
92
|
+
|
|
93
|
+
**`[slates]`** — the app composes this for you. `composeReferences()` cites each canonical character or environment reference inline as `Name (image N)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*"Zhang San (corresponding to image 1)"* `[:1976]`). You never hand-write role labels or index numbers.
|
|
94
|
+
|
|
95
|
+
## Action description `[official :1602-1621]`
|
|
96
|
+
|
|
97
|
+
- **Body-part specificity + quantified degree.** Name hands, legs, head, shoulders, back — and supplement **range, speed, and force**. *"slowly raise a hand", "quickly turn the head", "push hard off the ground", "slightly lower the head."*
|
|
98
|
+
- **Prioritize slow, gentle, continuous small movements.** Avoid high-burst, large-dynamic actions — sprinting, big jumps, violent rolls. *"walk slowly", "gently raise a hand", "sit down naturally with the motion."* **This is the official basis for the folk rule that "fast" degrades quality** — it is not a banned token, it is a class of motion the model handles badly.
|
|
99
|
+
- **Supplement transitions between actions.** Specify inertia and continuity between consecutive beats so movement reads coherent: *"use the inertia of turning around to naturally raise a hand", "naturally transition from a pause into raising a hand."*
|
|
100
|
+
|
|
101
|
+
## Externalize emotion `[official :1623-1636]`
|
|
102
|
+
|
|
103
|
+
Replace abstract emotion words ("very sad", "extremely angry") with **specific physical detail**. This is the highest-leverage single habit in the official guide:
|
|
104
|
+
|
|
105
|
+
| Abstract | Externalized as actions and details |
|
|
106
|
+
|---|---|
|
|
107
|
+
| **Sadness** | head lowering, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the corner of clothing, tears welling but not falling |
|
|
108
|
+
| **Joy** | corners of the mouth rising uncontrollably, brows and eyes relaxing, steps becoming light, unconsciously humming a tune |
|
|
109
|
+
| **Nervousness / anxiety** | frequently checking the watch, fingers constantly tapping the tabletop, rapid breathing, eyes darting away |
|
|
110
|
+
| **Anger** | both fists clenched, jawline tense, chest heaving, eyes sharp, squeezing words out through gritted teeth |
|
|
111
|
+
| **Relief** | letting out a long breath, tense shoulders completely relaxing, a faint smile appearing, looking up toward the distance |
|
|
112
|
+
|
|
113
|
+
## Camera `[official :1643-1648]`
|
|
114
|
+
|
|
115
|
+
> *"The model has a **strong understanding of camera movement terms**, so you can **directly use standard camera movement terminology**, such as 'medium shot, close-up, wide shot, slow push-in, smooth lateral tracking, fixed shot.'"*
|
|
116
|
+
|
|
117
|
+
This is an **open vocabulary, not a fixed list** — and it explicitly includes **shot size** (close-up / medium / wide / long shot), which is as much a camera instruction as the move itself.
|
|
118
|
+
|
|
119
|
+
> ⚠️ *"Try to specify only 1 type of camera movement in a single shot. Do not require push, pull, pan, and move at the same time, as this will increase image instability."* `[:1648]`
|
|
120
|
+
|
|
121
|
+
## Image quality, style, and constraints `[official :1656-1679]`
|
|
122
|
+
|
|
123
|
+
These three slots "define creative boundaries for the model, unify image quality and artistic tone, and avoid visual flaws and random deviations."
|
|
124
|
+
|
|
125
|
+
**1. Image quality** — define clarity, texture detail, and lighting quality. Official vocabulary: `HD` · `rich details` · `cinematic texture` · `natural colors` · `soft lighting`.
|
|
126
|
+
|
|
127
|
+
> ⚠️ This is a **real slot with real vocabulary** — do not confuse it with Stable-Diffusion-era quality incantations. `8K` / `masterpiece` / `trending on artstation` remain banned slop tokens (see Part 3); *"cinematic texture, rich details, natural colors"* is the officially sanctioned way to ask for the same thing.
|
|
128
|
+
|
|
129
|
+
**2. Style** — the overall art style and visual tone: `cyberpunk cool blue-purple tone` · `retro film` · `fresh Japanese style`.
|
|
130
|
+
|
|
131
|
+
**3. Constraint words** — *"Constraint words are very important. They can effectively avoid visual flaws, deformities, breakdowns, and unreasonable elements."* Official templates, verbatim:
|
|
132
|
+
|
|
133
|
+
- **No subtitles** — "keep it subtitle-free" / "avoid generating any text or subtitles"
|
|
134
|
+
- **No logo** — "do not generate a logo"
|
|
135
|
+
- **No watermark** — "do not generate a watermark"
|
|
136
|
+
|
|
137
|
+
Seedance has **no `negativePrompt` field** — constraints go inline in this slot. See Part 3 for the wider inline-negative kit.
|
|
138
|
+
|
|
139
|
+
## 🔴 Duplicated characters — the twin problem `[official :1948-1994]`
|
|
140
|
+
|
|
141
|
+
**Symptom:** in frames with **many characters**, where **three-view / multi-view character images** are supplied as references, two identical characters appear in the same generated frame.
|
|
142
|
+
|
|
143
|
+
**Root causes** `[:1954-1959]`:
|
|
144
|
+
1. Character subjects are not clearly defined in the prompt, so the model cannot distinguish roles.
|
|
145
|
+
2. *"When character **three-view / multi-view images** are used as reference assets, it is easy to confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."*
|
|
146
|
+
|
|
147
|
+
**Official fixes, in their order** `[:1971-1994]` — ByteDance is explicit that these *reduce probability*, not eliminate it:
|
|
148
|
+
|
|
149
|
+
1. **Bind each character to its image explicitly**, in a consistent format. Official example: *"Zhang San (corresponding to image 1) throws the green passbook toward Li Si (corresponding to image 2), who is standing."*
|
|
150
|
+
2. **Append the global constraint verbatim** at the end of the prompt `[:1982]`:
|
|
151
|
+
> *"Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect. Keep only a single corresponding character in the same frame, and do not reproduce repeated copies of characters."*
|
|
152
|
+
3. **Optimize reference assets** `[:1988]` — *"For character reference images, prioritize independent single-person photos. Three-view or multi-view assets are not recommended."*
|
|
153
|
+
4. **Simplify the prompt** — do not paste a whole script; redundant copy confuses the model.
|
|
154
|
+
|
|
155
|
+
**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on identity sheets. Practical rule for Slates:
|
|
156
|
+
|
|
157
|
+
- **Multi-character Seedance shot** → bind every character to its image, append the anti-twin constraint, and prefer single-person / dominant-portrait references over multi-view sheets.
|
|
158
|
+
- **Single-character shot** → the standard character-sheet flow is fine.
|
|
159
|
+
|
|
160
|
+
**Too many reference people** `[official :2048-2052]` — past **4 reference people**, output stability drops (wrong headcount, duplicates). Official workaround: group the cast into images of ≤4 people each, generate those stills first, then drive the video from them.
|
|
161
|
+
|
|
162
|
+
## Worked examples `[official :1689-1745]`
|
|
163
|
+
|
|
164
|
+
These are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere.
|
|
165
|
+
|
|
166
|
+
**Example 1 — dormitory emotional short drama (dialogue-focused).** Assets: `@Image 1` half-body photo of the female lead · `@Image 2` dormitory scene reference · `@Video 1` camera-movement reference · `@Audio 1` indoor ambience.
|
|
167
|
+
|
|
168
|
+
> Use the girl in @Image 1 as the main character, use @Image 2 as the dormitory scene style reference, and refer to the camera movement in @Video 1.
|
|
169
|
+
>
|
|
170
|
+
> **Shot 1**: At dusk, **girl @Image 1** walks briskly to the **dormitory entrance @Image 2**. The camera follows steadily in a medium shot. Warm yellow sunlight spills into the hallway from the window. She pauses at the doorway, takes a deep breath, and looks slightly nervous.
|
|
171
|
+
>
|
|
172
|
+
> **Shot 2**: **Girl @Image 1** pushes the door open and enters the dormitory. The camera cuts to an indoor medium shot. Her roommates look up at her while organizing their books. One of them smiles and asks {How did the exam go? Did you pass?}. The camera slowly cuts between half-body close-ups of several people.
|
|
173
|
+
>
|
|
174
|
+
> **Shot 3**: **Girl @Image 1** first lowers her head with a dejected expression. The camera gives her a close-up. Then she raises her head, unable to hold back a smile, laughs out loud, and says {I was kidding}. Her roommates start chasing and play-fighting with her. The camera slowly pulls back and freezes on a wide shot of the dormitory filled with laughter.
|
|
175
|
+
>
|
|
176
|
+
> The entire video should have a high-definition cinematic documentary style, with warm tones and soft lighting. The character's face remains stable without deformation; movements are natural and smooth, with no stutter or flicker. The ambient sound blends naturally with @Audio 1.
|
|
177
|
+
|
|
178
|
+
**Example 2 — ancient-style cliff confrontation (action/atmosphere-focused).** Assets: `@Image 1` female lead in red · `@Image 2` assassin in black · `@Image 3` cliff and bamboo forest · `@Video 1` martial-arts camera reference · `@Audio 1` drum beats.
|
|
179
|
+
|
|
180
|
+
> Use the woman in red from @Image 1 as the female lead, use the woman in black from @Image 2 as the opponent, use the cliff and bamboo forest environment in @Image 3 as the scene reference, refer to the overall camera movement and action rhythm in @Video 1, and synchronize the background sound effects with @Audio 1.
|
|
181
|
+
>
|
|
182
|
+
> **Shot 1**: At dusk, the camera slowly pushes in from a side medium shot of **woman in red @Image 1**. She stands at the edge of the cliff and lifts a wine flask to drink. Her sleeves and robe hem sway gently in the mountain wind. The camera circles halfway around her, moving from the front to her back. In the distance, a figure in black is faintly visible in the bamboo forest.
|
|
183
|
+
>
|
|
184
|
+
> **Shot 2**: The camera zooms and fades into a long shot. From a drone perspective, it overlooks the entire cliff and bamboo forest. The two characters stand at opposite ends of the cliff. The mountain wind lifts their robe hems and dust, and the rhythm slightly accelerates with the drum beats.
|
|
185
|
+
>
|
|
186
|
+
> **Shot 3**: The camera cuts back to a ground-level close shot. The two slowly draw their swords and face off. **Woman in red @Image 1** shifts from a careless expression to a cold gaze. **Woman in black @Image 2** looks determined, and the sword tip trembles slightly. The camera steadily follows the two as they circle each other, finally freezing on a close-up of the instant before the two swords meet.
|
|
187
|
+
>
|
|
188
|
+
> The overall visual style should feel like a cinematic wuxia world in misty rain, with cool tones, low saturation, a film-grain texture, and rich light-and-shadow layers. The characters' faces and body proportions remain stable without deformation. Movements are continuous and natural, not stiff, with no clipping or stutter.
|
|
189
|
+
|
|
190
|
+
## Other official notes
|
|
191
|
+
|
|
192
|
+
- **On-screen text** `[official :1758]` — Seedance can render common text (ad slogans, subtitles, speech bubbles) and will auto-match style/colour from context, or take an explicit colour / style / timing / position. Prefer **common characters**; avoid rare glyphs and special symbols. (For *guaranteed* legible text, the start-frame route in Part 3 is still safer.)
|
|
193
|
+
- **Extension degrades quality** `[official :2004-2024]` — using a generated video as the input for extension compounds degradation, with mottled colour blocks in face regions. Limit repeated continuations; prefer HD assets as input.
|
|
194
|
+
- **Special effects that miss** `[official :2031-2044]` — when a described effect comes out wrong (a countdown that scrolls randomly), define it with a **reference video** instead of words: *"the way the number '2999' appears should reference video 1."*
|
|
195
|
+
|
|
196
|
+
---
|
|
197
|
+
|
|
198
|
+
# Part 2 — Slates-specific `[slates]`
|
|
199
|
+
|
|
200
|
+
## Reference media — caps and transport
|
|
201
|
+
|
|
202
|
+
Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
|
|
203
|
+
|
|
204
|
+
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
|
|
205
|
+
|
|
206
|
+
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
207
|
+
|
|
208
|
+
These aren't separate Seedance features — they're prompting strategies over reference media.
|
|
209
|
+
|
|
210
|
+
- **Motion transfer:** subject image as a reference + the driving clip (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`
|
|
211
|
+
- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`
|
|
212
|
+
- **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script → their likeness AND voice deliver the new line.
|
|
213
|
+
|
|
214
|
+
## Reference rules (the verified ones)
|
|
215
|
+
|
|
216
|
+
<!-- @inject:references-read-literally -->
|
|
217
|
+
> **The general law: the model reads a reference literally.**
|
|
218
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
219
|
+
|
|
220
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
221
|
+
|
|
222
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
223
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
224
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
225
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
226
|
+
|
|
227
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
228
|
+
<!-- @end:references-read-literally -->
|
|
229
|
+
|
|
230
|
+
<!-- @inject:reference-rules-core -->
|
|
231
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
232
|
+
|
|
233
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
234
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
235
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
236
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
237
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
238
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
239
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
240
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
241
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
242
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
243
|
+
<!-- @end:reference-rules-core -->
|
|
244
|
+
|
|
245
|
+
### For Seedance specifically
|
|
246
|
+
|
|
247
|
+
- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say "still / scene / from a movie / from the image." The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one).
|
|
248
|
+
- **Seedance's own idiom for rule 2 is `Reference <Subject_N> in <Image_N>`** `[official :1389]` — `Image_N` indexes the order the refs are attached, so the name plus the index carries the role. The full binding grammar is in Part 1 (Subject binding).
|
|
249
|
+
- **Rule 3 has an official ceiling here.** The trend is MORE references (video and audio into Seedance), all addressed by name — but for **multi-character frames** see the twin-problem section above: bind every character to its image, append the anti-twin constraint, and prefer single-person references. Past 4 reference people, stability drops `[official :2048-2052]`.
|
|
250
|
+
- **Rule 8 holds even though Seedance can render common text natively** `[official :1758]`. A baked NB2 start frame is still the reliable route for text that must be legible.
|
|
251
|
+
- **Rule 5 pairs with the first/last-frame exclusion** — frames and reference images are mutually exclusive on this model (see Reference media above), so an environment you must match exactly costs you the frame lane.
|
|
252
|
+
|
|
253
|
+
---
|
|
254
|
+
|
|
255
|
+
# Part 3 — Community field notes `[community]`
|
|
256
|
+
|
|
257
|
+
Third-party guides and Slates field experience. Useful heuristics — but if one of these ever appears to contradict Part 1, **Part 1 wins**.
|
|
258
|
+
|
|
259
|
+
## Length
|
|
260
|
+
|
|
261
|
+
**Sweet spot 60-150 words** for a single shot (not 150-300 — that's the upper bound). Multi-shot storyboards run longer; official Example 1 above is ~230 words across three shots.
|
|
262
|
+
|
|
263
|
+
## Pin the subject in the first 20-30 words
|
|
264
|
+
|
|
265
|
+
The opening sentence is the **identity anchor**. If the subject isn't locked early, the model hallucinates new subjects mid-clip. (Compatible with Part 1: the binding preamble comes before `Shot 1`.)
|
|
266
|
+
|
|
267
|
+
```
|
|
268
|
+
A matte black earbud case sits on a polished obsidian surface...
|
|
269
|
+
```
|
|
270
|
+
|
|
271
|
+
## Lighting is a top quality lever
|
|
272
|
+
|
|
273
|
+
Lighting has an outsized impact on output quality — which is why it has its own slot in the official 8-slot formula. Describe it before or alongside the subject.
|
|
274
|
+
|
|
275
|
+
```
|
|
276
|
+
A cool-white diagonal beam from upper left, dust particles drifting through.
|
|
277
|
+
Soft golden hour lighting from low west angle.
|
|
278
|
+
Dramatic rim light against dark background.
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
## Camera and subject motion — separate sentences
|
|
282
|
+
|
|
283
|
+
Mixing them is a common cause of glitchy / shaky output.
|
|
284
|
+
|
|
285
|
+
❌ "The camera speed ramps as the earbud rises."
|
|
286
|
+
✅ "The earbud rises smoothly. The camera tracks upward."
|
|
287
|
+
|
|
288
|
+
## Slow-motion works; "fast" is a known bad token
|
|
289
|
+
|
|
290
|
+
Speed ramps and slow-motion are supported in natural language, and `fast` is widely reported as a quality-degrading keyword. **The official version of this rule is stronger and better founded** — prioritize slow, gentle, continuous small movements and avoid high-burst action (Part 1, Action description `[:1611-1615]`). Prompt the motion class, not the adjective.
|
|
291
|
+
|
|
292
|
+
```
|
|
293
|
+
the lid opens in slow-motion · the blade whips through the air
|
|
294
|
+
```
|
|
295
|
+
|
|
296
|
+
**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`. These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
|
|
297
|
+
|
|
298
|
+
## Style block at the end
|
|
299
|
+
|
|
300
|
+
One primary anchor + 2-3 supporting details, as the trailing paragraph (both official examples do exactly this). End with `Single continuous take` if you want one shot with no cuts. **Never** write `no cut` or `seamless transition` — not in the training vocabulary.
|
|
301
|
+
|
|
302
|
+
## ⚠️ Don't cross-pollinate image-model syntax
|
|
303
|
+
|
|
304
|
+
Named **lenses, apertures, film stocks, and camera bodies** — `85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`, `shot on Sony A7S3` — are an **image-model lever** (correct and encouraged in `reference-nano-banana.md`) and a **Seedance anti-pattern**. ByteDance's guide uses shot sizes, camera moves, pacing words, and the image-quality/style vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.
|
|
305
|
+
|
|
306
|
+
If you are carrying a look over from an NB2 start frame, translate it: `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`.
|
|
307
|
+
|
|
308
|
+
## Negative prompting — inline only
|
|
309
|
+
|
|
310
|
+
Seedance has **no `negativePrompt` field**. Put negatives in the constraints slot, led by the three official templates (Part 1):
|
|
311
|
+
|
|
312
|
+
```
|
|
313
|
+
keep it subtitle-free · do not generate a logo · do not generate a watermark
|
|
314
|
+
avoid jitter and bent limbs
|
|
315
|
+
avoid temporal flicker
|
|
316
|
+
avoid identity drift
|
|
317
|
+
no distortion, no stretching
|
|
318
|
+
```
|
|
319
|
+
|
|
320
|
+
Also fine: positive reframing ("empty street" not "no cars").
|
|
321
|
+
|
|
322
|
+
## Image-to-video / first-frame guidance
|
|
323
|
+
|
|
324
|
+
**Describe motion, not image.** The model already sees the visual; tokens spent re-describing appearance are wasted.
|
|
325
|
+
|
|
326
|
+
Stability phrases that help:
|
|
327
|
+
- `preserve composition and colors`
|
|
328
|
+
- `maintain exact appearance from reference image`
|
|
329
|
+
- `consistent character throughout, no deformation or drift`
|
|
330
|
+
|
|
331
|
+
**Cap I2V prompts under 60 words** when possible. Over 100 words frequently triggers silent generation failure.
|
|
332
|
+
|
|
333
|
+
## Common failure modes + fixes
|
|
334
|
+
|
|
335
|
+
| Failure | Fix |
|
|
336
|
+
|---|---|
|
|
337
|
+
| Hallucinated subject mid-clip | First 20-30 words = identity anchor |
|
|
338
|
+
| Bent limbs / extra fingers | `avoid jitter and bent limbs` in Constraints |
|
|
339
|
+
| Identity drift across multi-shot | Re-name the bound subject in **every** `Shot N` block `[official :1537]` |
|
|
340
|
+
| Two identical characters in one frame | The twin fix in Part 1 — bind each character to its image + append the global anti-twin constraint |
|
|
341
|
+
| Silent generation failure on I2V | Cut prompt under 100 words, single primary camera move |
|
|
342
|
+
| Speech / motion conflict | Limit dialogue to one line per action shot |
|
|
343
|
+
| Erratic/random pacing | You second-stamped. Remove all time markers and use `Shot N` `[official :1586]` |
|
|
344
|
+
|
|
345
|
+
## Sources
|
|
346
|
+
|
|
347
|
+
**Official (authoritative):**
|
|
348
|
+
- BytePlus ModelArk — Seedance 2.0 prompting guide, archived at `research/seedance-2-modelark-docs.md` (all `:NNNN` refs above)
|
|
349
|
+
|
|
350
|
+
**Community (secondary):**
|
|
351
|
+
- [fal.ai — How to Use Seedance 2.0](https://fal.ai/learn/tools/how-to-use-seedance-2-0)
|
|
352
|
+
- [apiyi.com — Seedance 2.0 Prompt Guide](https://help.apiyi.com/en/seedance-2-0-prompt-guide-video-generation-camera-style-tips-en.html)
|
|
353
|
+
- [atlabs.ai — Ultimate Seedance 2.0 Prompting Guide](https://www.atlabs.ai/blog/the-ultimate-seedance-2.0-prompting-guide-47-prompts-2026)
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
{
|
|
2
|
+
"schemaVersion": 1,
|
|
3
|
+
"source": "@slatesvideo/shared",
|
|
4
|
+
"sourceFiles": [
|
|
5
|
+
{
|
|
6
|
+
"path": "exports/slates-prompt-builder/prompt-builder.md",
|
|
7
|
+
"sha256": "d7510c28be9453ab1bb1be062ff04a79a75ede5150ac4564a978d4a80d6315c5"
|
|
8
|
+
},
|
|
9
|
+
{
|
|
10
|
+
"path": "skills/slates-character-identity.md",
|
|
11
|
+
"sha256": "cc0387b21c83bbf0840524e5b124db889d5ee584cbacd19ffaf5fd2ec132bd20"
|
|
12
|
+
},
|
|
13
|
+
{
|
|
14
|
+
"path": "skills/slates-prompting-seedance.md",
|
|
15
|
+
"sha256": "a67f8c72bf187b726d39d2efb3c72a44740f43d44d5166d9767b93c55b4dd0a2"
|
|
16
|
+
},
|
|
17
|
+
{
|
|
18
|
+
"path": "skills/slates-prompting-kling-v3.md",
|
|
19
|
+
"sha256": "629b9490103a1a7961b584b7e125faa45375265dd941d8aae466cf6b195d5884"
|
|
20
|
+
},
|
|
21
|
+
{
|
|
22
|
+
"path": "skills/slates-prompting-nano-banana-2.md",
|
|
23
|
+
"sha256": "fd47d89ffa5ac16c2d06e2d3c0e891813996dc15636a8988f7da6e8ee1816565"
|
|
24
|
+
},
|
|
25
|
+
{
|
|
26
|
+
"path": "skills/slates-content-policy.md",
|
|
27
|
+
"sha256": "ee5afd3e0d276d75c08b1747f551878be28d22151a369c969c55e58fda3f5710"
|
|
28
|
+
},
|
|
29
|
+
{
|
|
30
|
+
"path": "src/prompts/model-facts.ts",
|
|
31
|
+
"sha256": "3bdc599838204334e0dcbdb076c1d866b5eba2eed17c5de8f5077cfaa5f6fe0b"
|
|
32
|
+
}
|
|
33
|
+
],
|
|
34
|
+
"outputs": [
|
|
35
|
+
{
|
|
36
|
+
"path": "SKILL.md",
|
|
37
|
+
"bytes": 4969,
|
|
38
|
+
"sha256": "51bb840b158c84e5d9060af9849ec23b94031d0ff0413f13da9a97358b40ec66"
|
|
39
|
+
},
|
|
40
|
+
{
|
|
41
|
+
"path": "reference-character.md",
|
|
42
|
+
"bytes": 7950,
|
|
43
|
+
"sha256": "3f9af66752811a355225db039aeda9c724b685ede071900c51e61624c9e05622"
|
|
44
|
+
},
|
|
45
|
+
{
|
|
46
|
+
"path": "reference-seedance.md",
|
|
47
|
+
"bytes": 31667,
|
|
48
|
+
"sha256": "53a77f7bcfaa35be875df221f99843a74481ef6c6d8f7813b97a6e65d12f9dfe"
|
|
49
|
+
},
|
|
50
|
+
{
|
|
51
|
+
"path": "reference-kling.md",
|
|
52
|
+
"bytes": 13592,
|
|
53
|
+
"sha256": "3689d7a285cc8845bc41c5ffbed7b727a43cd3c108a2a46b8028c3528ef44ed5"
|
|
54
|
+
},
|
|
55
|
+
{
|
|
56
|
+
"path": "reference-nano-banana.md",
|
|
57
|
+
"bytes": 16437,
|
|
58
|
+
"sha256": "69fe636c780d6ddf9fe97e23b7d82d8961cac0037b985e3500a9485f295e211d"
|
|
59
|
+
},
|
|
60
|
+
{
|
|
61
|
+
"path": "reference-content-policy.md",
|
|
62
|
+
"bytes": 6785,
|
|
63
|
+
"sha256": "4c0fe7cd7d3d6b854af6702133fd81a3338a5cb01f48e4a626165c4eb9667b06"
|
|
64
|
+
}
|
|
65
|
+
],
|
|
66
|
+
"archive": {
|
|
67
|
+
"path": "slates-prompt-builder.skill",
|
|
68
|
+
"bytes": 37054,
|
|
69
|
+
"sha256": "823cbb640eb6f64afceec8665f20685cf851527d2b886e1443bccca0ead4afeb",
|
|
70
|
+
"entries": [
|
|
71
|
+
"SKILL.md",
|
|
72
|
+
"reference-character.md",
|
|
73
|
+
"reference-content-policy.md",
|
|
74
|
+
"reference-kling.md",
|
|
75
|
+
"reference-nano-banana.md",
|
|
76
|
+
"reference-seedance.md"
|
|
77
|
+
]
|
|
78
|
+
}
|
|
79
|
+
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@slatesvideo/shared",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.5",
|
|
4
4
|
"description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -24,11 +24,15 @@
|
|
|
24
24
|
"dist",
|
|
25
25
|
"!dist/**/*.map",
|
|
26
26
|
"skills",
|
|
27
|
+
"exports/slates-prompt-builder/generated",
|
|
27
28
|
"README.md"
|
|
28
29
|
],
|
|
29
30
|
"scripts": {
|
|
30
|
-
"
|
|
31
|
-
"
|
|
31
|
+
"sync-partials": "node scripts/sync-partials.mjs",
|
|
32
|
+
"build-prompt-builder": "node scripts/build-prompt-builder.mjs",
|
|
33
|
+
"check-prompt-builder": "node scripts/build-prompt-builder.mjs --check",
|
|
34
|
+
"build": "node scripts/sync-partials.mjs --check && node scripts/build-prompt-builder.mjs --check && node scripts/embed-skills.mjs && tsc",
|
|
35
|
+
"typecheck": "node scripts/sync-partials.mjs --check && node scripts/build-prompt-builder.mjs --check && node scripts/embed-skills.mjs && tsc --noEmit",
|
|
32
36
|
"prepublishOnly": "npm run build"
|
|
33
37
|
},
|
|
34
38
|
"repository": {
|
|
@@ -59,6 +63,7 @@
|
|
|
59
63
|
"zod": "^3.23.0"
|
|
60
64
|
},
|
|
61
65
|
"devDependencies": {
|
|
66
|
+
"fflate": "^0.8.3",
|
|
62
67
|
"typescript": "^5.7.0"
|
|
63
68
|
}
|
|
64
69
|
}
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify:
|
|
2
|
+
|
|
3
|
+
```
|
|
4
|
+
source phrase or declared default → what you wrote → what it resolves
|
|
5
|
+
"in a diner" → chrome-and-vinyl booth, 3/4 on the counter → fixes the anchor so blocking is repeatable
|
|
6
|
+
(no time of day) → late afternoon, low warm key → default; say the word and it changes
|
|
7
|
+
(no camera) → slow push-in, single move → one move per shot; stacking increases instability
|
|
8
|
+
```
|
|
9
|
+
|
|
10
|
+
**Hard rule: never silently add weather, props, style, or camera movement.** If it wasn't in the brief and you added it, it goes in the log. This is the "why did you add that?" affordance — for an agent that writes prompts on the user's behalf and spends their credits, it is what keeps the model in assembly and the user in the director's chair.
|
|
11
|
+
|
|
12
|
+
> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
2
|
+
|
|
3
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
4
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
5
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
6
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
7
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
8
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
9
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
10
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
11
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
12
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
@@ -0,0 +1,2 @@
|
|
|
1
|
+
<!-- consumer:ts -->
|
|
2
|
+
Name each reference inline; never write role essays. Slates does this for you: `@mention` a subject or environment and it composes `Marcus (image 1) in the cafe (image 2)`, citing them in the exact order it sends them. One canonical identity image avoids competing facial renderings; a "Reference Image Instructions" block drags reference lighting into your scene. Start with 2-3 focused refs.
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
> **The general law: the model reads a reference literally.**
|
|
2
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
3
|
+
|
|
4
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
5
|
+
|
|
6
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
7
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
8
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
9
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
10
|
+
|
|
11
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
@@ -0,0 +1,3 @@
|
|
|
1
|
+
**A visible defect in the still is already a STOP.** Do not animate it. Fix the frame first, then move to motion — and go to motion only when the crop passes the still scan and you genuinely need movement to confirm an uncertain edge, reflection, or object.
|
|
2
|
+
|
|
3
|
+
This is a **cost** rule as much as a craft rule: a 1080p/10s premium video generation costs many multiples of an image re-roll, and video is where a defect stops being fixable. Anything wrong in the still gets worse in motion — soft geometry mushes, broken-but-plausible objects fall apart, oily textures start crawling. **Animating a known-bad frame is the single most expensive mistake in the pipeline.** Re-rolling the image is the cheap move; re-rolling the video is not.
|