@slatesvideo/shared 0.5.4 → 0.5.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +13 -0
- package/dist/operations/index.d.ts +3 -8
- package/dist/operations/index.js +21 -38
- package/dist/prompts/character-sheet.d.ts +10 -21
- package/dist/prompts/character-sheet.js +50 -54
- package/dist/prompts/partials.generated.js +2 -2
- package/dist/prompts/prompting-tips.js +2 -2
- package/dist/prompts/reference-composer.d.ts +1 -1
- package/dist/prompts/reference-composer.js +3 -4
- package/dist/prompts/reference-rules.js +3 -3
- package/dist/skills/content.js +10 -10
- package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
- package/exports/slates-prompt-builder/generated/reference-character.md +78 -0
- package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
- package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
- package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +7 -3
- package/skills/_partials/reference-rules-core.md +1 -1
- package/skills/_partials/reference-tips-short.md +1 -1
- package/skills/{slates-character-turnaround.md → slates-character-identity.md} +32 -22
- package/skills/slates-edit-and-iterate.md +1 -1
- package/skills/slates-one-prompt-film.md +3 -3
- package/skills/slates-prompting-flux-2-max.md +1 -1
- package/skills/slates-prompting-gpt-image-2.md +1 -1
- package/skills/slates-prompting-kling-v3.md +8 -6
- package/skills/slates-prompting-nano-banana-2.md +7 -5
- package/skills/slates-prompting-omni-flash.md +1 -1
- package/skills/slates-prompting-seedance.md +15 -9
- package/skills/slates-prompting-veo-3.md +1 -1
|
@@ -47,7 +47,7 @@ The user's request is one of:
|
|
|
47
47
|
### 4. Generate, evaluate, decide
|
|
48
48
|
- Estimate cost first.
|
|
49
49
|
- After generation, the result is inline. Compare side-by-side with the original (`slates_get_asset_image` again).
|
|
50
|
-
- If the delta is correct: bind to the same
|
|
50
|
+
- If the delta is correct: bind to the same role (frame, character identity, etc.) the original was bound to.
|
|
51
51
|
- If the delta missed: one focused refinement, then regenerate. Cap at 3 tries.
|
|
52
52
|
|
|
53
53
|
### 5. Hand back
|
|
@@ -33,7 +33,7 @@ A 4-10 shot script is where you invent the most on the user's behalf — time of
|
|
|
33
33
|
|
|
34
34
|
### 2. Set up the project
|
|
35
35
|
- `slates_create_project` named for the piece.
|
|
36
|
-
- Recurring character? Build it properly — `slates_create_character` + the `slates-character-
|
|
36
|
+
- Recurring character? Build it properly — `slates_create_character` + the `slates-character-identity` recipe — so every frame references the same identity.
|
|
37
37
|
- Recurring location? `slates_create_environment`.
|
|
38
38
|
- One-off shots don't need character/environment records; skip the ceremony.
|
|
39
39
|
|
|
@@ -49,7 +49,7 @@ Price the whole batch before the first generation: frame images (count × model
|
|
|
49
49
|
Per `slates-cost-discipline` 3b: that single OK authorizes `confirm=true` for **every enumerated call in the batch** — no per-call re-asking. Re-confirm only if a call's price overruns the plan >25% or new calls get added (extra retakes, new shots).
|
|
50
50
|
|
|
51
51
|
### 5. Generate frame images
|
|
52
|
-
Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the character
|
|
52
|
+
Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the character identity / environment / prior frames for consistency (Slates names each reference inline as "image N" — you don't hand-write role labels; reuse the same subject name across shots). Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`.
|
|
53
53
|
|
|
54
54
|
**Multi-take where it matters:** for the hook shot and any shot the whole film hangs on, generate 2-4 variants (cheap model or 1k), pull them back with `slates_get_assets_batch`, pick the strongest on composition + identity, discard the rest. Don't multi-take filler shots.
|
|
55
55
|
|
|
@@ -83,4 +83,4 @@ Shots delivered, total spent vs. approved plan, the export path, and the single
|
|
|
83
83
|
- **Skeleton before spend.** Project + storyboard structure are free; generation isn't.
|
|
84
84
|
- **Look at everything.** Every image inline, every video via `slates_get_asset_video_frames` if a clip seems off. Never assemble a timeline from clips you haven't evaluated.
|
|
85
85
|
- **3-strike rule per shot.** Three failed takes on one shot = stop, show the user what you tried, ask.
|
|
86
|
-
- **Consistency comes from references, not luck.** Same
|
|
86
|
+
- **Consistency comes from references, not luck.** Same identity asset on every character frame; same environment refs across a location's shots.
|
|
@@ -103,7 +103,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
103
103
|
|
|
104
104
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
105
105
|
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
106
|
-
3. **One identity sheet per character
|
|
106
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
107
107
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
108
108
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
109
109
|
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
@@ -29,7 +29,7 @@ Never rely on the provider default (it's high — the priciest tier). The Slates
|
|
|
29
29
|
|
|
30
30
|
- State the grid explicitly and number the cells: "a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom".
|
|
31
31
|
- Give each cell ONE content clause: "Panel 3: the character mid-jump, side view".
|
|
32
|
-
- Character sheets:
|
|
32
|
+
- Character identity sheets: GPT Image 2 holds structured panel layouts; the Banana line holds the *face* better. Prefer NB2/NB Pro for identity-critical sheets and GPT Image 2 when labels or annotations are the main requirement.
|
|
33
33
|
|
|
34
34
|
## References & editing
|
|
35
35
|
|
|
@@ -125,7 +125,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
125
125
|
|
|
126
126
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
127
127
|
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
128
|
-
3. **One identity sheet per character
|
|
128
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
129
129
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
130
130
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
131
131
|
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
@@ -164,7 +164,7 @@ Layer scene-specific suppressions on top.
|
|
|
164
164
|
- **Pro**: higher visual quality, no audio
|
|
165
165
|
- **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references
|
|
166
166
|
|
|
167
|
-
Pick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models
|
|
167
|
+
Pick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.
|
|
168
168
|
|
|
169
169
|
## Benchmark prompt structure
|
|
170
170
|
|
|
@@ -179,6 +179,7 @@ Cinematic example (paraphrasing fal blog patterns):
|
|
|
179
179
|
> Shot 2: Medium shot of a detective in a trench coat ducking under an awning, water dripping from his hat brim. [Detective: weary, raspy]: 'I knew she'd come back.' Ambient noise: distant traffic, rain on metal.
|
|
180
180
|
> Shot 3: Close-up on his eyes, narrowing as headlights flash across his face."
|
|
181
181
|
|
|
182
|
+
<!-- slates-only -->
|
|
182
183
|
## Pre-flight: references arrive inline, refer by code
|
|
183
184
|
|
|
184
185
|
When you call `slates_generate_video` with `firstFrameAssetId` or `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside cost + `requires_confirm: true`. Look at them, revise prompt if needed, then re-call with `confirm=true`. Kling Omni multi-character with several ingredient images especially benefits — confirm each character image lands cleanly before spending.
|
|
@@ -187,14 +188,15 @@ When talking to the user about the gen, refer to each reference by its short cod
|
|
|
187
188
|
|
|
188
189
|
- ✅ "I'm anchoring on **IMG-A12** as the detective and **IMG-A18** as the alleyway environment — Omni will handle the line delivery in EN."
|
|
189
190
|
- ❌ "I'm using the detective image and the alley one..." (which alley? Three exist.)
|
|
191
|
+
<!-- /slates-only -->
|
|
190
192
|
|
|
191
|
-
## Video-to-video EDIT (`slates_edit_video`) — @Video1 / @ElementN / @ImageN
|
|
193
|
+
## Video-to-video EDIT<!-- slates-only --> (`slates_edit_video`)<!-- /slates-only --> — @Video1 / @ElementN / @ImageN
|
|
192
194
|
|
|
193
195
|
Kling O3 edit takes an EXISTING 3-15s clip and changes only what the prompt names — character swap, environment change, style transfer — in one pass, no masking. Original motion, camera, and audio are preserved by default. Its notation is Kling's own, different from the "image N" naming used everywhere else:
|
|
194
196
|
|
|
195
197
|
- **`@Video1`** — the source clip (always; the transport anchors the instruction to it).
|
|
196
|
-
- **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically)
|
|
197
|
-
- **`@Image1..`** — style/appearance references (pass as `styleAssetIds`)
|
|
198
|
+
- **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images<!-- slates-only --> (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically)<!-- /slates-only -->.
|
|
199
|
+
- **`@Image1..`** — style/appearance references<!-- slates-only --> (pass as `styleAssetIds`)<!-- /slates-only -->.
|
|
198
200
|
- Max **4 combined** element + image refs per edit.
|
|
199
201
|
|
|
200
202
|
**Prompt shape — the change, not the whole scene:**
|
|
@@ -212,7 +214,7 @@ Rules:
|
|
|
212
214
|
- One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).
|
|
213
215
|
- Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.
|
|
214
216
|
- Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.
|
|
215
|
-
- Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings — see `slates-model-selection
|
|
217
|
+
- Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.
|
|
216
218
|
|
|
217
219
|
## Sources
|
|
218
220
|
|
|
@@ -5,7 +5,7 @@ description: How to write prompts that produce cinematic, photorealistic results
|
|
|
5
5
|
|
|
6
6
|
# Nano Banana 2 — cinematic & photorealistic prompting
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Nano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the default model behind `slates_generate_image` — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
|
|
9
9
|
|
|
10
10
|
Knowledge cutoff: January 2025. Anything after needs explicit reference images.
|
|
11
11
|
|
|
@@ -100,7 +100,7 @@ Default to #1. Reach for #2 only when positive framing can't suppress the unwant
|
|
|
100
100
|
## Reference images
|
|
101
101
|
|
|
102
102
|
- **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.
|
|
103
|
-
- **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style (or pass `referenceAssetIds`)
|
|
103
|
+
- **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
104
104
|
|
|
105
105
|
### Reference rules (the verified ones)
|
|
106
106
|
|
|
@@ -123,7 +123,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
123
123
|
|
|
124
124
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
125
125
|
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
126
|
-
3. **One identity sheet per character
|
|
126
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
127
127
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
128
128
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
129
129
|
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
@@ -135,7 +135,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
135
135
|
|
|
136
136
|
### For Nano Banana 2 specifically
|
|
137
137
|
|
|
138
|
-
- **NB2's own consistency lever is "assign a distinct name to each character/object."** That is Google's phrasing for rule 3 —
|
|
138
|
+
- **NB2's own consistency lever is "assign a distinct name to each character/object."** That is Google's phrasing for rule 3 — cite each canonical identity inline by name.
|
|
139
139
|
- **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.
|
|
140
140
|
- **Character consistency is officially "not 100% perfect"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.
|
|
141
141
|
- **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.
|
|
@@ -156,7 +156,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
156
156
|
|
|
157
157
|
## Resolution tactics
|
|
158
158
|
|
|
159
|
-
- Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change —
|
|
159
|
+
- Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — check current numbers<!-- slates-only --> by calling `slates_estimate_generation_cost`<!-- /slates-only -->. Pick the cheapest resolution that serves the use case.
|
|
160
160
|
- **At 2K and above, the model allocates more tokens to surface detail** — explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.
|
|
161
161
|
- 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.
|
|
162
162
|
- 2K generations vary 20-60s+. Don't time-budget tightly.
|
|
@@ -182,4 +182,6 @@ Everything in this skill applies to the whole Nano Banana family; two variants t
|
|
|
182
182
|
- **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.
|
|
183
183
|
- **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.
|
|
184
184
|
|
|
185
|
+
<!-- slates-only -->
|
|
185
186
|
Routing between them (and vs GPT Image 2 / FLUX / Seedream): `slates-model-selection`.
|
|
187
|
+
<!-- /slates-only -->
|
|
@@ -30,7 +30,7 @@ Google's fast video generation + editing model ("Nano Banana Pro for video" in c
|
|
|
30
30
|
|
|
31
31
|
- **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
|
|
32
32
|
- Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.
|
|
33
|
-
- **Name references inline** the standard Slates way ("Marcus (
|
|
33
|
+
- **Name references inline** the standard Slates way ("Marcus (image 1) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
|
|
34
34
|
- **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language ("rain patters on the tin roof"). Negative direction as plain instructions ("Do not show text").
|
|
35
35
|
- Duration is an explicit 3–10s integer param; cost scales linearly per second.
|
|
36
36
|
|
|
@@ -55,7 +55,7 @@ Seedance classifies your request from the phrasing. Use the pattern that matches
|
|
|
55
55
|
|
|
56
56
|
> *"For edit / extend video tasks, directly use `<Video_N>` to refer to the video. **Do not use "reference `<Video_N>`"**, to avoid being incorrectly identified as a reference task."*
|
|
57
57
|
|
|
58
|
-
This is easy to trip: Slates has an edit lane (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)
|
|
58
|
+
This is easy to trip: Slates has an edit lane<!-- slates-only --> (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)<!-- /slates-only -->. Writing *"reference video 1 and change the jacket to red"* gets classified as a **reference** task — the model generates a brand-new clip inspired by the source instead of editing it. Write *"Strictly edit video 1, and modify the blue jacket to red."*
|
|
59
59
|
|
|
60
60
|
## Shot structure — "Shot 1 / Shot 2 / Shot 3" `[official :1563-1598]`
|
|
61
61
|
|
|
@@ -91,7 +91,7 @@ Every time a subject appears, it must be **explicitly referred to**. Two support
|
|
|
91
91
|
|
|
92
92
|
Also official: keep descriptions concise, avoid redundancy, avoid semantic conflicts (contradictory traits for one subject), and prefer expressing spatial relationships through reference images rather than dense text. `[:1550-1556]`
|
|
93
93
|
|
|
94
|
-
**`[slates]`** — the app composes this for you. `composeReferences()` cites each reference inline as `Name (image N)`
|
|
94
|
+
**`[slates]`** — the app composes this for you. `composeReferences()` cites each canonical character or environment reference inline as `Name (image N)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*"Zhang San (corresponding to image 1)"* `[:1976]`). You never hand-write role labels or index numbers.
|
|
95
95
|
|
|
96
96
|
## Action description `[official :1602-1621]`
|
|
97
97
|
|
|
@@ -153,7 +153,7 @@ Seedance has **no `negativePrompt` field** — constraints go inline in this slo
|
|
|
153
153
|
3. **Optimize reference assets** `[:1988]` — *"For character reference images, prioritize independent single-person photos. Three-view or multi-view assets are not recommended."*
|
|
154
154
|
4. **Simplify the prompt** — do not paste a whole script; redundant copy confuses the model.
|
|
155
155
|
|
|
156
|
-
**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on
|
|
156
|
+
**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on identity sheets. Practical rule for Slates:
|
|
157
157
|
|
|
158
158
|
- **Multi-character Seedance shot** → bind every character to its image, append the anti-twin constraint, and prefer single-person / dominant-portrait references over multi-view sheets.
|
|
159
159
|
- **Single-character shot** → the standard character-sheet flow is fine.
|
|
@@ -206,13 +206,16 @@ Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 aud
|
|
|
206
206
|
|
|
207
207
|
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
208
208
|
|
|
209
|
-
These aren't separate Seedance features — they're prompting strategies over reference media
|
|
209
|
+
These aren't separate Seedance features — they're prompting strategies over reference media.<!-- slates-only --> The Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:<!-- /slates-only -->
|
|
210
210
|
|
|
211
|
-
- **Motion transfer:** subject image as a reference + the driving clip via `videoReferenceAssetId
|
|
212
|
-
- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with
|
|
211
|
+
- **Motion transfer:** subject image as a reference + the driving clip<!-- slates-only --> via `videoReferenceAssetId`<!-- /slates-only --> (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`
|
|
212
|
+
- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`
|
|
213
213
|
- **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script → their likeness AND voice deliver the new line.
|
|
214
|
+
<!-- slates-only -->
|
|
214
215
|
- **Billing:** a reference VIDEO switches the cost key to `seedance-2*-vref-{res}-{T}s` where T = clip seconds + output seconds — quote before confirming. Audio references are free (audio is included on every route).
|
|
216
|
+
<!-- /slates-only -->
|
|
215
217
|
|
|
218
|
+
<!-- slates-only -->
|
|
216
219
|
## Faces — set `seedanceFace` for AI-character faces
|
|
217
220
|
|
|
218
221
|
Seedance routes through **three tiers** depending on the face in the reference, exposed as the "Face in Reference" toggle plus the real-face params on `slates_generate_video`:
|
|
@@ -223,8 +226,9 @@ Seedance routes through **three tiers** depending on the face in the reference,
|
|
|
223
226
|
|
|
224
227
|
Rules:
|
|
225
228
|
- **The real-vs-AI call is the PROVIDER'S, not yours.** ByteDance's classifier is probabilistic — some real photos pass the standard face route (billed at the cheap rate; fine), others get rejected with `[REAL_FACE_DETECTED]` (auto-refunded). Don't preemptively route to the real-face tier just because a photo looks real; try `seedanceFace: true` first and escalate only on the marked rejection. Public figures / celebrities fail on every route.
|
|
226
|
-
- It's about the **reference, not the output.** If your character
|
|
229
|
+
- It's about the **reference, not the output.** If your character identity or generated portrait shows a face, turn it on. A product shot with no person stays off.
|
|
227
230
|
- Don't toggle it on "just in case" — a faceless gen on the face route burns ~45% extra for nothing.
|
|
231
|
+
<!-- /slates-only -->
|
|
228
232
|
|
|
229
233
|
## Reference rules (the verified ones)
|
|
230
234
|
|
|
@@ -247,7 +251,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
247
251
|
|
|
248
252
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
249
253
|
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
250
|
-
3. **One identity sheet per character
|
|
254
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
251
255
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
252
256
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
253
257
|
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
@@ -259,12 +263,13 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
259
263
|
|
|
260
264
|
### For Seedance specifically
|
|
261
265
|
|
|
262
|
-
- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say "still / scene / from a movie / from the image." The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one — see slates-cost-discipline).
|
|
266
|
+
- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say "still / scene / from a movie / from the image." The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one<!-- slates-only --> — see slates-cost-discipline<!-- /slates-only -->).
|
|
263
267
|
- **Seedance's own idiom for rule 2 is `Reference <Subject_N> in <Image_N>`** `[official :1389]` — `Image_N` indexes the order the refs are attached, so the name plus the index carries the role. The full binding grammar is in Part 1 (Subject binding).
|
|
264
268
|
- **Rule 3 has an official ceiling here.** The trend is MORE references (video and audio into Seedance), all addressed by name — but for **multi-character frames** see the twin-problem section above: bind every character to its image, append the anti-twin constraint, and prefer single-person references. Past 4 reference people, stability drops `[official :2048-2052]`.
|
|
265
269
|
- **Rule 8 holds even though Seedance can render common text natively** `[official :1758]`. A baked NB2 start frame is still the reliable route for text that must be legible.
|
|
266
270
|
- **Rule 5 pairs with the first/last-frame exclusion** — frames and reference images are mutually exclusive on this model (see Reference media above), so an environment you must match exactly costs you the frame lane.
|
|
267
271
|
|
|
272
|
+
<!-- slates-only -->
|
|
268
273
|
## Pre-flight: references arrive inline, refer by code
|
|
269
274
|
|
|
270
275
|
When you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** — if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.
|
|
@@ -273,6 +278,7 @@ When talking to the user about the gen, refer to each reference by its short cod
|
|
|
273
278
|
|
|
274
279
|
- ✅ "I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame — the camera move is going to be a slow dolly forward through the gap."
|
|
275
280
|
- ❌ "I'm using the first beach image and the last one..." (which? They have four.)
|
|
281
|
+
<!-- /slates-only -->
|
|
276
282
|
|
|
277
283
|
---
|
|
278
284
|
|
|
@@ -117,7 +117,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
117
117
|
|
|
118
118
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
119
119
|
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
120
|
-
3. **One identity sheet per character
|
|
120
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
121
121
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
122
122
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
123
123
|
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|