@slatesvideo/shared 0.5.6 → 0.5.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.d.ts +1 -1
- package/dist/index.js +4 -1
- package/dist/operations/index.d.ts +43 -36
- package/dist/operations/index.js +289 -262
- package/dist/prompts/model-facts.d.ts +42 -0
- package/dist/prompts/model-facts.js +99 -18
- package/dist/prompts/partials.generated.js +2 -0
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +130 -100
- package/dist/prompts/reference-composer.d.ts +21 -3
- package/dist/prompts/reference-composer.js +80 -10
- package/dist/skills/content.js +7 -7
- package/exports/slates-prompt-builder/generated/SKILL.md +2 -2
- package/exports/slates-prompt-builder/generated/reference-seedance.md +13 -2
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +8 -8
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +1 -1
- package/skills/_partials/seedance-25-timestamps-short.md +2 -0
- package/skills/_partials/seedance-25-timestamps.md +34 -0
- package/skills/slates-model-selection.md +21 -19
- package/skills/slates-prompting-elevenlabs.md +16 -78
- package/skills/slates-prompting-lip-sync.md +12 -14
- package/skills/slates-prompting-motion-transfer.md +18 -14
- package/skills/slates-prompting-seed-audio.md +2 -2
- package/skills/slates-prompting-seedance-2-5.md +290 -0
- package/skills/slates-prompting-seedance.md +19 -3
- package/skills/slates-prompting-suno.md +0 -110
|
@@ -0,0 +1,290 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-seedance-2-5
|
|
3
|
+
description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is a SECOND SEAT next to 2.0, not an upgrade — it buys 30-second takes, 30 image references, audio-only references and INTEGER-SECOND TIMESTAMPS, and it gives up 1080p and 4K entirely. Timestamps are the one grammar difference that matters: 2.0 ignores them and answers only to shot numbers, 2.5 acts on them. Otherwise it shares 2.0's grammar (read slates-prompting-seedance for subject binding, camera and constraint vocabulary); this file covers what is different, plus the two hazards unique to 2.5 — the prompt-intent task classifier and the cost trap at 720p.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Seedance 2.5 — prompting
|
|
7
|
+
|
|
8
|
+
**Read `slates-prompting-seedance` first.** The prompt GRAMMAR is the same model family: the
|
|
9
|
+
8-slot advanced formula, subject binding by `<Subject_N>@<Image_N>`, camera vocabulary,
|
|
10
|
+
externalised emotion, inline constraint words, the anti-twin fix. None of it is restated here.
|
|
11
|
+
This file is only what 2.5 changes — and the biggest change is that **2.5 acts on timestamps
|
|
12
|
+
where 2.0 ignores them.**
|
|
13
|
+
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
## The one fact that decides whether you use it at all
|
|
17
|
+
|
|
18
|
+
**Seedance 2.5 is 480p or 720p. There is no 1080p and no 4K, on any provider.**
|
|
19
|
+
|
|
20
|
+
That is not a Slates limitation or a tier gate — the model does not produce those resolutions.
|
|
21
|
+
So 2.5 does not replace 2.0; it sits beside it:
|
|
22
|
+
|
|
23
|
+
| | Seedance 2.0 | Seedance 2.5 |
|
|
24
|
+
|---|---|---|
|
|
25
|
+
| Resolution | 480p / 720p / 1080p / **native 4K** | **480p / 720p only** |
|
|
26
|
+
| Length | 4–15s | **4–30s in one take** |
|
|
27
|
+
| Reference budget | 15 (9 image + 3 video + 3 audio) | **50 (30 image + 10 video + 10 audio)** |
|
|
28
|
+
| Combined reference video/audio | ≤15s | **≤30s** |
|
|
29
|
+
| Audio-only reference | ✗ (needs an image or video alongside) | **✓** |
|
|
30
|
+
| **Timestamps in the prompt** | **✗ — ignored; shot numbers only** | **✓ — integer seconds, acted on** |
|
|
31
|
+
| Multi-view image as ONE subject reference | ✗ (not recommended) | **✓ (up to 5 subjects)** |
|
|
32
|
+
| Video edit as its own task type | ✗ | **✓ (`seedance-2.5-edit`)** |
|
|
33
|
+
| Default video model | **yes** | no |
|
|
34
|
+
|
|
35
|
+
**Route to 2.5 when the shot needs LENGTH, MANY REFERENCES, or an audio-only reference.
|
|
36
|
+
Route to 2.0 when resolution matters at all** — which, for anything a client will see full-screen,
|
|
37
|
+
is most of the time.
|
|
38
|
+
|
|
39
|
+
---
|
|
40
|
+
|
|
41
|
+
## 🚨 Hazard 1 — the prompt-intent task classifier
|
|
42
|
+
|
|
43
|
+
This is the one that costs money and time, and it has no equivalent on 2.0.
|
|
44
|
+
|
|
45
|
+
**Seedance 2.5 sorts every request into one of five task types** — text-to-video,
|
|
46
|
+
reference-to-video, first/last-frame, **video edit**, **video extend** — from the reference roles
|
|
47
|
+
attached **plus the intent of your sentence**. Each type then has its own parameter constraints,
|
|
48
|
+
and a violation comes back **asynchronously**: the task queues, credits are reserved, and only then
|
|
49
|
+
does it fail.
|
|
50
|
+
|
|
51
|
+
The trigger words are ordinary English:
|
|
52
|
+
|
|
53
|
+
| Reclassified as | Words that do it (ByteDance's own list) |
|
|
54
|
+
|---|---|
|
|
55
|
+
| **video edit** | `edit video` · `add` · `insert` · `remove` · `delete` · `modify` · `replace` · `change to` |
|
|
56
|
+
| **video extend** | `extend forward` · `extend backward` · `continue` · `continue from` · `extend the story` |
|
|
57
|
+
|
|
58
|
+
So a perfectly legitimate reference-to-video prompt — *"a wide shot of the workshop, **remove** the
|
|
59
|
+
tripod from frame"* — gets classified as an edit and fails on constraints it never set.
|
|
60
|
+
|
|
61
|
+
**What to do:**
|
|
62
|
+
|
|
63
|
+
1. **If you mean to edit an existing clip, say so with the MODEL, not the sentence.** Call
|
|
64
|
+
`slates_edit_video` with `model: 'seedance-2.5-edit'`. That routes to a dedicated
|
|
65
|
+
task-typed endpoint and the classifier never has to guess.
|
|
66
|
+
2. **If you mean a fresh shot, describe the finished frame rather than an instruction to change
|
|
67
|
+
one.** Not *"remove the tripod"* → *"the workshop bench, clear and uncluttered"*. Not
|
|
68
|
+
*"add rain"* → *"heavy rain falling through the streetlight"*. This is better prompting anyway:
|
|
69
|
+
the model renders what you describe, it does not take edits to an imagined draft.
|
|
70
|
+
3. The trigger only fires when **references are attached**. A plain text-to-video prompt is safe
|
|
71
|
+
however it is worded.
|
|
72
|
+
|
|
73
|
+
**Slates will warn you, and it will never rewrite your prompt.** When a 2.5 reference generation's
|
|
74
|
+
prompt contains one of these words, the composer shows a warning that NAMES the words and the agent
|
|
75
|
+
route returns the same string. Silently editing the user's sentence to dodge a provider classifier
|
|
76
|
+
is forbidden — the words that reach the model are always the words the user can see.
|
|
77
|
+
|
|
78
|
+
---
|
|
79
|
+
|
|
80
|
+
## 🚨 Hazard 2 — 720p is not "the cheap one" any more
|
|
81
|
+
|
|
82
|
+
Every other model in Slates trains the habit that lower resolution means lower cost. 2.5 breaks it,
|
|
83
|
+
because the thing that moves the bill is **length**, and 2.5's length ceiling is double 2.0's.
|
|
84
|
+
|
|
85
|
+
Worked, at the shipped rates:
|
|
86
|
+
|
|
87
|
+
| Generation | Credits |
|
|
88
|
+
|---|---|
|
|
89
|
+
| 2.5 · 480p · 5s · faceless | 28 |
|
|
90
|
+
| 2.5 · 720p · 5s · faceless | 59 |
|
|
91
|
+
| 2.5 · 720p · 30s · faceless | 355 |
|
|
92
|
+
| 2.5 · 720p · 30s · AI-face route | **484** |
|
|
93
|
+
| 2.5 · 720p · 30s · consented real-face route | **710** |
|
|
94
|
+
| *(for scale)* 2.0 · 1080p · 15s · AI-face route | 411 |
|
|
95
|
+
|
|
96
|
+
**A 30-second 720p clip can cost more than a 15-second 1080p one** — and a base licence starts
|
|
97
|
+
with 1,000 credits. Someone who reads "720p" as "cheap" and asks for a 30-second take on the
|
|
98
|
+
real-face route has spent 71% of their welcome grant on one clip.
|
|
99
|
+
|
|
100
|
+
**Discipline:**
|
|
101
|
+
|
|
102
|
+
- **Always quote with `slates_estimate_generation_cost` before a take over ~10 seconds,** and say
|
|
103
|
+
the number out loud before generating.
|
|
104
|
+
- **Draft at 480p and 4–8 seconds.** Prove the composition, the motion and the identity first;
|
|
105
|
+
spend the length only on the take you already know works.
|
|
106
|
+
- **Length is a creative decision, not a default.** 30 seconds is available; it is rarely the right
|
|
107
|
+
answer for a single shot. Multi-shot storyboards inside one 30s generation are what the length is
|
|
108
|
+
actually for.
|
|
109
|
+
- Read `slates-cost-discipline` — all of it applies, more sharply here.
|
|
110
|
+
|
|
111
|
+
---
|
|
112
|
+
|
|
113
|
+
## Timestamps — the one grammar change
|
|
114
|
+
|
|
115
|
+
<!-- @inject:seedance-25-timestamps -->
|
|
116
|
+
**2.0 does not respond to timestamps and answers only to shot numbers. 2.5 responds to
|
|
117
|
+
integer-second timestamps.** That is ByteDance's own first line under "Differences from Seedance
|
|
118
|
+
2.0", and it is why a 30-second take is usable at all: the length is only worth buying if you can
|
|
119
|
+
say *when* things happen inside it.
|
|
120
|
+
|
|
121
|
+
Both formats are valid on 2.5, and you can mix them — `Shot N` blocks for a storyboard whose
|
|
122
|
+
pacing you are happy to leave to the model, timestamps when a beat has to land at a moment.
|
|
123
|
+
|
|
124
|
+
**Three ways to control time, all first-party:**
|
|
125
|
+
|
|
126
|
+
| Form | Write it like |
|
|
127
|
+
|---|---|
|
|
128
|
+
| **Interval** | `0-3 seconds… 3-7 seconds… 7-15 seconds` or `[1s-4s]… [4s-8s]… [8s-12s]` |
|
|
129
|
+
| **Time point** | *"Quick left sideways transition at the 5-second mark."* |
|
|
130
|
+
| **Relative** | *"After 3 seconds, everyone around him shakes their head."* · *"The frame freezes for 1 second after he presses the shutter."* |
|
|
131
|
+
|
|
132
|
+
**The rules that come with them:**
|
|
133
|
+
|
|
134
|
+
- **One second is the smallest unit.** Integers only — no `2.5s`, no frames.
|
|
135
|
+
- **No gaps in the timeline.** `0-3s… 5-6s…` leaves 3-5s unspecified and the model fills it however
|
|
136
|
+
it likes. Intervals must abut: `0-3s`, `3-7s`, `7-15s`.
|
|
137
|
+
- **Budget the plot to the seconds.** Too little content in a range and the model improvises to
|
|
138
|
+
fill it; too much and you get extra cuts or dropped beats. This is the actual craft of a 30s take.
|
|
139
|
+
- **Never time-code a high-frequency action.** *"Shake your head three times per second"* is
|
|
140
|
+
explicitly called out as a misuse — timestamps schedule beats, they don't choreograph frames.
|
|
141
|
+
- **Transitions want both halves:** the moment AND the method — *"At the 5-second mark, the camera
|
|
142
|
+
transitions leftward with a left wipe into a natural dissolve."*
|
|
143
|
+
- **Timestamps work on an EDIT too**, and that is where they earn the most: they scope a change in
|
|
144
|
+
time as well as in content — *"Change the man's action from drinking coffee to mopping the floor
|
|
145
|
+
from 4-6 seconds in Video 1, and leave the rest of the content unchanged."* Without a range, a
|
|
146
|
+
whole-clip instruction is applied to the whole clip.
|
|
147
|
+
|
|
148
|
+
Do **not** carry this back to 2.0, and do not carry Veo's `[00:00-00:02]` bracket syntax into
|
|
149
|
+
either — 2.0 ignores time entirely, and the cross-model syntax swap is its own known failure.
|
|
150
|
+
<!-- @end:seedance-25-timestamps -->
|
|
151
|
+
|
|
152
|
+
---
|
|
153
|
+
|
|
154
|
+
## What the extra reference budget is actually for
|
|
155
|
+
|
|
156
|
+
30 image references (up from 9) does **not** mean "attach 30 images". Every rule in
|
|
157
|
+
`slates-prompting-seedance` about references still holds — 2–4 strong references beat both
|
|
158
|
+
extremes, and one reference per role.
|
|
159
|
+
|
|
160
|
+
**Where 2.5 moves the ceiling, per ByteDance's own input recommendations:**
|
|
161
|
+
|
|
162
|
+
| | Stable | Works, but expect re-rolls |
|
|
163
|
+
|---|---|---|
|
|
164
|
+
| Subjects bound by IMAGE reference | 1–8 | 9–12 |
|
|
165
|
+
| Subjects bound by VIDEO or AUDIO reference | 1–5 | 6–10 |
|
|
166
|
+
| Reference clip length, per subject | 5–10s | longer drops stability |
|
|
167
|
+
|
|
168
|
+
**Multi-view images of one subject are supported on 2.5** — a turnaround sheet can be a single
|
|
169
|
+
reference image, where 2.0 wanted one authoritative rendering per subject. Past **5 subjects**,
|
|
170
|
+
go back to single-view images, one per view, rather than one image carrying several viewpoints.
|
|
171
|
+
|
|
172
|
+
The larger budget earns its keep in exactly two places:
|
|
173
|
+
|
|
174
|
+
- **A long multi-shot take** where different shots need different subjects and locations bound —
|
|
175
|
+
the budget is spread across the storyboard, not stacked on one frame.
|
|
176
|
+
- **Video and audio references alongside images**, which is where 2.5's 10 + 10 matters far more
|
|
177
|
+
than the image count.
|
|
178
|
+
|
|
179
|
+
### Audio-only references — the genuinely new input
|
|
180
|
+
|
|
181
|
+
2.0 required an image or video alongside any audio reference. **2.5 accepts audio on its own.**
|
|
182
|
+
That makes one recipe possible that was not before: drive a scene's timing, voice or ambience from
|
|
183
|
+
a recording with no visual anchor at all — a voice line, a music bed, a room tone — and let the
|
|
184
|
+
model build the picture to it. Cite it the same way as any other reference
|
|
185
|
+
(`Reference the timbre in <Audio_N> to generate…`), and remember that audio references carry **no
|
|
186
|
+
billing dimension** on any Seedance route: audio is included.
|
|
187
|
+
|
|
188
|
+
### Video references
|
|
189
|
+
|
|
190
|
+
Up to 10 clips, ≤30s combined (2.0: 3 clips, ≤15s). A reference VIDEO switches the cost key to
|
|
191
|
+
`seedance-2.5*-vref-{res}-{T}s`, where **T = Σ input seconds + output seconds** — the sum is across
|
|
192
|
+
**every** clip attached, not just the longest. Three 6-second references on a 12-second output bills
|
|
193
|
+
30 seconds, not 12 and not 18. Quote before confirming.
|
|
194
|
+
|
|
195
|
+
**The cap is a refusal, not a trim.** Attach an eleventh clip, or push past 30 combined seconds, and
|
|
196
|
+
the composition is rejected before anything uploads. That asymmetry is deliberate: reference images
|
|
197
|
+
warn-and-trim because dropping one doesn't change the price, and a dropped reference VIDEO would be
|
|
198
|
+
one you were quoted for and the model never saw.
|
|
199
|
+
|
|
200
|
+
### Mixing all three in one call
|
|
201
|
+
|
|
202
|
+
50 files total (30 image + 10 video + 10 audio) is a shared budget. Everything is cited positionally
|
|
203
|
+
by type — `image 1`, `video 2`, `audio 1` — in attachment order, so reordering the attachments
|
|
204
|
+
renumbers the citations. Write the prompt against those numbers:
|
|
205
|
+
|
|
206
|
+
```
|
|
207
|
+
Marcus (image 1) performs the motion from video 1 in the workshop from image 2,
|
|
208
|
+
speaking the line in audio 1. Preserve his identity, appearance and outfit.
|
|
209
|
+
```
|
|
210
|
+
|
|
211
|
+
Frames and reference media stay mutually exclusive, in every combination — the reference endpoint
|
|
212
|
+
has no first/last-frame parameters at all, so this is a shape mismatch rather than a preference.
|
|
213
|
+
|
|
214
|
+
---
|
|
215
|
+
|
|
216
|
+
## Seedance 2.5 Edit (`slates_edit_video`, `model: 'seedance-2.5-edit'`)
|
|
217
|
+
|
|
218
|
+
Its own picker row and its own op call, deliberately: the task type is **the model you chose**,
|
|
219
|
+
never something inferred from your sentence.
|
|
220
|
+
|
|
221
|
+
**Why route here at all:** it is the **only edit engine in Slates that accepts a clip longer than 15
|
|
222
|
+
seconds** (4–30s, versus Kling O3 Edit's 3–15s and Omni Flash Edit's 3–10s). For a clip inside the
|
|
223
|
+
others' range, choose on fidelity instead — Omni Flash Edit won the prompt-only head-to-head, and
|
|
224
|
+
Kling O3 Edit is the one that takes element and style reference images.
|
|
225
|
+
|
|
226
|
+
**How it behaves:**
|
|
227
|
+
|
|
228
|
+
- **Output length follows the SOURCE clip**, and the bill is the ceiled source length. The provider
|
|
229
|
+
requires an automatic duration on this task type, so there is no length knob — the clip you attach
|
|
230
|
+
is the quote. The returned clip can differ from the source by up to ~0.3s, which only compresses
|
|
231
|
+
transition frames; a clip that 2.5 itself generated comes back at exactly its input length.
|
|
232
|
+
- **Source clips under 20 seconds edit more reliably.** 4–30s is what the task type accepts;
|
|
233
|
+
ByteDance's own recommendation is to stay inside 20 for quality. A 28-second source is legal and
|
|
234
|
+
will need more attempts.
|
|
235
|
+
- **The aspect ratio follows the source clip too.** No ratio control; the frame is the clip's frame.
|
|
236
|
+
- **480p or 720p output**, native audio.
|
|
237
|
+
- **Prompt and source clip only** on this op. The MODEL takes reference images on an edit
|
|
238
|
+
(ByteDance recommends 1–5 — *"replace the man in dark clothing in @Video 1 with @Image 2"*);
|
|
239
|
+
**Slates has not wired that path**, so today an edit that must lock an identity from a photo
|
|
240
|
+
goes to Kling O3 Edit. Constraint of our build, not of the model — worth revisiting.
|
|
241
|
+
- **An edit bills roughly DOUBLE a plain 2.5 generation of the same length**, because every provider
|
|
242
|
+
charges an edit on input + output seconds. Read the confirm gate's number; do not reason from the
|
|
243
|
+
generation rate.
|
|
244
|
+
- **Set `seedanceFace: true` when a character's face is visible in the clip.** The faceless provider
|
|
245
|
+
blocks faces outright — this is not a price optimisation, it is whether the job runs at all.
|
|
246
|
+
- **There is no consented-real-face route for editing.** Real-person footage that the AI-face route
|
|
247
|
+
rejects has to go to Kling O3 Edit.
|
|
248
|
+
|
|
249
|
+
**Prompting an edit** — the same discipline as every other edit engine: **describe only what
|
|
250
|
+
changes.** The source already carries its composition, motion, timing and performance; re-describing
|
|
251
|
+
them fights the model. Use Seedance's own edit grammar from `slates-prompting-seedance`
|
|
252
|
+
(*"Strictly edit `<Video_1>`, and modify `<Original_Characteristic>` to `<New_Characteristic>`"*) and
|
|
253
|
+
**never** write *"reference video 1"* in an edit — the official guide is explicit that this phrasing
|
|
254
|
+
gets the request reclassified as a reference task, which is the same landmine as Hazard 1.
|
|
255
|
+
|
|
256
|
+
Two things sharpen an edit prompt, both first-party:
|
|
257
|
+
|
|
258
|
+
- **Say it as A → B, not as an outcome.** *"Change the man's action from drinking coffee to mopping
|
|
259
|
+
the floor"* beats *"the man mops the floor"* — naming what it currently is tells the model what
|
|
260
|
+
to overwrite.
|
|
261
|
+
- **Timestamp a partial edit** — the edit task type reads the same integer-second timestamps the
|
|
262
|
+
generation path does. Rules and forms are in § Timestamps above; this is the single most useful
|
|
263
|
+
thing they buy.
|
|
264
|
+
|
|
265
|
+
**Audio is editable too, and it is the least obvious use of this row.** The same op rewrites what
|
|
266
|
+
is heard while the picture stays put: change a spoken line, change the accent, translate the
|
|
267
|
+
dialogue and re-fit the lip movement, strip or replace the BGM or a sound effect. *"Only edit the
|
|
268
|
+
man's dialogue in Video 1: change it to 'Don't come over here,' in an American accent"* is an edit,
|
|
269
|
+
not a lip-sync job. Bill it like any other edit — on the source clip's length.
|
|
270
|
+
|
|
271
|
+
---
|
|
272
|
+
|
|
273
|
+
## Faces, and what does NOT change
|
|
274
|
+
|
|
275
|
+
The three-tier face routing is identical to 2.0 — faceless → default route, an AI character's face →
|
|
276
|
+
`seedanceFace: true` (the relaxed provider, a real cost premium), a real person's photo → the
|
|
277
|
+
consent-gated premium route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`
|
|
278
|
+
set **only** after the user explicitly confirms they hold the rights to the likeness. The full rules,
|
|
279
|
+
including why the real-vs-AI call is the provider's and not yours, are in
|
|
280
|
+
`slates-prompting-seedance`.
|
|
281
|
+
|
|
282
|
+
Also unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:
|
|
283
|
+
|
|
284
|
+
- **One primary camera move per shot.**
|
|
285
|
+
- **No lens / aperture / film-stock vocabulary.** That is image-model syntax and a Seedance
|
|
286
|
+
anti-pattern.
|
|
287
|
+
- **No `negativePrompt` field** — constraints go inline, and 2.5 acts on negative phrasing in
|
|
288
|
+
exactly two dimensions: subtitles (*"no subtitles"*) and audio (*"no BGM; environmental and
|
|
289
|
+
action sounds only"*, *"no audio"*). Everywhere else, describe what you want, not what you don't.
|
|
290
|
+
- **Legible in-shot text still belongs in a baked start frame**, not in the video prompt.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-seedance
|
|
3
|
-
description: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance structures multi-beat prompts as a "Shot 1 / Shot 2 / Shot 3" storyboard against an 8-slot advanced formula — never per-second time stamps. Its syntax differs from Kling, Veo and the image models; don't cross-pollinate (in particular, no lens / aperture / film-stock vocabulary).
|
|
3
|
+
description: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance 2.0 structures multi-beat prompts as a "Shot 1 / Shot 2 / Shot 3" storyboard against an 8-slot advanced formula — never per-second time stamps, which 2.0 does not respond to (Seedance 2.5 does; see slates-prompting-seedance-2-5). Its syntax differs from Kling, Veo and the image models; don't cross-pollinate (in particular, no lens / aperture / film-stock vocabulary).
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Seedance 2.0 — prompting
|
|
@@ -162,7 +162,7 @@ Seedance has **no `negativePrompt` field** — constraints go inline in this slo
|
|
|
162
162
|
|
|
163
163
|
## Worked examples `[official :1689-1745]`
|
|
164
164
|
|
|
165
|
-
These are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere.
|
|
165
|
+
These are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere — 2.0 does not respond to them at all, which is version-scoped and reverses on 2.5.
|
|
166
166
|
|
|
167
167
|
**Example 1 — dormitory emotional short drama (dialogue-focused).** Assets: `@Image 1` half-body photo of the female lead · `@Image 2` dormitory scene reference · `@Video 1` camera-movement reference · `@Audio 1` indoor ambience.
|
|
168
168
|
|
|
@@ -202,7 +202,23 @@ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding pre
|
|
|
202
202
|
|
|
203
203
|
Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
|
|
204
204
|
|
|
205
|
-
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
|
|
205
|
+
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.
|
|
206
|
+
|
|
207
|
+
### All three modalities go in ONE call
|
|
208
|
+
|
|
209
|
+
The caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.
|
|
210
|
+
|
|
211
|
+
Cite each by type and index, in the order they were attached — `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.
|
|
212
|
+
|
|
213
|
+
```
|
|
214
|
+
Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
|
|
215
|
+
speaking the line in audio 1. Preserve his identity, appearance and outfit.
|
|
216
|
+
```
|
|
217
|
+
<!-- slates-only -->
|
|
218
|
+
**Attaching a clip is NOT the same as editing it.** "Add as reference" puts it in the composer alongside everything else and wipes nothing; "Edit with AI" makes the clip the canvas and clears the tray for a fresh instruction. Two different jobs, two different menu entries — never infer one from the other.
|
|
219
|
+
|
|
220
|
+
**Over the cap is REFUSED, never trimmed.** A reference video is priced into the quote before it is sent, so a clip silently dropped after the quote would be a clip you paid for and the model never saw. Remove one and retry.
|
|
221
|
+
<!-- /slates-only -->
|
|
206
222
|
|
|
207
223
|
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
208
224
|
|
|
@@ -1,110 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: slates-prompting-suno
|
|
3
|
-
description: How to prompt Suno in Slates. Read before calling slates_generate_audio with model suno. Full music tracks - EVERY call returns TWO variations for one flat price and duration is FREE up to 360 seconds. The rule that decides everything - in CUSTOM mode the prompt field is the EXACT LYRICS (sung as written), in DESCRIPTION mode it is a description and the lyrics get written for you. Covers the mode matrix, style vs prompt steering, instrumental scoring, negative tags, and the character caps.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Suno — prompting
|
|
7
|
-
|
|
8
|
-
Full music generation, reached through the sunoapi.org wrapper. Models `V4`, `V4_5`, `V4_5PLUS`, `V4_5ALL`, `V5`, `V5_5`.
|
|
9
|
-
|
|
10
|
-
**Two facts that should shape every decision:**
|
|
11
|
-
|
|
12
|
-
1. **Every call returns TWO songs** — two genuinely different takes on the same brief, for one flat price. Audition both before re-rolling.
|
|
13
|
-
2. **Length is free.** A 360-second track costs exactly what a default one costs (measured against the live provider balance 2026-07-31: a default call and a `duration: 240` call both debited the same). There is never a reason to generate a bed shorter than your edit.
|
|
14
|
-
|
|
15
|
-
## Where it routes
|
|
16
|
-
|
|
17
|
-
- **Anything a listener would call a song or a score** — theme, underscore, needle-drop, montage bed, end-card sting.
|
|
18
|
-
- **NOT** ambience or room tone — that is `seed-audio`, which is cheaper and better at it.
|
|
19
|
-
- **NOT** a single effect — that is `eleven-sfx`.
|
|
20
|
-
- **AUDIO-ONLY.**
|
|
21
|
-
|
|
22
|
-
## 🚨 THE RULE: which mode you are in changes what `prompt` means
|
|
23
|
-
|
|
24
|
-
| `customMode` | `instrumental` | Required | What `prompt` means |
|
|
25
|
-
|---|---|---|---|
|
|
26
|
-
| `false` | either | `prompt` only (≤500 chars) | **A description.** Lyrics get written for you. |
|
|
27
|
-
| `true` | `true` | `style`, `title` | **Unused.** Style + title do all the steering. |
|
|
28
|
-
| `true` | `false` | `style`, `title`, `prompt` | **THE EXACT LYRICS**, sung as written. |
|
|
29
|
-
|
|
30
|
-
Putting a description in the prompt field while `customMode: true` and `instrumental: false` gets your description **sung back at you**. This is the single most common Suno mistake and it costs a full generation every time.
|
|
31
|
-
|
|
32
|
-
## Description mode — the fast path
|
|
33
|
-
|
|
34
|
-
```
|
|
35
|
-
customMode: false
|
|
36
|
-
prompt: "brooding synthwave for a night drive, analog bass, no vocals, 90 bpm"
|
|
37
|
-
```
|
|
38
|
-
|
|
39
|
-
Use it when you need a mood and do not care about specific words. 500-character cap. This is the right default for background beds.
|
|
40
|
-
|
|
41
|
-
## Custom mode — when the words matter
|
|
42
|
-
|
|
43
|
-
```
|
|
44
|
-
customMode: true
|
|
45
|
-
instrumental: false
|
|
46
|
-
style: "dream pop, hazy, reverb-heavy guitars, female vocal, 100 bpm"
|
|
47
|
-
title: "Blue Hour"
|
|
48
|
-
prompt: "[Verse 1]\nThe lights come on before we're ready\n..."
|
|
49
|
-
```
|
|
50
|
-
|
|
51
|
-
Structure tags (`[Verse]`, `[Chorus]`, `[Bridge]`, `[Outro]`) inside the lyrics are how you control the arrangement. Everything that is not a structure tag will be sung.
|
|
52
|
-
|
|
53
|
-
## Instrumental scoring
|
|
54
|
-
|
|
55
|
-
```
|
|
56
|
-
customMode: true
|
|
57
|
-
instrumental: true
|
|
58
|
-
style: "tense orchestral strings, low brass swells, no percussion"
|
|
59
|
-
title: "Approach"
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
`instrumental: true` is the right answer for almost every film bed — a vocal you did not ask for will fight your dialogue.
|
|
63
|
-
|
|
64
|
-
## Steer with `style`, not with adjective piles
|
|
65
|
-
|
|
66
|
-
Genre + era + instrumentation + tempo belong in `style`, not stuffed into `prompt`.
|
|
67
|
-
|
|
68
|
-
```
|
|
69
|
-
✓ style: "90s trip-hop, dusty breakbeat, Rhodes piano, upright bass, 85 bpm"
|
|
70
|
-
✗ prompt: "a really cool dusty 90s trip hop song with a Rhodes and..."
|
|
71
|
-
```
|
|
72
|
-
|
|
73
|
-
`negativeTags` removes what keeps creeping in: `"brass, EDM drop, male vocal"`.
|
|
74
|
-
|
|
75
|
-
## The steering knobs
|
|
76
|
-
|
|
77
|
-
| Param | Range | Reach for it when |
|
|
78
|
-
|---|---|---|
|
|
79
|
-
| `duration` | 10–360s (**V5_5 + custom mode only**) | Always, when the bed must outlast the cut. It is free. |
|
|
80
|
-
| `vocalGender` | `m` / `f` — **the wire values, not "male"/"female"** | A specific voice is required. |
|
|
81
|
-
| `styleWeight` | 0–1 | The style field is being ignored (raise) or strangling the song (lower). |
|
|
82
|
-
| `weirdnessConstraint` | 0–1 | Takes are too safe (raise) or falling apart (lower). |
|
|
83
|
-
| `audioWeight` | 0–1 | Balancing an audio input against the prompt. |
|
|
84
|
-
| `personaId` / `personaModel` | — | A series needs the same voice/sound across episodes. |
|
|
85
|
-
|
|
86
|
-
## Character caps
|
|
87
|
-
|
|
88
|
-
| Field | V4 | V4_5 / V4_5PLUS / V5 / V5_5 | V4_5ALL |
|
|
89
|
-
|---|---|---|---|
|
|
90
|
-
| prompt (custom = literal lyrics) | 3000 | 5000 | 5000 |
|
|
91
|
-
| prompt (non-custom = description) | 500 | 500 | 500 |
|
|
92
|
-
| style | 200 | 1000 | 1000 |
|
|
93
|
-
| title | 80 | 100 | 80 |
|
|
94
|
-
|
|
95
|
-
## Iterating
|
|
96
|
-
|
|
97
|
-
- **Audition both returned songs first.** A re-roll costs a full generation; the second variation is already paid for.
|
|
98
|
-
- Wrong genre → fix `style`. Wrong words → you are in the wrong mode, check the matrix above.
|
|
99
|
-
- Something keeps appearing that you do not want → `negativeTags`, not more prompt.
|
|
100
|
-
- Three failed generations on the same brief means the style field is too vague, not that the seed is unlucky.
|
|
101
|
-
|
|
102
|
-
## Ops notes
|
|
103
|
-
|
|
104
|
-
- Tracks take 2–3 minutes; a streamable preview exists ~30–40s in. Use `background: true` and poll.
|
|
105
|
-
- The provider hosts files for a limited window — **Slates downloads and stores them locally as soon as the track finishes**, so nothing expires out from under a project.
|
|
106
|
-
- Suno has **no official public API**; this rides an unofficial wrapper. Treat availability as best-effort and do not build a deadline around it.
|
|
107
|
-
|
|
108
|
-
## Content notes
|
|
109
|
-
|
|
110
|
-
Provider-side moderation rejects lyrics and style prompts naming real artists or protected material (`SENSITIVE_WORD_ERROR`). Describe the sound, not the artist. See slates-content-policy.
|