@slatesvideo/shared 0.5.6 → 0.5.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,290 @@
1
+ ---
2
+ name: slates-prompting-seedance-2-5
3
+ description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is a SECOND SEAT next to 2.0, not an upgrade — it buys 30-second takes, 30 image references, audio-only references and INTEGER-SECOND TIMESTAMPS, and it gives up 1080p and 4K entirely. Timestamps are the one grammar difference that matters: 2.0 ignores them and answers only to shot numbers, 2.5 acts on them. Otherwise it shares 2.0's grammar (read slates-prompting-seedance for subject binding, camera and constraint vocabulary); this file covers what is different, plus the two hazards unique to 2.5 — the prompt-intent task classifier and the cost trap at 720p.
4
+ ---
5
+
6
+ # Seedance 2.5 — prompting
7
+
8
+ **Read `slates-prompting-seedance` first.** The prompt GRAMMAR is the same model family: the
9
+ 8-slot advanced formula, subject binding by `<Subject_N>@<Image_N>`, camera vocabulary,
10
+ externalised emotion, inline constraint words, the anti-twin fix. None of it is restated here.
11
+ This file is only what 2.5 changes — and the biggest change is that **2.5 acts on timestamps
12
+ where 2.0 ignores them.**
13
+
14
+ ---
15
+
16
+ ## The one fact that decides whether you use it at all
17
+
18
+ **Seedance 2.5 is 480p or 720p. There is no 1080p and no 4K, on any provider.**
19
+
20
+ That is not a Slates limitation or a tier gate — the model does not produce those resolutions.
21
+ So 2.5 does not replace 2.0; it sits beside it:
22
+
23
+ | | Seedance 2.0 | Seedance 2.5 |
24
+ |---|---|---|
25
+ | Resolution | 480p / 720p / 1080p / **native 4K** | **480p / 720p only** |
26
+ | Length | 4–15s | **4–30s in one take** |
27
+ | Reference budget | 15 (9 image + 3 video + 3 audio) | **50 (30 image + 10 video + 10 audio)** |
28
+ | Combined reference video/audio | ≤15s | **≤30s** |
29
+ | Audio-only reference | ✗ (needs an image or video alongside) | **✓** |
30
+ | **Timestamps in the prompt** | **✗ — ignored; shot numbers only** | **✓ — integer seconds, acted on** |
31
+ | Multi-view image as ONE subject reference | ✗ (not recommended) | **✓ (up to 5 subjects)** |
32
+ | Video edit as its own task type | ✗ | **✓ (`seedance-2.5-edit`)** |
33
+ | Default video model | **yes** | no |
34
+
35
+ **Route to 2.5 when the shot needs LENGTH, MANY REFERENCES, or an audio-only reference.
36
+ Route to 2.0 when resolution matters at all** — which, for anything a client will see full-screen,
37
+ is most of the time.
38
+
39
+ ---
40
+
41
+ ## 🚨 Hazard 1 — the prompt-intent task classifier
42
+
43
+ This is the one that costs money and time, and it has no equivalent on 2.0.
44
+
45
+ **Seedance 2.5 sorts every request into one of five task types** — text-to-video,
46
+ reference-to-video, first/last-frame, **video edit**, **video extend** — from the reference roles
47
+ attached **plus the intent of your sentence**. Each type then has its own parameter constraints,
48
+ and a violation comes back **asynchronously**: the task queues, credits are reserved, and only then
49
+ does it fail.
50
+
51
+ The trigger words are ordinary English:
52
+
53
+ | Reclassified as | Words that do it (ByteDance's own list) |
54
+ |---|---|
55
+ | **video edit** | `edit video` · `add` · `insert` · `remove` · `delete` · `modify` · `replace` · `change to` |
56
+ | **video extend** | `extend forward` · `extend backward` · `continue` · `continue from` · `extend the story` |
57
+
58
+ So a perfectly legitimate reference-to-video prompt — *"a wide shot of the workshop, **remove** the
59
+ tripod from frame"* — gets classified as an edit and fails on constraints it never set.
60
+
61
+ **What to do:**
62
+
63
+ 1. **If you mean to edit an existing clip, say so with the MODEL, not the sentence.** Call
64
+ `slates_edit_video` with `model: 'seedance-2.5-edit'`. That routes to a dedicated
65
+ task-typed endpoint and the classifier never has to guess.
66
+ 2. **If you mean a fresh shot, describe the finished frame rather than an instruction to change
67
+ one.** Not *"remove the tripod"* → *"the workshop bench, clear and uncluttered"*. Not
68
+ *"add rain"* → *"heavy rain falling through the streetlight"*. This is better prompting anyway:
69
+ the model renders what you describe, it does not take edits to an imagined draft.
70
+ 3. The trigger only fires when **references are attached**. A plain text-to-video prompt is safe
71
+ however it is worded.
72
+
73
+ **Slates will warn you, and it will never rewrite your prompt.** When a 2.5 reference generation's
74
+ prompt contains one of these words, the composer shows a warning that NAMES the words and the agent
75
+ route returns the same string. Silently editing the user's sentence to dodge a provider classifier
76
+ is forbidden — the words that reach the model are always the words the user can see.
77
+
78
+ ---
79
+
80
+ ## 🚨 Hazard 2 — 720p is not "the cheap one" any more
81
+
82
+ Every other model in Slates trains the habit that lower resolution means lower cost. 2.5 breaks it,
83
+ because the thing that moves the bill is **length**, and 2.5's length ceiling is double 2.0's.
84
+
85
+ Worked, at the shipped rates:
86
+
87
+ | Generation | Credits |
88
+ |---|---|
89
+ | 2.5 · 480p · 5s · faceless | 28 |
90
+ | 2.5 · 720p · 5s · faceless | 59 |
91
+ | 2.5 · 720p · 30s · faceless | 355 |
92
+ | 2.5 · 720p · 30s · AI-face route | **484** |
93
+ | 2.5 · 720p · 30s · consented real-face route | **710** |
94
+ | *(for scale)* 2.0 · 1080p · 15s · AI-face route | 411 |
95
+
96
+ **A 30-second 720p clip can cost more than a 15-second 1080p one** — and a base licence starts
97
+ with 1,000 credits. Someone who reads "720p" as "cheap" and asks for a 30-second take on the
98
+ real-face route has spent 71% of their welcome grant on one clip.
99
+
100
+ **Discipline:**
101
+
102
+ - **Always quote with `slates_estimate_generation_cost` before a take over ~10 seconds,** and say
103
+ the number out loud before generating.
104
+ - **Draft at 480p and 4–8 seconds.** Prove the composition, the motion and the identity first;
105
+ spend the length only on the take you already know works.
106
+ - **Length is a creative decision, not a default.** 30 seconds is available; it is rarely the right
107
+ answer for a single shot. Multi-shot storyboards inside one 30s generation are what the length is
108
+ actually for.
109
+ - Read `slates-cost-discipline` — all of it applies, more sharply here.
110
+
111
+ ---
112
+
113
+ ## Timestamps — the one grammar change
114
+
115
+ <!-- @inject:seedance-25-timestamps -->
116
+ **2.0 does not respond to timestamps and answers only to shot numbers. 2.5 responds to
117
+ integer-second timestamps.** That is ByteDance's own first line under "Differences from Seedance
118
+ 2.0", and it is why a 30-second take is usable at all: the length is only worth buying if you can
119
+ say *when* things happen inside it.
120
+
121
+ Both formats are valid on 2.5, and you can mix them — `Shot N` blocks for a storyboard whose
122
+ pacing you are happy to leave to the model, timestamps when a beat has to land at a moment.
123
+
124
+ **Three ways to control time, all first-party:**
125
+
126
+ | Form | Write it like |
127
+ |---|---|
128
+ | **Interval** | `0-3 seconds… 3-7 seconds… 7-15 seconds` or `[1s-4s]… [4s-8s]… [8s-12s]` |
129
+ | **Time point** | *"Quick left sideways transition at the 5-second mark."* |
130
+ | **Relative** | *"After 3 seconds, everyone around him shakes their head."* · *"The frame freezes for 1 second after he presses the shutter."* |
131
+
132
+ **The rules that come with them:**
133
+
134
+ - **One second is the smallest unit.** Integers only — no `2.5s`, no frames.
135
+ - **No gaps in the timeline.** `0-3s… 5-6s…` leaves 3-5s unspecified and the model fills it however
136
+ it likes. Intervals must abut: `0-3s`, `3-7s`, `7-15s`.
137
+ - **Budget the plot to the seconds.** Too little content in a range and the model improvises to
138
+ fill it; too much and you get extra cuts or dropped beats. This is the actual craft of a 30s take.
139
+ - **Never time-code a high-frequency action.** *"Shake your head three times per second"* is
140
+ explicitly called out as a misuse — timestamps schedule beats, they don't choreograph frames.
141
+ - **Transitions want both halves:** the moment AND the method — *"At the 5-second mark, the camera
142
+ transitions leftward with a left wipe into a natural dissolve."*
143
+ - **Timestamps work on an EDIT too**, and that is where they earn the most: they scope a change in
144
+ time as well as in content — *"Change the man's action from drinking coffee to mopping the floor
145
+ from 4-6 seconds in Video 1, and leave the rest of the content unchanged."* Without a range, a
146
+ whole-clip instruction is applied to the whole clip.
147
+
148
+ Do **not** carry this back to 2.0, and do not carry Veo's `[00:00-00:02]` bracket syntax into
149
+ either — 2.0 ignores time entirely, and the cross-model syntax swap is its own known failure.
150
+ <!-- @end:seedance-25-timestamps -->
151
+
152
+ ---
153
+
154
+ ## What the extra reference budget is actually for
155
+
156
+ 30 image references (up from 9) does **not** mean "attach 30 images". Every rule in
157
+ `slates-prompting-seedance` about references still holds — 2–4 strong references beat both
158
+ extremes, and one reference per role.
159
+
160
+ **Where 2.5 moves the ceiling, per ByteDance's own input recommendations:**
161
+
162
+ | | Stable | Works, but expect re-rolls |
163
+ |---|---|---|
164
+ | Subjects bound by IMAGE reference | 1–8 | 9–12 |
165
+ | Subjects bound by VIDEO or AUDIO reference | 1–5 | 6–10 |
166
+ | Reference clip length, per subject | 5–10s | longer drops stability |
167
+
168
+ **Multi-view images of one subject are supported on 2.5** — a turnaround sheet can be a single
169
+ reference image, where 2.0 wanted one authoritative rendering per subject. Past **5 subjects**,
170
+ go back to single-view images, one per view, rather than one image carrying several viewpoints.
171
+
172
+ The larger budget earns its keep in exactly two places:
173
+
174
+ - **A long multi-shot take** where different shots need different subjects and locations bound —
175
+ the budget is spread across the storyboard, not stacked on one frame.
176
+ - **Video and audio references alongside images**, which is where 2.5's 10 + 10 matters far more
177
+ than the image count.
178
+
179
+ ### Audio-only references — the genuinely new input
180
+
181
+ 2.0 required an image or video alongside any audio reference. **2.5 accepts audio on its own.**
182
+ That makes one recipe possible that was not before: drive a scene's timing, voice or ambience from
183
+ a recording with no visual anchor at all — a voice line, a music bed, a room tone — and let the
184
+ model build the picture to it. Cite it the same way as any other reference
185
+ (`Reference the timbre in <Audio_N> to generate…`), and remember that audio references carry **no
186
+ billing dimension** on any Seedance route: audio is included.
187
+
188
+ ### Video references
189
+
190
+ Up to 10 clips, ≤30s combined (2.0: 3 clips, ≤15s). A reference VIDEO switches the cost key to
191
+ `seedance-2.5*-vref-{res}-{T}s`, where **T = Σ input seconds + output seconds** — the sum is across
192
+ **every** clip attached, not just the longest. Three 6-second references on a 12-second output bills
193
+ 30 seconds, not 12 and not 18. Quote before confirming.
194
+
195
+ **The cap is a refusal, not a trim.** Attach an eleventh clip, or push past 30 combined seconds, and
196
+ the composition is rejected before anything uploads. That asymmetry is deliberate: reference images
197
+ warn-and-trim because dropping one doesn't change the price, and a dropped reference VIDEO would be
198
+ one you were quoted for and the model never saw.
199
+
200
+ ### Mixing all three in one call
201
+
202
+ 50 files total (30 image + 10 video + 10 audio) is a shared budget. Everything is cited positionally
203
+ by type — `image 1`, `video 2`, `audio 1` — in attachment order, so reordering the attachments
204
+ renumbers the citations. Write the prompt against those numbers:
205
+
206
+ ```
207
+ Marcus (image 1) performs the motion from video 1 in the workshop from image 2,
208
+ speaking the line in audio 1. Preserve his identity, appearance and outfit.
209
+ ```
210
+
211
+ Frames and reference media stay mutually exclusive, in every combination — the reference endpoint
212
+ has no first/last-frame parameters at all, so this is a shape mismatch rather than a preference.
213
+
214
+ ---
215
+
216
+ ## Seedance 2.5 Edit (`slates_edit_video`, `model: 'seedance-2.5-edit'`)
217
+
218
+ Its own picker row and its own op call, deliberately: the task type is **the model you chose**,
219
+ never something inferred from your sentence.
220
+
221
+ **Why route here at all:** it is the **only edit engine in Slates that accepts a clip longer than 15
222
+ seconds** (4–30s, versus Kling O3 Edit's 3–15s and Omni Flash Edit's 3–10s). For a clip inside the
223
+ others' range, choose on fidelity instead — Omni Flash Edit won the prompt-only head-to-head, and
224
+ Kling O3 Edit is the one that takes element and style reference images.
225
+
226
+ **How it behaves:**
227
+
228
+ - **Output length follows the SOURCE clip**, and the bill is the ceiled source length. The provider
229
+ requires an automatic duration on this task type, so there is no length knob — the clip you attach
230
+ is the quote. The returned clip can differ from the source by up to ~0.3s, which only compresses
231
+ transition frames; a clip that 2.5 itself generated comes back at exactly its input length.
232
+ - **Source clips under 20 seconds edit more reliably.** 4–30s is what the task type accepts;
233
+ ByteDance's own recommendation is to stay inside 20 for quality. A 28-second source is legal and
234
+ will need more attempts.
235
+ - **The aspect ratio follows the source clip too.** No ratio control; the frame is the clip's frame.
236
+ - **480p or 720p output**, native audio.
237
+ - **Prompt and source clip only** on this op. The MODEL takes reference images on an edit
238
+ (ByteDance recommends 1–5 — *"replace the man in dark clothing in @Video 1 with @Image 2"*);
239
+ **Slates has not wired that path**, so today an edit that must lock an identity from a photo
240
+ goes to Kling O3 Edit. Constraint of our build, not of the model — worth revisiting.
241
+ - **An edit bills roughly DOUBLE a plain 2.5 generation of the same length**, because every provider
242
+ charges an edit on input + output seconds. Read the confirm gate's number; do not reason from the
243
+ generation rate.
244
+ - **Set `seedanceFace: true` when a character's face is visible in the clip.** The faceless provider
245
+ blocks faces outright — this is not a price optimisation, it is whether the job runs at all.
246
+ - **There is no consented-real-face route for editing.** Real-person footage that the AI-face route
247
+ rejects has to go to Kling O3 Edit.
248
+
249
+ **Prompting an edit** — the same discipline as every other edit engine: **describe only what
250
+ changes.** The source already carries its composition, motion, timing and performance; re-describing
251
+ them fights the model. Use Seedance's own edit grammar from `slates-prompting-seedance`
252
+ (*"Strictly edit `<Video_1>`, and modify `<Original_Characteristic>` to `<New_Characteristic>`"*) and
253
+ **never** write *"reference video 1"* in an edit — the official guide is explicit that this phrasing
254
+ gets the request reclassified as a reference task, which is the same landmine as Hazard 1.
255
+
256
+ Two things sharpen an edit prompt, both first-party:
257
+
258
+ - **Say it as A → B, not as an outcome.** *"Change the man's action from drinking coffee to mopping
259
+ the floor"* beats *"the man mops the floor"* — naming what it currently is tells the model what
260
+ to overwrite.
261
+ - **Timestamp a partial edit** — the edit task type reads the same integer-second timestamps the
262
+ generation path does. Rules and forms are in § Timestamps above; this is the single most useful
263
+ thing they buy.
264
+
265
+ **Audio is editable too, and it is the least obvious use of this row.** The same op rewrites what
266
+ is heard while the picture stays put: change a spoken line, change the accent, translate the
267
+ dialogue and re-fit the lip movement, strip or replace the BGM or a sound effect. *"Only edit the
268
+ man's dialogue in Video 1: change it to 'Don't come over here,' in an American accent"* is an edit,
269
+ not a lip-sync job. Bill it like any other edit — on the source clip's length.
270
+
271
+ ---
272
+
273
+ ## Faces, and what does NOT change
274
+
275
+ The three-tier face routing is identical to 2.0 — faceless → default route, an AI character's face →
276
+ `seedanceFace: true` (the relaxed provider, a real cost premium), a real person's photo → the
277
+ consent-gated premium route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`
278
+ set **only** after the user explicitly confirms they hold the rights to the likeness. The full rules,
279
+ including why the real-vs-AI call is the provider's and not yours, are in
280
+ `slates-prompting-seedance`.
281
+
282
+ Also unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:
283
+
284
+ - **One primary camera move per shot.**
285
+ - **No lens / aperture / film-stock vocabulary.** That is image-model syntax and a Seedance
286
+ anti-pattern.
287
+ - **No `negativePrompt` field** — constraints go inline, and 2.5 acts on negative phrasing in
288
+ exactly two dimensions: subtitles (*"no subtitles"*) and audio (*"no BGM; environmental and
289
+ action sounds only"*, *"no audio"*). Everywhere else, describe what you want, not what you don't.
290
+ - **Legible in-shot text still belongs in a baked start frame**, not in the video prompt.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-seedance
3
- description: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance structures multi-beat prompts as a "Shot 1 / Shot 2 / Shot 3" storyboard against an 8-slot advanced formula — never per-second time stamps. Its syntax differs from Kling, Veo and the image models; don't cross-pollinate (in particular, no lens / aperture / film-stock vocabulary).
3
+ description: How to prompt Seedance 2.0 (ByteDance video model). Read before calling slates_generate_video with model seedance-2. Seedance 2.0 structures multi-beat prompts as a "Shot 1 / Shot 2 / Shot 3" storyboard against an 8-slot advanced formula — never per-second time stamps, which 2.0 does not respond to (Seedance 2.5 does; see slates-prompting-seedance-2-5). Its syntax differs from Kling, Veo and the image models; don't cross-pollinate (in particular, no lens / aperture / film-stock vocabulary).
4
4
  ---
5
5
 
6
6
  # Seedance 2.0 — prompting
@@ -162,7 +162,7 @@ Seedance has **no `negativePrompt` field** — constraints go inline in this slo
162
162
 
163
163
  ## Worked examples `[official :1689-1745]`
164
164
 
165
- These are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere.
165
+ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere — 2.0 does not respond to them at all, which is version-scoped and reverses on 2.5.
166
166
 
167
167
  **Example 1 — dormitory emotional short drama (dialogue-focused).** Assets: `@Image 1` half-body photo of the female lead · `@Image 2` dormitory scene reference · `@Video 1` camera-movement reference · `@Audio 1` indoor ambience.
168
168
 
@@ -202,7 +202,23 @@ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding pre
202
202
 
203
203
  Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
204
204
 
205
- **Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
205
+ **Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.
206
+
207
+ ### All three modalities go in ONE call
208
+
209
+ The caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.
210
+
211
+ Cite each by type and index, in the order they were attached — `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.
212
+
213
+ ```
214
+ Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
215
+ speaking the line in audio 1. Preserve his identity, appearance and outfit.
216
+ ```
217
+ <!-- slates-only -->
218
+ **Attaching a clip is NOT the same as editing it.** "Add as reference" puts it in the composer alongside everything else and wipes nothing; "Edit with AI" makes the clip the canvas and clears the tray for a fresh instruction. Two different jobs, two different menu entries — never infer one from the other.
219
+
220
+ **Over the cap is REFUSED, never trimmed.** A reference video is priced into the quote before it is sent, so a clip silently dropped after the quote would be a clip you paid for and the model never saw. Remove one and retry.
221
+ <!-- /slates-only -->
206
222
 
207
223
  ### Motion transfer & lip-sync recipes (reference video / audio)
208
224
 
@@ -1,110 +0,0 @@
1
- ---
2
- name: slates-prompting-suno
3
- description: How to prompt Suno in Slates. Read before calling slates_generate_audio with model suno. Full music tracks - EVERY call returns TWO variations for one flat price and duration is FREE up to 360 seconds. The rule that decides everything - in CUSTOM mode the prompt field is the EXACT LYRICS (sung as written), in DESCRIPTION mode it is a description and the lyrics get written for you. Covers the mode matrix, style vs prompt steering, instrumental scoring, negative tags, and the character caps.
4
- ---
5
-
6
- # Suno — prompting
7
-
8
- Full music generation, reached through the sunoapi.org wrapper. Models `V4`, `V4_5`, `V4_5PLUS`, `V4_5ALL`, `V5`, `V5_5`.
9
-
10
- **Two facts that should shape every decision:**
11
-
12
- 1. **Every call returns TWO songs** — two genuinely different takes on the same brief, for one flat price. Audition both before re-rolling.
13
- 2. **Length is free.** A 360-second track costs exactly what a default one costs (measured against the live provider balance 2026-07-31: a default call and a `duration: 240` call both debited the same). There is never a reason to generate a bed shorter than your edit.
14
-
15
- ## Where it routes
16
-
17
- - **Anything a listener would call a song or a score** — theme, underscore, needle-drop, montage bed, end-card sting.
18
- - **NOT** ambience or room tone — that is `seed-audio`, which is cheaper and better at it.
19
- - **NOT** a single effect — that is `eleven-sfx`.
20
- - **AUDIO-ONLY.**
21
-
22
- ## 🚨 THE RULE: which mode you are in changes what `prompt` means
23
-
24
- | `customMode` | `instrumental` | Required | What `prompt` means |
25
- |---|---|---|---|
26
- | `false` | either | `prompt` only (≤500 chars) | **A description.** Lyrics get written for you. |
27
- | `true` | `true` | `style`, `title` | **Unused.** Style + title do all the steering. |
28
- | `true` | `false` | `style`, `title`, `prompt` | **THE EXACT LYRICS**, sung as written. |
29
-
30
- Putting a description in the prompt field while `customMode: true` and `instrumental: false` gets your description **sung back at you**. This is the single most common Suno mistake and it costs a full generation every time.
31
-
32
- ## Description mode — the fast path
33
-
34
- ```
35
- customMode: false
36
- prompt: "brooding synthwave for a night drive, analog bass, no vocals, 90 bpm"
37
- ```
38
-
39
- Use it when you need a mood and do not care about specific words. 500-character cap. This is the right default for background beds.
40
-
41
- ## Custom mode — when the words matter
42
-
43
- ```
44
- customMode: true
45
- instrumental: false
46
- style: "dream pop, hazy, reverb-heavy guitars, female vocal, 100 bpm"
47
- title: "Blue Hour"
48
- prompt: "[Verse 1]\nThe lights come on before we're ready\n..."
49
- ```
50
-
51
- Structure tags (`[Verse]`, `[Chorus]`, `[Bridge]`, `[Outro]`) inside the lyrics are how you control the arrangement. Everything that is not a structure tag will be sung.
52
-
53
- ## Instrumental scoring
54
-
55
- ```
56
- customMode: true
57
- instrumental: true
58
- style: "tense orchestral strings, low brass swells, no percussion"
59
- title: "Approach"
60
- ```
61
-
62
- `instrumental: true` is the right answer for almost every film bed — a vocal you did not ask for will fight your dialogue.
63
-
64
- ## Steer with `style`, not with adjective piles
65
-
66
- Genre + era + instrumentation + tempo belong in `style`, not stuffed into `prompt`.
67
-
68
- ```
69
- ✓ style: "90s trip-hop, dusty breakbeat, Rhodes piano, upright bass, 85 bpm"
70
- ✗ prompt: "a really cool dusty 90s trip hop song with a Rhodes and..."
71
- ```
72
-
73
- `negativeTags` removes what keeps creeping in: `"brass, EDM drop, male vocal"`.
74
-
75
- ## The steering knobs
76
-
77
- | Param | Range | Reach for it when |
78
- |---|---|---|
79
- | `duration` | 10–360s (**V5_5 + custom mode only**) | Always, when the bed must outlast the cut. It is free. |
80
- | `vocalGender` | `m` / `f` — **the wire values, not "male"/"female"** | A specific voice is required. |
81
- | `styleWeight` | 0–1 | The style field is being ignored (raise) or strangling the song (lower). |
82
- | `weirdnessConstraint` | 0–1 | Takes are too safe (raise) or falling apart (lower). |
83
- | `audioWeight` | 0–1 | Balancing an audio input against the prompt. |
84
- | `personaId` / `personaModel` | — | A series needs the same voice/sound across episodes. |
85
-
86
- ## Character caps
87
-
88
- | Field | V4 | V4_5 / V4_5PLUS / V5 / V5_5 | V4_5ALL |
89
- |---|---|---|---|
90
- | prompt (custom = literal lyrics) | 3000 | 5000 | 5000 |
91
- | prompt (non-custom = description) | 500 | 500 | 500 |
92
- | style | 200 | 1000 | 1000 |
93
- | title | 80 | 100 | 80 |
94
-
95
- ## Iterating
96
-
97
- - **Audition both returned songs first.** A re-roll costs a full generation; the second variation is already paid for.
98
- - Wrong genre → fix `style`. Wrong words → you are in the wrong mode, check the matrix above.
99
- - Something keeps appearing that you do not want → `negativeTags`, not more prompt.
100
- - Three failed generations on the same brief means the style field is too vague, not that the seed is unlucky.
101
-
102
- ## Ops notes
103
-
104
- - Tracks take 2–3 minutes; a streamable preview exists ~30–40s in. Use `background: true` and poll.
105
- - The provider hosts files for a limited window — **Slates downloads and stores them locally as soon as the track finishes**, so nothing expires out from under a project.
106
- - Suno has **no official public API**; this rides an unofficial wrapper. Treat availability as best-effort and do not build a deadline around it.
107
-
108
- ## Content notes
109
-
110
- Provider-side moderation rejects lyrics and style prompts naming real artists or protected material (`SENSITIVE_WORD_ERROR`). Describe the sound, not the artist. See slates-content-policy.