@slatesvideo/shared 0.5.4 → 0.5.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +13 -0
- package/dist/index.d.ts +1 -1
- package/dist/index.js +2 -2
- package/dist/operations/index.d.ts +84 -9
- package/dist/operations/index.js +455 -47
- package/dist/prompts/character-sheet.d.ts +11 -21
- package/dist/prompts/character-sheet.js +124 -59
- package/dist/prompts/model-facts.d.ts +1 -1
- package/dist/prompts/model-facts.js +32 -0
- package/dist/prompts/partials.generated.js +2 -2
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +196 -2
- package/dist/prompts/reference-composer.d.ts +1 -1
- package/dist/prompts/reference-composer.js +3 -4
- package/dist/prompts/reference-rules.d.ts +19 -2
- package/dist/prompts/reference-rules.js +21 -4
- package/dist/skills/content.js +14 -11
- package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
- package/{skills/slates-character-turnaround.md → exports/slates-prompt-builder/generated/reference-character.md} +26 -33
- package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
- package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
- package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +7 -3
- package/skills/_partials/reference-rules-core.md +1 -1
- package/skills/_partials/reference-tips-short.md +1 -1
- package/skills/slates-character-identity.md +105 -0
- package/skills/slates-edit-and-iterate.md +1 -1
- package/skills/slates-model-selection.md +26 -0
- package/skills/slates-one-prompt-film.md +3 -3
- package/skills/slates-prompting-elevenlabs.md +131 -0
- package/skills/slates-prompting-flux-2-max.md +1 -1
- package/skills/slates-prompting-gpt-image-2.md +1 -1
- package/skills/slates-prompting-kling-v3.md +8 -6
- package/skills/slates-prompting-nano-banana-2.md +7 -5
- package/skills/slates-prompting-omni-flash.md +1 -1
- package/skills/slates-prompting-seed-audio.md +110 -0
- package/skills/slates-prompting-seedance.md +15 -9
- package/skills/slates-prompting-suno.md +110 -0
- package/skills/slates-prompting-veo-3.md +1 -1
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-seed-audio
|
|
3
|
+
description: How to prompt Seed Audio 1.0 (ByteDance, via fal). Read before calling slates_generate_audio with model seed-audio. The one-pass audio SCENE model — dialogue, SFX and ambience together from ONE plain sentence. CRITICAL - it has NO duration parameter, so length must be named IN THE PROMPT TEXT and Slates bills the duration you request. Covers the one-sentence doctrine, the crowd-size rule, why Kling "SFX:" syntax hurts here, and the audio-refs-XOR-image input rule.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Seed Audio 1.0 — prompting
|
|
7
|
+
|
|
8
|
+
ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence. 1–120 seconds. It is the default audio model in Slates and the workhorse for continuity beds.
|
|
9
|
+
|
|
10
|
+
## Where it routes
|
|
11
|
+
|
|
12
|
+
- **Scene audio, room tone, ambience beds, crowd/nature soundscapes** — anything where several sounds share a space. One generation, not three layered ones.
|
|
13
|
+
- **Fast scratch dialogue** when the exact wording is still moving. Once the script locks and the read has to be repeatable, switch to `eleven-v3`.
|
|
14
|
+
- **NOT** a single effect that must land on a known frame — that is `eleven-sfx`, which takes an exact duration.
|
|
15
|
+
- **NOT** music — that is `suno`.
|
|
16
|
+
- **AUDIO-ONLY.** It cannot produce images or video.
|
|
17
|
+
|
|
18
|
+
## THE RULES
|
|
19
|
+
|
|
20
|
+
### 1. 🚨 There is no duration parameter — the words set the length
|
|
21
|
+
|
|
22
|
+
This is the single most important fact about this model. Output length is driven by the prompt text ("… 15 seconds"), capped at 120s.
|
|
23
|
+
|
|
24
|
+
**Slates handles this for you:** the `durationSeconds` param appends the duration to the prompt and **bills that number of seconds**. So:
|
|
25
|
+
|
|
26
|
+
- Set `durationSeconds` to what you actually want.
|
|
27
|
+
- **Do not also write a different length into your sentence.** Two numbers fight, and you pay for the one you selected, not the one you got.
|
|
28
|
+
- If the returned clip is shorter than requested you still paid for the request — that is the deal that keeps the displayed price equal to the charge. Ask for what you need.
|
|
29
|
+
|
|
30
|
+
<!-- slates-only -->
|
|
31
|
+
The server re-derives the billed key from `durationSeconds` (a client cannot under-bill), probes the returned `audio.duration` after completion, and logs `SEED AUDIO BILLING DRIFT` if the model overshot. No auto-charge, no refund — the request is the contract.
|
|
32
|
+
<!-- /slates-only -->
|
|
33
|
+
|
|
34
|
+
### 2. One plain sentence. No production jargon.
|
|
35
|
+
|
|
36
|
+
Field-proven (Higgsfield sprint, 2026-07-27/28). Working prompts look like this:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
tiny applause of 2 or 3 people at an open mic. 15 seconds
|
|
40
|
+
nature soundscape, wide open field cicadas and birds and a loon.
|
|
41
|
+
a diner at 2am, one coffee machine hissing, cutlery somewhere behind the counter
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Not this:
|
|
45
|
+
|
|
46
|
+
```
|
|
47
|
+
✗ AMBIENCE: interior diner, night. SFX: espresso machine (hiss, 2s), cutlery.
|
|
48
|
+
✗ Wide shot of a diner. Slow push in. Warm tungsten. Ambient noise: ...
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Shot language, camera moves and lighting belong to video prompts. Here they are just words the model has to ignore.
|
|
52
|
+
|
|
53
|
+
### 3. Never bring Kling's audio syntax to this model
|
|
54
|
+
|
|
55
|
+
`SFX:` and `Ambient noise:` prefixes and `Background music:` labels are **Kling 3.0 video** syntax. Seed Audio has no parser for them — it reads them as text in the scene and the output gets measurably worse. Describe the sounds directly instead.
|
|
56
|
+
|
|
57
|
+
### 4. Name the crowd size, the room size, the distance
|
|
58
|
+
|
|
59
|
+
The highest-leverage single edit on any bed. Unqualified nouns default big:
|
|
60
|
+
|
|
61
|
+
| Vague | What it returns | Fixed |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| `applause` | a full auditorium | `tiny applause of 2 or 3 people` |
|
|
64
|
+
| `traffic` | a highway | `one car passing on a wet residential street` |
|
|
65
|
+
| `crowd` | a stadium | `four people talking at the next table` |
|
|
66
|
+
|
|
67
|
+
Distance words (`far off`, `muffled through a wall`, `right next to the mic`) work the same way and are how you build depth in one sentence.
|
|
68
|
+
|
|
69
|
+
### 5. Beds must outlast the cut
|
|
70
|
+
|
|
71
|
+
Ask for a few seconds more than the clip needs so the edit has handles to fade through. A bed that ends exactly on the cut always sounds clipped. This is a product requirement, not a preference — it is why the duration control exists at all.
|
|
72
|
+
|
|
73
|
+
### 6. Dialogue goes in quotes, inside the same sentence as the room
|
|
74
|
+
|
|
75
|
+
```
|
|
76
|
+
a tired bartender says, "we closed twenty minutes ago", glasses clinking behind him
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
Pick a preset voice when a specific speaker matters. Leave `voice` unset and the scene casts itself — which is usually right for crowd and background dialogue.
|
|
80
|
+
|
|
81
|
+
Preset voices (20): `vivi_mixed_en_zh_ja_es_id`, `mindy_en_es_id_pt_zh`, `kian_en_zh`, `cedric_en_zh`, `sophie_en_zh`, `jean_en_zh`, `magnus_en_zh`, `mabel_en_zh`, `nadia_en_zh`, `opal_en_zh`, `pearl_en_zh`, `quentin_en_zh`, `corinne_mixed_en_zh`, `esther_mixed_en_zh`, `lyla_mixed_en_zh`, `tracy_es_zh`, `sandy_es_mixed_en_zh`, `felix_zh`, `celeste_zh`, `monkey_king_zh`.
|
|
82
|
+
|
|
83
|
+
Set `multilingual: true` for non-English or mixed-language lines.
|
|
84
|
+
|
|
85
|
+
### 7. Inputs: up to 3 audio clips **XOR** one image. Never both.
|
|
86
|
+
|
|
87
|
+
- **Audio references** — up to 3 clips, each ≤30s and ≤10MB (wav/mp3/pcm/ogg_opus). Refer to them in the prompt as `@Audio1`, `@Audio2`, `@Audio3`: *"match the room tone of @Audio1"*.
|
|
88
|
+
- **Image reference** — one image (jpeg/png/webp ≤10MB). The model scores what it sees.
|
|
89
|
+
- Sending both is rejected by the API. Pick the one that carries the intent.
|
|
90
|
+
|
|
91
|
+
### 8. The knobs, and when to touch them
|
|
92
|
+
|
|
93
|
+
| Param | Range | Reach for it when |
|
|
94
|
+
|---|---|---|
|
|
95
|
+
| `speed` | 0.5–2.0 | Dialogue is racing or dragging against picture. |
|
|
96
|
+
| `volume` | 0.5–2.0 | Rarely — normalize on the timeline instead. |
|
|
97
|
+
| `pitch` | −12…+12 semitones | Ageing or shifting a voice. Small moves only; ±3 is already a lot. |
|
|
98
|
+
| `multilingual` | bool | Non-English or code-switched lines. |
|
|
99
|
+
| `sampleRate` | 8k–48k | Leave at 24000 unless you are matching an existing stem. |
|
|
100
|
+
| `outputFormat` | mp3 / wav / pcm / ogg_opus | wav when this is going into a mix; mp3 otherwise. |
|
|
101
|
+
|
|
102
|
+
## Iterating
|
|
103
|
+
|
|
104
|
+
- A bed that came back wrong is almost always a **scale** problem (crowd/room too big) or a **jargon** problem (the sentence reads like a spec). Fix those two before touching `speed`/`pitch`.
|
|
105
|
+
- Three failed takes on the same sentence means the sentence is wrong, not the seed. Rewrite it the way you would say it out loud.
|
|
106
|
+
- Generations are cheap enough at short durations that auditioning two phrasings beats agonizing over one.
|
|
107
|
+
|
|
108
|
+
## Content notes
|
|
109
|
+
|
|
110
|
+
Provider-side moderation applies to voices and to recognizable real people. See slates-content-policy.
|
|
@@ -55,7 +55,7 @@ Seedance classifies your request from the phrasing. Use the pattern that matches
|
|
|
55
55
|
|
|
56
56
|
> *"For edit / extend video tasks, directly use `<Video_N>` to refer to the video. **Do not use "reference `<Video_N>`"**, to avoid being incorrectly identified as a reference task."*
|
|
57
57
|
|
|
58
|
-
This is easy to trip: Slates has an edit lane (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)
|
|
58
|
+
This is easy to trip: Slates has an edit lane<!-- slates-only --> (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)<!-- /slates-only -->. Writing *"reference video 1 and change the jacket to red"* gets classified as a **reference** task — the model generates a brand-new clip inspired by the source instead of editing it. Write *"Strictly edit video 1, and modify the blue jacket to red."*
|
|
59
59
|
|
|
60
60
|
## Shot structure — "Shot 1 / Shot 2 / Shot 3" `[official :1563-1598]`
|
|
61
61
|
|
|
@@ -91,7 +91,7 @@ Every time a subject appears, it must be **explicitly referred to**. Two support
|
|
|
91
91
|
|
|
92
92
|
Also official: keep descriptions concise, avoid redundancy, avoid semantic conflicts (contradictory traits for one subject), and prefer expressing spatial relationships through reference images rather than dense text. `[:1550-1556]`
|
|
93
93
|
|
|
94
|
-
**`[slates]`** — the app composes this for you. `composeReferences()` cites each reference inline as `Name (image N)`
|
|
94
|
+
**`[slates]`** — the app composes this for you. `composeReferences()` cites each canonical character or environment reference inline as `Name (image N)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*"Zhang San (corresponding to image 1)"* `[:1976]`). You never hand-write role labels or index numbers.
|
|
95
95
|
|
|
96
96
|
## Action description `[official :1602-1621]`
|
|
97
97
|
|
|
@@ -153,7 +153,7 @@ Seedance has **no `negativePrompt` field** — constraints go inline in this slo
|
|
|
153
153
|
3. **Optimize reference assets** `[:1988]` — *"For character reference images, prioritize independent single-person photos. Three-view or multi-view assets are not recommended."*
|
|
154
154
|
4. **Simplify the prompt** — do not paste a whole script; redundant copy confuses the model.
|
|
155
155
|
|
|
156
|
-
**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on
|
|
156
|
+
**Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on identity sheets. Practical rule for Slates:
|
|
157
157
|
|
|
158
158
|
- **Multi-character Seedance shot** → bind every character to its image, append the anti-twin constraint, and prefer single-person / dominant-portrait references over multi-view sheets.
|
|
159
159
|
- **Single-character shot** → the standard character-sheet flow is fine.
|
|
@@ -206,13 +206,16 @@ Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 aud
|
|
|
206
206
|
|
|
207
207
|
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
208
208
|
|
|
209
|
-
These aren't separate Seedance features — they're prompting strategies over reference media
|
|
209
|
+
These aren't separate Seedance features — they're prompting strategies over reference media.<!-- slates-only --> The Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:<!-- /slates-only -->
|
|
210
210
|
|
|
211
|
-
- **Motion transfer:** subject image as a reference + the driving clip via `videoReferenceAssetId
|
|
212
|
-
- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with
|
|
211
|
+
- **Motion transfer:** subject image as a reference + the driving clip<!-- slates-only --> via `videoReferenceAssetId`<!-- /slates-only --> (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`
|
|
212
|
+
- **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`
|
|
213
213
|
- **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script → their likeness AND voice deliver the new line.
|
|
214
|
+
<!-- slates-only -->
|
|
214
215
|
- **Billing:** a reference VIDEO switches the cost key to `seedance-2*-vref-{res}-{T}s` where T = clip seconds + output seconds — quote before confirming. Audio references are free (audio is included on every route).
|
|
216
|
+
<!-- /slates-only -->
|
|
215
217
|
|
|
218
|
+
<!-- slates-only -->
|
|
216
219
|
## Faces — set `seedanceFace` for AI-character faces
|
|
217
220
|
|
|
218
221
|
Seedance routes through **three tiers** depending on the face in the reference, exposed as the "Face in Reference" toggle plus the real-face params on `slates_generate_video`:
|
|
@@ -223,8 +226,9 @@ Seedance routes through **three tiers** depending on the face in the reference,
|
|
|
223
226
|
|
|
224
227
|
Rules:
|
|
225
228
|
- **The real-vs-AI call is the PROVIDER'S, not yours.** ByteDance's classifier is probabilistic — some real photos pass the standard face route (billed at the cheap rate; fine), others get rejected with `[REAL_FACE_DETECTED]` (auto-refunded). Don't preemptively route to the real-face tier just because a photo looks real; try `seedanceFace: true` first and escalate only on the marked rejection. Public figures / celebrities fail on every route.
|
|
226
|
-
- It's about the **reference, not the output.** If your character
|
|
229
|
+
- It's about the **reference, not the output.** If your character identity or generated portrait shows a face, turn it on. A product shot with no person stays off.
|
|
227
230
|
- Don't toggle it on "just in case" — a faceless gen on the face route burns ~45% extra for nothing.
|
|
231
|
+
<!-- /slates-only -->
|
|
228
232
|
|
|
229
233
|
## Reference rules (the verified ones)
|
|
230
234
|
|
|
@@ -247,7 +251,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
247
251
|
|
|
248
252
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
249
253
|
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
250
|
-
3. **One identity sheet per character
|
|
254
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
251
255
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
252
256
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
253
257
|
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
@@ -259,12 +263,13 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
259
263
|
|
|
260
264
|
### For Seedance specifically
|
|
261
265
|
|
|
262
|
-
- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say "still / scene / from a movie / from the image." The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one — see slates-cost-discipline).
|
|
266
|
+
- **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say "still / scene / from a movie / from the image." The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one<!-- slates-only --> — see slates-cost-discipline<!-- /slates-only -->).
|
|
263
267
|
- **Seedance's own idiom for rule 2 is `Reference <Subject_N> in <Image_N>`** `[official :1389]` — `Image_N` indexes the order the refs are attached, so the name plus the index carries the role. The full binding grammar is in Part 1 (Subject binding).
|
|
264
268
|
- **Rule 3 has an official ceiling here.** The trend is MORE references (video and audio into Seedance), all addressed by name — but for **multi-character frames** see the twin-problem section above: bind every character to its image, append the anti-twin constraint, and prefer single-person references. Past 4 reference people, stability drops `[official :2048-2052]`.
|
|
265
269
|
- **Rule 8 holds even though Seedance can render common text natively** `[official :1758]`. A baked NB2 start frame is still the reliable route for text that must be legible.
|
|
266
270
|
- **Rule 5 pairs with the first/last-frame exclusion** — frames and reference images are mutually exclusive on this model (see Reference media above), so an environment you must match exactly costs you the frame lane.
|
|
267
271
|
|
|
272
|
+
<!-- slates-only -->
|
|
268
273
|
## Pre-flight: references arrive inline, refer by code
|
|
269
274
|
|
|
270
275
|
When you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** — if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.
|
|
@@ -273,6 +278,7 @@ When talking to the user about the gen, refer to each reference by its short cod
|
|
|
273
278
|
|
|
274
279
|
- ✅ "I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame — the camera move is going to be a slow dolly forward through the gap."
|
|
275
280
|
- ❌ "I'm using the first beach image and the last one..." (which? They have four.)
|
|
281
|
+
<!-- /slates-only -->
|
|
276
282
|
|
|
277
283
|
---
|
|
278
284
|
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-suno
|
|
3
|
+
description: How to prompt Suno in Slates. Read before calling slates_generate_audio with model suno. Full music tracks - EVERY call returns TWO variations for one flat price and duration is FREE up to 360 seconds. The rule that decides everything - in CUSTOM mode the prompt field is the EXACT LYRICS (sung as written), in DESCRIPTION mode it is a description and the lyrics get written for you. Covers the mode matrix, style vs prompt steering, instrumental scoring, negative tags, and the character caps.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Suno — prompting
|
|
7
|
+
|
|
8
|
+
Full music generation, reached through the sunoapi.org wrapper. Models `V4`, `V4_5`, `V4_5PLUS`, `V4_5ALL`, `V5`, `V5_5`.
|
|
9
|
+
|
|
10
|
+
**Two facts that should shape every decision:**
|
|
11
|
+
|
|
12
|
+
1. **Every call returns TWO songs** — two genuinely different takes on the same brief, for one flat price. Audition both before re-rolling.
|
|
13
|
+
2. **Length is free.** A 360-second track costs exactly what a default one costs (measured against the live provider balance 2026-07-31: a default call and a `duration: 240` call both debited the same). There is never a reason to generate a bed shorter than your edit.
|
|
14
|
+
|
|
15
|
+
## Where it routes
|
|
16
|
+
|
|
17
|
+
- **Anything a listener would call a song or a score** — theme, underscore, needle-drop, montage bed, end-card sting.
|
|
18
|
+
- **NOT** ambience or room tone — that is `seed-audio`, which is cheaper and better at it.
|
|
19
|
+
- **NOT** a single effect — that is `eleven-sfx`.
|
|
20
|
+
- **AUDIO-ONLY.**
|
|
21
|
+
|
|
22
|
+
## 🚨 THE RULE: which mode you are in changes what `prompt` means
|
|
23
|
+
|
|
24
|
+
| `customMode` | `instrumental` | Required | What `prompt` means |
|
|
25
|
+
|---|---|---|---|
|
|
26
|
+
| `false` | either | `prompt` only (≤500 chars) | **A description.** Lyrics get written for you. |
|
|
27
|
+
| `true` | `true` | `style`, `title` | **Unused.** Style + title do all the steering. |
|
|
28
|
+
| `true` | `false` | `style`, `title`, `prompt` | **THE EXACT LYRICS**, sung as written. |
|
|
29
|
+
|
|
30
|
+
Putting a description in the prompt field while `customMode: true` and `instrumental: false` gets your description **sung back at you**. This is the single most common Suno mistake and it costs a full generation every time.
|
|
31
|
+
|
|
32
|
+
## Description mode — the fast path
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
customMode: false
|
|
36
|
+
prompt: "brooding synthwave for a night drive, analog bass, no vocals, 90 bpm"
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Use it when you need a mood and do not care about specific words. 500-character cap. This is the right default for background beds.
|
|
40
|
+
|
|
41
|
+
## Custom mode — when the words matter
|
|
42
|
+
|
|
43
|
+
```
|
|
44
|
+
customMode: true
|
|
45
|
+
instrumental: false
|
|
46
|
+
style: "dream pop, hazy, reverb-heavy guitars, female vocal, 100 bpm"
|
|
47
|
+
title: "Blue Hour"
|
|
48
|
+
prompt: "[Verse 1]\nThe lights come on before we're ready\n..."
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Structure tags (`[Verse]`, `[Chorus]`, `[Bridge]`, `[Outro]`) inside the lyrics are how you control the arrangement. Everything that is not a structure tag will be sung.
|
|
52
|
+
|
|
53
|
+
## Instrumental scoring
|
|
54
|
+
|
|
55
|
+
```
|
|
56
|
+
customMode: true
|
|
57
|
+
instrumental: true
|
|
58
|
+
style: "tense orchestral strings, low brass swells, no percussion"
|
|
59
|
+
title: "Approach"
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
`instrumental: true` is the right answer for almost every film bed — a vocal you did not ask for will fight your dialogue.
|
|
63
|
+
|
|
64
|
+
## Steer with `style`, not with adjective piles
|
|
65
|
+
|
|
66
|
+
Genre + era + instrumentation + tempo belong in `style`, not stuffed into `prompt`.
|
|
67
|
+
|
|
68
|
+
```
|
|
69
|
+
✓ style: "90s trip-hop, dusty breakbeat, Rhodes piano, upright bass, 85 bpm"
|
|
70
|
+
✗ prompt: "a really cool dusty 90s trip hop song with a Rhodes and..."
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
`negativeTags` removes what keeps creeping in: `"brass, EDM drop, male vocal"`.
|
|
74
|
+
|
|
75
|
+
## The steering knobs
|
|
76
|
+
|
|
77
|
+
| Param | Range | Reach for it when |
|
|
78
|
+
|---|---|---|
|
|
79
|
+
| `duration` | 10–360s (**V5_5 + custom mode only**) | Always, when the bed must outlast the cut. It is free. |
|
|
80
|
+
| `vocalGender` | `m` / `f` — **the wire values, not "male"/"female"** | A specific voice is required. |
|
|
81
|
+
| `styleWeight` | 0–1 | The style field is being ignored (raise) or strangling the song (lower). |
|
|
82
|
+
| `weirdnessConstraint` | 0–1 | Takes are too safe (raise) or falling apart (lower). |
|
|
83
|
+
| `audioWeight` | 0–1 | Balancing an audio input against the prompt. |
|
|
84
|
+
| `personaId` / `personaModel` | — | A series needs the same voice/sound across episodes. |
|
|
85
|
+
|
|
86
|
+
## Character caps
|
|
87
|
+
|
|
88
|
+
| Field | V4 | V4_5 / V4_5PLUS / V5 / V5_5 | V4_5ALL |
|
|
89
|
+
|---|---|---|---|
|
|
90
|
+
| prompt (custom = literal lyrics) | 3000 | 5000 | 5000 |
|
|
91
|
+
| prompt (non-custom = description) | 500 | 500 | 500 |
|
|
92
|
+
| style | 200 | 1000 | 1000 |
|
|
93
|
+
| title | 80 | 100 | 80 |
|
|
94
|
+
|
|
95
|
+
## Iterating
|
|
96
|
+
|
|
97
|
+
- **Audition both returned songs first.** A re-roll costs a full generation; the second variation is already paid for.
|
|
98
|
+
- Wrong genre → fix `style`. Wrong words → you are in the wrong mode, check the matrix above.
|
|
99
|
+
- Something keeps appearing that you do not want → `negativeTags`, not more prompt.
|
|
100
|
+
- Three failed generations on the same brief means the style field is too vague, not that the seed is unlucky.
|
|
101
|
+
|
|
102
|
+
## Ops notes
|
|
103
|
+
|
|
104
|
+
- Tracks take 2–3 minutes; a streamable preview exists ~30–40s in. Use `background: true` and poll.
|
|
105
|
+
- The provider hosts files for a limited window — **Slates downloads and stores them locally as soon as the track finishes**, so nothing expires out from under a project.
|
|
106
|
+
- Suno has **no official public API**; this rides an unofficial wrapper. Treat availability as best-effort and do not build a deadline around it.
|
|
107
|
+
|
|
108
|
+
## Content notes
|
|
109
|
+
|
|
110
|
+
Provider-side moderation rejects lyrics and style prompts naming real artists or protected material (`SENSITIVE_WORD_ERROR`). Describe the sound, not the artist. See slates-content-policy.
|
|
@@ -117,7 +117,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
117
117
|
|
|
118
118
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
119
119
|
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
120
|
-
3. **One identity sheet per character
|
|
120
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
121
121
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
122
122
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
123
123
|
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|