@slatesvideo/shared 0.5.4 → 0.5.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/README.md +13 -0
  2. package/dist/index.d.ts +1 -1
  3. package/dist/index.js +2 -2
  4. package/dist/operations/index.d.ts +84 -9
  5. package/dist/operations/index.js +455 -47
  6. package/dist/prompts/character-sheet.d.ts +11 -21
  7. package/dist/prompts/character-sheet.js +124 -59
  8. package/dist/prompts/model-facts.d.ts +1 -1
  9. package/dist/prompts/model-facts.js +32 -0
  10. package/dist/prompts/partials.generated.js +2 -2
  11. package/dist/prompts/prompting-tips.d.ts +1 -1
  12. package/dist/prompts/prompting-tips.js +196 -2
  13. package/dist/prompts/reference-composer.d.ts +1 -1
  14. package/dist/prompts/reference-composer.js +3 -4
  15. package/dist/prompts/reference-rules.d.ts +19 -2
  16. package/dist/prompts/reference-rules.js +21 -4
  17. package/dist/skills/content.js +14 -11
  18. package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
  19. package/{skills/slates-character-turnaround.md → exports/slates-prompt-builder/generated/reference-character.md} +26 -33
  20. package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
  21. package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
  22. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
  23. package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
  24. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
  25. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  26. package/package.json +7 -3
  27. package/skills/_partials/reference-rules-core.md +1 -1
  28. package/skills/_partials/reference-tips-short.md +1 -1
  29. package/skills/slates-character-identity.md +105 -0
  30. package/skills/slates-edit-and-iterate.md +1 -1
  31. package/skills/slates-model-selection.md +26 -0
  32. package/skills/slates-one-prompt-film.md +3 -3
  33. package/skills/slates-prompting-elevenlabs.md +131 -0
  34. package/skills/slates-prompting-flux-2-max.md +1 -1
  35. package/skills/slates-prompting-gpt-image-2.md +1 -1
  36. package/skills/slates-prompting-kling-v3.md +8 -6
  37. package/skills/slates-prompting-nano-banana-2.md +7 -5
  38. package/skills/slates-prompting-omni-flash.md +1 -1
  39. package/skills/slates-prompting-seed-audio.md +110 -0
  40. package/skills/slates-prompting-seedance.md +15 -9
  41. package/skills/slates-prompting-suno.md +110 -0
  42. package/skills/slates-prompting-veo-3.md +1 -1
@@ -0,0 +1,110 @@
1
+ ---
2
+ name: slates-prompting-seed-audio
3
+ description: How to prompt Seed Audio 1.0 (ByteDance, via fal). Read before calling slates_generate_audio with model seed-audio. The one-pass audio SCENE model — dialogue, SFX and ambience together from ONE plain sentence. CRITICAL - it has NO duration parameter, so length must be named IN THE PROMPT TEXT and Slates bills the duration you request. Covers the one-sentence doctrine, the crowd-size rule, why Kling "SFX:" syntax hurts here, and the audio-refs-XOR-image input rule.
4
+ ---
5
+
6
+ # Seed Audio 1.0 — prompting
7
+
8
+ ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence. 1–120 seconds. It is the default audio model in Slates and the workhorse for continuity beds.
9
+
10
+ ## Where it routes
11
+
12
+ - **Scene audio, room tone, ambience beds, crowd/nature soundscapes** — anything where several sounds share a space. One generation, not three layered ones.
13
+ - **Fast scratch dialogue** when the exact wording is still moving. Once the script locks and the read has to be repeatable, switch to `eleven-v3`.
14
+ - **NOT** a single effect that must land on a known frame — that is `eleven-sfx`, which takes an exact duration.
15
+ - **NOT** music — that is `suno`.
16
+ - **AUDIO-ONLY.** It cannot produce images or video.
17
+
18
+ ## THE RULES
19
+
20
+ ### 1. 🚨 There is no duration parameter — the words set the length
21
+
22
+ This is the single most important fact about this model. Output length is driven by the prompt text ("… 15 seconds"), capped at 120s.
23
+
24
+ **Slates handles this for you:** the `durationSeconds` param appends the duration to the prompt and **bills that number of seconds**. So:
25
+
26
+ - Set `durationSeconds` to what you actually want.
27
+ - **Do not also write a different length into your sentence.** Two numbers fight, and you pay for the one you selected, not the one you got.
28
+ - If the returned clip is shorter than requested you still paid for the request — that is the deal that keeps the displayed price equal to the charge. Ask for what you need.
29
+
30
+ <!-- slates-only -->
31
+ The server re-derives the billed key from `durationSeconds` (a client cannot under-bill), probes the returned `audio.duration` after completion, and logs `SEED AUDIO BILLING DRIFT` if the model overshot. No auto-charge, no refund — the request is the contract.
32
+ <!-- /slates-only -->
33
+
34
+ ### 2. One plain sentence. No production jargon.
35
+
36
+ Field-proven (Higgsfield sprint, 2026-07-27/28). Working prompts look like this:
37
+
38
+ ```
39
+ tiny applause of 2 or 3 people at an open mic. 15 seconds
40
+ nature soundscape, wide open field cicadas and birds and a loon.
41
+ a diner at 2am, one coffee machine hissing, cutlery somewhere behind the counter
42
+ ```
43
+
44
+ Not this:
45
+
46
+ ```
47
+ ✗ AMBIENCE: interior diner, night. SFX: espresso machine (hiss, 2s), cutlery.
48
+ ✗ Wide shot of a diner. Slow push in. Warm tungsten. Ambient noise: ...
49
+ ```
50
+
51
+ Shot language, camera moves and lighting belong to video prompts. Here they are just words the model has to ignore.
52
+
53
+ ### 3. Never bring Kling's audio syntax to this model
54
+
55
+ `SFX:` and `Ambient noise:` prefixes and `Background music:` labels are **Kling 3.0 video** syntax. Seed Audio has no parser for them — it reads them as text in the scene and the output gets measurably worse. Describe the sounds directly instead.
56
+
57
+ ### 4. Name the crowd size, the room size, the distance
58
+
59
+ The highest-leverage single edit on any bed. Unqualified nouns default big:
60
+
61
+ | Vague | What it returns | Fixed |
62
+ |---|---|---|
63
+ | `applause` | a full auditorium | `tiny applause of 2 or 3 people` |
64
+ | `traffic` | a highway | `one car passing on a wet residential street` |
65
+ | `crowd` | a stadium | `four people talking at the next table` |
66
+
67
+ Distance words (`far off`, `muffled through a wall`, `right next to the mic`) work the same way and are how you build depth in one sentence.
68
+
69
+ ### 5. Beds must outlast the cut
70
+
71
+ Ask for a few seconds more than the clip needs so the edit has handles to fade through. A bed that ends exactly on the cut always sounds clipped. This is a product requirement, not a preference — it is why the duration control exists at all.
72
+
73
+ ### 6. Dialogue goes in quotes, inside the same sentence as the room
74
+
75
+ ```
76
+ a tired bartender says, "we closed twenty minutes ago", glasses clinking behind him
77
+ ```
78
+
79
+ Pick a preset voice when a specific speaker matters. Leave `voice` unset and the scene casts itself — which is usually right for crowd and background dialogue.
80
+
81
+ Preset voices (20): `vivi_mixed_en_zh_ja_es_id`, `mindy_en_es_id_pt_zh`, `kian_en_zh`, `cedric_en_zh`, `sophie_en_zh`, `jean_en_zh`, `magnus_en_zh`, `mabel_en_zh`, `nadia_en_zh`, `opal_en_zh`, `pearl_en_zh`, `quentin_en_zh`, `corinne_mixed_en_zh`, `esther_mixed_en_zh`, `lyla_mixed_en_zh`, `tracy_es_zh`, `sandy_es_mixed_en_zh`, `felix_zh`, `celeste_zh`, `monkey_king_zh`.
82
+
83
+ Set `multilingual: true` for non-English or mixed-language lines.
84
+
85
+ ### 7. Inputs: up to 3 audio clips **XOR** one image. Never both.
86
+
87
+ - **Audio references** — up to 3 clips, each ≤30s and ≤10MB (wav/mp3/pcm/ogg_opus). Refer to them in the prompt as `@Audio1`, `@Audio2`, `@Audio3`: *"match the room tone of @Audio1"*.
88
+ - **Image reference** — one image (jpeg/png/webp ≤10MB). The model scores what it sees.
89
+ - Sending both is rejected by the API. Pick the one that carries the intent.
90
+
91
+ ### 8. The knobs, and when to touch them
92
+
93
+ | Param | Range | Reach for it when |
94
+ |---|---|---|
95
+ | `speed` | 0.5–2.0 | Dialogue is racing or dragging against picture. |
96
+ | `volume` | 0.5–2.0 | Rarely — normalize on the timeline instead. |
97
+ | `pitch` | −12…+12 semitones | Ageing or shifting a voice. Small moves only; ±3 is already a lot. |
98
+ | `multilingual` | bool | Non-English or code-switched lines. |
99
+ | `sampleRate` | 8k–48k | Leave at 24000 unless you are matching an existing stem. |
100
+ | `outputFormat` | mp3 / wav / pcm / ogg_opus | wav when this is going into a mix; mp3 otherwise. |
101
+
102
+ ## Iterating
103
+
104
+ - A bed that came back wrong is almost always a **scale** problem (crowd/room too big) or a **jargon** problem (the sentence reads like a spec). Fix those two before touching `speed`/`pitch`.
105
+ - Three failed takes on the same sentence means the sentence is wrong, not the seed. Rewrite it the way you would say it out loud.
106
+ - Generations are cheap enough at short durations that auditioning two phrasings beats agonizing over one.
107
+
108
+ ## Content notes
109
+
110
+ Provider-side moderation applies to voices and to recognizable real people. See slates-content-policy.
@@ -55,7 +55,7 @@ Seedance classifies your request from the phrasing. Use the pattern that matches
55
55
 
56
56
  > *"For edit / extend video tasks, directly use `<Video_N>` to refer to the video. **Do not use "reference `<Video_N>`"**, to avoid being incorrectly identified as a reference task."*
57
57
 
58
- This is easy to trip: Slates has an edit lane (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes). Writing *"reference video 1 and change the jacket to red"* gets classified as a **reference** task — the model generates a brand-new clip inspired by the source instead of editing it. Write *"Strictly edit video 1, and modify the blue jacket to red."*
58
+ This is easy to trip: Slates has an edit lane<!-- slates-only --> (`slates_generate_video` with `videoReferenceAssetId`, plus the Seedance edit/relocate routes)<!-- /slates-only -->. Writing *"reference video 1 and change the jacket to red"* gets classified as a **reference** task — the model generates a brand-new clip inspired by the source instead of editing it. Write *"Strictly edit video 1, and modify the blue jacket to red."*
59
59
 
60
60
  ## Shot structure — "Shot 1 / Shot 2 / Shot 3" `[official :1563-1598]`
61
61
 
@@ -91,7 +91,7 @@ Every time a subject appears, it must be **explicitly referred to**. Two support
91
91
 
92
92
  Also official: keep descriptions concise, avoid redundancy, avoid semantic conflicts (contradictory traits for one subject), and prefer expressing spatial relationships through reference images rather than dense text. `[:1550-1556]`
93
93
 
94
- **`[slates]`** — the app composes this for you. `composeReferences()` cites each reference inline as `Name (image N)` / `Name (images 1 and 2)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*"Zhang San (corresponding to image 1)"* `[:1976]`). You never hand-write role labels or index numbers.
94
+ **`[slates]`** — the app composes this for you. `composeReferences()` cites each canonical character or environment reference inline as `Name (image N)` in the exact order it sends them, which is ByteDance's own duplicate-character format (*"Zhang San (corresponding to image 1)"* `[:1976]`). You never hand-write role labels or index numbers.
95
95
 
96
96
  ## Action description `[official :1602-1621]`
97
97
 
@@ -153,7 +153,7 @@ Seedance has **no `negativePrompt` field** — constraints go inline in this slo
153
153
  3. **Optimize reference assets** `[:1988]` — *"For character reference images, prioritize independent single-person photos. Three-view or multi-view assets are not recommended."*
154
154
  4. **Simplify the prompt** — do not paste a whole script; redundant copy confuses the model.
155
155
 
156
- **Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on turnaround sheets. Practical rule for Slates:
156
+ **Scope this honestly.** This is troubleshooting for the twin problem in **multi-character frames**, not a blanket verdict on identity sheets. Practical rule for Slates:
157
157
 
158
158
  - **Multi-character Seedance shot** → bind every character to its image, append the anti-twin constraint, and prefer single-person / dominant-portrait references over multi-view sheets.
159
159
  - **Single-character shot** → the standard character-sheet flow is fine.
@@ -206,13 +206,16 @@ Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 aud
206
206
 
207
207
  ### Motion transfer & lip-sync recipes (reference video / audio)
208
208
 
209
- These aren't separate Seedance features — they're prompting strategies over reference media, and the Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:
209
+ These aren't separate Seedance features — they're prompting strategies over reference media.<!-- slates-only --> The Slates tools (`slates_generate_motion_transfer` / `slates_generate_lip_sync` with the seedance engine) compose them for you. When driving them by hand through `slates_generate_video`:<!-- /slates-only -->
210
210
 
211
- - **Motion transfer:** subject image as a reference + the driving clip via `videoReferenceAssetId` (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`
212
- - **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with `generate_audio` on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (`audioReferenceAssetId`, ≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`
211
+ - **Motion transfer:** subject image as a reference + the driving clip<!-- slates-only --> via `videoReferenceAssetId`<!-- /slates-only --> (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`
212
+ - **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`
213
213
  - **Voice + face from one clip (the talking-head recipe):** ONE unedited 2–15s clip of the person speaking (clear voice, no music, no cuts) as the video reference + prompt with the new script → their likeness AND voice deliver the new line.
214
+ <!-- slates-only -->
214
215
  - **Billing:** a reference VIDEO switches the cost key to `seedance-2*-vref-{res}-{T}s` where T = clip seconds + output seconds — quote before confirming. Audio references are free (audio is included on every route).
216
+ <!-- /slates-only -->
215
217
 
218
+ <!-- slates-only -->
216
219
  ## Faces — set `seedanceFace` for AI-character faces
217
220
 
218
221
  Seedance routes through **three tiers** depending on the face in the reference, exposed as the "Face in Reference" toggle plus the real-face params on `slates_generate_video`:
@@ -223,8 +226,9 @@ Seedance routes through **three tiers** depending on the face in the reference,
223
226
 
224
227
  Rules:
225
228
  - **The real-vs-AI call is the PROVIDER'S, not yours.** ByteDance's classifier is probabilistic — some real photos pass the standard face route (billed at the cheap rate; fine), others get rejected with `[REAL_FACE_DETECTED]` (auto-refunded). Don't preemptively route to the real-face tier just because a photo looks real; try `seedanceFace: true` first and escalate only on the marked rejection. Public figures / celebrities fail on every route.
226
- - It's about the **reference, not the output.** If your character refs (turnaround, expression sheet, a generated portrait) show a face, turn it on. A product shot with no person stays off.
229
+ - It's about the **reference, not the output.** If your character identity or generated portrait shows a face, turn it on. A product shot with no person stays off.
227
230
  - Don't toggle it on "just in case" — a faceless gen on the face route burns ~45% extra for nothing.
231
+ <!-- /slates-only -->
228
232
 
229
233
  ## Reference rules (the verified ones)
230
234
 
@@ -247,7 +251,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
247
251
 
248
252
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
249
253
  2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
250
- 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
254
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
251
255
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
252
256
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
253
257
  6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
@@ -259,12 +263,13 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
259
263
 
260
264
  ### For Seedance specifically
261
265
 
262
- - **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say "still / scene / from a movie / from the image." The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one — see slates-cost-discipline).
266
+ - **Describe the ACTION, never the reference's content.** With refs attached, prompt only what is *happening* — motion, change, camera. Never re-describe what's in the reference, and never say "still / scene / from a movie / from the image." The model already sees the refs; narrating them wastes tokens and induces drift. Injection is stochastic — if a roll misses, **re-roll, don't re-engineer** (and a slow gen is not a failed one<!-- slates-only --> — see slates-cost-discipline<!-- /slates-only -->).
263
267
  - **Seedance's own idiom for rule 2 is `Reference <Subject_N> in <Image_N>`** `[official :1389]` — `Image_N` indexes the order the refs are attached, so the name plus the index carries the role. The full binding grammar is in Part 1 (Subject binding).
264
268
  - **Rule 3 has an official ceiling here.** The trend is MORE references (video and audio into Seedance), all addressed by name — but for **multi-character frames** see the twin-problem section above: bind every character to its image, append the anti-twin constraint, and prefer single-person references. Past 4 reference people, stability drops `[official :2048-2052]`.
265
269
  - **Rule 8 holds even though Seedance can render common text natively** `[official :1758]`. A baked NB2 start frame is still the reliable route for text that must be legible.
266
270
  - **Rule 5 pairs with the first/last-frame exclusion** — frames and reference images are mutually exclusive on this model (see Reference media above), so an environment you must match exactly costs you the frame lane.
267
271
 
272
+ <!-- slates-only -->
268
273
  ## Pre-flight: references arrive inline, refer by code
269
274
 
270
275
  When you call `slates_generate_video` with reference asset IDs (firstFrameAssetId, lastFrameAssetId, ingredientAssetIds), the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. **Look at the references** — if they suggest a different framing, lighting, or motion than your current prompt captures, revise the prompt before re-calling with `confirm=true`.
@@ -273,6 +278,7 @@ When talking to the user about the gen, refer to each reference by its short cod
273
278
 
274
279
  - ✅ "I'm using **IMG-A12** as the first frame and **IMG-A15** as the last frame — the camera move is going to be a slow dolly forward through the gap."
275
280
  - ❌ "I'm using the first beach image and the last one..." (which? They have four.)
281
+ <!-- /slates-only -->
276
282
 
277
283
  ---
278
284
 
@@ -0,0 +1,110 @@
1
+ ---
2
+ name: slates-prompting-suno
3
+ description: How to prompt Suno in Slates. Read before calling slates_generate_audio with model suno. Full music tracks - EVERY call returns TWO variations for one flat price and duration is FREE up to 360 seconds. The rule that decides everything - in CUSTOM mode the prompt field is the EXACT LYRICS (sung as written), in DESCRIPTION mode it is a description and the lyrics get written for you. Covers the mode matrix, style vs prompt steering, instrumental scoring, negative tags, and the character caps.
4
+ ---
5
+
6
+ # Suno — prompting
7
+
8
+ Full music generation, reached through the sunoapi.org wrapper. Models `V4`, `V4_5`, `V4_5PLUS`, `V4_5ALL`, `V5`, `V5_5`.
9
+
10
+ **Two facts that should shape every decision:**
11
+
12
+ 1. **Every call returns TWO songs** — two genuinely different takes on the same brief, for one flat price. Audition both before re-rolling.
13
+ 2. **Length is free.** A 360-second track costs exactly what a default one costs (measured against the live provider balance 2026-07-31: a default call and a `duration: 240` call both debited the same). There is never a reason to generate a bed shorter than your edit.
14
+
15
+ ## Where it routes
16
+
17
+ - **Anything a listener would call a song or a score** — theme, underscore, needle-drop, montage bed, end-card sting.
18
+ - **NOT** ambience or room tone — that is `seed-audio`, which is cheaper and better at it.
19
+ - **NOT** a single effect — that is `eleven-sfx`.
20
+ - **AUDIO-ONLY.**
21
+
22
+ ## 🚨 THE RULE: which mode you are in changes what `prompt` means
23
+
24
+ | `customMode` | `instrumental` | Required | What `prompt` means |
25
+ |---|---|---|---|
26
+ | `false` | either | `prompt` only (≤500 chars) | **A description.** Lyrics get written for you. |
27
+ | `true` | `true` | `style`, `title` | **Unused.** Style + title do all the steering. |
28
+ | `true` | `false` | `style`, `title`, `prompt` | **THE EXACT LYRICS**, sung as written. |
29
+
30
+ Putting a description in the prompt field while `customMode: true` and `instrumental: false` gets your description **sung back at you**. This is the single most common Suno mistake and it costs a full generation every time.
31
+
32
+ ## Description mode — the fast path
33
+
34
+ ```
35
+ customMode: false
36
+ prompt: "brooding synthwave for a night drive, analog bass, no vocals, 90 bpm"
37
+ ```
38
+
39
+ Use it when you need a mood and do not care about specific words. 500-character cap. This is the right default for background beds.
40
+
41
+ ## Custom mode — when the words matter
42
+
43
+ ```
44
+ customMode: true
45
+ instrumental: false
46
+ style: "dream pop, hazy, reverb-heavy guitars, female vocal, 100 bpm"
47
+ title: "Blue Hour"
48
+ prompt: "[Verse 1]\nThe lights come on before we're ready\n..."
49
+ ```
50
+
51
+ Structure tags (`[Verse]`, `[Chorus]`, `[Bridge]`, `[Outro]`) inside the lyrics are how you control the arrangement. Everything that is not a structure tag will be sung.
52
+
53
+ ## Instrumental scoring
54
+
55
+ ```
56
+ customMode: true
57
+ instrumental: true
58
+ style: "tense orchestral strings, low brass swells, no percussion"
59
+ title: "Approach"
60
+ ```
61
+
62
+ `instrumental: true` is the right answer for almost every film bed — a vocal you did not ask for will fight your dialogue.
63
+
64
+ ## Steer with `style`, not with adjective piles
65
+
66
+ Genre + era + instrumentation + tempo belong in `style`, not stuffed into `prompt`.
67
+
68
+ ```
69
+ ✓ style: "90s trip-hop, dusty breakbeat, Rhodes piano, upright bass, 85 bpm"
70
+ ✗ prompt: "a really cool dusty 90s trip hop song with a Rhodes and..."
71
+ ```
72
+
73
+ `negativeTags` removes what keeps creeping in: `"brass, EDM drop, male vocal"`.
74
+
75
+ ## The steering knobs
76
+
77
+ | Param | Range | Reach for it when |
78
+ |---|---|---|
79
+ | `duration` | 10–360s (**V5_5 + custom mode only**) | Always, when the bed must outlast the cut. It is free. |
80
+ | `vocalGender` | `m` / `f` — **the wire values, not "male"/"female"** | A specific voice is required. |
81
+ | `styleWeight` | 0–1 | The style field is being ignored (raise) or strangling the song (lower). |
82
+ | `weirdnessConstraint` | 0–1 | Takes are too safe (raise) or falling apart (lower). |
83
+ | `audioWeight` | 0–1 | Balancing an audio input against the prompt. |
84
+ | `personaId` / `personaModel` | — | A series needs the same voice/sound across episodes. |
85
+
86
+ ## Character caps
87
+
88
+ | Field | V4 | V4_5 / V4_5PLUS / V5 / V5_5 | V4_5ALL |
89
+ |---|---|---|---|
90
+ | prompt (custom = literal lyrics) | 3000 | 5000 | 5000 |
91
+ | prompt (non-custom = description) | 500 | 500 | 500 |
92
+ | style | 200 | 1000 | 1000 |
93
+ | title | 80 | 100 | 80 |
94
+
95
+ ## Iterating
96
+
97
+ - **Audition both returned songs first.** A re-roll costs a full generation; the second variation is already paid for.
98
+ - Wrong genre → fix `style`. Wrong words → you are in the wrong mode, check the matrix above.
99
+ - Something keeps appearing that you do not want → `negativeTags`, not more prompt.
100
+ - Three failed generations on the same brief means the style field is too vague, not that the seed is unlucky.
101
+
102
+ ## Ops notes
103
+
104
+ - Tracks take 2–3 minutes; a streamable preview exists ~30–40s in. Use `background: true` and poll.
105
+ - The provider hosts files for a limited window — **Slates downloads and stores them locally as soon as the track finishes**, so nothing expires out from under a project.
106
+ - Suno has **no official public API**; this rides an unofficial wrapper. Treat availability as best-effort and do not build a deadline around it.
107
+
108
+ ## Content notes
109
+
110
+ Provider-side moderation rejects lyrics and style prompts naming real artists or protected material (`SENSITIVE_WORD_ERROR`). Describe the sound, not the artist. See slates-content-policy.
@@ -117,7 +117,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
117
117
 
118
118
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
119
119
  2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
120
- 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
120
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
121
121
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
122
122
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
123
123
  6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.