@slatesvideo/shared 0.5.5 → 0.5.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.d.ts +1 -1
- package/dist/index.js +2 -2
- package/dist/operations/index.d.ts +81 -1
- package/dist/operations/index.js +435 -10
- package/dist/prompts/character-sheet.d.ts +2 -1
- package/dist/prompts/character-sheet.js +82 -13
- package/dist/prompts/model-facts.d.ts +1 -1
- package/dist/prompts/model-facts.js +32 -0
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +194 -0
- package/dist/prompts/reference-rules.d.ts +19 -2
- package/dist/prompts/reference-rules.js +18 -1
- package/dist/skills/content.js +5 -2
- package/exports/slates-prompt-builder/generated/reference-character.md +9 -4
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +6 -6
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +1 -1
- package/skills/slates-character-identity.md +9 -4
- package/skills/slates-model-selection.md +26 -0
- package/skills/slates-prompting-elevenlabs.md +131 -0
- package/skills/slates-prompting-seed-audio.md +110 -0
- package/skills/slates-prompting-suno.md +110 -0
|
@@ -27,12 +27,14 @@ Slates generates **one identity sheet per character**, bound as the character's
|
|
|
27
27
|
| Panel | What it carries |
|
|
28
28
|
|---|---|
|
|
29
29
|
| **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
|
|
30
|
-
| **Full-body front, relaxed A-pose —
|
|
30
|
+
| **Full-body front, relaxed A-pose — cropped at the collarbone, just the face cropped out** | Build, proportion, wardrobe. The face is cropped off on purpose: a front-facing body panel renders a ~40px face that can't match the portrait's, so the sheet would carry two competing identities and the model averages them. **Only the face** — neck, arms and hands render as skin. |
|
|
31
31
|
| **Full-body back, head and hair visible** | Hair fall and the back of the outfit — the only panel where either reads. Keeps its head because there's no face to compete with. |
|
|
32
32
|
|
|
33
33
|
The rule is **kill every competing rendering of the FACE, not every head** — which is why exactly one body panel is headless.
|
|
34
34
|
|
|
35
|
-
On a deep neutral-grey plate (
|
|
35
|
+
On a deep neutral-grey plate (hex `3a3a3c`, emitted without the `#` — see the sigil warning in Don'ts), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look. Expression is **a slight natural smile with the teeth just visible** — a closed mouth carries no dental information, so every downstream smiling shot invents teeth, and teeth are person-specific.
|
|
36
|
+
|
|
37
|
+
**Two carve-outs, scoped differently on purpose.** Non-human characters get a natural neutral expression instead of a smile — that one is scoped by *having a human mouth*, so a bipedal robot or humanoid alien is covered. Quadrupeds and non-bipedal characters get a natural standing stance with the head shown on both body panels — that one is *anatomical*. **Both are conditionals the image model evaluates against your reference; neither is a code branch, because the op has no character-kind input.**
|
|
36
38
|
|
|
37
39
|
**The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
|
|
38
40
|
|
|
@@ -53,7 +55,8 @@ If text only: generate from prompt-only — less consistent, so warn the user.
|
|
|
53
55
|
- Default to Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted spend.
|
|
54
56
|
- When the result returns inline, **evaluate it before binding**:
|
|
55
57
|
- Is the portrait clearly the largest panel, and is it off-frontal?
|
|
56
|
-
- **Is the front body panel cleanly headless** — an empty collar
|
|
58
|
+
- **Is the front body panel cleanly headless** — an empty collar above a normally rendered body, no partial face, no floating jaw, no smeared neck stump? A botched crop is worse than no crop.
|
|
59
|
+
- **Is the body still there?** Neck, forearms and hands rendered as skin, not an empty outfit floating on nothing. A hollow garment means the invisible-mannequin genre ran unbounded.
|
|
57
60
|
- Do the body panels read as the same build, wardrobe and hair as the portrait?
|
|
58
61
|
- Catchlights present, irises readable rather than black holes?
|
|
59
62
|
- Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
|
|
@@ -73,6 +76,8 @@ Critically, the app injects **no** wardrobe, expression, or lighting directive.
|
|
|
73
76
|
- **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
|
|
74
77
|
- **Don't** skip binding. An unbound asset doesn't help downstream.
|
|
75
78
|
- **Don't** invent character details. Stick to what's in the reference image and the user's description.
|
|
76
|
-
- **Don't** describe the
|
|
79
|
+
- **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on gpt-image-2** — fal returns `content_policy_violation` with `loc: ["body","prompt"]`, so the text is rejected before any image is read, because an anatomical absence reads as gore to OpenAI's classifier. It passed NB2, which is why the original receipt looked safe: **it was model-scoped.** State an exclusion as a framing choice, never as a missing body part.
|
|
80
|
+
- **Don't** invoke the invisible-mannequin genre without bounding it to the face. **"an invisible-mannequin presentation where the clothing holds its own shape" removed all the skin** — no neck, no hands, no forearms, a garment floating on nothing — because that *is* the e-commerce genre in full: an empty outfit. **"with just the face cropped out"** keeps the anchor and bounds it. Generalises: a genre anchor imports the whole genre, so name what STAYS, not only what goes.
|
|
81
|
+
- **Don't** put `#` or `@` anywhere in prompt text. Both are reference-token sigils in the desktop prompt composer and an unresolved one is **silently deleted** — no error, no log, just missing words. `#3a3a3c` reached fal as `background ()` on a real 2026-07-30 request, meaning the plate value had never been delivered to any model since the composer shipped. Write hex values bare.
|
|
77
82
|
- **Don't** use 4K — wastes credits, no quality gain at sheet scale.
|
|
78
83
|
- **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `reference-seedance.md`.
|
|
@@ -8,7 +8,7 @@
|
|
|
8
8
|
},
|
|
9
9
|
{
|
|
10
10
|
"path": "skills/slates-character-identity.md",
|
|
11
|
-
"sha256": "
|
|
11
|
+
"sha256": "52191db058ada8096297e97958268548b352a5f81739d1bfd01da340867f2d35"
|
|
12
12
|
},
|
|
13
13
|
{
|
|
14
14
|
"path": "skills/slates-prompting-seedance.md",
|
|
@@ -28,7 +28,7 @@
|
|
|
28
28
|
},
|
|
29
29
|
{
|
|
30
30
|
"path": "src/prompts/model-facts.ts",
|
|
31
|
-
"sha256": "
|
|
31
|
+
"sha256": "f3ac5db676a96d1dbdbe7aaa364ca5df3c4739637ef7d61840ad10a585931a62"
|
|
32
32
|
}
|
|
33
33
|
],
|
|
34
34
|
"outputs": [
|
|
@@ -39,8 +39,8 @@
|
|
|
39
39
|
},
|
|
40
40
|
{
|
|
41
41
|
"path": "reference-character.md",
|
|
42
|
-
"bytes":
|
|
43
|
-
"sha256": "
|
|
42
|
+
"bytes": 10081,
|
|
43
|
+
"sha256": "7dd5b693f3095b6678106c1a41c1303d57b3d7a30da0eeffc867fc5000fd6c74"
|
|
44
44
|
},
|
|
45
45
|
{
|
|
46
46
|
"path": "reference-seedance.md",
|
|
@@ -65,8 +65,8 @@
|
|
|
65
65
|
],
|
|
66
66
|
"archive": {
|
|
67
67
|
"path": "slates-prompt-builder.skill",
|
|
68
|
-
"bytes":
|
|
69
|
-
"sha256": "
|
|
68
|
+
"bytes": 37938,
|
|
69
|
+
"sha256": "48331704f54f422542f7a1a164915f3739b925e3f225a3710dbd8ecb2963ce68",
|
|
70
70
|
"entries": [
|
|
71
71
|
"SKILL.md",
|
|
72
72
|
"reference-character.md",
|
|
Binary file
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@slatesvideo/shared",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.6",
|
|
4
4
|
"description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -28,12 +28,14 @@ Slates generates **one identity sheet per character**, bound as the character's
|
|
|
28
28
|
| Panel | What it carries |
|
|
29
29
|
|---|---|
|
|
30
30
|
| **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
|
|
31
|
-
| **Full-body front, relaxed A-pose —
|
|
31
|
+
| **Full-body front, relaxed A-pose — cropped at the collarbone, just the face cropped out** | Build, proportion, wardrobe. The face is cropped off on purpose: a front-facing body panel renders a ~40px face that can't match the portrait's, so the sheet would carry two competing identities and the model averages them. **Only the face** — neck, arms and hands render as skin. |
|
|
32
32
|
| **Full-body back, head and hair visible** | Hair fall and the back of the outfit — the only panel where either reads. Keeps its head because there's no face to compete with. |
|
|
33
33
|
|
|
34
34
|
The rule is **kill every competing rendering of the FACE, not every head** — which is why exactly one body panel is headless.
|
|
35
35
|
|
|
36
|
-
On a deep neutral-grey plate (
|
|
36
|
+
On a deep neutral-grey plate (hex `3a3a3c`, emitted without the `#` — see the sigil warning in Don'ts), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look. Expression is **a slight natural smile with the teeth just visible** — a closed mouth carries no dental information, so every downstream smiling shot invents teeth, and teeth are person-specific.
|
|
37
|
+
|
|
38
|
+
**Two carve-outs, scoped differently on purpose.** Non-human characters get a natural neutral expression instead of a smile — that one is scoped by *having a human mouth*, so a bipedal robot or humanoid alien is covered. Quadrupeds and non-bipedal characters get a natural standing stance with the head shown on both body panels — that one is *anatomical*. **Both are conditionals the image model evaluates against your reference; neither is a code branch, because the op has no character-kind input.**
|
|
37
39
|
|
|
38
40
|
**The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
|
|
39
41
|
|
|
@@ -69,7 +71,8 @@ If text only: generate from prompt-only — less consistent, so warn the user.
|
|
|
69
71
|
- Default to Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted spend.
|
|
70
72
|
- When the result returns inline, **evaluate it before binding**:
|
|
71
73
|
- Is the portrait clearly the largest panel, and is it off-frontal?
|
|
72
|
-
- **Is the front body panel cleanly headless** — an empty collar
|
|
74
|
+
- **Is the front body panel cleanly headless** — an empty collar above a normally rendered body, no partial face, no floating jaw, no smeared neck stump? A botched crop is worse than no crop.
|
|
75
|
+
- **Is the body still there?** Neck, forearms and hands rendered as skin, not an empty outfit floating on nothing. A hollow garment means the invisible-mannequin genre ran unbounded.
|
|
73
76
|
- Do the body panels read as the same build, wardrobe and hair as the portrait?
|
|
74
77
|
- Catchlights present, irises readable rather than black holes?
|
|
75
78
|
- Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
|
|
@@ -95,6 +98,8 @@ Critically, the app injects **no** wardrobe, expression, or lighting directive.
|
|
|
95
98
|
- **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
|
|
96
99
|
- **Don't** skip binding. An unbound asset doesn't help downstream.
|
|
97
100
|
- **Don't** invent character details. Stick to what's in the reference image and the user's description.
|
|
98
|
-
- **Don't** describe the
|
|
101
|
+
- **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on gpt-image-2** — fal returns `content_policy_violation` with `loc: ["body","prompt"]`, so the text is rejected before any image is read, because an anatomical absence reads as gore to OpenAI's classifier. It passed NB2, which is why the original receipt looked safe: **it was model-scoped.** State an exclusion as a framing choice, never as a missing body part.
|
|
102
|
+
- **Don't** invoke the invisible-mannequin genre without bounding it to the face. **"an invisible-mannequin presentation where the clothing holds its own shape" removed all the skin** — no neck, no hands, no forearms, a garment floating on nothing — because that *is* the e-commerce genre in full: an empty outfit. **"with just the face cropped out"** keeps the anchor and bounds it. Generalises: a genre anchor imports the whole genre, so name what STAYS, not only what goes.
|
|
103
|
+
- **Don't** put `#` or `@` anywhere in prompt text. Both are reference-token sigils in the desktop prompt composer and an unresolved one is **silently deleted** — no error, no log, just missing words. `#3a3a3c` reached fal as `background ()` on a real 2026-07-30 request, meaning the plate value had never been delivered to any model since the composer shipped. Write hex values bare.
|
|
99
104
|
- **Don't** use 4K — wastes credits, no quality gain at sheet scale.
|
|
100
105
|
- **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `slates-prompting-seedance`.
|
|
@@ -91,6 +91,32 @@ Both tools have a cheap Kling utility lane and a premium Seedance lane. The capa
|
|
|
91
91
|
|
|
92
92
|
**Split rule of thumb:** readable text / panels / UI → GPT Image 2; photoreal, character-locked, widescreen, or edit-heavy → the Banana line; drafts → NB2 Lite; uncensored or odd resolutions → Seedream/FLUX.
|
|
93
93
|
|
|
94
|
+
## Audio routing
|
|
95
|
+
|
|
96
|
+
**Image and video models cannot generate standalone audio, and none of the four audio models can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.
|
|
97
|
+
|
|
98
|
+
| Job | Model | Why |
|
|
99
|
+
|---|---|---|
|
|
100
|
+
| **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. 1–120s. The continuity-bed workhorse. |
|
|
101
|
+
| **The exact words, in a repeatable named voice** — ad reads, narration, character lines to lip-sync against | **Eleven v3** (`eleven-v3`) | Verbatim text, 20 preset voices, re-renderable after a copy tweak without the performance drifting. Billed per 100 characters. |
|
|
102
|
+
| **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control (0.5–22s) and a real loop mode. |
|
|
103
|
+
| **A song or a score** | **Suno** (`suno`) | Full music with structure. Two variations per call for one flat price, and duration is free to 360s. |
|
|
104
|
+
|
|
105
|
+
### Named audio escalation triggers
|
|
106
|
+
|
|
107
|
+
- **"It needs to sound like a place"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.
|
|
108
|
+
- **"Read this line"** with copy that a client can still change → Eleven v3. Scratch dialogue while the script is moving can stay on Seed Audio.
|
|
109
|
+
- **"That needs a thump right there"** → Sound Effects, with the duration set to roughly the length of the event.
|
|
110
|
+
- **"Give it a track"** → Suno, `instrumental: true` unless a vocal is genuinely wanted (an unasked-for vocal fights dialogue).
|
|
111
|
+
|
|
112
|
+
**Rules:**
|
|
113
|
+
|
|
114
|
+
- **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.
|
|
115
|
+
- **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.
|
|
116
|
+
- **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — on Suno the extra seconds are literally free.
|
|
117
|
+
- **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.
|
|
118
|
+
- Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`, `slates-prompting-suno`.
|
|
119
|
+
|
|
94
120
|
## Cost is a tiebreaker, not the router
|
|
95
121
|
|
|
96
122
|
Route by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.
|
|
@@ -0,0 +1,131 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-elevenlabs
|
|
3
|
+
description: How to prompt the two ElevenLabs surfaces in Slates. Read before calling slates_generate_audio with model eleven-v3 (Eleven v3 text-to-speech - controlled, repeatable, named-voice voiceover, billed per 100 characters) or eleven-sfx (Sound Effects v2 - one short effect with an EXACT duration, 0.5-22s, billed per second). Covers the "the text field is spoken verbatim" rule, punctuation as the only timing control, stability, describing an effect by its physical cause, and when to use Seed Audio instead.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# ElevenLabs — prompting (Eleven v3 TTS + Sound Effects v2)
|
|
7
|
+
|
|
8
|
+
Two separate surfaces from the same vendor, carried on fal (`fal-ai/elevenlabs/tts/eleven-v3`, `fal-ai/elevenlabs/sound-effects/v2`). They share nothing but a bill — treat them as different tools.
|
|
9
|
+
|
|
10
|
+
## Where they route
|
|
11
|
+
|
|
12
|
+
- **`eleven-v3`** — the exact words matter and the read must be **repeatable**: ad reads, narration, character lines you will lip-sync against, anything a client will ask you to re-render after a copy tweak. Billed per 100 characters of text, rounded up.
|
|
13
|
+
- **`eleven-sfx`** — a single sound that has to land on a known frame, or a seamless loop. Billed per second, 0.5–22s.
|
|
14
|
+
- **Neither** for layered scenes. A room with dialogue *and* clatter *and* ambience is one `seed-audio` pass, not three ElevenLabs generations.
|
|
15
|
+
- **AUDIO-ONLY.** Neither can produce images or video.
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
## Eleven v3 (`eleven-v3`) — THE RULES
|
|
20
|
+
|
|
21
|
+
### 1. 🚨 The text field is the script. Every character is spoken.
|
|
22
|
+
|
|
23
|
+
```
|
|
24
|
+
✗ (excited) Read this fast — "Grab yours today!"
|
|
25
|
+
✓ Grab yours today!
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
Stage directions, speaker names, bracketed emotion tags and markdown all get read out loud. There is no instruction channel — direction lives in `stability` and in how you punctuate.
|
|
29
|
+
|
|
30
|
+
### 2. Punctuation is the only timing control
|
|
31
|
+
|
|
32
|
+
| You want | Write |
|
|
33
|
+
|---|---|
|
|
34
|
+
| a hard stop | `It works. Every time.` |
|
|
35
|
+
| a beat, not a stop | `It works — every time.` |
|
|
36
|
+
| a trailing hesitation | `It works… mostly.` |
|
|
37
|
+
| a list rhythm | `Faster, cheaper, and yours.` |
|
|
38
|
+
|
|
39
|
+
Rewrite the punctuation before you touch a setting. It moves the read more than `stability` does.
|
|
40
|
+
|
|
41
|
+
### 3. Pick a voice and keep it
|
|
42
|
+
|
|
43
|
+
20 presets: Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill. Default `Rachel`.
|
|
44
|
+
|
|
45
|
+
One voice per character, one per piece. A series that swaps voices between shots reads as an accident. There is **no voice cloning on this route** — a cloned ElevenLabs voice ID is not supported and must not be assumed to pass through.
|
|
46
|
+
|
|
47
|
+
### 4. Stability
|
|
48
|
+
|
|
49
|
+
| Value | Behavior | Use for |
|
|
50
|
+
|---|---|---|
|
|
51
|
+
| ~0.3 | expressive, varies take-to-take | one dramatic line, character dialogue |
|
|
52
|
+
| 0.5 (default) | balanced | most reads |
|
|
53
|
+
| ~0.8 | flat, highly repeatable | long narration, anything you will re-render |
|
|
54
|
+
|
|
55
|
+
Raise it when re-rolls keep giving you a different performance. Lower it when the read is lifeless.
|
|
56
|
+
|
|
57
|
+
### 5. Spell out what TTS gets wrong
|
|
58
|
+
|
|
59
|
+
Acronyms, product names, prices, years and URLs are where it embarrasses itself. Write the pronunciation:
|
|
60
|
+
|
|
61
|
+
```
|
|
62
|
+
SKU → "ess kay you"
|
|
63
|
+
2026 → "twenty twenty six"
|
|
64
|
+
$19.99 → "nineteen ninety nine"
|
|
65
|
+
slates.video → "slates dot video"
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
Set `languageCode` (ISO 639-1) to force a language when the text is ambiguous or code-switched.
|
|
69
|
+
|
|
70
|
+
### 6. Length is money, honestly
|
|
71
|
+
|
|
72
|
+
1–5000 characters, billed in 100-character buckets rounded up. A tightened sentence costs less; a pasted stray paragraph costs more. This is the one Slates surface where editing the copy is also a cost control.
|
|
73
|
+
|
|
74
|
+
### 7. Timestamps are free — leave them on
|
|
75
|
+
|
|
76
|
+
Word-level timestamps come back with every generation at no extra charge. They are exactly what a caption/subtitle pass consumes. There is no reason to disable them.
|
|
77
|
+
|
|
78
|
+
---
|
|
79
|
+
|
|
80
|
+
## Sound Effects v2 (`eleven-sfx`) — THE RULES
|
|
81
|
+
|
|
82
|
+
### 1. Describe the physical CAUSE, not the label
|
|
83
|
+
|
|
84
|
+
```
|
|
85
|
+
✗ door sound
|
|
86
|
+
✓ heavy oak door slams shut in a stone hallway
|
|
87
|
+
|
|
88
|
+
✗ whoosh
|
|
89
|
+
✓ a thick rope swung fast past a microphone, low air displacement
|
|
90
|
+
|
|
91
|
+
✗ footsteps
|
|
92
|
+
✓ boots on wet gravel, slow, one person
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
Material + weight + surface + room. Naming all four is the difference between a usable effect and a stock-library shrug. Cap is 450 characters — you will not need them.
|
|
96
|
+
|
|
97
|
+
### 2. One sound per generation
|
|
98
|
+
|
|
99
|
+
This surface makes a single event. A door, then footsteps, then a siren is three generations layered on the timeline — or one `seed-audio` scene, which is usually cheaper and always more coherent.
|
|
100
|
+
|
|
101
|
+
### 3. Duration is always explicit, and it is the price
|
|
102
|
+
|
|
103
|
+
`durationSeconds` is 0.5–22 and Slates **always sends it**. (Left null the model picks, which makes the charge non-deterministic — so it is never left null.)
|
|
104
|
+
|
|
105
|
+
| Kind of sound | Ask for |
|
|
106
|
+
|---|---|
|
|
107
|
+
| impact, hit, click | 0.5–1s |
|
|
108
|
+
| whoosh, riser, transition | 2–4s |
|
|
109
|
+
| loopable bed | 8–22s + `loop: true` |
|
|
110
|
+
|
|
111
|
+
Over-asking pads the tail with room tone you then trim. Under-asking clips the decay.
|
|
112
|
+
|
|
113
|
+
### 4. Loops
|
|
114
|
+
|
|
115
|
+
`loop: true` tiles without a seam — rain, engine hum, crowd murmur, machine noise. Combine with a longer duration so the loop point is not obvious.
|
|
116
|
+
|
|
117
|
+
### 5. Prompt influence
|
|
118
|
+
|
|
119
|
+
`promptInfluence` 0–1, default 0.3. Higher hugs your wording with less variation between takes; lower explores. Raise it when a re-roll keeps wandering off the brief; lower it when every take sounds like the same take.
|
|
120
|
+
|
|
121
|
+
---
|
|
122
|
+
|
|
123
|
+
## Iterating on either surface
|
|
124
|
+
|
|
125
|
+
- TTS re-rolls that keep drifting = raise `stability`. TTS reads that sound robotic = lower it, then fix the punctuation.
|
|
126
|
+
- SFX re-rolls that keep missing = the prompt named a label instead of a cause. Rewrite it as a physical event.
|
|
127
|
+
- Three failed takes means the prompt is wrong, not the seed.
|
|
128
|
+
|
|
129
|
+
## Content notes
|
|
130
|
+
|
|
131
|
+
ElevenLabs applies its own moderation, and voice likeness of real people is restricted by their terms. See slates-content-policy.
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-seed-audio
|
|
3
|
+
description: How to prompt Seed Audio 1.0 (ByteDance, via fal). Read before calling slates_generate_audio with model seed-audio. The one-pass audio SCENE model — dialogue, SFX and ambience together from ONE plain sentence. CRITICAL - it has NO duration parameter, so length must be named IN THE PROMPT TEXT and Slates bills the duration you request. Covers the one-sentence doctrine, the crowd-size rule, why Kling "SFX:" syntax hurts here, and the audio-refs-XOR-image input rule.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Seed Audio 1.0 — prompting
|
|
7
|
+
|
|
8
|
+
ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence. 1–120 seconds. It is the default audio model in Slates and the workhorse for continuity beds.
|
|
9
|
+
|
|
10
|
+
## Where it routes
|
|
11
|
+
|
|
12
|
+
- **Scene audio, room tone, ambience beds, crowd/nature soundscapes** — anything where several sounds share a space. One generation, not three layered ones.
|
|
13
|
+
- **Fast scratch dialogue** when the exact wording is still moving. Once the script locks and the read has to be repeatable, switch to `eleven-v3`.
|
|
14
|
+
- **NOT** a single effect that must land on a known frame — that is `eleven-sfx`, which takes an exact duration.
|
|
15
|
+
- **NOT** music — that is `suno`.
|
|
16
|
+
- **AUDIO-ONLY.** It cannot produce images or video.
|
|
17
|
+
|
|
18
|
+
## THE RULES
|
|
19
|
+
|
|
20
|
+
### 1. 🚨 There is no duration parameter — the words set the length
|
|
21
|
+
|
|
22
|
+
This is the single most important fact about this model. Output length is driven by the prompt text ("… 15 seconds"), capped at 120s.
|
|
23
|
+
|
|
24
|
+
**Slates handles this for you:** the `durationSeconds` param appends the duration to the prompt and **bills that number of seconds**. So:
|
|
25
|
+
|
|
26
|
+
- Set `durationSeconds` to what you actually want.
|
|
27
|
+
- **Do not also write a different length into your sentence.** Two numbers fight, and you pay for the one you selected, not the one you got.
|
|
28
|
+
- If the returned clip is shorter than requested you still paid for the request — that is the deal that keeps the displayed price equal to the charge. Ask for what you need.
|
|
29
|
+
|
|
30
|
+
<!-- slates-only -->
|
|
31
|
+
The server re-derives the billed key from `durationSeconds` (a client cannot under-bill), probes the returned `audio.duration` after completion, and logs `SEED AUDIO BILLING DRIFT` if the model overshot. No auto-charge, no refund — the request is the contract.
|
|
32
|
+
<!-- /slates-only -->
|
|
33
|
+
|
|
34
|
+
### 2. One plain sentence. No production jargon.
|
|
35
|
+
|
|
36
|
+
Field-proven (Higgsfield sprint, 2026-07-27/28). Working prompts look like this:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
tiny applause of 2 or 3 people at an open mic. 15 seconds
|
|
40
|
+
nature soundscape, wide open field cicadas and birds and a loon.
|
|
41
|
+
a diner at 2am, one coffee machine hissing, cutlery somewhere behind the counter
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Not this:
|
|
45
|
+
|
|
46
|
+
```
|
|
47
|
+
✗ AMBIENCE: interior diner, night. SFX: espresso machine (hiss, 2s), cutlery.
|
|
48
|
+
✗ Wide shot of a diner. Slow push in. Warm tungsten. Ambient noise: ...
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Shot language, camera moves and lighting belong to video prompts. Here they are just words the model has to ignore.
|
|
52
|
+
|
|
53
|
+
### 3. Never bring Kling's audio syntax to this model
|
|
54
|
+
|
|
55
|
+
`SFX:` and `Ambient noise:` prefixes and `Background music:` labels are **Kling 3.0 video** syntax. Seed Audio has no parser for them — it reads them as text in the scene and the output gets measurably worse. Describe the sounds directly instead.
|
|
56
|
+
|
|
57
|
+
### 4. Name the crowd size, the room size, the distance
|
|
58
|
+
|
|
59
|
+
The highest-leverage single edit on any bed. Unqualified nouns default big:
|
|
60
|
+
|
|
61
|
+
| Vague | What it returns | Fixed |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| `applause` | a full auditorium | `tiny applause of 2 or 3 people` |
|
|
64
|
+
| `traffic` | a highway | `one car passing on a wet residential street` |
|
|
65
|
+
| `crowd` | a stadium | `four people talking at the next table` |
|
|
66
|
+
|
|
67
|
+
Distance words (`far off`, `muffled through a wall`, `right next to the mic`) work the same way and are how you build depth in one sentence.
|
|
68
|
+
|
|
69
|
+
### 5. Beds must outlast the cut
|
|
70
|
+
|
|
71
|
+
Ask for a few seconds more than the clip needs so the edit has handles to fade through. A bed that ends exactly on the cut always sounds clipped. This is a product requirement, not a preference — it is why the duration control exists at all.
|
|
72
|
+
|
|
73
|
+
### 6. Dialogue goes in quotes, inside the same sentence as the room
|
|
74
|
+
|
|
75
|
+
```
|
|
76
|
+
a tired bartender says, "we closed twenty minutes ago", glasses clinking behind him
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
Pick a preset voice when a specific speaker matters. Leave `voice` unset and the scene casts itself — which is usually right for crowd and background dialogue.
|
|
80
|
+
|
|
81
|
+
Preset voices (20): `vivi_mixed_en_zh_ja_es_id`, `mindy_en_es_id_pt_zh`, `kian_en_zh`, `cedric_en_zh`, `sophie_en_zh`, `jean_en_zh`, `magnus_en_zh`, `mabel_en_zh`, `nadia_en_zh`, `opal_en_zh`, `pearl_en_zh`, `quentin_en_zh`, `corinne_mixed_en_zh`, `esther_mixed_en_zh`, `lyla_mixed_en_zh`, `tracy_es_zh`, `sandy_es_mixed_en_zh`, `felix_zh`, `celeste_zh`, `monkey_king_zh`.
|
|
82
|
+
|
|
83
|
+
Set `multilingual: true` for non-English or mixed-language lines.
|
|
84
|
+
|
|
85
|
+
### 7. Inputs: up to 3 audio clips **XOR** one image. Never both.
|
|
86
|
+
|
|
87
|
+
- **Audio references** — up to 3 clips, each ≤30s and ≤10MB (wav/mp3/pcm/ogg_opus). Refer to them in the prompt as `@Audio1`, `@Audio2`, `@Audio3`: *"match the room tone of @Audio1"*.
|
|
88
|
+
- **Image reference** — one image (jpeg/png/webp ≤10MB). The model scores what it sees.
|
|
89
|
+
- Sending both is rejected by the API. Pick the one that carries the intent.
|
|
90
|
+
|
|
91
|
+
### 8. The knobs, and when to touch them
|
|
92
|
+
|
|
93
|
+
| Param | Range | Reach for it when |
|
|
94
|
+
|---|---|---|
|
|
95
|
+
| `speed` | 0.5–2.0 | Dialogue is racing or dragging against picture. |
|
|
96
|
+
| `volume` | 0.5–2.0 | Rarely — normalize on the timeline instead. |
|
|
97
|
+
| `pitch` | −12…+12 semitones | Ageing or shifting a voice. Small moves only; ±3 is already a lot. |
|
|
98
|
+
| `multilingual` | bool | Non-English or code-switched lines. |
|
|
99
|
+
| `sampleRate` | 8k–48k | Leave at 24000 unless you are matching an existing stem. |
|
|
100
|
+
| `outputFormat` | mp3 / wav / pcm / ogg_opus | wav when this is going into a mix; mp3 otherwise. |
|
|
101
|
+
|
|
102
|
+
## Iterating
|
|
103
|
+
|
|
104
|
+
- A bed that came back wrong is almost always a **scale** problem (crowd/room too big) or a **jargon** problem (the sentence reads like a spec). Fix those two before touching `speed`/`pitch`.
|
|
105
|
+
- Three failed takes on the same sentence means the sentence is wrong, not the seed. Rewrite it the way you would say it out loud.
|
|
106
|
+
- Generations are cheap enough at short durations that auditioning two phrasings beats agonizing over one.
|
|
107
|
+
|
|
108
|
+
## Content notes
|
|
109
|
+
|
|
110
|
+
Provider-side moderation applies to voices and to recognizable real people. See slates-content-policy.
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-suno
|
|
3
|
+
description: How to prompt Suno in Slates. Read before calling slates_generate_audio with model suno. Full music tracks - EVERY call returns TWO variations for one flat price and duration is FREE up to 360 seconds. The rule that decides everything - in CUSTOM mode the prompt field is the EXACT LYRICS (sung as written), in DESCRIPTION mode it is a description and the lyrics get written for you. Covers the mode matrix, style vs prompt steering, instrumental scoring, negative tags, and the character caps.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Suno — prompting
|
|
7
|
+
|
|
8
|
+
Full music generation, reached through the sunoapi.org wrapper. Models `V4`, `V4_5`, `V4_5PLUS`, `V4_5ALL`, `V5`, `V5_5`.
|
|
9
|
+
|
|
10
|
+
**Two facts that should shape every decision:**
|
|
11
|
+
|
|
12
|
+
1. **Every call returns TWO songs** — two genuinely different takes on the same brief, for one flat price. Audition both before re-rolling.
|
|
13
|
+
2. **Length is free.** A 360-second track costs exactly what a default one costs (measured against the live provider balance 2026-07-31: a default call and a `duration: 240` call both debited the same). There is never a reason to generate a bed shorter than your edit.
|
|
14
|
+
|
|
15
|
+
## Where it routes
|
|
16
|
+
|
|
17
|
+
- **Anything a listener would call a song or a score** — theme, underscore, needle-drop, montage bed, end-card sting.
|
|
18
|
+
- **NOT** ambience or room tone — that is `seed-audio`, which is cheaper and better at it.
|
|
19
|
+
- **NOT** a single effect — that is `eleven-sfx`.
|
|
20
|
+
- **AUDIO-ONLY.**
|
|
21
|
+
|
|
22
|
+
## 🚨 THE RULE: which mode you are in changes what `prompt` means
|
|
23
|
+
|
|
24
|
+
| `customMode` | `instrumental` | Required | What `prompt` means |
|
|
25
|
+
|---|---|---|---|
|
|
26
|
+
| `false` | either | `prompt` only (≤500 chars) | **A description.** Lyrics get written for you. |
|
|
27
|
+
| `true` | `true` | `style`, `title` | **Unused.** Style + title do all the steering. |
|
|
28
|
+
| `true` | `false` | `style`, `title`, `prompt` | **THE EXACT LYRICS**, sung as written. |
|
|
29
|
+
|
|
30
|
+
Putting a description in the prompt field while `customMode: true` and `instrumental: false` gets your description **sung back at you**. This is the single most common Suno mistake and it costs a full generation every time.
|
|
31
|
+
|
|
32
|
+
## Description mode — the fast path
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
customMode: false
|
|
36
|
+
prompt: "brooding synthwave for a night drive, analog bass, no vocals, 90 bpm"
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Use it when you need a mood and do not care about specific words. 500-character cap. This is the right default for background beds.
|
|
40
|
+
|
|
41
|
+
## Custom mode — when the words matter
|
|
42
|
+
|
|
43
|
+
```
|
|
44
|
+
customMode: true
|
|
45
|
+
instrumental: false
|
|
46
|
+
style: "dream pop, hazy, reverb-heavy guitars, female vocal, 100 bpm"
|
|
47
|
+
title: "Blue Hour"
|
|
48
|
+
prompt: "[Verse 1]\nThe lights come on before we're ready\n..."
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Structure tags (`[Verse]`, `[Chorus]`, `[Bridge]`, `[Outro]`) inside the lyrics are how you control the arrangement. Everything that is not a structure tag will be sung.
|
|
52
|
+
|
|
53
|
+
## Instrumental scoring
|
|
54
|
+
|
|
55
|
+
```
|
|
56
|
+
customMode: true
|
|
57
|
+
instrumental: true
|
|
58
|
+
style: "tense orchestral strings, low brass swells, no percussion"
|
|
59
|
+
title: "Approach"
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
`instrumental: true` is the right answer for almost every film bed — a vocal you did not ask for will fight your dialogue.
|
|
63
|
+
|
|
64
|
+
## Steer with `style`, not with adjective piles
|
|
65
|
+
|
|
66
|
+
Genre + era + instrumentation + tempo belong in `style`, not stuffed into `prompt`.
|
|
67
|
+
|
|
68
|
+
```
|
|
69
|
+
✓ style: "90s trip-hop, dusty breakbeat, Rhodes piano, upright bass, 85 bpm"
|
|
70
|
+
✗ prompt: "a really cool dusty 90s trip hop song with a Rhodes and..."
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
`negativeTags` removes what keeps creeping in: `"brass, EDM drop, male vocal"`.
|
|
74
|
+
|
|
75
|
+
## The steering knobs
|
|
76
|
+
|
|
77
|
+
| Param | Range | Reach for it when |
|
|
78
|
+
|---|---|---|
|
|
79
|
+
| `duration` | 10–360s (**V5_5 + custom mode only**) | Always, when the bed must outlast the cut. It is free. |
|
|
80
|
+
| `vocalGender` | `m` / `f` — **the wire values, not "male"/"female"** | A specific voice is required. |
|
|
81
|
+
| `styleWeight` | 0–1 | The style field is being ignored (raise) or strangling the song (lower). |
|
|
82
|
+
| `weirdnessConstraint` | 0–1 | Takes are too safe (raise) or falling apart (lower). |
|
|
83
|
+
| `audioWeight` | 0–1 | Balancing an audio input against the prompt. |
|
|
84
|
+
| `personaId` / `personaModel` | — | A series needs the same voice/sound across episodes. |
|
|
85
|
+
|
|
86
|
+
## Character caps
|
|
87
|
+
|
|
88
|
+
| Field | V4 | V4_5 / V4_5PLUS / V5 / V5_5 | V4_5ALL |
|
|
89
|
+
|---|---|---|---|
|
|
90
|
+
| prompt (custom = literal lyrics) | 3000 | 5000 | 5000 |
|
|
91
|
+
| prompt (non-custom = description) | 500 | 500 | 500 |
|
|
92
|
+
| style | 200 | 1000 | 1000 |
|
|
93
|
+
| title | 80 | 100 | 80 |
|
|
94
|
+
|
|
95
|
+
## Iterating
|
|
96
|
+
|
|
97
|
+
- **Audition both returned songs first.** A re-roll costs a full generation; the second variation is already paid for.
|
|
98
|
+
- Wrong genre → fix `style`. Wrong words → you are in the wrong mode, check the matrix above.
|
|
99
|
+
- Something keeps appearing that you do not want → `negativeTags`, not more prompt.
|
|
100
|
+
- Three failed generations on the same brief means the style field is too vague, not that the seed is unlucky.
|
|
101
|
+
|
|
102
|
+
## Ops notes
|
|
103
|
+
|
|
104
|
+
- Tracks take 2–3 minutes; a streamable preview exists ~30–40s in. Use `background: true` and poll.
|
|
105
|
+
- The provider hosts files for a limited window — **Slates downloads and stores them locally as soon as the track finishes**, so nothing expires out from under a project.
|
|
106
|
+
- Suno has **no official public API**; this rides an unofficial wrapper. Treat availability as best-effort and do not build a deadline around it.
|
|
107
|
+
|
|
108
|
+
## Content notes
|
|
109
|
+
|
|
110
|
+
Provider-side moderation rejects lyrics and style prompts naming real artists or protected material (`SENSITIVE_WORD_ERROR`). Describe the sound, not the artist. See slates-content-policy.
|