@slatesvideo/shared 0.5.5 → 0.5.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.d.ts +2 -2
- package/dist/index.js +6 -3
- package/dist/operations/index.d.ts +106 -19
- package/dist/operations/index.js +624 -172
- package/dist/prompts/character-sheet.d.ts +2 -1
- package/dist/prompts/character-sheet.js +82 -13
- package/dist/prompts/model-facts.d.ts +43 -1
- package/dist/prompts/model-facts.js +104 -2
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +208 -5
- package/dist/prompts/reference-composer.d.ts +21 -3
- package/dist/prompts/reference-composer.js +80 -10
- package/dist/prompts/reference-rules.d.ts +19 -2
- package/dist/prompts/reference-rules.js +18 -1
- package/dist/skills/content.js +8 -5
- package/exports/slates-prompt-builder/generated/SKILL.md +2 -2
- package/exports/slates-prompt-builder/generated/reference-character.md +9 -4
- package/exports/slates-prompt-builder/generated/reference-seedance.md +12 -1
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +11 -11
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +1 -1
- package/skills/slates-character-identity.md +9 -4
- package/skills/slates-model-selection.md +39 -11
- package/skills/slates-prompting-elevenlabs.md +69 -0
- package/skills/slates-prompting-lip-sync.md +12 -14
- package/skills/slates-prompting-motion-transfer.md +18 -14
- package/skills/slates-prompting-seed-audio.md +110 -0
- package/skills/slates-prompting-seedance-2-5.md +215 -0
- package/skills/slates-prompting-seedance.md +17 -1
|
@@ -16,8 +16,8 @@ This portable skill is deliberately thin. Its reference files are generated dire
|
|
|
16
16
|
<!-- @generated:model-routing -->
|
|
17
17
|
| Model | Canonical route | Guide |
|
|
18
18
|
|---|---|---|
|
|
19
|
-
| **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics.
|
|
20
|
-
| **Seedance 2.0** | PREMIUM video tier — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those).
|
|
19
|
+
| **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics. Kling is also the ONLY engine behind the Motion Transfer and Lip Sync tools (MC std/pro, lip-sync, avatar) — those two tools are Kling-only. | `reference-kling.md` |
|
|
20
|
+
| **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those). 4-15s, up to 9 ingredient images. Strong I2V / own-footage restyle. Native 4K, but 4K VIDEO is a Pro-only tier gate (base maxes at 1080p; server returns PRO_REQUIRED) — default 1080p unless the user is on Pro. Attaching a clip as a video reference (own-footage restyle, motion or dialogue conditioning) bills combined input+output seconds. 2.0 STAYS THE DEFAULT over 2.5 because it is the only Seedance with 1080p and 4K. | `reference-seedance.md` |
|
|
21
21
|
| **Nano Banana 2 (Gemini 3.1 Flash Image)** | Default image model. 14 refs hard cap (10 object + 4 character). Brief it like a creative director, not tag soup. No negativePrompt field — use positive reframing. Best image start-frame for legible text. Knowledge cutoff Jan 2025. | `reference-nano-banana.md` |
|
|
22
22
|
<!-- @end:model-routing -->
|
|
23
23
|
|
|
@@ -27,12 +27,14 @@ Slates generates **one identity sheet per character**, bound as the character's
|
|
|
27
27
|
| Panel | What it carries |
|
|
28
28
|
|---|---|
|
|
29
29
|
| **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
|
|
30
|
-
| **Full-body front, relaxed A-pose —
|
|
30
|
+
| **Full-body front, relaxed A-pose — cropped at the collarbone, just the face cropped out** | Build, proportion, wardrobe. The face is cropped off on purpose: a front-facing body panel renders a ~40px face that can't match the portrait's, so the sheet would carry two competing identities and the model averages them. **Only the face** — neck, arms and hands render as skin. |
|
|
31
31
|
| **Full-body back, head and hair visible** | Hair fall and the back of the outfit — the only panel where either reads. Keeps its head because there's no face to compete with. |
|
|
32
32
|
|
|
33
33
|
The rule is **kill every competing rendering of the FACE, not every head** — which is why exactly one body panel is headless.
|
|
34
34
|
|
|
35
|
-
On a deep neutral-grey plate (
|
|
35
|
+
On a deep neutral-grey plate (hex `3a3a3c`, emitted without the `#` — see the sigil warning in Don'ts), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look. Expression is **a slight natural smile with the teeth just visible** — a closed mouth carries no dental information, so every downstream smiling shot invents teeth, and teeth are person-specific.
|
|
36
|
+
|
|
37
|
+
**Two carve-outs, scoped differently on purpose.** Non-human characters get a natural neutral expression instead of a smile — that one is scoped by *having a human mouth*, so a bipedal robot or humanoid alien is covered. Quadrupeds and non-bipedal characters get a natural standing stance with the head shown on both body panels — that one is *anatomical*. **Both are conditionals the image model evaluates against your reference; neither is a code branch, because the op has no character-kind input.**
|
|
36
38
|
|
|
37
39
|
**The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
|
|
38
40
|
|
|
@@ -53,7 +55,8 @@ If text only: generate from prompt-only — less consistent, so warn the user.
|
|
|
53
55
|
- Default to Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted spend.
|
|
54
56
|
- When the result returns inline, **evaluate it before binding**:
|
|
55
57
|
- Is the portrait clearly the largest panel, and is it off-frontal?
|
|
56
|
-
- **Is the front body panel cleanly headless** — an empty collar
|
|
58
|
+
- **Is the front body panel cleanly headless** — an empty collar above a normally rendered body, no partial face, no floating jaw, no smeared neck stump? A botched crop is worse than no crop.
|
|
59
|
+
- **Is the body still there?** Neck, forearms and hands rendered as skin, not an empty outfit floating on nothing. A hollow garment means the invisible-mannequin genre ran unbounded.
|
|
57
60
|
- Do the body panels read as the same build, wardrobe and hair as the portrait?
|
|
58
61
|
- Catchlights present, irises readable rather than black holes?
|
|
59
62
|
- Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
|
|
@@ -73,6 +76,8 @@ Critically, the app injects **no** wardrobe, expression, or lighting directive.
|
|
|
73
76
|
- **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
|
|
74
77
|
- **Don't** skip binding. An unbound asset doesn't help downstream.
|
|
75
78
|
- **Don't** invent character details. Stick to what's in the reference image and the user's description.
|
|
76
|
-
- **Don't** describe the
|
|
79
|
+
- **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on gpt-image-2** — fal returns `content_policy_violation` with `loc: ["body","prompt"]`, so the text is rejected before any image is read, because an anatomical absence reads as gore to OpenAI's classifier. It passed NB2, which is why the original receipt looked safe: **it was model-scoped.** State an exclusion as a framing choice, never as a missing body part.
|
|
80
|
+
- **Don't** invoke the invisible-mannequin genre without bounding it to the face. **"an invisible-mannequin presentation where the clothing holds its own shape" removed all the skin** — no neck, no hands, no forearms, a garment floating on nothing — because that *is* the e-commerce genre in full: an empty outfit. **"with just the face cropped out"** keeps the anchor and bounds it. Generalises: a genre anchor imports the whole genre, so name what STAYS, not only what goes.
|
|
81
|
+
- **Don't** put `#` or `@` anywhere in prompt text. Both are reference-token sigils in the desktop prompt composer and an unresolved one is **silently deleted** — no error, no log, just missing words. `#3a3a3c` reached fal as `background ()` on a real 2026-07-30 request, meaning the plate value had never been delivered to any model since the composer shipped. Write hex values bare.
|
|
77
82
|
- **Don't** use 4K — wastes credits, no quality gain at sheet scale.
|
|
78
83
|
- **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `reference-seedance.md`.
|
|
@@ -201,7 +201,18 @@ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding pre
|
|
|
201
201
|
|
|
202
202
|
Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
|
|
203
203
|
|
|
204
|
-
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
|
|
204
|
+
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.
|
|
205
|
+
|
|
206
|
+
### All three modalities go in ONE call
|
|
207
|
+
|
|
208
|
+
The caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.
|
|
209
|
+
|
|
210
|
+
Cite each by type and index, in the order they were attached — `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.
|
|
211
|
+
|
|
212
|
+
```
|
|
213
|
+
Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
|
|
214
|
+
speaking the line in audio 1. Preserve his identity, appearance and outfit.
|
|
215
|
+
```
|
|
205
216
|
|
|
206
217
|
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
207
218
|
|
|
@@ -8,11 +8,11 @@
|
|
|
8
8
|
},
|
|
9
9
|
{
|
|
10
10
|
"path": "skills/slates-character-identity.md",
|
|
11
|
-
"sha256": "
|
|
11
|
+
"sha256": "52191db058ada8096297e97958268548b352a5f81739d1bfd01da340867f2d35"
|
|
12
12
|
},
|
|
13
13
|
{
|
|
14
14
|
"path": "skills/slates-prompting-seedance.md",
|
|
15
|
-
"sha256": "
|
|
15
|
+
"sha256": "81ba0d25c52ef503500d532d9376898400c0a192c63c6ce2af566a5176c92050"
|
|
16
16
|
},
|
|
17
17
|
{
|
|
18
18
|
"path": "skills/slates-prompting-kling-v3.md",
|
|
@@ -28,24 +28,24 @@
|
|
|
28
28
|
},
|
|
29
29
|
{
|
|
30
30
|
"path": "src/prompts/model-facts.ts",
|
|
31
|
-
"sha256": "
|
|
31
|
+
"sha256": "b55054947d0286024f0bf0565968cf5b0a5c5e0d786aa5b2b78aed7e7be9706b"
|
|
32
32
|
}
|
|
33
33
|
],
|
|
34
34
|
"outputs": [
|
|
35
35
|
{
|
|
36
36
|
"path": "SKILL.md",
|
|
37
|
-
"bytes":
|
|
38
|
-
"sha256": "
|
|
37
|
+
"bytes": 4954,
|
|
38
|
+
"sha256": "8d161adfb51a6da1800d9cacd93181e13966807a3490e75bd89a0093d0311cde"
|
|
39
39
|
},
|
|
40
40
|
{
|
|
41
41
|
"path": "reference-character.md",
|
|
42
|
-
"bytes":
|
|
43
|
-
"sha256": "
|
|
42
|
+
"bytes": 10081,
|
|
43
|
+
"sha256": "7dd5b693f3095b6678106c1a41c1303d57b3d7a30da0eeffc867fc5000fd6c74"
|
|
44
44
|
},
|
|
45
45
|
{
|
|
46
46
|
"path": "reference-seedance.md",
|
|
47
|
-
"bytes":
|
|
48
|
-
"sha256": "
|
|
47
|
+
"bytes": 32473,
|
|
48
|
+
"sha256": "0175afc387d27ca47439dfa2a33d6f96394a369044b4539928ff4f36fbf8126d"
|
|
49
49
|
},
|
|
50
50
|
{
|
|
51
51
|
"path": "reference-kling.md",
|
|
@@ -65,8 +65,8 @@
|
|
|
65
65
|
],
|
|
66
66
|
"archive": {
|
|
67
67
|
"path": "slates-prompt-builder.skill",
|
|
68
|
-
"bytes":
|
|
69
|
-
"sha256": "
|
|
68
|
+
"bytes": 38258,
|
|
69
|
+
"sha256": "66c29297ee18dab6f429b48b75f47bb53d8180e35674adf03f3290ee84a7779b",
|
|
70
70
|
"entries": [
|
|
71
71
|
"SKILL.md",
|
|
72
72
|
"reference-character.md",
|
|
Binary file
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@slatesvideo/shared",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.7",
|
|
4
4
|
"description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -28,12 +28,14 @@ Slates generates **one identity sheet per character**, bound as the character's
|
|
|
28
28
|
| Panel | What it carries |
|
|
29
29
|
|---|---|
|
|
30
30
|
| **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
|
|
31
|
-
| **Full-body front, relaxed A-pose —
|
|
31
|
+
| **Full-body front, relaxed A-pose — cropped at the collarbone, just the face cropped out** | Build, proportion, wardrobe. The face is cropped off on purpose: a front-facing body panel renders a ~40px face that can't match the portrait's, so the sheet would carry two competing identities and the model averages them. **Only the face** — neck, arms and hands render as skin. |
|
|
32
32
|
| **Full-body back, head and hair visible** | Hair fall and the back of the outfit — the only panel where either reads. Keeps its head because there's no face to compete with. |
|
|
33
33
|
|
|
34
34
|
The rule is **kill every competing rendering of the FACE, not every head** — which is why exactly one body panel is headless.
|
|
35
35
|
|
|
36
|
-
On a deep neutral-grey plate (
|
|
36
|
+
On a deep neutral-grey plate (hex `3a3a3c`, emitted without the `#` — see the sigil warning in Don'ts), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look. Expression is **a slight natural smile with the teeth just visible** — a closed mouth carries no dental information, so every downstream smiling shot invents teeth, and teeth are person-specific.
|
|
37
|
+
|
|
38
|
+
**Two carve-outs, scoped differently on purpose.** Non-human characters get a natural neutral expression instead of a smile — that one is scoped by *having a human mouth*, so a bipedal robot or humanoid alien is covered. Quadrupeds and non-bipedal characters get a natural standing stance with the head shown on both body panels — that one is *anatomical*. **Both are conditionals the image model evaluates against your reference; neither is a code branch, because the op has no character-kind input.**
|
|
37
39
|
|
|
38
40
|
**The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
|
|
39
41
|
|
|
@@ -69,7 +71,8 @@ If text only: generate from prompt-only — less consistent, so warn the user.
|
|
|
69
71
|
- Default to Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted spend.
|
|
70
72
|
- When the result returns inline, **evaluate it before binding**:
|
|
71
73
|
- Is the portrait clearly the largest panel, and is it off-frontal?
|
|
72
|
-
- **Is the front body panel cleanly headless** — an empty collar
|
|
74
|
+
- **Is the front body panel cleanly headless** — an empty collar above a normally rendered body, no partial face, no floating jaw, no smeared neck stump? A botched crop is worse than no crop.
|
|
75
|
+
- **Is the body still there?** Neck, forearms and hands rendered as skin, not an empty outfit floating on nothing. A hollow garment means the invisible-mannequin genre ran unbounded.
|
|
73
76
|
- Do the body panels read as the same build, wardrobe and hair as the portrait?
|
|
74
77
|
- Catchlights present, irises readable rather than black holes?
|
|
75
78
|
- Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
|
|
@@ -95,6 +98,8 @@ Critically, the app injects **no** wardrobe, expression, or lighting directive.
|
|
|
95
98
|
- **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
|
|
96
99
|
- **Don't** skip binding. An unbound asset doesn't help downstream.
|
|
97
100
|
- **Don't** invent character details. Stick to what's in the reference image and the user's description.
|
|
98
|
-
- **Don't** describe the
|
|
101
|
+
- **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on gpt-image-2** — fal returns `content_policy_violation` with `loc: ["body","prompt"]`, so the text is rejected before any image is read, because an anatomical absence reads as gore to OpenAI's classifier. It passed NB2, which is why the original receipt looked safe: **it was model-scoped.** State an exclusion as a framing choice, never as a missing body part.
|
|
102
|
+
- **Don't** invoke the invisible-mannequin genre without bounding it to the face. **"an invisible-mannequin presentation where the clothing holds its own shape" removed all the skin** — no neck, no hands, no forearms, a garment floating on nothing — because that *is* the e-commerce genre in full: an empty outfit. **"with just the face cropped out"** keeps the anchor and bounds it. Generalises: a genre anchor imports the whole genre, so name what STAYS, not only what goes.
|
|
103
|
+
- **Don't** put `#` or `@` anywhere in prompt text. Both are reference-token sigils in the desktop prompt composer and an unresolved one is **silently deleted** — no error, no log, just missing words. `#3a3a3c` reached fal as `background ()` on a real 2026-07-30 request, meaning the plate value had never been delivered to any model since the composer shipped. Write hex values bare.
|
|
99
104
|
- **Don't** use 4K — wastes credits, no quality gain at sheet scale.
|
|
100
105
|
- **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `slates-prompting-seedance`.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-model-selection
|
|
3
|
-
description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 only) and never the default.
|
|
3
|
+
description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Seedance 2.5 is a SECOND SEAT beside 2.0 (30s takes and 30 references, but 480p/720p only — never an upgrade); Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 only) and never the default.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Model selection — the routing doctrine
|
|
@@ -28,6 +28,7 @@ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni F
|
|
|
28
28
|
| Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
|
|
29
29
|
| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
|
|
30
30
|
| The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
|
|
31
|
+
| **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs — and **480p/720p ONLY, no 1080p and no 4K on any provider**. If resolution matters at all, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) 720p is NOT the cheap seat here — a 30s 720p face gen is 484 credits, more than a 15s 1080p Seedance 2.0 face gen (411), against a 1,000-credit welcome grant. Quote before any take over ~10s. |
|
|
31
32
|
|
|
32
33
|
### Named Seedance escalation triggers
|
|
33
34
|
|
|
@@ -49,26 +50,27 @@ Concrete beats route better than an abstract category. Cost stays a tiebreaker,
|
|
|
49
50
|
| **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
|
|
50
51
|
| **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
|
|
51
52
|
| Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
|
|
52
|
-
|
|
|
53
|
+
| **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |
|
|
54
|
+
| AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
|
|
53
55
|
|
|
54
56
|
- **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
|
|
55
57
|
- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.
|
|
56
58
|
- **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
|
|
57
59
|
- Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
|
|
58
60
|
|
|
59
|
-
## Motion Transfer & Lip Sync routing (
|
|
61
|
+
## Motion Transfer & Lip Sync routing (Kling-only tools)
|
|
60
62
|
|
|
61
|
-
Both tools
|
|
63
|
+
Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that bolts motion or lip movement onto a finished source as a dedicated post-process.
|
|
62
64
|
|
|
63
|
-
| Job |
|
|
65
|
+
| Job | Tool | Why |
|
|
64
66
|
|---|---|---|
|
|
65
|
-
|
|
|
66
|
-
|
|
|
67
|
-
|
|
68
|
-
|
|
67
|
+
| Motion retarget onto a still character | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |
|
|
68
|
+
| Re-voice a clip, or animate a still portrait | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |
|
|
69
|
+
|
|
70
|
+
**Want the Seedance version of either?** It is not a switch on these tools — it is a normal `slates_generate_video` on `seedance-2` with the clip attached as a **video reference** and the motion or dialogue written into the prompt ("the character from image 1 performs the exact motion from video 1"). That routes to the same endpoint the tool would have called, with the prompt visible and editable instead of ghost-written. Single-pass conditioning genuinely beats post-hoc retargeting on fast choreography, contact, cloth and hair — and it carries native audio — so escalate there whenever fidelity matters.
|
|
69
71
|
|
|
70
|
-
-
|
|
71
|
-
-
|
|
72
|
+
- Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane.
|
|
73
|
+
- Faces on that route go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person (premium realface pricing).
|
|
72
74
|
|
|
73
75
|
**Rules:**
|
|
74
76
|
|
|
@@ -91,6 +93,32 @@ Both tools have a cheap Kling utility lane and a premium Seedance lane. The capa
|
|
|
91
93
|
|
|
92
94
|
**Split rule of thumb:** readable text / panels / UI → GPT Image 2; photoreal, character-locked, widescreen, or edit-heavy → the Banana line; drafts → NB2 Lite; uncensored or odd resolutions → Seedream/FLUX.
|
|
93
95
|
|
|
96
|
+
## Audio routing
|
|
97
|
+
|
|
98
|
+
**Image and video models cannot generate standalone audio, and neither audio model can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.
|
|
99
|
+
|
|
100
|
+
| Job | Model | Why |
|
|
101
|
+
|---|---|---|
|
|
102
|
+
| **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. 1–120s. The continuity-bed workhorse and the only speech surface. |
|
|
103
|
+
| **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control (0.5–22s) and a real loop mode. |
|
|
104
|
+
|
|
105
|
+
**There is no music model and no cast-voiceover model.** A song is imported (Slates reads audio files and puts them on the timeline), not generated. A line that has to be spoken is generated on Seed Audio and lip-synced against.
|
|
106
|
+
|
|
107
|
+
### Named audio escalation triggers
|
|
108
|
+
|
|
109
|
+
- **"It needs to sound like a place"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.
|
|
110
|
+
- **"Read this line"** → Seed Audio, with the line in quotes inside the scene sentence. Re-roll until the take is right, then lip-sync against it.
|
|
111
|
+
- **"That needs a thump right there"** → Sound Effects, with the duration set to roughly the length of the event.
|
|
112
|
+
- **"Give it a track"** → there is no music generation. Say so and offer to lay an imported track on an audio track.
|
|
113
|
+
|
|
114
|
+
**Rules:**
|
|
115
|
+
|
|
116
|
+
- **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.
|
|
117
|
+
- **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.
|
|
118
|
+
- **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — and remember those extra seconds are billed on both surfaces.
|
|
119
|
+
- **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.
|
|
120
|
+
- Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`.
|
|
121
|
+
|
|
94
122
|
## Cost is a tiebreaker, not the router
|
|
95
123
|
|
|
96
124
|
Route by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.
|
|
@@ -0,0 +1,69 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-elevenlabs
|
|
3
|
+
description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before calling slates_generate_audio with model eleven-sfx — ONE short effect with an EXACT duration (0.5-22s), or a seamless loop, billed per second. Covers describing an effect by its physical cause, the one-sound-per-generation rule, picking a duration, loops, prompt_influence, and when to use Seed Audio instead.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# ElevenLabs Sound Effects v2 — prompting
|
|
7
|
+
|
|
8
|
+
One short sound with an exact length, carried on fal (`fal-ai/elevenlabs/sound-effects/v2`). This is the only Slates audio surface with a real duration control and a real loop mode.
|
|
9
|
+
|
|
10
|
+
## Where it routes
|
|
11
|
+
|
|
12
|
+
- **A single hit that has to land on a known frame** — door slam, whoosh, impact, UI blip, riser.
|
|
13
|
+
- **A seamless loop** you can lay under a whole scene — rain, engine hum, crowd murmur, machine noise.
|
|
14
|
+
- **NOT** layered scenes. A room with dialogue *and* clatter *and* ambience is one `seed-audio` pass, not three SFX generations.
|
|
15
|
+
- **NOT** speech. Dialogue, narration and scratch VO are `seed-audio` — it casts and performs the line inside the scene.
|
|
16
|
+
- **AUDIO-ONLY.** It cannot produce images or video.
|
|
17
|
+
|
|
18
|
+
## THE RULES
|
|
19
|
+
|
|
20
|
+
### 1. Describe the physical CAUSE, not the label
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
✗ door sound
|
|
24
|
+
✓ heavy oak door slams shut in a stone hallway
|
|
25
|
+
|
|
26
|
+
✗ whoosh
|
|
27
|
+
✓ a thick rope swung fast past a microphone, low air displacement
|
|
28
|
+
|
|
29
|
+
✗ footsteps
|
|
30
|
+
✓ boots on wet gravel, slow, one person
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
Material + weight + surface + room. Naming all four is the difference between a usable effect and a stock-library shrug. Cap is 450 characters — you will not need them.
|
|
34
|
+
|
|
35
|
+
### 2. One sound per generation
|
|
36
|
+
|
|
37
|
+
This surface makes a single event. A door, then footsteps, then a siren is three generations layered on the timeline — or one `seed-audio` scene, which is usually cheaper and always more coherent.
|
|
38
|
+
|
|
39
|
+
### 3. Duration is always explicit, and it is the price
|
|
40
|
+
|
|
41
|
+
`durationSeconds` is 0.5–22 and Slates **always sends it**. (Left null the model picks, which makes the charge non-deterministic — so it is never left null.)
|
|
42
|
+
|
|
43
|
+
| Kind of sound | Ask for |
|
|
44
|
+
|---|---|
|
|
45
|
+
| impact, hit, click | 0.5–1s |
|
|
46
|
+
| whoosh, riser, transition | 2–4s |
|
|
47
|
+
| loopable bed | 8–22s + `loop: true` |
|
|
48
|
+
|
|
49
|
+
Over-asking pads the tail with room tone you then trim. Under-asking clips the decay.
|
|
50
|
+
|
|
51
|
+
### 4. Loops
|
|
52
|
+
|
|
53
|
+
`loop: true` tiles without a seam — rain, engine hum, crowd murmur, machine noise. Combine with a longer duration so the loop point is not obvious.
|
|
54
|
+
|
|
55
|
+
For a bed longer than 22s, this is the wrong surface: `seed-audio` runs to 120s in one pass.
|
|
56
|
+
|
|
57
|
+
### 5. Prompt influence
|
|
58
|
+
|
|
59
|
+
`promptInfluence` 0–1, default 0.3. Higher hugs your wording with less variation between takes; lower explores. Raise it when a re-roll keeps wandering off the brief; lower it when every take sounds like the same take.
|
|
60
|
+
|
|
61
|
+
## Iterating
|
|
62
|
+
|
|
63
|
+
- Re-rolls that keep missing = the prompt named a **label** instead of a **cause**. Rewrite it as a physical event.
|
|
64
|
+
- A hit that lands but sounds wrong in the scene is usually a *room* problem — name the space ("in a stone hallway", "in a padded studio", "outdoors, no reflections").
|
|
65
|
+
- Three failed takes means the prompt is wrong, not the seed.
|
|
66
|
+
|
|
67
|
+
## Content notes
|
|
68
|
+
|
|
69
|
+
ElevenLabs applies its own moderation. See slates-content-policy.
|
|
@@ -1,33 +1,31 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-lip-sync
|
|
3
|
-
description: How to set up lip-sync — Kling (
|
|
3
|
+
description: How to set up lip-sync — Kling-only (dedicated lip-sync and avatar endpoints, 5-second outputs). Read before calling slates_generate_lip_sync. Two flows — video→video re-dub and image→video avatar — with different inputs, pricing, and gotchas. Voice catalog, framing rules, audio file constraints, and which tier to pick. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Lip-sync — setup guide
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
**This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; every entry is a real endpoint and every output is 5 seconds.
|
|
9
9
|
|
|
10
|
-
| Flow | Source |
|
|
10
|
+
| Flow | Source | Model | Cost | Use case |
|
|
11
11
|
|------|--------|-------|-----------|----------|
|
|
12
12
|
| Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |
|
|
13
13
|
| Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |
|
|
14
14
|
| Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |
|
|
15
|
-
| **Seedance native** | image or video | `engine=seedance-2` | per second (`seedance-2-face-*`; video sources bill input+output seconds) | **Premium**: natural delivery, whole-body performance, voice cloned from a video source, audio included |
|
|
16
15
|
|
|
17
|
-
Pick
|
|
16
|
+
Pick `sourceType` deliberately — it decides the pricing tier and the underlying endpoint.
|
|
18
17
|
|
|
19
|
-
## Seedance
|
|
18
|
+
## Want Seedance instead? That is a video generation, not a mode here
|
|
20
19
|
|
|
21
|
-
|
|
20
|
+
Seedance can generate the performance rather than bolting a mouth onto finished pixels — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal, and a video source keeps its own voice. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the clip (or portrait) attached as a video/ingredient reference and the dialogue written into the prompt yourself.
|
|
22
21
|
|
|
23
|
-
|
|
24
|
-
- **`audioMethod=upload`** drives the speech from a ≤15s audio file instead (a reference-audio input, no billing surcharge).
|
|
25
|
-
- **Sources:** image (any style; same framing rules as the avatar flow below) or a 2–15s video clip. Output duration follows the source/audio/line length (4–15s), not a fixed 5s.
|
|
26
|
-
- **Billing:** image sources bill the normal `seedance-2-face-{res}-{N}s` keys; video sources bill combined input+output seconds (`-vref-` keys) — pass `sourceSeconds` and quote via the confirm gate.
|
|
27
|
-
- **Faces:** `seedanceFace` defaults true. A REAL person → `[REAL_FACE_DETECTED]` → confirm consent → retry with `seedanceRealFace=true, realFaceConsent=true` (premium realface pricing).
|
|
28
|
-
- When to pick it: hero dialogue shots, natural delivery, "make this clip's person say X in their own voice". Stay on Kling for cheap utility re-dubs and long clips.
|
|
22
|
+
That is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a "video 1" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.
|
|
29
23
|
|
|
30
|
-
|
|
24
|
+
- Driving clips must be 2–15s; output duration is whatever you set (4–15s).
|
|
25
|
+
- Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming.
|
|
26
|
+
- Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.
|
|
27
|
+
|
|
28
|
+
Everything below is about the Kling tool.
|
|
31
29
|
|
|
32
30
|
## Choosing video vs avatar
|
|
33
31
|
|
|
@@ -1,32 +1,36 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-motion-transfer
|
|
3
|
-
description: How to set up motion transfer — Kling Motion Control (
|
|
3
|
+
description: How to set up motion transfer — Kling Motion Control only (std and pro tiers, 5-second outputs). Read before calling slates_generate_motion_transfer. Reference image (character) + driving video (motion source) → new video of the character performing the motion. Asset selection rules, character_orientation, tiers, and prompt usage. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Motion transfer — setup guide
|
|
7
7
|
|
|
8
|
-
Take a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion.
|
|
8
|
+
Take a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion. **This tool is Kling-only** — it wraps Kling Motion Control and nothing else.
|
|
9
9
|
|
|
10
|
-
|
|
|
10
|
+
| Tier | Cost | Use case |
|
|
11
11
|
|------|-----------|----------|
|
|
12
12
|
| Kling std (`kling-mc-std-5s`) | ~32 credits / 5s | General motion transfer, budget lane |
|
|
13
13
|
| Kling pro (`kling-mc-pro-5s`) | ~42 credits / 5s | Cleaner anatomy, better identity preservation |
|
|
14
|
-
| **Seedance 2.0** (`motionModel=seedance-2`) | per second of input+output (`seedance-2-face-vref-*`) | **Premium lane** — single-pass generation with the driving clip as a native conditioning signal: better motion fidelity, native audio, prompt-directed |
|
|
15
14
|
|
|
16
|
-
|
|
15
|
+
Both tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)
|
|
17
16
|
|
|
18
|
-
## Seedance
|
|
17
|
+
## Want Seedance instead? That is a video generation, not a mode here
|
|
19
18
|
|
|
20
|
-
Kling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion.
|
|
19
|
+
Kling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion, and the output carries native audio. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the driving clip attached as a video reference and the character image as an ingredient, then write the prompt yourself:
|
|
21
20
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
- **Native audio included** — the output can carry sound from the prompt (or the driving clip's vibe); no audio surcharge.
|
|
27
|
-
- `characterOrientation` is Kling-only; Seedance framing follows the prompt + `aspectRatio`.
|
|
21
|
+
```
|
|
22
|
+
The character from image 1 performs the exact motion, choreography, and camera
|
|
23
|
+
movement from video 1. Preserve the character's identity, appearance, and outfit.
|
|
24
|
+
```
|
|
28
25
|
|
|
29
|
-
|
|
26
|
+
That is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.
|
|
27
|
+
|
|
28
|
+
- **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).
|
|
29
|
+
- **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.
|
|
30
|
+
- **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
|
|
31
|
+
- `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.
|
|
32
|
+
|
|
33
|
+
Everything below is about the Kling tool.
|
|
30
34
|
|
|
31
35
|
## Inputs
|
|
32
36
|
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-seed-audio
|
|
3
|
+
description: How to prompt Seed Audio 1.0 (ByteDance, via fal). Read before calling slates_generate_audio with model seed-audio. The one-pass audio SCENE model — dialogue, SFX and ambience together from ONE plain sentence. CRITICAL - it has NO duration parameter, so length must be named IN THE PROMPT TEXT and Slates bills the duration you request. Covers the one-sentence doctrine, the crowd-size rule, why Kling "SFX:" syntax hurts here, and the audio-refs-XOR-image input rule.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Seed Audio 1.0 — prompting
|
|
7
|
+
|
|
8
|
+
ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence. 1–120 seconds. It is the default audio model in Slates and the workhorse for continuity beds.
|
|
9
|
+
|
|
10
|
+
## Where it routes
|
|
11
|
+
|
|
12
|
+
- **Scene audio, room tone, ambience beds, crowd/nature soundscapes** — anything where several sounds share a space. One generation, not three layered ones.
|
|
13
|
+
- **Dialogue and scratch VO.** This is the only speech surface in Slates: the line is performed inside the scene. Lock the read by re-rolling until a take is right, then lip-sync against it with `slates_generate_lip_sync`.
|
|
14
|
+
- **NOT** a single effect that must land on a known frame — that is `eleven-sfx`, which takes an exact duration.
|
|
15
|
+
- **NOT** music. Slates has no music model; import a track and drop it on an audio track.
|
|
16
|
+
- **AUDIO-ONLY.** It cannot produce images or video.
|
|
17
|
+
|
|
18
|
+
## THE RULES
|
|
19
|
+
|
|
20
|
+
### 1. 🚨 There is no duration parameter — the words set the length
|
|
21
|
+
|
|
22
|
+
This is the single most important fact about this model. Output length is driven by the prompt text ("… 15 seconds"), capped at 120s.
|
|
23
|
+
|
|
24
|
+
**Slates handles this for you:** the `durationSeconds` param appends the duration to the prompt and **bills that number of seconds**. So:
|
|
25
|
+
|
|
26
|
+
- Set `durationSeconds` to what you actually want.
|
|
27
|
+
- **Do not also write a different length into your sentence.** Two numbers fight, and you pay for the one you selected, not the one you got.
|
|
28
|
+
- If the returned clip is shorter than requested you still paid for the request — that is the deal that keeps the displayed price equal to the charge. Ask for what you need.
|
|
29
|
+
|
|
30
|
+
<!-- slates-only -->
|
|
31
|
+
The server re-derives the billed key from `durationSeconds` (a client cannot under-bill), probes the returned `audio.duration` after completion, and logs `SEED AUDIO BILLING DRIFT` if the model overshot. No auto-charge, no refund — the request is the contract.
|
|
32
|
+
<!-- /slates-only -->
|
|
33
|
+
|
|
34
|
+
### 2. One plain sentence. No production jargon.
|
|
35
|
+
|
|
36
|
+
Field-proven (Higgsfield sprint, 2026-07-27/28). Working prompts look like this:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
tiny applause of 2 or 3 people at an open mic. 15 seconds
|
|
40
|
+
nature soundscape, wide open field cicadas and birds and a loon.
|
|
41
|
+
a diner at 2am, one coffee machine hissing, cutlery somewhere behind the counter
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Not this:
|
|
45
|
+
|
|
46
|
+
```
|
|
47
|
+
✗ AMBIENCE: interior diner, night. SFX: espresso machine (hiss, 2s), cutlery.
|
|
48
|
+
✗ Wide shot of a diner. Slow push in. Warm tungsten. Ambient noise: ...
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Shot language, camera moves and lighting belong to video prompts. Here they are just words the model has to ignore.
|
|
52
|
+
|
|
53
|
+
### 3. Never bring Kling's audio syntax to this model
|
|
54
|
+
|
|
55
|
+
`SFX:` and `Ambient noise:` prefixes and `Background music:` labels are **Kling 3.0 video** syntax. Seed Audio has no parser for them — it reads them as text in the scene and the output gets measurably worse. Describe the sounds directly instead.
|
|
56
|
+
|
|
57
|
+
### 4. Name the crowd size, the room size, the distance
|
|
58
|
+
|
|
59
|
+
The highest-leverage single edit on any bed. Unqualified nouns default big:
|
|
60
|
+
|
|
61
|
+
| Vague | What it returns | Fixed |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| `applause` | a full auditorium | `tiny applause of 2 or 3 people` |
|
|
64
|
+
| `traffic` | a highway | `one car passing on a wet residential street` |
|
|
65
|
+
| `crowd` | a stadium | `four people talking at the next table` |
|
|
66
|
+
|
|
67
|
+
Distance words (`far off`, `muffled through a wall`, `right next to the mic`) work the same way and are how you build depth in one sentence.
|
|
68
|
+
|
|
69
|
+
### 5. Beds must outlast the cut
|
|
70
|
+
|
|
71
|
+
Ask for a few seconds more than the clip needs so the edit has handles to fade through. A bed that ends exactly on the cut always sounds clipped. This is a product requirement, not a preference — it is why the duration control exists at all.
|
|
72
|
+
|
|
73
|
+
### 6. Dialogue goes in quotes, inside the same sentence as the room
|
|
74
|
+
|
|
75
|
+
```
|
|
76
|
+
a tired bartender says, "we closed twenty minutes ago", glasses clinking behind him
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
Pick a preset voice when a specific speaker matters. Leave `voice` unset and the scene casts itself — which is usually right for crowd and background dialogue.
|
|
80
|
+
|
|
81
|
+
Preset voices (20): `vivi_mixed_en_zh_ja_es_id`, `mindy_en_es_id_pt_zh`, `kian_en_zh`, `cedric_en_zh`, `sophie_en_zh`, `jean_en_zh`, `magnus_en_zh`, `mabel_en_zh`, `nadia_en_zh`, `opal_en_zh`, `pearl_en_zh`, `quentin_en_zh`, `corinne_mixed_en_zh`, `esther_mixed_en_zh`, `lyla_mixed_en_zh`, `tracy_es_zh`, `sandy_es_mixed_en_zh`, `felix_zh`, `celeste_zh`, `monkey_king_zh`.
|
|
82
|
+
|
|
83
|
+
Set `multilingual: true` for non-English or mixed-language lines.
|
|
84
|
+
|
|
85
|
+
### 7. Inputs: up to 3 audio clips **XOR** one image. Never both.
|
|
86
|
+
|
|
87
|
+
- **Audio references** — up to 3 clips, each ≤30s and ≤10MB (wav/mp3/pcm/ogg_opus). Refer to them in the prompt as `@Audio1`, `@Audio2`, `@Audio3`: *"match the room tone of @Audio1"*.
|
|
88
|
+
- **Image reference** — one image (jpeg/png/webp ≤10MB). The model scores what it sees.
|
|
89
|
+
- Sending both is rejected by the API. Pick the one that carries the intent.
|
|
90
|
+
|
|
91
|
+
### 8. The knobs, and when to touch them
|
|
92
|
+
|
|
93
|
+
| Param | Range | Reach for it when |
|
|
94
|
+
|---|---|---|
|
|
95
|
+
| `speed` | 0.5–2.0 | Dialogue is racing or dragging against picture. |
|
|
96
|
+
| `volume` | 0.5–2.0 | Rarely — normalize on the timeline instead. |
|
|
97
|
+
| `pitch` | −12…+12 semitones | Ageing or shifting a voice. Small moves only; ±3 is already a lot. |
|
|
98
|
+
| `multilingual` | bool | Non-English or code-switched lines. |
|
|
99
|
+
| `sampleRate` | 8k–48k | Leave at 24000 unless you are matching an existing stem. |
|
|
100
|
+
| `outputFormat` | mp3 / wav / pcm / ogg_opus | wav when this is going into a mix; mp3 otherwise. |
|
|
101
|
+
|
|
102
|
+
## Iterating
|
|
103
|
+
|
|
104
|
+
- A bed that came back wrong is almost always a **scale** problem (crowd/room too big) or a **jargon** problem (the sentence reads like a spec). Fix those two before touching `speed`/`pitch`.
|
|
105
|
+
- Three failed takes on the same sentence means the sentence is wrong, not the seed. Rewrite it the way you would say it out loud.
|
|
106
|
+
- Generations are cheap enough at short durations that auditioning two phrasings beats agonizing over one.
|
|
107
|
+
|
|
108
|
+
## Content notes
|
|
109
|
+
|
|
110
|
+
Provider-side moderation applies to voices and to recognizable real people. See slates-content-policy.
|