@slatesvideo/shared 0.5.6 → 0.5.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.d.ts +1 -1
- package/dist/index.js +4 -1
- package/dist/operations/index.d.ts +43 -36
- package/dist/operations/index.js +289 -262
- package/dist/prompts/model-facts.d.ts +42 -0
- package/dist/prompts/model-facts.js +88 -18
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +108 -99
- package/dist/prompts/reference-composer.d.ts +21 -3
- package/dist/prompts/reference-composer.js +80 -10
- package/dist/skills/content.js +7 -7
- package/exports/slates-prompt-builder/generated/SKILL.md +2 -2
- package/exports/slates-prompt-builder/generated/reference-seedance.md +12 -1
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +8 -8
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +1 -1
- package/skills/slates-model-selection.md +21 -19
- package/skills/slates-prompting-elevenlabs.md +16 -78
- package/skills/slates-prompting-lip-sync.md +12 -14
- package/skills/slates-prompting-motion-transfer.md +18 -14
- package/skills/slates-prompting-seed-audio.md +2 -2
- package/skills/slates-prompting-seedance-2-5.md +215 -0
- package/skills/slates-prompting-seedance.md +17 -1
- package/skills/slates-prompting-suno.md +0 -110
|
@@ -16,8 +16,8 @@ This portable skill is deliberately thin. Its reference files are generated dire
|
|
|
16
16
|
<!-- @generated:model-routing -->
|
|
17
17
|
| Model | Canonical route | Guide |
|
|
18
18
|
|---|---|---|
|
|
19
|
-
| **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics.
|
|
20
|
-
| **Seedance 2.0** | PREMIUM video tier — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those).
|
|
19
|
+
| **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics. Kling is also the ONLY engine behind the Motion Transfer and Lip Sync tools (MC std/pro, lip-sync, avatar) — those two tools are Kling-only. | `reference-kling.md` |
|
|
20
|
+
| **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those). 4-15s, up to 9 ingredient images. Strong I2V / own-footage restyle. Native 4K, but 4K VIDEO is a Pro-only tier gate (base maxes at 1080p; server returns PRO_REQUIRED) — default 1080p unless the user is on Pro. Attaching a clip as a video reference (own-footage restyle, motion or dialogue conditioning) bills combined input+output seconds. 2.0 STAYS THE DEFAULT over 2.5 because it is the only Seedance with 1080p and 4K. | `reference-seedance.md` |
|
|
21
21
|
| **Nano Banana 2 (Gemini 3.1 Flash Image)** | Default image model. 14 refs hard cap (10 object + 4 character). Brief it like a creative director, not tag soup. No negativePrompt field — use positive reframing. Best image start-frame for legible text. Knowledge cutoff Jan 2025. | `reference-nano-banana.md` |
|
|
22
22
|
<!-- @end:model-routing -->
|
|
23
23
|
|
|
@@ -201,7 +201,18 @@ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding pre
|
|
|
201
201
|
|
|
202
202
|
Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
|
|
203
203
|
|
|
204
|
-
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
|
|
204
|
+
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.
|
|
205
|
+
|
|
206
|
+
### All three modalities go in ONE call
|
|
207
|
+
|
|
208
|
+
The caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.
|
|
209
|
+
|
|
210
|
+
Cite each by type and index, in the order they were attached — `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.
|
|
211
|
+
|
|
212
|
+
```
|
|
213
|
+
Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
|
|
214
|
+
speaking the line in audio 1. Preserve his identity, appearance and outfit.
|
|
215
|
+
```
|
|
205
216
|
|
|
206
217
|
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
207
218
|
|
|
@@ -12,7 +12,7 @@
|
|
|
12
12
|
},
|
|
13
13
|
{
|
|
14
14
|
"path": "skills/slates-prompting-seedance.md",
|
|
15
|
-
"sha256": "
|
|
15
|
+
"sha256": "81ba0d25c52ef503500d532d9376898400c0a192c63c6ce2af566a5176c92050"
|
|
16
16
|
},
|
|
17
17
|
{
|
|
18
18
|
"path": "skills/slates-prompting-kling-v3.md",
|
|
@@ -28,14 +28,14 @@
|
|
|
28
28
|
},
|
|
29
29
|
{
|
|
30
30
|
"path": "src/prompts/model-facts.ts",
|
|
31
|
-
"sha256": "
|
|
31
|
+
"sha256": "b55054947d0286024f0bf0565968cf5b0a5c5e0d786aa5b2b78aed7e7be9706b"
|
|
32
32
|
}
|
|
33
33
|
],
|
|
34
34
|
"outputs": [
|
|
35
35
|
{
|
|
36
36
|
"path": "SKILL.md",
|
|
37
|
-
"bytes":
|
|
38
|
-
"sha256": "
|
|
37
|
+
"bytes": 4954,
|
|
38
|
+
"sha256": "8d161adfb51a6da1800d9cacd93181e13966807a3490e75bd89a0093d0311cde"
|
|
39
39
|
},
|
|
40
40
|
{
|
|
41
41
|
"path": "reference-character.md",
|
|
@@ -44,8 +44,8 @@
|
|
|
44
44
|
},
|
|
45
45
|
{
|
|
46
46
|
"path": "reference-seedance.md",
|
|
47
|
-
"bytes":
|
|
48
|
-
"sha256": "
|
|
47
|
+
"bytes": 32473,
|
|
48
|
+
"sha256": "0175afc387d27ca47439dfa2a33d6f96394a369044b4539928ff4f36fbf8126d"
|
|
49
49
|
},
|
|
50
50
|
{
|
|
51
51
|
"path": "reference-kling.md",
|
|
@@ -65,8 +65,8 @@
|
|
|
65
65
|
],
|
|
66
66
|
"archive": {
|
|
67
67
|
"path": "slates-prompt-builder.skill",
|
|
68
|
-
"bytes":
|
|
69
|
-
"sha256": "
|
|
68
|
+
"bytes": 38258,
|
|
69
|
+
"sha256": "66c29297ee18dab6f429b48b75f47bb53d8180e35674adf03f3290ee84a7779b",
|
|
70
70
|
"entries": [
|
|
71
71
|
"SKILL.md",
|
|
72
72
|
"reference-character.md",
|
|
Binary file
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@slatesvideo/shared",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.7",
|
|
4
4
|
"description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-model-selection
|
|
3
|
-
description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 only) and never the default.
|
|
3
|
+
description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Seedance 2.5 is a SECOND SEAT beside 2.0 (30s takes and 30 references, but 480p/720p only — never an upgrade); Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 only) and never the default.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Model selection — the routing doctrine
|
|
@@ -28,6 +28,7 @@ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni F
|
|
|
28
28
|
| Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
|
|
29
29
|
| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
|
|
30
30
|
| The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
|
|
31
|
+
| **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs — and **480p/720p ONLY, no 1080p and no 4K on any provider**. If resolution matters at all, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) 720p is NOT the cheap seat here — a 30s 720p face gen is 484 credits, more than a 15s 1080p Seedance 2.0 face gen (411), against a 1,000-credit welcome grant. Quote before any take over ~10s. |
|
|
31
32
|
|
|
32
33
|
### Named Seedance escalation triggers
|
|
33
34
|
|
|
@@ -49,26 +50,27 @@ Concrete beats route better than an abstract category. Cost stays a tiebreaker,
|
|
|
49
50
|
| **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
|
|
50
51
|
| **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
|
|
51
52
|
| Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
|
|
52
|
-
|
|
|
53
|
+
| **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |
|
|
54
|
+
| AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
|
|
53
55
|
|
|
54
56
|
- **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
|
|
55
57
|
- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.
|
|
56
58
|
- **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
|
|
57
59
|
- Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
|
|
58
60
|
|
|
59
|
-
## Motion Transfer & Lip Sync routing (
|
|
61
|
+
## Motion Transfer & Lip Sync routing (Kling-only tools)
|
|
60
62
|
|
|
61
|
-
Both tools
|
|
63
|
+
Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that bolts motion or lip movement onto a finished source as a dedicated post-process.
|
|
62
64
|
|
|
63
|
-
| Job |
|
|
65
|
+
| Job | Tool | Why |
|
|
64
66
|
|---|---|---|
|
|
65
|
-
|
|
|
66
|
-
|
|
|
67
|
-
|
|
68
|
-
|
|
67
|
+
| Motion retarget onto a still character | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |
|
|
68
|
+
| Re-voice a clip, or animate a still portrait | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |
|
|
69
|
+
|
|
70
|
+
**Want the Seedance version of either?** It is not a switch on these tools — it is a normal `slates_generate_video` on `seedance-2` with the clip attached as a **video reference** and the motion or dialogue written into the prompt ("the character from image 1 performs the exact motion from video 1"). That routes to the same endpoint the tool would have called, with the prompt visible and editable instead of ghost-written. Single-pass conditioning genuinely beats post-hoc retargeting on fast choreography, contact, cloth and hair — and it carries native audio — so escalate there whenever fidelity matters.
|
|
69
71
|
|
|
70
|
-
-
|
|
71
|
-
-
|
|
72
|
+
- Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane.
|
|
73
|
+
- Faces on that route go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person (premium realface pricing).
|
|
72
74
|
|
|
73
75
|
**Rules:**
|
|
74
76
|
|
|
@@ -93,29 +95,29 @@ Both tools have a cheap Kling utility lane and a premium Seedance lane. The capa
|
|
|
93
95
|
|
|
94
96
|
## Audio routing
|
|
95
97
|
|
|
96
|
-
**Image and video models cannot generate standalone audio, and
|
|
98
|
+
**Image and video models cannot generate standalone audio, and neither audio model can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.
|
|
97
99
|
|
|
98
100
|
| Job | Model | Why |
|
|
99
101
|
|---|---|---|
|
|
100
|
-
| **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. 1–120s. The continuity-bed workhorse. |
|
|
101
|
-
| **The exact words, in a repeatable named voice** — ad reads, narration, character lines to lip-sync against | **Eleven v3** (`eleven-v3`) | Verbatim text, 20 preset voices, re-renderable after a copy tweak without the performance drifting. Billed per 100 characters. |
|
|
102
|
+
| **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. 1–120s. The continuity-bed workhorse and the only speech surface. |
|
|
102
103
|
| **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control (0.5–22s) and a real loop mode. |
|
|
103
|
-
|
|
104
|
+
|
|
105
|
+
**There is no music model and no cast-voiceover model.** A song is imported (Slates reads audio files and puts them on the timeline), not generated. A line that has to be spoken is generated on Seed Audio and lip-synced against.
|
|
104
106
|
|
|
105
107
|
### Named audio escalation triggers
|
|
106
108
|
|
|
107
109
|
- **"It needs to sound like a place"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.
|
|
108
|
-
- **"Read this line"**
|
|
110
|
+
- **"Read this line"** → Seed Audio, with the line in quotes inside the scene sentence. Re-roll until the take is right, then lip-sync against it.
|
|
109
111
|
- **"That needs a thump right there"** → Sound Effects, with the duration set to roughly the length of the event.
|
|
110
|
-
- **"Give it a track"** →
|
|
112
|
+
- **"Give it a track"** → there is no music generation. Say so and offer to lay an imported track on an audio track.
|
|
111
113
|
|
|
112
114
|
**Rules:**
|
|
113
115
|
|
|
114
116
|
- **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.
|
|
115
117
|
- **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.
|
|
116
|
-
- **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles —
|
|
118
|
+
- **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — and remember those extra seconds are billed on both surfaces.
|
|
117
119
|
- **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.
|
|
118
|
-
- Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs
|
|
120
|
+
- Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`.
|
|
119
121
|
|
|
120
122
|
## Cost is a tiebreaker, not the router
|
|
121
123
|
|
|
@@ -1,83 +1,21 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-elevenlabs
|
|
3
|
-
description: How to prompt
|
|
3
|
+
description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before calling slates_generate_audio with model eleven-sfx — ONE short effect with an EXACT duration (0.5-22s), or a seamless loop, billed per second. Covers describing an effect by its physical cause, the one-sound-per-generation rule, picking a duration, loops, prompt_influence, and when to use Seed Audio instead.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# ElevenLabs
|
|
6
|
+
# ElevenLabs Sound Effects v2 — prompting
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
One short sound with an exact length, carried on fal (`fal-ai/elevenlabs/sound-effects/v2`). This is the only Slates audio surface with a real duration control and a real loop mode.
|
|
9
9
|
|
|
10
|
-
## Where
|
|
10
|
+
## Where it routes
|
|
11
11
|
|
|
12
|
-
-
|
|
13
|
-
-
|
|
14
|
-
- **
|
|
15
|
-
- **
|
|
12
|
+
- **A single hit that has to land on a known frame** — door slam, whoosh, impact, UI blip, riser.
|
|
13
|
+
- **A seamless loop** you can lay under a whole scene — rain, engine hum, crowd murmur, machine noise.
|
|
14
|
+
- **NOT** layered scenes. A room with dialogue *and* clatter *and* ambience is one `seed-audio` pass, not three SFX generations.
|
|
15
|
+
- **NOT** speech. Dialogue, narration and scratch VO are `seed-audio` — it casts and performs the line inside the scene.
|
|
16
|
+
- **AUDIO-ONLY.** It cannot produce images or video.
|
|
16
17
|
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
## Eleven v3 (`eleven-v3`) — THE RULES
|
|
20
|
-
|
|
21
|
-
### 1. 🚨 The text field is the script. Every character is spoken.
|
|
22
|
-
|
|
23
|
-
```
|
|
24
|
-
✗ (excited) Read this fast — "Grab yours today!"
|
|
25
|
-
✓ Grab yours today!
|
|
26
|
-
```
|
|
27
|
-
|
|
28
|
-
Stage directions, speaker names, bracketed emotion tags and markdown all get read out loud. There is no instruction channel — direction lives in `stability` and in how you punctuate.
|
|
29
|
-
|
|
30
|
-
### 2. Punctuation is the only timing control
|
|
31
|
-
|
|
32
|
-
| You want | Write |
|
|
33
|
-
|---|---|
|
|
34
|
-
| a hard stop | `It works. Every time.` |
|
|
35
|
-
| a beat, not a stop | `It works — every time.` |
|
|
36
|
-
| a trailing hesitation | `It works… mostly.` |
|
|
37
|
-
| a list rhythm | `Faster, cheaper, and yours.` |
|
|
38
|
-
|
|
39
|
-
Rewrite the punctuation before you touch a setting. It moves the read more than `stability` does.
|
|
40
|
-
|
|
41
|
-
### 3. Pick a voice and keep it
|
|
42
|
-
|
|
43
|
-
20 presets: Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill. Default `Rachel`.
|
|
44
|
-
|
|
45
|
-
One voice per character, one per piece. A series that swaps voices between shots reads as an accident. There is **no voice cloning on this route** — a cloned ElevenLabs voice ID is not supported and must not be assumed to pass through.
|
|
46
|
-
|
|
47
|
-
### 4. Stability
|
|
48
|
-
|
|
49
|
-
| Value | Behavior | Use for |
|
|
50
|
-
|---|---|---|
|
|
51
|
-
| ~0.3 | expressive, varies take-to-take | one dramatic line, character dialogue |
|
|
52
|
-
| 0.5 (default) | balanced | most reads |
|
|
53
|
-
| ~0.8 | flat, highly repeatable | long narration, anything you will re-render |
|
|
54
|
-
|
|
55
|
-
Raise it when re-rolls keep giving you a different performance. Lower it when the read is lifeless.
|
|
56
|
-
|
|
57
|
-
### 5. Spell out what TTS gets wrong
|
|
58
|
-
|
|
59
|
-
Acronyms, product names, prices, years and URLs are where it embarrasses itself. Write the pronunciation:
|
|
60
|
-
|
|
61
|
-
```
|
|
62
|
-
SKU → "ess kay you"
|
|
63
|
-
2026 → "twenty twenty six"
|
|
64
|
-
$19.99 → "nineteen ninety nine"
|
|
65
|
-
slates.video → "slates dot video"
|
|
66
|
-
```
|
|
67
|
-
|
|
68
|
-
Set `languageCode` (ISO 639-1) to force a language when the text is ambiguous or code-switched.
|
|
69
|
-
|
|
70
|
-
### 6. Length is money, honestly
|
|
71
|
-
|
|
72
|
-
1–5000 characters, billed in 100-character buckets rounded up. A tightened sentence costs less; a pasted stray paragraph costs more. This is the one Slates surface where editing the copy is also a cost control.
|
|
73
|
-
|
|
74
|
-
### 7. Timestamps are free — leave them on
|
|
75
|
-
|
|
76
|
-
Word-level timestamps come back with every generation at no extra charge. They are exactly what a caption/subtitle pass consumes. There is no reason to disable them.
|
|
77
|
-
|
|
78
|
-
---
|
|
79
|
-
|
|
80
|
-
## Sound Effects v2 (`eleven-sfx`) — THE RULES
|
|
18
|
+
## THE RULES
|
|
81
19
|
|
|
82
20
|
### 1. Describe the physical CAUSE, not the label
|
|
83
21
|
|
|
@@ -114,18 +52,18 @@ Over-asking pads the tail with room tone you then trim. Under-asking clips the d
|
|
|
114
52
|
|
|
115
53
|
`loop: true` tiles without a seam — rain, engine hum, crowd murmur, machine noise. Combine with a longer duration so the loop point is not obvious.
|
|
116
54
|
|
|
55
|
+
For a bed longer than 22s, this is the wrong surface: `seed-audio` runs to 120s in one pass.
|
|
56
|
+
|
|
117
57
|
### 5. Prompt influence
|
|
118
58
|
|
|
119
59
|
`promptInfluence` 0–1, default 0.3. Higher hugs your wording with less variation between takes; lower explores. Raise it when a re-roll keeps wandering off the brief; lower it when every take sounds like the same take.
|
|
120
60
|
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
## Iterating on either surface
|
|
61
|
+
## Iterating
|
|
124
62
|
|
|
125
|
-
-
|
|
126
|
-
-
|
|
63
|
+
- Re-rolls that keep missing = the prompt named a **label** instead of a **cause**. Rewrite it as a physical event.
|
|
64
|
+
- A hit that lands but sounds wrong in the scene is usually a *room* problem — name the space ("in a stone hallway", "in a padded studio", "outdoors, no reflections").
|
|
127
65
|
- Three failed takes means the prompt is wrong, not the seed.
|
|
128
66
|
|
|
129
67
|
## Content notes
|
|
130
68
|
|
|
131
|
-
ElevenLabs applies its own moderation
|
|
69
|
+
ElevenLabs applies its own moderation. See slates-content-policy.
|
|
@@ -1,33 +1,31 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-lip-sync
|
|
3
|
-
description: How to set up lip-sync — Kling (
|
|
3
|
+
description: How to set up lip-sync — Kling-only (dedicated lip-sync and avatar endpoints, 5-second outputs). Read before calling slates_generate_lip_sync. Two flows — video→video re-dub and image→video avatar — with different inputs, pricing, and gotchas. Voice catalog, framing rules, audio file constraints, and which tier to pick. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Lip-sync — setup guide
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
**This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; every entry is a real endpoint and every output is 5 seconds.
|
|
9
9
|
|
|
10
|
-
| Flow | Source |
|
|
10
|
+
| Flow | Source | Model | Cost | Use case |
|
|
11
11
|
|------|--------|-------|-----------|----------|
|
|
12
12
|
| Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |
|
|
13
13
|
| Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |
|
|
14
14
|
| Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |
|
|
15
|
-
| **Seedance native** | image or video | `engine=seedance-2` | per second (`seedance-2-face-*`; video sources bill input+output seconds) | **Premium**: natural delivery, whole-body performance, voice cloned from a video source, audio included |
|
|
16
15
|
|
|
17
|
-
Pick
|
|
16
|
+
Pick `sourceType` deliberately — it decides the pricing tier and the underlying endpoint.
|
|
18
17
|
|
|
19
|
-
## Seedance
|
|
18
|
+
## Want Seedance instead? That is a video generation, not a mode here
|
|
20
19
|
|
|
21
|
-
|
|
20
|
+
Seedance can generate the performance rather than bolting a mouth onto finished pixels — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal, and a video source keeps its own voice. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the clip (or portrait) attached as a video/ingredient reference and the dialogue written into the prompt yourself.
|
|
22
21
|
|
|
23
|
-
|
|
24
|
-
- **`audioMethod=upload`** drives the speech from a ≤15s audio file instead (a reference-audio input, no billing surcharge).
|
|
25
|
-
- **Sources:** image (any style; same framing rules as the avatar flow below) or a 2–15s video clip. Output duration follows the source/audio/line length (4–15s), not a fixed 5s.
|
|
26
|
-
- **Billing:** image sources bill the normal `seedance-2-face-{res}-{N}s` keys; video sources bill combined input+output seconds (`-vref-` keys) — pass `sourceSeconds` and quote via the confirm gate.
|
|
27
|
-
- **Faces:** `seedanceFace` defaults true. A REAL person → `[REAL_FACE_DETECTED]` → confirm consent → retry with `seedanceRealFace=true, realFaceConsent=true` (premium realface pricing).
|
|
28
|
-
- When to pick it: hero dialogue shots, natural delivery, "make this clip's person say X in their own voice". Stay on Kling for cheap utility re-dubs and long clips.
|
|
22
|
+
That is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a "video 1" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.
|
|
29
23
|
|
|
30
|
-
|
|
24
|
+
- Driving clips must be 2–15s; output duration is whatever you set (4–15s).
|
|
25
|
+
- Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming.
|
|
26
|
+
- Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.
|
|
27
|
+
|
|
28
|
+
Everything below is about the Kling tool.
|
|
31
29
|
|
|
32
30
|
## Choosing video vs avatar
|
|
33
31
|
|
|
@@ -1,32 +1,36 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-motion-transfer
|
|
3
|
-
description: How to set up motion transfer — Kling Motion Control (
|
|
3
|
+
description: How to set up motion transfer — Kling Motion Control only (std and pro tiers, 5-second outputs). Read before calling slates_generate_motion_transfer. Reference image (character) + driving video (motion source) → new video of the character performing the motion. Asset selection rules, character_orientation, tiers, and prompt usage. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Motion transfer — setup guide
|
|
7
7
|
|
|
8
|
-
Take a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion.
|
|
8
|
+
Take a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion. **This tool is Kling-only** — it wraps Kling Motion Control and nothing else.
|
|
9
9
|
|
|
10
|
-
|
|
|
10
|
+
| Tier | Cost | Use case |
|
|
11
11
|
|------|-----------|----------|
|
|
12
12
|
| Kling std (`kling-mc-std-5s`) | ~32 credits / 5s | General motion transfer, budget lane |
|
|
13
13
|
| Kling pro (`kling-mc-pro-5s`) | ~42 credits / 5s | Cleaner anatomy, better identity preservation |
|
|
14
|
-
| **Seedance 2.0** (`motionModel=seedance-2`) | per second of input+output (`seedance-2-face-vref-*`) | **Premium lane** — single-pass generation with the driving clip as a native conditioning signal: better motion fidelity, native audio, prompt-directed |
|
|
15
14
|
|
|
16
|
-
|
|
15
|
+
Both tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)
|
|
17
16
|
|
|
18
|
-
## Seedance
|
|
17
|
+
## Want Seedance instead? That is a video generation, not a mode here
|
|
19
18
|
|
|
20
|
-
Kling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion.
|
|
19
|
+
Kling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion, and the output carries native audio. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the driving clip attached as a video reference and the character image as an ingredient, then write the prompt yourself:
|
|
21
20
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
- **Native audio included** — the output can carry sound from the prompt (or the driving clip's vibe); no audio surcharge.
|
|
27
|
-
- `characterOrientation` is Kling-only; Seedance framing follows the prompt + `aspectRatio`.
|
|
21
|
+
```
|
|
22
|
+
The character from image 1 performs the exact motion, choreography, and camera
|
|
23
|
+
movement from video 1. Preserve the character's identity, appearance, and outfit.
|
|
24
|
+
```
|
|
28
25
|
|
|
29
|
-
|
|
26
|
+
That is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.
|
|
27
|
+
|
|
28
|
+
- **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).
|
|
29
|
+
- **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.
|
|
30
|
+
- **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
|
|
31
|
+
- `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.
|
|
32
|
+
|
|
33
|
+
Everything below is about the Kling tool.
|
|
30
34
|
|
|
31
35
|
## Inputs
|
|
32
36
|
|
|
@@ -10,9 +10,9 @@ ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.
|
|
|
10
10
|
## Where it routes
|
|
11
11
|
|
|
12
12
|
- **Scene audio, room tone, ambience beds, crowd/nature soundscapes** — anything where several sounds share a space. One generation, not three layered ones.
|
|
13
|
-
- **
|
|
13
|
+
- **Dialogue and scratch VO.** This is the only speech surface in Slates: the line is performed inside the scene. Lock the read by re-rolling until a take is right, then lip-sync against it with `slates_generate_lip_sync`.
|
|
14
14
|
- **NOT** a single effect that must land on a known frame — that is `eleven-sfx`, which takes an exact duration.
|
|
15
|
-
- **NOT** music
|
|
15
|
+
- **NOT** music. Slates has no music model; import a track and drop it on an audio track.
|
|
16
16
|
- **AUDIO-ONLY.** It cannot produce images or video.
|
|
17
17
|
|
|
18
18
|
## THE RULES
|