@slatesvideo/shared 0.5.6 → 0.5.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -16,8 +16,8 @@ This portable skill is deliberately thin. Its reference files are generated dire
16
16
  <!-- @generated:model-routing -->
17
17
  | Model | Canonical route | Guide |
18
18
  |---|---|---|
19
- | **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics. In the Motion Transfer / Lip Sync tools, Kling (MC / lip-sync / avatar) is the cheap utility lane; Seedance is the premium single-pass lane. | `reference-kling.md` |
20
- | **Seedance 2.0** | PREMIUM video tier — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those). Up to 9 ingredient images. Strong I2V / own-footage restyle. Native 4K, but 4K VIDEO is a Pro-only tier gate (base maxes at 1080p; server returns PRO_REQUIRED) — default 1080p unless the user is on Pro. Also the PREMIUM engine inside the Motion Transfer and Lip Sync tools (single-pass: driving video / dialogue are native conditioning signals — better motion fidelity, natural speech, voice cloned from a video source; video references bill input+output seconds). | `reference-seedance.md` |
19
+ | **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics. Kling is also the ONLY engine behind the Motion Transfer and Lip Sync tools (MC std/pro, lip-sync, avatar) — those two tools are Kling-only. | `reference-kling.md` |
20
+ | **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those). 4-15s, up to 9 ingredient images. Strong I2V / own-footage restyle. Native 4K, but 4K VIDEO is a Pro-only tier gate (base maxes at 1080p; server returns PRO_REQUIRED) — default 1080p unless the user is on Pro. Attaching a clip as a video reference (own-footage restyle, motion or dialogue conditioning) bills combined input+output seconds. 2.0 STAYS THE DEFAULT over 2.5 because it is the only Seedance with 1080p and 4K. | `reference-seedance.md` |
21
21
  | **Nano Banana 2 (Gemini 3.1 Flash Image)** | Default image model. 14 refs hard cap (10 object + 4 character). Brief it like a creative director, not tag soup. No negativePrompt field — use positive reframing. Best image start-frame for legible text. Knowledge cutoff Jan 2025. | `reference-nano-banana.md` |
22
22
  <!-- @end:model-routing -->
23
23
 
@@ -161,7 +161,7 @@ Seedance has **no `negativePrompt` field** — constraints go inline in this slo
161
161
 
162
162
  ## Worked examples `[official :1689-1745]`
163
163
 
164
- These are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere.
164
+ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding preamble, then `Shot N` blocks in event order, then a trailing style + stability paragraph. No time stamps anywhere — 2.0 does not respond to them at all, which is version-scoped and reverses on 2.5.
165
165
 
166
166
  **Example 1 — dormitory emotional short drama (dialogue-focused).** Assets: `@Image 1` half-body photo of the female lead · `@Image 2` dormitory scene reference · `@Video 1` camera-movement reference · `@Audio 1` indoor ambience.
167
167
 
@@ -201,7 +201,18 @@ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding pre
201
201
 
202
202
  Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
203
203
 
204
- **Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
204
+ **Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.
205
+
206
+ ### All three modalities go in ONE call
207
+
208
+ The caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.
209
+
210
+ Cite each by type and index, in the order they were attached — `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.
211
+
212
+ ```
213
+ Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
214
+ speaking the line in audio 1. Preserve his identity, appearance and outfit.
215
+ ```
205
216
 
206
217
  ### Motion transfer & lip-sync recipes (reference video / audio)
207
218
 
@@ -12,7 +12,7 @@
12
12
  },
13
13
  {
14
14
  "path": "skills/slates-prompting-seedance.md",
15
- "sha256": "a67f8c72bf187b726d39d2efb3c72a44740f43d44d5166d9767b93c55b4dd0a2"
15
+ "sha256": "ad9f3d5a20ef3bf52ebe916c81f67b9570fe1a43b54d64819da8b73972eb1a8e"
16
16
  },
17
17
  {
18
18
  "path": "skills/slates-prompting-kling-v3.md",
@@ -28,14 +28,14 @@
28
28
  },
29
29
  {
30
30
  "path": "src/prompts/model-facts.ts",
31
- "sha256": "f3ac5db676a96d1dbdbe7aaa364ca5df3c4739637ef7d61840ad10a585931a62"
31
+ "sha256": "b047ebd693dcfbeea352598a56ba405315df0b5287f6b8773388496b20efa3a2"
32
32
  }
33
33
  ],
34
34
  "outputs": [
35
35
  {
36
36
  "path": "SKILL.md",
37
- "bytes": 4969,
38
- "sha256": "51bb840b158c84e5d9060af9849ec23b94031d0ff0413f13da9a97358b40ec66"
37
+ "bytes": 4954,
38
+ "sha256": "8d161adfb51a6da1800d9cacd93181e13966807a3490e75bd89a0093d0311cde"
39
39
  },
40
40
  {
41
41
  "path": "reference-character.md",
@@ -44,8 +44,8 @@
44
44
  },
45
45
  {
46
46
  "path": "reference-seedance.md",
47
- "bytes": 31667,
48
- "sha256": "53a77f7bcfaa35be875df221f99843a74481ef6c6d8f7813b97a6e65d12f9dfe"
47
+ "bytes": 32558,
48
+ "sha256": "29b9408feec4ba91e5cd6f048ff6e6771fd2d96120378d4b6d52dec2a35eacd3"
49
49
  },
50
50
  {
51
51
  "path": "reference-kling.md",
@@ -65,8 +65,8 @@
65
65
  ],
66
66
  "archive": {
67
67
  "path": "slates-prompt-builder.skill",
68
- "bytes": 37938,
69
- "sha256": "48331704f54f422542f7a1a164915f3739b925e3f225a3710dbd8ecb2963ce68",
68
+ "bytes": 38295,
69
+ "sha256": "2cfe2ee4f80b0f64b53704a377c35d95ad15bedd5af3314f6b5242c8f6fb9bf1",
70
70
  "entries": [
71
71
  "SKILL.md",
72
72
  "reference-character.md",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@slatesvideo/shared",
3
- "version": "0.5.6",
3
+ "version": "0.5.8",
4
4
  "description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -0,0 +1,2 @@
1
+ <!-- consumer:ts -->
2
+ Seedance 2.0 ignores timing and answers only to "Shot 1 / Shot 2"; 2.5 acts on whole-second timestamps, and that is what makes a 30-second take controllable rather than just long. Three forms work: intervals ("0-3 seconds…3-7 seconds"), a point ("at the 5-second mark"), or relative ("after 3 seconds"). Whole seconds only, no gaps between intervals, and never to choreograph fast repeated motion. They work on edits too, where a range scopes the change: "…from 4-6 seconds…".
@@ -0,0 +1,34 @@
1
+ **2.0 does not respond to timestamps and answers only to shot numbers. 2.5 responds to
2
+ integer-second timestamps.** That is ByteDance's own first line under "Differences from Seedance
3
+ 2.0", and it is why a 30-second take is usable at all: the length is only worth buying if you can
4
+ say *when* things happen inside it.
5
+
6
+ Both formats are valid on 2.5, and you can mix them — `Shot N` blocks for a storyboard whose
7
+ pacing you are happy to leave to the model, timestamps when a beat has to land at a moment.
8
+
9
+ **Three ways to control time, all first-party:**
10
+
11
+ | Form | Write it like |
12
+ |---|---|
13
+ | **Interval** | `0-3 seconds… 3-7 seconds… 7-15 seconds` or `[1s-4s]… [4s-8s]… [8s-12s]` |
14
+ | **Time point** | *"Quick left sideways transition at the 5-second mark."* |
15
+ | **Relative** | *"After 3 seconds, everyone around him shakes their head."* · *"The frame freezes for 1 second after he presses the shutter."* |
16
+
17
+ **The rules that come with them:**
18
+
19
+ - **One second is the smallest unit.** Integers only — no `2.5s`, no frames.
20
+ - **No gaps in the timeline.** `0-3s… 5-6s…` leaves 3-5s unspecified and the model fills it however
21
+ it likes. Intervals must abut: `0-3s`, `3-7s`, `7-15s`.
22
+ - **Budget the plot to the seconds.** Too little content in a range and the model improvises to
23
+ fill it; too much and you get extra cuts or dropped beats. This is the actual craft of a 30s take.
24
+ - **Never time-code a high-frequency action.** *"Shake your head three times per second"* is
25
+ explicitly called out as a misuse — timestamps schedule beats, they don't choreograph frames.
26
+ - **Transitions want both halves:** the moment AND the method — *"At the 5-second mark, the camera
27
+ transitions leftward with a left wipe into a natural dissolve."*
28
+ - **Timestamps work on an EDIT too**, and that is where they earn the most: they scope a change in
29
+ time as well as in content — *"Change the man's action from drinking coffee to mopping the floor
30
+ from 4-6 seconds in Video 1, and leave the rest of the content unchanged."* Without a range, a
31
+ whole-clip instruction is applied to the whole clip.
32
+
33
+ Do **not** carry this back to 2.0, and do not carry Veo's `[00:00-00:02]` bracket syntax into
34
+ either — 2.0 ignores time entirely, and the cross-model syntax swap is its own known failure.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-model-selection
3
- description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 only) and never the default.
3
+ description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Seedance 2.5 is a SECOND SEAT beside 2.0 (30s takes, 30 references and timestamp control, but 480p/720p only — never an upgrade); Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 only) and never the default.
4
4
  ---
5
5
 
6
6
  # Model selection — the routing doctrine
@@ -28,6 +28,7 @@ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni F
28
28
  | Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
29
29
  | **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
30
30
  | The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
31
+ | **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs, and the only Seedance seat that **acts on timestamps** (rules in `slates-prompting-seedance-2-5` § Timestamps) — and **480p/720p ONLY, no 1080p and no 4K on any provider**. If resolution matters at all, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) 720p is NOT the cheap seat here — a 30s 720p face gen is 484 credits, more than a 15s 1080p Seedance 2.0 face gen (411), against a 1,000-credit welcome grant. Quote before any take over ~10s. |
31
32
 
32
33
  ### Named Seedance escalation triggers
33
34
 
@@ -49,26 +50,27 @@ Concrete beats route better than an abstract category. Cost stays a tiebreaker,
49
50
  | **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
50
51
  | **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
51
52
  | Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
52
- | AI-edit the user's OWN footage | Omni Flash Edit (3–10s) or Kling O3 Edit (3–15s, 720–3840px) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
53
+ | **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |
54
+ | AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
53
55
 
54
56
  - **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
55
57
  - **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.
56
58
  - **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
57
59
  - Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
58
60
 
59
- ## Motion Transfer & Lip Sync routing (two engines per tool)
61
+ ## Motion Transfer & Lip Sync routing (Kling-only tools)
60
62
 
61
- Both tools have a cheap Kling utility lane and a premium Seedance lane. The capability is the same; the execution model differs: Kling bolts motion/lip onto the source as a dedicated post-process; Seedance generates in a single pass with the driving clip / dialogue as native conditioning signals — better motion fidelity, natural speech delivery, audio included.
63
+ Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that bolts motion or lip movement onto a finished source as a dedicated post-process.
62
64
 
63
- | Job | Engine | Why |
65
+ | Job | Tool | Why |
64
66
  |---|---|---|
65
- | Quick motion retarget, budget lane, or driving clip >15s | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |
66
- | **Motion transfer where fidelity or audio matters** — dance, choreography, cinematic action | **Seedance 2.0** (`motionModel=seedance-2`) | Single-pass conditioning beats post-hoc retargeting; prompt-driven; native audio. Driving clip 2–15s; bills input+output seconds (vref keys). |
67
- | Cheap lip-sync utility (re-voice a clip, simple avatar) | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |
68
- | **Natural speech, voice cloned from the source clip, premium delivery** | **Seedance 2.0** (`engine=seedance-2`) | The line is spoken IN the generation (no TTS layer); a video source keeps its own voice; uploaded ≤15s audio can drive it. |
67
+ | Motion retarget onto a still character | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |
68
+ | Re-voice a clip, or animate a still portrait | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |
69
+
70
+ **Want the Seedance version of either?** It is not a switch on these tools — it is a normal `slates_generate_video` on `seedance-2` with the clip attached as a **video reference** and the motion or dialogue written into the prompt ("the character from image 1 performs the exact motion from video 1"). That routes to the same endpoint the tool would have called, with the prompt visible and editable instead of ghost-written. Single-pass conditioning genuinely beats post-hoc retargeting on fast choreography, contact, cloth and hair — and it carries native audio — so escalate there whenever fidelity matters.
69
71
 
70
- - Faces: Seedance tool gens default `seedanceFace=true` (sources are people). A REAL person triggers the consent cascade (`[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent`, premium realface pricing).
71
- - Billing: any Seedance gen with a video reference bills COMBINED input+output seconds (`seedance-2*-vref-*` keys) — always pass the clip duration and quote before confirming.
72
+ - Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane.
73
+ - Faces on that route go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person (premium realface pricing).
72
74
 
73
75
  **Rules:**
74
76
 
@@ -93,29 +95,29 @@ Both tools have a cheap Kling utility lane and a premium Seedance lane. The capa
93
95
 
94
96
  ## Audio routing
95
97
 
96
- **Image and video models cannot generate standalone audio, and none of the four audio models can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.
98
+ **Image and video models cannot generate standalone audio, and neither audio model can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.
97
99
 
98
100
  | Job | Model | Why |
99
101
  |---|---|---|
100
- | **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. 1–120s. The continuity-bed workhorse. |
101
- | **The exact words, in a repeatable named voice** — ad reads, narration, character lines to lip-sync against | **Eleven v3** (`eleven-v3`) | Verbatim text, 20 preset voices, re-renderable after a copy tweak without the performance drifting. Billed per 100 characters. |
102
+ | **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. 1–120s. The continuity-bed workhorse and the only speech surface. |
102
103
  | **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control (0.5–22s) and a real loop mode. |
103
- | **A song or a score** | **Suno** (`suno`) | Full music with structure. Two variations per call for one flat price, and duration is free to 360s. |
104
+
105
+ **There is no music model and no cast-voiceover model.** A song is imported (Slates reads audio files and puts them on the timeline), not generated. A line that has to be spoken is generated on Seed Audio and lip-synced against.
104
106
 
105
107
  ### Named audio escalation triggers
106
108
 
107
109
  - **"It needs to sound like a place"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.
108
- - **"Read this line"** with copy that a client can still change → Eleven v3. Scratch dialogue while the script is moving can stay on Seed Audio.
110
+ - **"Read this line"** → Seed Audio, with the line in quotes inside the scene sentence. Re-roll until the take is right, then lip-sync against it.
109
111
  - **"That needs a thump right there"** → Sound Effects, with the duration set to roughly the length of the event.
110
- - **"Give it a track"** → Suno, `instrumental: true` unless a vocal is genuinely wanted (an unasked-for vocal fights dialogue).
112
+ - **"Give it a track"** → there is no music generation. Say so and offer to lay an imported track on an audio track.
111
113
 
112
114
  **Rules:**
113
115
 
114
116
  - **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.
115
117
  - **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.
116
- - **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — on Suno the extra seconds are literally free.
118
+ - **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — and remember those extra seconds are billed on both surfaces.
117
119
  - **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.
118
- - Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`, `slates-prompting-suno`.
120
+ - Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`.
119
121
 
120
122
  ## Cost is a tiebreaker, not the router
121
123
 
@@ -1,83 +1,21 @@
1
1
  ---
2
2
  name: slates-prompting-elevenlabs
3
- description: How to prompt the two ElevenLabs surfaces in Slates. Read before calling slates_generate_audio with model eleven-v3 (Eleven v3 text-to-speech - controlled, repeatable, named-voice voiceover, billed per 100 characters) or eleven-sfx (Sound Effects v2 - one short effect with an EXACT duration, 0.5-22s, billed per second). Covers the "the text field is spoken verbatim" rule, punctuation as the only timing control, stability, describing an effect by its physical cause, and when to use Seed Audio instead.
3
+ description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before calling slates_generate_audio with model eleven-sfx — ONE short effect with an EXACT duration (0.5-22s), or a seamless loop, billed per second. Covers describing an effect by its physical cause, the one-sound-per-generation rule, picking a duration, loops, prompt_influence, and when to use Seed Audio instead.
4
4
  ---
5
5
 
6
- # ElevenLabs — prompting (Eleven v3 TTS + Sound Effects v2)
6
+ # ElevenLabs Sound Effects v2 — prompting
7
7
 
8
- Two separate surfaces from the same vendor, carried on fal (`fal-ai/elevenlabs/tts/eleven-v3`, `fal-ai/elevenlabs/sound-effects/v2`). They share nothing but a bill — treat them as different tools.
8
+ One short sound with an exact length, carried on fal (`fal-ai/elevenlabs/sound-effects/v2`). This is the only Slates audio surface with a real duration control and a real loop mode.
9
9
 
10
- ## Where they route
10
+ ## Where it routes
11
11
 
12
- - **`eleven-v3`** — the exact words matter and the read must be **repeatable**: ad reads, narration, character lines you will lip-sync against, anything a client will ask you to re-render after a copy tweak. Billed per 100 characters of text, rounded up.
13
- - **`eleven-sfx`** — a single sound that has to land on a known frame, or a seamless loop. Billed per second, 0.5–22s.
14
- - **Neither** for layered scenes. A room with dialogue *and* clatter *and* ambience is one `seed-audio` pass, not three ElevenLabs generations.
15
- - **AUDIO-ONLY.** Neither can produce images or video.
12
+ - **A single hit that has to land on a known frame** — door slam, whoosh, impact, UI blip, riser.
13
+ - **A seamless loop** you can lay under a whole scene — rain, engine hum, crowd murmur, machine noise.
14
+ - **NOT** layered scenes. A room with dialogue *and* clatter *and* ambience is one `seed-audio` pass, not three SFX generations.
15
+ - **NOT** speech. Dialogue, narration and scratch VO are `seed-audio` — it casts and performs the line inside the scene.
16
+ - **AUDIO-ONLY.** It cannot produce images or video.
16
17
 
17
- ---
18
-
19
- ## Eleven v3 (`eleven-v3`) — THE RULES
20
-
21
- ### 1. 🚨 The text field is the script. Every character is spoken.
22
-
23
- ```
24
- ✗ (excited) Read this fast — "Grab yours today!"
25
- ✓ Grab yours today!
26
- ```
27
-
28
- Stage directions, speaker names, bracketed emotion tags and markdown all get read out loud. There is no instruction channel — direction lives in `stability` and in how you punctuate.
29
-
30
- ### 2. Punctuation is the only timing control
31
-
32
- | You want | Write |
33
- |---|---|
34
- | a hard stop | `It works. Every time.` |
35
- | a beat, not a stop | `It works — every time.` |
36
- | a trailing hesitation | `It works… mostly.` |
37
- | a list rhythm | `Faster, cheaper, and yours.` |
38
-
39
- Rewrite the punctuation before you touch a setting. It moves the read more than `stability` does.
40
-
41
- ### 3. Pick a voice and keep it
42
-
43
- 20 presets: Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill. Default `Rachel`.
44
-
45
- One voice per character, one per piece. A series that swaps voices between shots reads as an accident. There is **no voice cloning on this route** — a cloned ElevenLabs voice ID is not supported and must not be assumed to pass through.
46
-
47
- ### 4. Stability
48
-
49
- | Value | Behavior | Use for |
50
- |---|---|---|
51
- | ~0.3 | expressive, varies take-to-take | one dramatic line, character dialogue |
52
- | 0.5 (default) | balanced | most reads |
53
- | ~0.8 | flat, highly repeatable | long narration, anything you will re-render |
54
-
55
- Raise it when re-rolls keep giving you a different performance. Lower it when the read is lifeless.
56
-
57
- ### 5. Spell out what TTS gets wrong
58
-
59
- Acronyms, product names, prices, years and URLs are where it embarrasses itself. Write the pronunciation:
60
-
61
- ```
62
- SKU → "ess kay you"
63
- 2026 → "twenty twenty six"
64
- $19.99 → "nineteen ninety nine"
65
- slates.video → "slates dot video"
66
- ```
67
-
68
- Set `languageCode` (ISO 639-1) to force a language when the text is ambiguous or code-switched.
69
-
70
- ### 6. Length is money, honestly
71
-
72
- 1–5000 characters, billed in 100-character buckets rounded up. A tightened sentence costs less; a pasted stray paragraph costs more. This is the one Slates surface where editing the copy is also a cost control.
73
-
74
- ### 7. Timestamps are free — leave them on
75
-
76
- Word-level timestamps come back with every generation at no extra charge. They are exactly what a caption/subtitle pass consumes. There is no reason to disable them.
77
-
78
- ---
79
-
80
- ## Sound Effects v2 (`eleven-sfx`) — THE RULES
18
+ ## THE RULES
81
19
 
82
20
  ### 1. Describe the physical CAUSE, not the label
83
21
 
@@ -114,18 +52,18 @@ Over-asking pads the tail with room tone you then trim. Under-asking clips the d
114
52
 
115
53
  `loop: true` tiles without a seam — rain, engine hum, crowd murmur, machine noise. Combine with a longer duration so the loop point is not obvious.
116
54
 
55
+ For a bed longer than 22s, this is the wrong surface: `seed-audio` runs to 120s in one pass.
56
+
117
57
  ### 5. Prompt influence
118
58
 
119
59
  `promptInfluence` 0–1, default 0.3. Higher hugs your wording with less variation between takes; lower explores. Raise it when a re-roll keeps wandering off the brief; lower it when every take sounds like the same take.
120
60
 
121
- ---
122
-
123
- ## Iterating on either surface
61
+ ## Iterating
124
62
 
125
- - TTS re-rolls that keep drifting = raise `stability`. TTS reads that sound robotic = lower it, then fix the punctuation.
126
- - SFX re-rolls that keep missing = the prompt named a label instead of a cause. Rewrite it as a physical event.
63
+ - Re-rolls that keep missing = the prompt named a **label** instead of a **cause**. Rewrite it as a physical event.
64
+ - A hit that lands but sounds wrong in the scene is usually a *room* problem — name the space ("in a stone hallway", "in a padded studio", "outdoors, no reflections").
127
65
  - Three failed takes means the prompt is wrong, not the seed.
128
66
 
129
67
  ## Content notes
130
68
 
131
- ElevenLabs applies its own moderation, and voice likeness of real people is restricted by their terms. See slates-content-policy.
69
+ ElevenLabs applies its own moderation. See slates-content-policy.
@@ -1,33 +1,31 @@
1
1
  ---
2
2
  name: slates-prompting-lip-sync
3
- description: How to set up lip-sync — Kling (cheap utility lane) or Seedance 2.0 (premium single-pass lane). Read before calling slates_generate_lip_sync. Flows — video→video re-dub, image→video avatar, and Seedance native speech — with different inputs, pricing, and gotchas. Voice catalog, framing rules, audio file constraints, and which engine/tier to pick.
3
+ description: How to set up lip-sync — Kling-only (dedicated lip-sync and avatar endpoints, 5-second outputs). Read before calling slates_generate_lip_sync. Two flows — video→video re-dub and image→video avatar — with different inputs, pricing, and gotchas. Voice catalog, framing rules, audio file constraints, and which tier to pick. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.
4
4
  ---
5
5
 
6
6
  # Lip-sync — setup guide
7
7
 
8
- Two engines. Kling is the cheap utility lane (dedicated lip-sync endpoints, 5-second outputs); Seedance is the premium lane (speech generated IN the video itself, single pass):
8
+ **This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; every entry is a real endpoint and every output is 5 seconds.
9
9
 
10
- | Flow | Source | Engine/Model | Cost | Use case |
10
+ | Flow | Source | Model | Cost | Use case |
11
11
  |------|--------|-------|-----------|----------|
12
12
  | Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |
13
13
  | Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |
14
14
  | Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |
15
- | **Seedance native** | image or video | `engine=seedance-2` | per second (`seedance-2-face-*`; video sources bill input+output seconds) | **Premium**: natural delivery, whole-body performance, voice cloned from a video source, audio included |
16
15
 
17
- Pick engine + `sourceType` deliberately — they decide the pricing tier and the underlying endpoint.
16
+ Pick `sourceType` deliberately — it decides the pricing tier and the underlying endpoint.
18
17
 
19
- ## Seedance engine (premium single-pass)
18
+ ## Want Seedance instead? That is a video generation, not a mode here
20
19
 
21
- Kling lip-sync moves the mouth on finished pixels; Seedance *generates* the performance — head movement, gesture, delivery energy — with the dialogue as a native conditioning signal. Key facts:
20
+ Seedance can generate the performance rather than bolting a mouth onto finished pixels — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal, and a video source keeps its own voice. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the clip (or portrait) attached as a video/ingredient reference and the dialogue written into the prompt yourself.
22
21
 
23
- - **`ttsText` becomes the spoken line, natively.** No TTS voice/speed params — a VIDEO source keeps its **own voice** (the model clones it from the clip's audio track); for an image source the voice follows the character's look, or describe it in the line's context.
24
- - **`audioMethod=upload`** drives the speech from a ≤15s audio file instead (a reference-audio input, no billing surcharge).
25
- - **Sources:** image (any style; same framing rules as the avatar flow below) or a 2–15s video clip. Output duration follows the source/audio/line length (4–15s), not a fixed 5s.
26
- - **Billing:** image sources bill the normal `seedance-2-face-{res}-{N}s` keys; video sources bill combined input+output seconds (`-vref-` keys) — pass `sourceSeconds` and quote via the confirm gate.
27
- - **Faces:** `seedanceFace` defaults true. A REAL person → `[REAL_FACE_DETECTED]` → confirm consent → retry with `seedanceRealFace=true, realFaceConsent=true` (premium realface pricing).
28
- - When to pick it: hero dialogue shots, natural delivery, "make this clip's person say X in their own voice". Stay on Kling for cheap utility re-dubs and long clips.
22
+ That is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a "video 1" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.
29
23
 
30
- Everything below applies to the **Kling** engine.
24
+ - Driving clips must be 2–15s; output duration is whatever you set (4–15s).
25
+ - Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming.
26
+ - Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.
27
+
28
+ Everything below is about the Kling tool.
31
29
 
32
30
  ## Choosing video vs avatar
33
31
 
@@ -1,32 +1,36 @@
1
1
  ---
2
2
  name: slates-prompting-motion-transfer
3
- description: How to set up motion transfer — Kling Motion Control (cheap utility lane) or Seedance 2.0 (premium single-pass lane). Read before calling slates_generate_motion_transfer. Reference image (character) + driving video (motion source) → new video of the character performing the motion. Asset selection rules, engine choice, character_orientation, tiers, and prompt usage.
3
+ description: How to set up motion transfer — Kling Motion Control only (std and pro tiers, 5-second outputs). Read before calling slates_generate_motion_transfer. Reference image (character) + driving video (motion source) → new video of the character performing the motion. Asset selection rules, character_orientation, tiers, and prompt usage. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.
4
4
  ---
5
5
 
6
6
  # Motion transfer — setup guide
7
7
 
8
- Take a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion. Two engines:
8
+ Take a still **target image** (your character) and a **source video** (the motion you want), produce a new video of your character performing the source video's motion. **This tool is Kling-only** — it wraps Kling Motion Control and nothing else.
9
9
 
10
- | Engine | Cost | Use case |
10
+ | Tier | Cost | Use case |
11
11
  |------|-----------|----------|
12
12
  | Kling std (`kling-mc-std-5s`) | ~32 credits / 5s | General motion transfer, budget lane |
13
13
  | Kling pro (`kling-mc-pro-5s`) | ~42 credits / 5s | Cleaner anatomy, better identity preservation |
14
- | **Seedance 2.0** (`motionModel=seedance-2`) | per second of input+output (`seedance-2-face-vref-*`) | **Premium lane** — single-pass generation with the driving clip as a native conditioning signal: better motion fidelity, native audio, prompt-directed |
15
14
 
16
- All tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)
15
+ Both tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)
17
16
 
18
- ## Seedance engine (premium single-pass)
17
+ ## Want Seedance instead? That is a video generation, not a mode here
19
18
 
20
- Kling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion. Key differences from Kling:
19
+ Kling MC retargets a skeleton onto a finished image; Seedance *generates* the shot with the motion as a conditioning input — the difference shows on fast choreography, physical contact, cloth/hair, and camera motion, and the output carries native audio. **It is not an engine switch on this tool.** Run a normal `slates_generate_video` on `seedance-2` with the driving clip attached as a video reference and the character image as an ingredient, then write the prompt yourself:
21
20
 
22
- - **Prompt-driven.** The prompt is the primary control, using ordinal references: `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.` Add style/setting/camera direction freely — Seedance re-generates the whole shot.
23
- - **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`character_orientation: 'video'` takes up to 30s).
24
- - **Billing = combined input+output seconds** (the vref keys). Pass `sourceVideoSeconds`; output `duration` defaults to the clip length (4–15s). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.
25
- - **Faces route through the face cascade**: `seedanceFace` defaults true (driving clips contain people). A REAL person → `[REAL_FACE_DETECTED]` → confirm consent with the user → retry `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
26
- - **Native audio included** — the output can carry sound from the prompt (or the driving clip's vibe); no audio surcharge.
27
- - `characterOrientation` is Kling-only; Seedance framing follows the prompt + `aspectRatio`.
21
+ ```
22
+ The character from image 1 performs the exact motion, choreography, and camera
23
+ movement from video 1. Preserve the character's identity, appearance, and outfit.
24
+ ```
28
25
 
29
- Everything below applies to the **Kling** engine.
26
+ That is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.
27
+
28
+ - **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).
29
+ - **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.
30
+ - **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
31
+ - `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.
32
+
33
+ Everything below is about the Kling tool.
30
34
 
31
35
  ## Inputs
32
36
 
@@ -10,9 +10,9 @@ ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.
10
10
  ## Where it routes
11
11
 
12
12
  - **Scene audio, room tone, ambience beds, crowd/nature soundscapes** — anything where several sounds share a space. One generation, not three layered ones.
13
- - **Fast scratch dialogue** when the exact wording is still moving. Once the script locks and the read has to be repeatable, switch to `eleven-v3`.
13
+ - **Dialogue and scratch VO.** This is the only speech surface in Slates: the line is performed inside the scene. Lock the read by re-rolling until a take is right, then lip-sync against it with `slates_generate_lip_sync`.
14
14
  - **NOT** a single effect that must land on a known frame — that is `eleven-sfx`, which takes an exact duration.
15
- - **NOT** music — that is `suno`.
15
+ - **NOT** music. Slates has no music model; import a track and drop it on an audio track.
16
16
  - **AUDIO-ONLY.** It cannot produce images or video.
17
17
 
18
18
  ## THE RULES