@slatesvideo/shared 0.5.6 → 0.5.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,215 @@
1
+ ---
2
+ name: slates-prompting-seedance-2-5
3
+ description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is a SECOND SEAT next to 2.0, not an upgrade — it buys 30-second takes, 30 image references and audio-only references, and it gives up 1080p and 4K entirely. Shares 2.0's prompt grammar (read slates-prompting-seedance for the shot structure, subject binding, camera and constraint vocabulary); this file covers only what is different, plus the two hazards unique to 2.5 — the prompt-intent task classifier and the cost trap at 720p.
4
+ ---
5
+
6
+ # Seedance 2.5 — prompting
7
+
8
+ **Read `slates-prompting-seedance` first.** The prompt GRAMMAR is the same model family: the
9
+ 8-slot advanced formula, the `Shot 1 / Shot 2 / Shot 3` storyboard, subject binding by
10
+ `<Subject_N>@<Image_N>`, camera vocabulary, externalised emotion, inline constraint words, the
11
+ anti-twin fix. None of it is restated here. This file is only what 2.5 changes.
12
+
13
+ ---
14
+
15
+ ## The one fact that decides whether you use it at all
16
+
17
+ **Seedance 2.5 is 480p or 720p. There is no 1080p and no 4K, on any provider.**
18
+
19
+ That is not a Slates limitation or a tier gate — the model does not produce those resolutions.
20
+ So 2.5 does not replace 2.0; it sits beside it:
21
+
22
+ | | Seedance 2.0 | Seedance 2.5 |
23
+ |---|---|---|
24
+ | Resolution | 480p / 720p / 1080p / **native 4K** | **480p / 720p only** |
25
+ | Length | 4–15s | **4–30s in one take** |
26
+ | Reference budget | 15 (9 image + 3 video + 3 audio) | **50 (30 image + 10 video + 10 audio)** |
27
+ | Combined reference video/audio | ≤15s | **≤30s** |
28
+ | Audio-only reference | ✗ (needs an image or video alongside) | **✓** |
29
+ | Video edit as its own task type | ✗ | **✓ (`seedance-2.5-edit`)** |
30
+ | Default video model | **yes** | no |
31
+
32
+ **Route to 2.5 when the shot needs LENGTH, MANY REFERENCES, or an audio-only reference.
33
+ Route to 2.0 when resolution matters at all** — which, for anything a client will see full-screen,
34
+ is most of the time.
35
+
36
+ ---
37
+
38
+ ## 🚨 Hazard 1 — the prompt-intent task classifier
39
+
40
+ This is the one that costs money and time, and it has no equivalent on 2.0.
41
+
42
+ **Seedance 2.5 sorts every request into one of five task types** — text-to-video,
43
+ reference-to-video, first/last-frame, **video edit**, **video extend** — from the reference roles
44
+ attached **plus the intent of your sentence**. Each type then has its own parameter constraints,
45
+ and a violation comes back **asynchronously**: the task queues, credits are reserved, and only then
46
+ does it fail.
47
+
48
+ The trigger words are ordinary English:
49
+
50
+ | Reclassified as | Words that do it |
51
+ |---|---|
52
+ | **video edit** | `add` · `remove` · `replace` · `change` · `edit the video` |
53
+ | **video extend** | `extend` · `continue` · `continue the story` |
54
+
55
+ So a perfectly legitimate reference-to-video prompt — *"a wide shot of the workshop, **remove** the
56
+ tripod from frame"* — gets classified as an edit and fails on constraints it never set.
57
+
58
+ **What to do:**
59
+
60
+ 1. **If you mean to edit an existing clip, say so with the MODEL, not the sentence.** Call
61
+ `slates_edit_video` with `model: 'seedance-2.5-edit'`. That routes to a dedicated
62
+ task-typed endpoint and the classifier never has to guess.
63
+ 2. **If you mean a fresh shot, describe the finished frame rather than an instruction to change
64
+ one.** Not *"remove the tripod"* → *"the workshop bench, clear and uncluttered"*. Not
65
+ *"add rain"* → *"heavy rain falling through the streetlight"*. This is better prompting anyway:
66
+ the model renders what you describe, it does not take edits to an imagined draft.
67
+ 3. The trigger only fires when **references are attached**. A plain text-to-video prompt is safe
68
+ however it is worded.
69
+
70
+ **Slates will warn you, and it will never rewrite your prompt.** When a 2.5 reference generation's
71
+ prompt contains one of these words, the composer shows a warning that NAMES the words and the agent
72
+ route returns the same string. Silently editing the user's sentence to dodge a provider classifier
73
+ is forbidden — the words that reach the model are always the words the user can see.
74
+
75
+ ---
76
+
77
+ ## 🚨 Hazard 2 — 720p is not "the cheap one" any more
78
+
79
+ Every other model in Slates trains the habit that lower resolution means lower cost. 2.5 breaks it,
80
+ because the thing that moves the bill is **length**, and 2.5's length ceiling is double 2.0's.
81
+
82
+ Worked, at the shipped rates:
83
+
84
+ | Generation | Credits |
85
+ |---|---|
86
+ | 2.5 · 480p · 5s · faceless | 28 |
87
+ | 2.5 · 720p · 5s · faceless | 59 |
88
+ | 2.5 · 720p · 30s · faceless | 355 |
89
+ | 2.5 · 720p · 30s · AI-face route | **484** |
90
+ | 2.5 · 720p · 30s · consented real-face route | **710** |
91
+ | *(for scale)* 2.0 · 1080p · 15s · AI-face route | 411 |
92
+
93
+ **A 30-second 720p clip can cost more than a 15-second 1080p one** — and a base licence starts
94
+ with 1,000 credits. Someone who reads "720p" as "cheap" and asks for a 30-second take on the
95
+ real-face route has spent 71% of their welcome grant on one clip.
96
+
97
+ **Discipline:**
98
+
99
+ - **Always quote with `slates_estimate_generation_cost` before a take over ~10 seconds,** and say
100
+ the number out loud before generating.
101
+ - **Draft at 480p and 4–8 seconds.** Prove the composition, the motion and the identity first;
102
+ spend the length only on the take you already know works.
103
+ - **Length is a creative decision, not a default.** 30 seconds is available; it is rarely the right
104
+ answer for a single shot. Multi-shot storyboards inside one 30s generation are what the length is
105
+ actually for.
106
+ - Read `slates-cost-discipline` — all of it applies, more sharply here.
107
+
108
+ ---
109
+
110
+ ## What the extra reference budget is actually for
111
+
112
+ 30 image references (up from 9) does **not** mean "attach 30 images". Every rule in
113
+ `slates-prompting-seedance` about references still holds — 2–4 strong references beat both extremes,
114
+ one reference per role, one authoritative rendering per subject, and past **4 reference people**
115
+ output stability drops regardless of the cap.
116
+
117
+ The larger budget earns its keep in exactly two places:
118
+
119
+ - **A long multi-shot take** where different shots need different subjects and locations bound —
120
+ the budget is spread across the storyboard, not stacked on one frame.
121
+ - **Video and audio references alongside images**, which is where 2.5's 10 + 10 matters far more
122
+ than the image count.
123
+
124
+ ### Audio-only references — the genuinely new input
125
+
126
+ 2.0 required an image or video alongside any audio reference. **2.5 accepts audio on its own.**
127
+ That makes one recipe possible that was not before: drive a scene's timing, voice or ambience from
128
+ a recording with no visual anchor at all — a voice line, a music bed, a room tone — and let the
129
+ model build the picture to it. Cite it the same way as any other reference
130
+ (`Reference the timbre in <Audio_N> to generate…`), and remember that audio references carry **no
131
+ billing dimension** on any Seedance route: audio is included.
132
+
133
+ ### Video references
134
+
135
+ Up to 10 clips, ≤30s combined (2.0: 3 clips, ≤15s). A reference VIDEO switches the cost key to
136
+ `seedance-2.5*-vref-{res}-{T}s`, where **T = Σ input seconds + output seconds** — the sum is across
137
+ **every** clip attached, not just the longest. Three 6-second references on a 12-second output bills
138
+ 30 seconds, not 12 and not 18. Quote before confirming.
139
+
140
+ **The cap is a refusal, not a trim.** Attach an eleventh clip, or push past 30 combined seconds, and
141
+ the composition is rejected before anything uploads. That asymmetry is deliberate: reference images
142
+ warn-and-trim because dropping one doesn't change the price, and a dropped reference VIDEO would be
143
+ one you were quoted for and the model never saw.
144
+
145
+ ### Mixing all three in one call
146
+
147
+ 50 files total (30 image + 10 video + 10 audio) is a shared budget. Everything is cited positionally
148
+ by type — `image 1`, `video 2`, `audio 1` — in attachment order, so reordering the attachments
149
+ renumbers the citations. Write the prompt against those numbers:
150
+
151
+ ```
152
+ Marcus (image 1) performs the motion from video 1 in the workshop from image 2,
153
+ speaking the line in audio 1. Preserve his identity, appearance and outfit.
154
+ ```
155
+
156
+ Frames and reference media stay mutually exclusive, in every combination — the reference endpoint
157
+ has no first/last-frame parameters at all, so this is a shape mismatch rather than a preference.
158
+
159
+ ---
160
+
161
+ ## Seedance 2.5 Edit (`slates_edit_video`, `model: 'seedance-2.5-edit'`)
162
+
163
+ Its own picker row and its own op call, deliberately: the task type is **the model you chose**,
164
+ never something inferred from your sentence.
165
+
166
+ **Why route here at all:** it is the **only edit engine in Slates that accepts a clip longer than 15
167
+ seconds** (4–30s, versus Kling O3 Edit's 3–15s and Omni Flash Edit's 3–10s). For a clip inside the
168
+ others' range, choose on fidelity instead — Omni Flash Edit won the prompt-only head-to-head, and
169
+ Kling O3 Edit is the one that takes element and style reference images.
170
+
171
+ **How it behaves:**
172
+
173
+ - **Output length follows the SOURCE clip**, and the bill is the ceiled source length. The provider
174
+ requires an automatic duration on this task type, so there is no length knob — the clip you attach
175
+ is the quote.
176
+ - **The aspect ratio follows the source clip too.** No ratio control; the frame is the clip's frame.
177
+ - **480p or 720p output**, native audio.
178
+ - **Prompt and source clip only** on this op — no character or style reference images. If the edit
179
+ needs a reference image to lock an identity, that is Kling O3 Edit's job.
180
+ - **An edit bills roughly DOUBLE a plain 2.5 generation of the same length**, because every provider
181
+ charges an edit on input + output seconds. Read the confirm gate's number; do not reason from the
182
+ generation rate.
183
+ - **Set `seedanceFace: true` when a character's face is visible in the clip.** The faceless provider
184
+ blocks faces outright — this is not a price optimisation, it is whether the job runs at all.
185
+ - **There is no consented-real-face route for editing.** Real-person footage that the AI-face route
186
+ rejects has to go to Kling O3 Edit.
187
+
188
+ **Prompting an edit** — the same discipline as every other edit engine: **describe only what
189
+ changes.** The source already carries its composition, motion, timing and performance; re-describing
190
+ them fights the model. Use Seedance's own edit grammar from `slates-prompting-seedance`
191
+ (*"Strictly edit `<Video_1>`, and modify `<Original_Characteristic>` to `<New_Characteristic>`"*) and
192
+ **never** write *"reference video 1"* in an edit — the official guide is explicit that this phrasing
193
+ gets the request reclassified as a reference task, which is the same landmine as Hazard 1.
194
+
195
+ ---
196
+
197
+ ## Faces, and what does NOT change
198
+
199
+ The three-tier face routing is identical to 2.0 — faceless → default route, an AI character's face →
200
+ `seedanceFace: true` (the relaxed provider, a real cost premium), a real person's photo → the
201
+ consent-gated premium route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`
202
+ set **only** after the user explicitly confirms they hold the rights to the likeness. The full rules,
203
+ including why the real-vs-AI call is the provider's and not yours, are in
204
+ `slates-prompting-seedance`.
205
+
206
+ Also unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:
207
+
208
+ - **No time stamps.** `Shot 1 / Shot 2 / Shot 3`, ordered by when events occur, pacing left to the
209
+ model. A 30-second take makes this rule matter more, not less — it is more shots, not longer
210
+ segment labels.
211
+ - **One primary camera move per shot.**
212
+ - **No lens / aperture / film-stock vocabulary.** That is image-model syntax and a Seedance
213
+ anti-pattern.
214
+ - **No `negativePrompt` field** — constraints go inline.
215
+ - **Legible in-shot text still belongs in a baked start frame**, not in the video prompt.
@@ -202,7 +202,23 @@ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding pre
202
202
 
203
203
  Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
204
204
 
205
- **Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
205
+ **Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.
206
+
207
+ ### All three modalities go in ONE call
208
+
209
+ The caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.
210
+
211
+ Cite each by type and index, in the order they were attached — `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.
212
+
213
+ ```
214
+ Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
215
+ speaking the line in audio 1. Preserve his identity, appearance and outfit.
216
+ ```
217
+ <!-- slates-only -->
218
+ **Attaching a clip is NOT the same as editing it.** "Add as reference" puts it in the composer alongside everything else and wipes nothing; "Edit with AI" makes the clip the canvas and clears the tray for a fresh instruction. Two different jobs, two different menu entries — never infer one from the other.
219
+
220
+ **Over the cap is REFUSED, never trimmed.** A reference video is priced into the quote before it is sent, so a clip silently dropped after the quote would be a clip you paid for and the model never saw. Remove one and retry.
221
+ <!-- /slates-only -->
206
222
 
207
223
  ### Motion transfer & lip-sync recipes (reference video / audio)
208
224
 
@@ -1,110 +0,0 @@
1
- ---
2
- name: slates-prompting-suno
3
- description: How to prompt Suno in Slates. Read before calling slates_generate_audio with model suno. Full music tracks - EVERY call returns TWO variations for one flat price and duration is FREE up to 360 seconds. The rule that decides everything - in CUSTOM mode the prompt field is the EXACT LYRICS (sung as written), in DESCRIPTION mode it is a description and the lyrics get written for you. Covers the mode matrix, style vs prompt steering, instrumental scoring, negative tags, and the character caps.
4
- ---
5
-
6
- # Suno — prompting
7
-
8
- Full music generation, reached through the sunoapi.org wrapper. Models `V4`, `V4_5`, `V4_5PLUS`, `V4_5ALL`, `V5`, `V5_5`.
9
-
10
- **Two facts that should shape every decision:**
11
-
12
- 1. **Every call returns TWO songs** — two genuinely different takes on the same brief, for one flat price. Audition both before re-rolling.
13
- 2. **Length is free.** A 360-second track costs exactly what a default one costs (measured against the live provider balance 2026-07-31: a default call and a `duration: 240` call both debited the same). There is never a reason to generate a bed shorter than your edit.
14
-
15
- ## Where it routes
16
-
17
- - **Anything a listener would call a song or a score** — theme, underscore, needle-drop, montage bed, end-card sting.
18
- - **NOT** ambience or room tone — that is `seed-audio`, which is cheaper and better at it.
19
- - **NOT** a single effect — that is `eleven-sfx`.
20
- - **AUDIO-ONLY.**
21
-
22
- ## 🚨 THE RULE: which mode you are in changes what `prompt` means
23
-
24
- | `customMode` | `instrumental` | Required | What `prompt` means |
25
- |---|---|---|---|
26
- | `false` | either | `prompt` only (≤500 chars) | **A description.** Lyrics get written for you. |
27
- | `true` | `true` | `style`, `title` | **Unused.** Style + title do all the steering. |
28
- | `true` | `false` | `style`, `title`, `prompt` | **THE EXACT LYRICS**, sung as written. |
29
-
30
- Putting a description in the prompt field while `customMode: true` and `instrumental: false` gets your description **sung back at you**. This is the single most common Suno mistake and it costs a full generation every time.
31
-
32
- ## Description mode — the fast path
33
-
34
- ```
35
- customMode: false
36
- prompt: "brooding synthwave for a night drive, analog bass, no vocals, 90 bpm"
37
- ```
38
-
39
- Use it when you need a mood and do not care about specific words. 500-character cap. This is the right default for background beds.
40
-
41
- ## Custom mode — when the words matter
42
-
43
- ```
44
- customMode: true
45
- instrumental: false
46
- style: "dream pop, hazy, reverb-heavy guitars, female vocal, 100 bpm"
47
- title: "Blue Hour"
48
- prompt: "[Verse 1]\nThe lights come on before we're ready\n..."
49
- ```
50
-
51
- Structure tags (`[Verse]`, `[Chorus]`, `[Bridge]`, `[Outro]`) inside the lyrics are how you control the arrangement. Everything that is not a structure tag will be sung.
52
-
53
- ## Instrumental scoring
54
-
55
- ```
56
- customMode: true
57
- instrumental: true
58
- style: "tense orchestral strings, low brass swells, no percussion"
59
- title: "Approach"
60
- ```
61
-
62
- `instrumental: true` is the right answer for almost every film bed — a vocal you did not ask for will fight your dialogue.
63
-
64
- ## Steer with `style`, not with adjective piles
65
-
66
- Genre + era + instrumentation + tempo belong in `style`, not stuffed into `prompt`.
67
-
68
- ```
69
- ✓ style: "90s trip-hop, dusty breakbeat, Rhodes piano, upright bass, 85 bpm"
70
- ✗ prompt: "a really cool dusty 90s trip hop song with a Rhodes and..."
71
- ```
72
-
73
- `negativeTags` removes what keeps creeping in: `"brass, EDM drop, male vocal"`.
74
-
75
- ## The steering knobs
76
-
77
- | Param | Range | Reach for it when |
78
- |---|---|---|
79
- | `duration` | 10–360s (**V5_5 + custom mode only**) | Always, when the bed must outlast the cut. It is free. |
80
- | `vocalGender` | `m` / `f` — **the wire values, not "male"/"female"** | A specific voice is required. |
81
- | `styleWeight` | 0–1 | The style field is being ignored (raise) or strangling the song (lower). |
82
- | `weirdnessConstraint` | 0–1 | Takes are too safe (raise) or falling apart (lower). |
83
- | `audioWeight` | 0–1 | Balancing an audio input against the prompt. |
84
- | `personaId` / `personaModel` | — | A series needs the same voice/sound across episodes. |
85
-
86
- ## Character caps
87
-
88
- | Field | V4 | V4_5 / V4_5PLUS / V5 / V5_5 | V4_5ALL |
89
- |---|---|---|---|
90
- | prompt (custom = literal lyrics) | 3000 | 5000 | 5000 |
91
- | prompt (non-custom = description) | 500 | 500 | 500 |
92
- | style | 200 | 1000 | 1000 |
93
- | title | 80 | 100 | 80 |
94
-
95
- ## Iterating
96
-
97
- - **Audition both returned songs first.** A re-roll costs a full generation; the second variation is already paid for.
98
- - Wrong genre → fix `style`. Wrong words → you are in the wrong mode, check the matrix above.
99
- - Something keeps appearing that you do not want → `negativeTags`, not more prompt.
100
- - Three failed generations on the same brief means the style field is too vague, not that the seed is unlucky.
101
-
102
- ## Ops notes
103
-
104
- - Tracks take 2–3 minutes; a streamable preview exists ~30–40s in. Use `background: true` and poll.
105
- - The provider hosts files for a limited window — **Slates downloads and stores them locally as soon as the track finishes**, so nothing expires out from under a project.
106
- - Suno has **no official public API**; this rides an unofficial wrapper. Treat availability as best-effort and do not build a deadline around it.
107
-
108
- ## Content notes
109
-
110
- Provider-side moderation rejects lyrics and style prompts naming real artists or protected material (`SENSITIVE_WORD_ERROR`). Describe the sound, not the artist. See slates-content-policy.