@slatesvideo/shared 0.5.5 → 0.5.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.d.ts +2 -2
- package/dist/index.js +6 -3
- package/dist/operations/index.d.ts +106 -19
- package/dist/operations/index.js +624 -172
- package/dist/prompts/character-sheet.d.ts +2 -1
- package/dist/prompts/character-sheet.js +82 -13
- package/dist/prompts/model-facts.d.ts +43 -1
- package/dist/prompts/model-facts.js +104 -2
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +208 -5
- package/dist/prompts/reference-composer.d.ts +21 -3
- package/dist/prompts/reference-composer.js +80 -10
- package/dist/prompts/reference-rules.d.ts +19 -2
- package/dist/prompts/reference-rules.js +18 -1
- package/dist/skills/content.js +8 -5
- package/exports/slates-prompt-builder/generated/SKILL.md +2 -2
- package/exports/slates-prompt-builder/generated/reference-character.md +9 -4
- package/exports/slates-prompt-builder/generated/reference-seedance.md +12 -1
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +11 -11
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +1 -1
- package/skills/slates-character-identity.md +9 -4
- package/skills/slates-model-selection.md +39 -11
- package/skills/slates-prompting-elevenlabs.md +69 -0
- package/skills/slates-prompting-lip-sync.md +12 -14
- package/skills/slates-prompting-motion-transfer.md +18 -14
- package/skills/slates-prompting-seed-audio.md +110 -0
- package/skills/slates-prompting-seedance-2-5.md +215 -0
- package/skills/slates-prompting-seedance.md +17 -1
|
@@ -0,0 +1,215 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-seedance-2-5
|
|
3
|
+
description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is a SECOND SEAT next to 2.0, not an upgrade — it buys 30-second takes, 30 image references and audio-only references, and it gives up 1080p and 4K entirely. Shares 2.0's prompt grammar (read slates-prompting-seedance for the shot structure, subject binding, camera and constraint vocabulary); this file covers only what is different, plus the two hazards unique to 2.5 — the prompt-intent task classifier and the cost trap at 720p.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Seedance 2.5 — prompting
|
|
7
|
+
|
|
8
|
+
**Read `slates-prompting-seedance` first.** The prompt GRAMMAR is the same model family: the
|
|
9
|
+
8-slot advanced formula, the `Shot 1 / Shot 2 / Shot 3` storyboard, subject binding by
|
|
10
|
+
`<Subject_N>@<Image_N>`, camera vocabulary, externalised emotion, inline constraint words, the
|
|
11
|
+
anti-twin fix. None of it is restated here. This file is only what 2.5 changes.
|
|
12
|
+
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
## The one fact that decides whether you use it at all
|
|
16
|
+
|
|
17
|
+
**Seedance 2.5 is 480p or 720p. There is no 1080p and no 4K, on any provider.**
|
|
18
|
+
|
|
19
|
+
That is not a Slates limitation or a tier gate — the model does not produce those resolutions.
|
|
20
|
+
So 2.5 does not replace 2.0; it sits beside it:
|
|
21
|
+
|
|
22
|
+
| | Seedance 2.0 | Seedance 2.5 |
|
|
23
|
+
|---|---|---|
|
|
24
|
+
| Resolution | 480p / 720p / 1080p / **native 4K** | **480p / 720p only** |
|
|
25
|
+
| Length | 4–15s | **4–30s in one take** |
|
|
26
|
+
| Reference budget | 15 (9 image + 3 video + 3 audio) | **50 (30 image + 10 video + 10 audio)** |
|
|
27
|
+
| Combined reference video/audio | ≤15s | **≤30s** |
|
|
28
|
+
| Audio-only reference | ✗ (needs an image or video alongside) | **✓** |
|
|
29
|
+
| Video edit as its own task type | ✗ | **✓ (`seedance-2.5-edit`)** |
|
|
30
|
+
| Default video model | **yes** | no |
|
|
31
|
+
|
|
32
|
+
**Route to 2.5 when the shot needs LENGTH, MANY REFERENCES, or an audio-only reference.
|
|
33
|
+
Route to 2.0 when resolution matters at all** — which, for anything a client will see full-screen,
|
|
34
|
+
is most of the time.
|
|
35
|
+
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
## 🚨 Hazard 1 — the prompt-intent task classifier
|
|
39
|
+
|
|
40
|
+
This is the one that costs money and time, and it has no equivalent on 2.0.
|
|
41
|
+
|
|
42
|
+
**Seedance 2.5 sorts every request into one of five task types** — text-to-video,
|
|
43
|
+
reference-to-video, first/last-frame, **video edit**, **video extend** — from the reference roles
|
|
44
|
+
attached **plus the intent of your sentence**. Each type then has its own parameter constraints,
|
|
45
|
+
and a violation comes back **asynchronously**: the task queues, credits are reserved, and only then
|
|
46
|
+
does it fail.
|
|
47
|
+
|
|
48
|
+
The trigger words are ordinary English:
|
|
49
|
+
|
|
50
|
+
| Reclassified as | Words that do it |
|
|
51
|
+
|---|---|
|
|
52
|
+
| **video edit** | `add` · `remove` · `replace` · `change` · `edit the video` |
|
|
53
|
+
| **video extend** | `extend` · `continue` · `continue the story` |
|
|
54
|
+
|
|
55
|
+
So a perfectly legitimate reference-to-video prompt — *"a wide shot of the workshop, **remove** the
|
|
56
|
+
tripod from frame"* — gets classified as an edit and fails on constraints it never set.
|
|
57
|
+
|
|
58
|
+
**What to do:**
|
|
59
|
+
|
|
60
|
+
1. **If you mean to edit an existing clip, say so with the MODEL, not the sentence.** Call
|
|
61
|
+
`slates_edit_video` with `model: 'seedance-2.5-edit'`. That routes to a dedicated
|
|
62
|
+
task-typed endpoint and the classifier never has to guess.
|
|
63
|
+
2. **If you mean a fresh shot, describe the finished frame rather than an instruction to change
|
|
64
|
+
one.** Not *"remove the tripod"* → *"the workshop bench, clear and uncluttered"*. Not
|
|
65
|
+
*"add rain"* → *"heavy rain falling through the streetlight"*. This is better prompting anyway:
|
|
66
|
+
the model renders what you describe, it does not take edits to an imagined draft.
|
|
67
|
+
3. The trigger only fires when **references are attached**. A plain text-to-video prompt is safe
|
|
68
|
+
however it is worded.
|
|
69
|
+
|
|
70
|
+
**Slates will warn you, and it will never rewrite your prompt.** When a 2.5 reference generation's
|
|
71
|
+
prompt contains one of these words, the composer shows a warning that NAMES the words and the agent
|
|
72
|
+
route returns the same string. Silently editing the user's sentence to dodge a provider classifier
|
|
73
|
+
is forbidden — the words that reach the model are always the words the user can see.
|
|
74
|
+
|
|
75
|
+
---
|
|
76
|
+
|
|
77
|
+
## 🚨 Hazard 2 — 720p is not "the cheap one" any more
|
|
78
|
+
|
|
79
|
+
Every other model in Slates trains the habit that lower resolution means lower cost. 2.5 breaks it,
|
|
80
|
+
because the thing that moves the bill is **length**, and 2.5's length ceiling is double 2.0's.
|
|
81
|
+
|
|
82
|
+
Worked, at the shipped rates:
|
|
83
|
+
|
|
84
|
+
| Generation | Credits |
|
|
85
|
+
|---|---|
|
|
86
|
+
| 2.5 · 480p · 5s · faceless | 28 |
|
|
87
|
+
| 2.5 · 720p · 5s · faceless | 59 |
|
|
88
|
+
| 2.5 · 720p · 30s · faceless | 355 |
|
|
89
|
+
| 2.5 · 720p · 30s · AI-face route | **484** |
|
|
90
|
+
| 2.5 · 720p · 30s · consented real-face route | **710** |
|
|
91
|
+
| *(for scale)* 2.0 · 1080p · 15s · AI-face route | 411 |
|
|
92
|
+
|
|
93
|
+
**A 30-second 720p clip can cost more than a 15-second 1080p one** — and a base licence starts
|
|
94
|
+
with 1,000 credits. Someone who reads "720p" as "cheap" and asks for a 30-second take on the
|
|
95
|
+
real-face route has spent 71% of their welcome grant on one clip.
|
|
96
|
+
|
|
97
|
+
**Discipline:**
|
|
98
|
+
|
|
99
|
+
- **Always quote with `slates_estimate_generation_cost` before a take over ~10 seconds,** and say
|
|
100
|
+
the number out loud before generating.
|
|
101
|
+
- **Draft at 480p and 4–8 seconds.** Prove the composition, the motion and the identity first;
|
|
102
|
+
spend the length only on the take you already know works.
|
|
103
|
+
- **Length is a creative decision, not a default.** 30 seconds is available; it is rarely the right
|
|
104
|
+
answer for a single shot. Multi-shot storyboards inside one 30s generation are what the length is
|
|
105
|
+
actually for.
|
|
106
|
+
- Read `slates-cost-discipline` — all of it applies, more sharply here.
|
|
107
|
+
|
|
108
|
+
---
|
|
109
|
+
|
|
110
|
+
## What the extra reference budget is actually for
|
|
111
|
+
|
|
112
|
+
30 image references (up from 9) does **not** mean "attach 30 images". Every rule in
|
|
113
|
+
`slates-prompting-seedance` about references still holds — 2–4 strong references beat both extremes,
|
|
114
|
+
one reference per role, one authoritative rendering per subject, and past **4 reference people**
|
|
115
|
+
output stability drops regardless of the cap.
|
|
116
|
+
|
|
117
|
+
The larger budget earns its keep in exactly two places:
|
|
118
|
+
|
|
119
|
+
- **A long multi-shot take** where different shots need different subjects and locations bound —
|
|
120
|
+
the budget is spread across the storyboard, not stacked on one frame.
|
|
121
|
+
- **Video and audio references alongside images**, which is where 2.5's 10 + 10 matters far more
|
|
122
|
+
than the image count.
|
|
123
|
+
|
|
124
|
+
### Audio-only references — the genuinely new input
|
|
125
|
+
|
|
126
|
+
2.0 required an image or video alongside any audio reference. **2.5 accepts audio on its own.**
|
|
127
|
+
That makes one recipe possible that was not before: drive a scene's timing, voice or ambience from
|
|
128
|
+
a recording with no visual anchor at all — a voice line, a music bed, a room tone — and let the
|
|
129
|
+
model build the picture to it. Cite it the same way as any other reference
|
|
130
|
+
(`Reference the timbre in <Audio_N> to generate…`), and remember that audio references carry **no
|
|
131
|
+
billing dimension** on any Seedance route: audio is included.
|
|
132
|
+
|
|
133
|
+
### Video references
|
|
134
|
+
|
|
135
|
+
Up to 10 clips, ≤30s combined (2.0: 3 clips, ≤15s). A reference VIDEO switches the cost key to
|
|
136
|
+
`seedance-2.5*-vref-{res}-{T}s`, where **T = Σ input seconds + output seconds** — the sum is across
|
|
137
|
+
**every** clip attached, not just the longest. Three 6-second references on a 12-second output bills
|
|
138
|
+
30 seconds, not 12 and not 18. Quote before confirming.
|
|
139
|
+
|
|
140
|
+
**The cap is a refusal, not a trim.** Attach an eleventh clip, or push past 30 combined seconds, and
|
|
141
|
+
the composition is rejected before anything uploads. That asymmetry is deliberate: reference images
|
|
142
|
+
warn-and-trim because dropping one doesn't change the price, and a dropped reference VIDEO would be
|
|
143
|
+
one you were quoted for and the model never saw.
|
|
144
|
+
|
|
145
|
+
### Mixing all three in one call
|
|
146
|
+
|
|
147
|
+
50 files total (30 image + 10 video + 10 audio) is a shared budget. Everything is cited positionally
|
|
148
|
+
by type — `image 1`, `video 2`, `audio 1` — in attachment order, so reordering the attachments
|
|
149
|
+
renumbers the citations. Write the prompt against those numbers:
|
|
150
|
+
|
|
151
|
+
```
|
|
152
|
+
Marcus (image 1) performs the motion from video 1 in the workshop from image 2,
|
|
153
|
+
speaking the line in audio 1. Preserve his identity, appearance and outfit.
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
Frames and reference media stay mutually exclusive, in every combination — the reference endpoint
|
|
157
|
+
has no first/last-frame parameters at all, so this is a shape mismatch rather than a preference.
|
|
158
|
+
|
|
159
|
+
---
|
|
160
|
+
|
|
161
|
+
## Seedance 2.5 Edit (`slates_edit_video`, `model: 'seedance-2.5-edit'`)
|
|
162
|
+
|
|
163
|
+
Its own picker row and its own op call, deliberately: the task type is **the model you chose**,
|
|
164
|
+
never something inferred from your sentence.
|
|
165
|
+
|
|
166
|
+
**Why route here at all:** it is the **only edit engine in Slates that accepts a clip longer than 15
|
|
167
|
+
seconds** (4–30s, versus Kling O3 Edit's 3–15s and Omni Flash Edit's 3–10s). For a clip inside the
|
|
168
|
+
others' range, choose on fidelity instead — Omni Flash Edit won the prompt-only head-to-head, and
|
|
169
|
+
Kling O3 Edit is the one that takes element and style reference images.
|
|
170
|
+
|
|
171
|
+
**How it behaves:**
|
|
172
|
+
|
|
173
|
+
- **Output length follows the SOURCE clip**, and the bill is the ceiled source length. The provider
|
|
174
|
+
requires an automatic duration on this task type, so there is no length knob — the clip you attach
|
|
175
|
+
is the quote.
|
|
176
|
+
- **The aspect ratio follows the source clip too.** No ratio control; the frame is the clip's frame.
|
|
177
|
+
- **480p or 720p output**, native audio.
|
|
178
|
+
- **Prompt and source clip only** on this op — no character or style reference images. If the edit
|
|
179
|
+
needs a reference image to lock an identity, that is Kling O3 Edit's job.
|
|
180
|
+
- **An edit bills roughly DOUBLE a plain 2.5 generation of the same length**, because every provider
|
|
181
|
+
charges an edit on input + output seconds. Read the confirm gate's number; do not reason from the
|
|
182
|
+
generation rate.
|
|
183
|
+
- **Set `seedanceFace: true` when a character's face is visible in the clip.** The faceless provider
|
|
184
|
+
blocks faces outright — this is not a price optimisation, it is whether the job runs at all.
|
|
185
|
+
- **There is no consented-real-face route for editing.** Real-person footage that the AI-face route
|
|
186
|
+
rejects has to go to Kling O3 Edit.
|
|
187
|
+
|
|
188
|
+
**Prompting an edit** — the same discipline as every other edit engine: **describe only what
|
|
189
|
+
changes.** The source already carries its composition, motion, timing and performance; re-describing
|
|
190
|
+
them fights the model. Use Seedance's own edit grammar from `slates-prompting-seedance`
|
|
191
|
+
(*"Strictly edit `<Video_1>`, and modify `<Original_Characteristic>` to `<New_Characteristic>`"*) and
|
|
192
|
+
**never** write *"reference video 1"* in an edit — the official guide is explicit that this phrasing
|
|
193
|
+
gets the request reclassified as a reference task, which is the same landmine as Hazard 1.
|
|
194
|
+
|
|
195
|
+
---
|
|
196
|
+
|
|
197
|
+
## Faces, and what does NOT change
|
|
198
|
+
|
|
199
|
+
The three-tier face routing is identical to 2.0 — faceless → default route, an AI character's face →
|
|
200
|
+
`seedanceFace: true` (the relaxed provider, a real cost premium), a real person's photo → the
|
|
201
|
+
consent-gated premium route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`
|
|
202
|
+
set **only** after the user explicitly confirms they hold the rights to the likeness. The full rules,
|
|
203
|
+
including why the real-vs-AI call is the provider's and not yours, are in
|
|
204
|
+
`slates-prompting-seedance`.
|
|
205
|
+
|
|
206
|
+
Also unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:
|
|
207
|
+
|
|
208
|
+
- **No time stamps.** `Shot 1 / Shot 2 / Shot 3`, ordered by when events occur, pacing left to the
|
|
209
|
+
model. A 30-second take makes this rule matter more, not less — it is more shots, not longer
|
|
210
|
+
segment labels.
|
|
211
|
+
- **One primary camera move per shot.**
|
|
212
|
+
- **No lens / aperture / film-stock vocabulary.** That is image-model syntax and a Seedance
|
|
213
|
+
anti-pattern.
|
|
214
|
+
- **No `negativePrompt` field** — constraints go inline.
|
|
215
|
+
- **Legible in-shot text still belongs in a baked start frame**, not in the video prompt.
|
|
@@ -202,7 +202,23 @@ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding pre
|
|
|
202
202
|
|
|
203
203
|
Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
|
|
204
204
|
|
|
205
|
-
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
|
|
205
|
+
**Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.
|
|
206
|
+
|
|
207
|
+
### All three modalities go in ONE call
|
|
208
|
+
|
|
209
|
+
The caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.
|
|
210
|
+
|
|
211
|
+
Cite each by type and index, in the order they were attached — `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.
|
|
212
|
+
|
|
213
|
+
```
|
|
214
|
+
Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
|
|
215
|
+
speaking the line in audio 1. Preserve his identity, appearance and outfit.
|
|
216
|
+
```
|
|
217
|
+
<!-- slates-only -->
|
|
218
|
+
**Attaching a clip is NOT the same as editing it.** "Add as reference" puts it in the composer alongside everything else and wipes nothing; "Edit with AI" makes the clip the canvas and clears the tray for a fresh instruction. Two different jobs, two different menu entries — never infer one from the other.
|
|
219
|
+
|
|
220
|
+
**Over the cap is REFUSED, never trimmed.** A reference video is priced into the quote before it is sent, so a clip silently dropped after the quote would be a clip you paid for and the model never saw. Remove one and retry.
|
|
221
|
+
<!-- /slates-only -->
|
|
206
222
|
|
|
207
223
|
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
208
224
|
|