@slatesvideo/shared 0.5.5 → 0.5.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (29) hide show
  1. package/dist/index.d.ts +2 -2
  2. package/dist/index.js +6 -3
  3. package/dist/operations/index.d.ts +106 -19
  4. package/dist/operations/index.js +624 -172
  5. package/dist/prompts/character-sheet.d.ts +2 -1
  6. package/dist/prompts/character-sheet.js +82 -13
  7. package/dist/prompts/model-facts.d.ts +43 -1
  8. package/dist/prompts/model-facts.js +104 -2
  9. package/dist/prompts/prompting-tips.d.ts +1 -1
  10. package/dist/prompts/prompting-tips.js +208 -5
  11. package/dist/prompts/reference-composer.d.ts +21 -3
  12. package/dist/prompts/reference-composer.js +80 -10
  13. package/dist/prompts/reference-rules.d.ts +19 -2
  14. package/dist/prompts/reference-rules.js +18 -1
  15. package/dist/skills/content.js +8 -5
  16. package/exports/slates-prompt-builder/generated/SKILL.md +2 -2
  17. package/exports/slates-prompt-builder/generated/reference-character.md +9 -4
  18. package/exports/slates-prompt-builder/generated/reference-seedance.md +12 -1
  19. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +11 -11
  20. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  21. package/package.json +1 -1
  22. package/skills/slates-character-identity.md +9 -4
  23. package/skills/slates-model-selection.md +39 -11
  24. package/skills/slates-prompting-elevenlabs.md +69 -0
  25. package/skills/slates-prompting-lip-sync.md +12 -14
  26. package/skills/slates-prompting-motion-transfer.md +18 -14
  27. package/skills/slates-prompting-seed-audio.md +110 -0
  28. package/skills/slates-prompting-seedance-2-5.md +215 -0
  29. package/skills/slates-prompting-seedance.md +17 -1
@@ -0,0 +1,215 @@
1
+ ---
2
+ name: slates-prompting-seedance-2-5
3
+ description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is a SECOND SEAT next to 2.0, not an upgrade — it buys 30-second takes, 30 image references and audio-only references, and it gives up 1080p and 4K entirely. Shares 2.0's prompt grammar (read slates-prompting-seedance for the shot structure, subject binding, camera and constraint vocabulary); this file covers only what is different, plus the two hazards unique to 2.5 — the prompt-intent task classifier and the cost trap at 720p.
4
+ ---
5
+
6
+ # Seedance 2.5 — prompting
7
+
8
+ **Read `slates-prompting-seedance` first.** The prompt GRAMMAR is the same model family: the
9
+ 8-slot advanced formula, the `Shot 1 / Shot 2 / Shot 3` storyboard, subject binding by
10
+ `<Subject_N>@<Image_N>`, camera vocabulary, externalised emotion, inline constraint words, the
11
+ anti-twin fix. None of it is restated here. This file is only what 2.5 changes.
12
+
13
+ ---
14
+
15
+ ## The one fact that decides whether you use it at all
16
+
17
+ **Seedance 2.5 is 480p or 720p. There is no 1080p and no 4K, on any provider.**
18
+
19
+ That is not a Slates limitation or a tier gate — the model does not produce those resolutions.
20
+ So 2.5 does not replace 2.0; it sits beside it:
21
+
22
+ | | Seedance 2.0 | Seedance 2.5 |
23
+ |---|---|---|
24
+ | Resolution | 480p / 720p / 1080p / **native 4K** | **480p / 720p only** |
25
+ | Length | 4–15s | **4–30s in one take** |
26
+ | Reference budget | 15 (9 image + 3 video + 3 audio) | **50 (30 image + 10 video + 10 audio)** |
27
+ | Combined reference video/audio | ≤15s | **≤30s** |
28
+ | Audio-only reference | ✗ (needs an image or video alongside) | **✓** |
29
+ | Video edit as its own task type | ✗ | **✓ (`seedance-2.5-edit`)** |
30
+ | Default video model | **yes** | no |
31
+
32
+ **Route to 2.5 when the shot needs LENGTH, MANY REFERENCES, or an audio-only reference.
33
+ Route to 2.0 when resolution matters at all** — which, for anything a client will see full-screen,
34
+ is most of the time.
35
+
36
+ ---
37
+
38
+ ## 🚨 Hazard 1 — the prompt-intent task classifier
39
+
40
+ This is the one that costs money and time, and it has no equivalent on 2.0.
41
+
42
+ **Seedance 2.5 sorts every request into one of five task types** — text-to-video,
43
+ reference-to-video, first/last-frame, **video edit**, **video extend** — from the reference roles
44
+ attached **plus the intent of your sentence**. Each type then has its own parameter constraints,
45
+ and a violation comes back **asynchronously**: the task queues, credits are reserved, and only then
46
+ does it fail.
47
+
48
+ The trigger words are ordinary English:
49
+
50
+ | Reclassified as | Words that do it |
51
+ |---|---|
52
+ | **video edit** | `add` · `remove` · `replace` · `change` · `edit the video` |
53
+ | **video extend** | `extend` · `continue` · `continue the story` |
54
+
55
+ So a perfectly legitimate reference-to-video prompt — *"a wide shot of the workshop, **remove** the
56
+ tripod from frame"* — gets classified as an edit and fails on constraints it never set.
57
+
58
+ **What to do:**
59
+
60
+ 1. **If you mean to edit an existing clip, say so with the MODEL, not the sentence.** Call
61
+ `slates_edit_video` with `model: 'seedance-2.5-edit'`. That routes to a dedicated
62
+ task-typed endpoint and the classifier never has to guess.
63
+ 2. **If you mean a fresh shot, describe the finished frame rather than an instruction to change
64
+ one.** Not *"remove the tripod"* → *"the workshop bench, clear and uncluttered"*. Not
65
+ *"add rain"* → *"heavy rain falling through the streetlight"*. This is better prompting anyway:
66
+ the model renders what you describe, it does not take edits to an imagined draft.
67
+ 3. The trigger only fires when **references are attached**. A plain text-to-video prompt is safe
68
+ however it is worded.
69
+
70
+ **Slates will warn you, and it will never rewrite your prompt.** When a 2.5 reference generation's
71
+ prompt contains one of these words, the composer shows a warning that NAMES the words and the agent
72
+ route returns the same string. Silently editing the user's sentence to dodge a provider classifier
73
+ is forbidden — the words that reach the model are always the words the user can see.
74
+
75
+ ---
76
+
77
+ ## 🚨 Hazard 2 — 720p is not "the cheap one" any more
78
+
79
+ Every other model in Slates trains the habit that lower resolution means lower cost. 2.5 breaks it,
80
+ because the thing that moves the bill is **length**, and 2.5's length ceiling is double 2.0's.
81
+
82
+ Worked, at the shipped rates:
83
+
84
+ | Generation | Credits |
85
+ |---|---|
86
+ | 2.5 · 480p · 5s · faceless | 28 |
87
+ | 2.5 · 720p · 5s · faceless | 59 |
88
+ | 2.5 · 720p · 30s · faceless | 355 |
89
+ | 2.5 · 720p · 30s · AI-face route | **484** |
90
+ | 2.5 · 720p · 30s · consented real-face route | **710** |
91
+ | *(for scale)* 2.0 · 1080p · 15s · AI-face route | 411 |
92
+
93
+ **A 30-second 720p clip can cost more than a 15-second 1080p one** — and a base licence starts
94
+ with 1,000 credits. Someone who reads "720p" as "cheap" and asks for a 30-second take on the
95
+ real-face route has spent 71% of their welcome grant on one clip.
96
+
97
+ **Discipline:**
98
+
99
+ - **Always quote with `slates_estimate_generation_cost` before a take over ~10 seconds,** and say
100
+ the number out loud before generating.
101
+ - **Draft at 480p and 4–8 seconds.** Prove the composition, the motion and the identity first;
102
+ spend the length only on the take you already know works.
103
+ - **Length is a creative decision, not a default.** 30 seconds is available; it is rarely the right
104
+ answer for a single shot. Multi-shot storyboards inside one 30s generation are what the length is
105
+ actually for.
106
+ - Read `slates-cost-discipline` — all of it applies, more sharply here.
107
+
108
+ ---
109
+
110
+ ## What the extra reference budget is actually for
111
+
112
+ 30 image references (up from 9) does **not** mean "attach 30 images". Every rule in
113
+ `slates-prompting-seedance` about references still holds — 2–4 strong references beat both extremes,
114
+ one reference per role, one authoritative rendering per subject, and past **4 reference people**
115
+ output stability drops regardless of the cap.
116
+
117
+ The larger budget earns its keep in exactly two places:
118
+
119
+ - **A long multi-shot take** where different shots need different subjects and locations bound —
120
+ the budget is spread across the storyboard, not stacked on one frame.
121
+ - **Video and audio references alongside images**, which is where 2.5's 10 + 10 matters far more
122
+ than the image count.
123
+
124
+ ### Audio-only references — the genuinely new input
125
+
126
+ 2.0 required an image or video alongside any audio reference. **2.5 accepts audio on its own.**
127
+ That makes one recipe possible that was not before: drive a scene's timing, voice or ambience from
128
+ a recording with no visual anchor at all — a voice line, a music bed, a room tone — and let the
129
+ model build the picture to it. Cite it the same way as any other reference
130
+ (`Reference the timbre in <Audio_N> to generate…`), and remember that audio references carry **no
131
+ billing dimension** on any Seedance route: audio is included.
132
+
133
+ ### Video references
134
+
135
+ Up to 10 clips, ≤30s combined (2.0: 3 clips, ≤15s). A reference VIDEO switches the cost key to
136
+ `seedance-2.5*-vref-{res}-{T}s`, where **T = Σ input seconds + output seconds** — the sum is across
137
+ **every** clip attached, not just the longest. Three 6-second references on a 12-second output bills
138
+ 30 seconds, not 12 and not 18. Quote before confirming.
139
+
140
+ **The cap is a refusal, not a trim.** Attach an eleventh clip, or push past 30 combined seconds, and
141
+ the composition is rejected before anything uploads. That asymmetry is deliberate: reference images
142
+ warn-and-trim because dropping one doesn't change the price, and a dropped reference VIDEO would be
143
+ one you were quoted for and the model never saw.
144
+
145
+ ### Mixing all three in one call
146
+
147
+ 50 files total (30 image + 10 video + 10 audio) is a shared budget. Everything is cited positionally
148
+ by type — `image 1`, `video 2`, `audio 1` — in attachment order, so reordering the attachments
149
+ renumbers the citations. Write the prompt against those numbers:
150
+
151
+ ```
152
+ Marcus (image 1) performs the motion from video 1 in the workshop from image 2,
153
+ speaking the line in audio 1. Preserve his identity, appearance and outfit.
154
+ ```
155
+
156
+ Frames and reference media stay mutually exclusive, in every combination — the reference endpoint
157
+ has no first/last-frame parameters at all, so this is a shape mismatch rather than a preference.
158
+
159
+ ---
160
+
161
+ ## Seedance 2.5 Edit (`slates_edit_video`, `model: 'seedance-2.5-edit'`)
162
+
163
+ Its own picker row and its own op call, deliberately: the task type is **the model you chose**,
164
+ never something inferred from your sentence.
165
+
166
+ **Why route here at all:** it is the **only edit engine in Slates that accepts a clip longer than 15
167
+ seconds** (4–30s, versus Kling O3 Edit's 3–15s and Omni Flash Edit's 3–10s). For a clip inside the
168
+ others' range, choose on fidelity instead — Omni Flash Edit won the prompt-only head-to-head, and
169
+ Kling O3 Edit is the one that takes element and style reference images.
170
+
171
+ **How it behaves:**
172
+
173
+ - **Output length follows the SOURCE clip**, and the bill is the ceiled source length. The provider
174
+ requires an automatic duration on this task type, so there is no length knob — the clip you attach
175
+ is the quote.
176
+ - **The aspect ratio follows the source clip too.** No ratio control; the frame is the clip's frame.
177
+ - **480p or 720p output**, native audio.
178
+ - **Prompt and source clip only** on this op — no character or style reference images. If the edit
179
+ needs a reference image to lock an identity, that is Kling O3 Edit's job.
180
+ - **An edit bills roughly DOUBLE a plain 2.5 generation of the same length**, because every provider
181
+ charges an edit on input + output seconds. Read the confirm gate's number; do not reason from the
182
+ generation rate.
183
+ - **Set `seedanceFace: true` when a character's face is visible in the clip.** The faceless provider
184
+ blocks faces outright — this is not a price optimisation, it is whether the job runs at all.
185
+ - **There is no consented-real-face route for editing.** Real-person footage that the AI-face route
186
+ rejects has to go to Kling O3 Edit.
187
+
188
+ **Prompting an edit** — the same discipline as every other edit engine: **describe only what
189
+ changes.** The source already carries its composition, motion, timing and performance; re-describing
190
+ them fights the model. Use Seedance's own edit grammar from `slates-prompting-seedance`
191
+ (*"Strictly edit `<Video_1>`, and modify `<Original_Characteristic>` to `<New_Characteristic>`"*) and
192
+ **never** write *"reference video 1"* in an edit — the official guide is explicit that this phrasing
193
+ gets the request reclassified as a reference task, which is the same landmine as Hazard 1.
194
+
195
+ ---
196
+
197
+ ## Faces, and what does NOT change
198
+
199
+ The three-tier face routing is identical to 2.0 — faceless → default route, an AI character's face →
200
+ `seedanceFace: true` (the relaxed provider, a real cost premium), a real person's photo → the
201
+ consent-gated premium route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`
202
+ set **only** after the user explicitly confirms they hold the rights to the likeness. The full rules,
203
+ including why the real-vs-AI call is the provider's and not yours, are in
204
+ `slates-prompting-seedance`.
205
+
206
+ Also unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:
207
+
208
+ - **No time stamps.** `Shot 1 / Shot 2 / Shot 3`, ordered by when events occur, pacing left to the
209
+ model. A 30-second take makes this rule matter more, not less — it is more shots, not longer
210
+ segment labels.
211
+ - **One primary camera move per shot.**
212
+ - **No lens / aperture / film-stock vocabulary.** That is image-model syntax and a Seedance
213
+ anti-pattern.
214
+ - **No `negativePrompt` field** — constraints go inline.
215
+ - **Legible in-shot text still belongs in a baked start frame**, not in the video prompt.
@@ -202,7 +202,23 @@ These are ByteDance's own end-to-end cases. Note the shape: an asset-binding pre
202
202
 
203
203
  Reference-to-video accepts up to **9 reference images, 3 reference videos, 3 audio clips** `[official :275-281]`. Text+audio-only and audio-only inputs are not supported.
204
204
 
205
- **Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)*
205
+ **Mutually exclusive:** first-frame/last-frame mode CANNOT be combined with reference images. The error reads `"first/last frame content cannot be mixed with reference media content."` Pick one or the other. *(Official note `[:284]`: you can approximate first/last frames via prompt wording inside a multimodal call, but if the frames must be exact, use the dedicated first/last-frame route.)* The same rule covers reference VIDEO and AUDIO: they ride the reference endpoint, which has no frame parameters at all.
206
+
207
+ ### All three modalities go in ONE call
208
+
209
+ The caps are a shared budget, not three separate features: **12 files total on 2.0** (9 image + 3 video + 3 audio), **15 seconds of reference video combined**, **15 seconds of audio combined**. On 2.0 an audio reference needs at least one image or video alongside it; 2.5 accepts audio on its own.
210
+
211
+ Cite each by type and index, in the order they were attached — `image 1`, `video 1`, `audio 1`. The index is positional: reorder the attachments and the numbers move with them.
212
+
213
+ ```
214
+ Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
215
+ speaking the line in audio 1. Preserve his identity, appearance and outfit.
216
+ ```
217
+ <!-- slates-only -->
218
+ **Attaching a clip is NOT the same as editing it.** "Add as reference" puts it in the composer alongside everything else and wipes nothing; "Edit with AI" makes the clip the canvas and clears the tray for a fresh instruction. Two different jobs, two different menu entries — never infer one from the other.
219
+
220
+ **Over the cap is REFUSED, never trimmed.** A reference video is priced into the quote before it is sent, so a clip silently dropped after the quote would be a clip you paid for and the model never saw. Remove one and retry.
221
+ <!-- /slates-only -->
206
222
 
207
223
  ### Motion transfer & lip-sync recipes (reference video / audio)
208
224