@slatesvideo/shared 0.7.1 → 0.7.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (88) hide show
  1. package/dist/clients/cloud.d.ts +4 -0
  2. package/dist/clients/cloud.js +11 -3
  3. package/dist/index.d.ts +2 -1
  4. package/dist/index.js +2 -1
  5. package/dist/manual/content.d.ts +1 -1
  6. package/dist/manual/content.js +1 -1
  7. package/dist/manual/index.d.ts +11 -2
  8. package/dist/manual/index.js +178 -14
  9. package/dist/operations/index.d.ts +292 -94
  10. package/dist/operations/index.js +870 -199
  11. package/dist/operations/surface.d.ts +6 -2
  12. package/dist/operations/surface.js +35 -5
  13. package/dist/prompts/agent-doctrine.d.ts +4 -4
  14. package/dist/prompts/agent-doctrine.js +18 -29
  15. package/dist/prompts/generation-policy.d.ts +1 -1
  16. package/dist/prompts/guide-discovery.d.ts +23 -0
  17. package/dist/prompts/guide-discovery.js +39 -0
  18. package/dist/prompts/guide-retrieval.js +1 -1
  19. package/dist/prompts/model-capabilities.d.ts +8 -9
  20. package/dist/prompts/model-capabilities.js +11 -51
  21. package/dist/prompts/model-facts.d.ts +2 -2
  22. package/dist/prompts/model-facts.js +15 -26
  23. package/dist/prompts/partials.generated.js +6 -3
  24. package/dist/prompts/prompting-tips.d.ts +1 -1
  25. package/dist/prompts/prompting-tips.js +21 -63
  26. package/dist/prompts/search-terms.d.ts +3 -0
  27. package/dist/prompts/search-terms.js +24 -0
  28. package/dist/skills/content.js +36 -37
  29. package/dist/skills/metadata.d.ts +7 -0
  30. package/dist/skills/metadata.js +29 -0
  31. package/exports/slates-chatgpt-images/generated/SKILL.md +7 -1
  32. package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
  33. package/exports/slates-prompt-builder/generated/SKILL.md +28 -16
  34. package/exports/slates-prompt-builder/generated/reference-character.md +12 -13
  35. package/exports/slates-prompt-builder/generated/reference-content-policy.md +2 -2
  36. package/exports/slates-prompt-builder/generated/reference-gpt-image-2-5.md +191 -0
  37. package/exports/slates-prompt-builder/generated/reference-kling.md +32 -11
  38. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +24 -6
  39. package/exports/slates-prompt-builder/generated/reference-omni-flash.md +65 -0
  40. package/exports/slates-prompt-builder/generated/reference-seedance-2-5.md +362 -0
  41. package/exports/slates-prompt-builder/generated/reference-seedance.md +34 -4
  42. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +77 -23
  43. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  44. package/package.json +2 -1
  45. package/skills/_partials/blender-action-curves.md +24 -0
  46. package/skills/_partials/cinematic-card.md +1 -1
  47. package/skills/_partials/iteration-diagnosis.md +5 -0
  48. package/skills/_partials/model-routing.md +35 -0
  49. package/skills/_partials/seedance-25-timestamps.md +2 -2
  50. package/skills/_partials/still-gate.md +2 -2
  51. package/skills/_partials/thresholds.md +1 -1
  52. package/skills/slates-blocking-to-prompt.md +15 -13
  53. package/skills/slates-camera-language.md +45 -7
  54. package/skills/slates-character-identity.md +8 -6
  55. package/skills/slates-chatgpt-images.md +7 -1
  56. package/skills/slates-cinematic-look.md +1 -1
  57. package/skills/slates-content-policy.md +4 -6
  58. package/skills/slates-cost-discipline.md +18 -12
  59. package/skills/slates-dialogue-blocking.md +6 -6
  60. package/skills/slates-direct-response-ad.md +1 -1
  61. package/skills/slates-edit-and-iterate.md +12 -4
  62. package/skills/slates-model-selection.md +82 -90
  63. package/skills/slates-one-prompt-film.md +1 -1
  64. package/skills/slates-previs-blocking.md +44 -13
  65. package/skills/slates-project-organization.md +2 -2
  66. package/skills/slates-prompting-elevenlabs.md +4 -4
  67. package/skills/slates-prompting-flux-2-max.md +2 -3
  68. package/skills/slates-prompting-gpt-image-2-5.md +2 -2
  69. package/skills/slates-prompting-inworld-tts.md +1 -1
  70. package/skills/slates-prompting-kling-v3.md +11 -9
  71. package/skills/slates-prompting-lip-sync.md +15 -15
  72. package/skills/slates-prompting-ltx-2-5.md +5 -6
  73. package/skills/slates-prompting-minimax-h3.md +11 -11
  74. package/skills/slates-prompting-motion-transfer.md +8 -8
  75. package/skills/slates-prompting-nano-banana-2.md +8 -4
  76. package/skills/slates-prompting-omni-flash.md +9 -9
  77. package/skills/slates-prompting-seed-audio.md +24 -4
  78. package/skills/slates-prompting-seedance-2-5.md +40 -30
  79. package/skills/slates-prompting-seedance.md +4 -4
  80. package/skills/slates-prompting-seedream-5-lite.md +6 -6
  81. package/skills/slates-restyle-from-blocking.md +2 -2
  82. package/skills/slates-script-craft.md +1 -1
  83. package/skills/slates-shot-variety.md +1 -1
  84. package/skills/slates-storyboard-from-script.md +1 -1
  85. package/skills/slates-style-prompting.md +8 -6
  86. package/skills/slates-ugc-influencer-ad.md +1 -1
  87. package/skills/slates-vision-feedback-loop.md +118 -110
  88. package/skills/slates-prompting-veo-3.md +0 -224
@@ -0,0 +1,65 @@
1
+ <!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
2
+
3
+ > Generated from the production Slates guide. Model-specific syntax and measured examples apply to the endpoints named below. For another generation tool, check its current schema and reference handling; its limits, billing and defaults may differ.
4
+
5
+ # Gemini Omni Flash — prompting
6
+
7
+ **Card — Gemini Omni Flash.** Two different jobs with OPPOSITE prompt rules, and getting them the wrong way round is the whole failure mode.
8
+
9
+ **The five levers**
10
+ 1. **Editing: short prompt, ONE change, nothing else.** Google's own doc says so and a 2026-07-09 receipt confirms it — a long "keep every frame identical" preamble produced WORSE drift than two sentences.
11
+ 2. **Editing: always end with `Keep everything else the same.`** — the one documented preservation lever.
12
+
13
+ 3. **Editing: describe the EFFECT, never a real object as a metaphor.** "Candle-like flame" rendered a literal candle in the subject's hand.
14
+ 4. **Editing: no chained stage directions.** Several beats cued to moments ("…flies onto his shoulder when he calls it, and perches as he walks…") hard-failed with `invalid_request`. One effect tied to an action already in the footage ("…when he snaps his fingers…") passed. Collapse to one continuous action; the model syncs to the footage's own motion.
15
+ 5. **Generation: the opposite — describe fully.** Subject, action, setting, `camera tracking alongside`, `overcast flat light`, tone. Audio is prompt-driven with no parameters: dialogue in quotes, sound in plain language — `rain patters on the tin roof`, `spray from tyres`, `a horn somewhere behind`.
16
+
17
+ **Examples**
18
+ - Edit: `Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.`
19
+ - Generate: `A courier in a yellow shell jacket weaves between stalled cars on a wet arterial road, camera tracking alongside at shoulder height. Overcast flat light, spray from tyres. Rain patters on car roofs, a horn somewhere behind.`
20
+
21
+ **Hard constraint:** it is a CHEAP DRAFT seat for generation and the EDIT-fidelity winner for footage-synced VFX — never a hero generation shot. Expect a possible jitter or doubled speech beat in the last half second of an edit: trim the tail rather than burning a re-roll.
22
+
23
+ **Never use in an EDIT prompt** (each one has a receipt above):
24
+ - a long preservation preamble — it produces WORSE drift than `Keep everything else the same.`
25
+ - a real object as a metaphor: `candle-like`, `flame-like`, `laser-like`
26
+ - several staged beats cued to moments (`when he calls it, and perches as he walks`): these hard-fail, they do not merely drift. One effect tied to an action already in the clip (`when he snaps his fingers`) passed
27
+ - harm-to-person framing: `ignite`, `catch fire`, `on fire` applied to a person trips the safety filter
28
+
29
+ Google's fast video generation + editing model ("Nano Banana Pro for video" in creator slang — a nickname; it is NOT the NB Pro image model). Carried on fal (`google/gemini-omni-flash*`). 720p only, 24fps, 3–10 second clips, 16:9 or 9:16. **Audio is native and included** — dialogue, SFX, and ambient generate WITH the video at no extra cost.
30
+
31
+ ## Where it routes
32
+
33
+ - **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: SKILL.md.
34
+ - **Drafts and iteration volume with sound**: an audio-native video seat (~6.4 cr/s at 720p).
35
+ - **Hero-generation quality is still unproven.** Use the current model-routing guide for the final generation seat; the edit-fidelity receipt does not establish generation quality.
36
+
37
+ ## Editing (model `omni-flash-edit`) — THE RULES (receipts, not theory)
38
+
39
+ 1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes."* Live receipt 2026-07-09: a long "keep every frame/word/movement identical…" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same."*
40
+ 2. **Always end with "Keep everything else the same."** — the one documented preservation lever.
41
+ 3. **Never name a real-world object as a metaphor.** "Candle-like flame" rendered a literal candle in his hand. Describe the effect itself ("small magical flames on his fingertips").
42
+ 3b. **No chained stage directions: they HARD-FAIL, not drift.** Receipt 2026-07-09: "a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…" → deterministic `invalid_request` (2×, "could not generate with the given inputs"); collapsing to one continuous action — "A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke." — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video. The winning flames prompt above ("when he snaps his fingers") shows one effect tied to an action the footage already contains is fine; which part of the dragon prompt triggered the refusal is untested beyond that.
43
+ 4. **Safety filter (Google's, strict about harm-to-person):** "fingertips ignite / catch fire" → `content_policy_violation`. Frame effects as magical/harmless VFX: "small magical flames appear on his fingertips" passed. See reference-content-policy.md §Gemini for the substitution patterns.
44
+ 5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.
45
+ 6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.
46
+ 7. Source clip 3–10s (trim longer clips first). Output length follows the source; billing per output second, rounded up. Voice editing unsupported — never ask it to change dialogue.
47
+ 8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).
48
+ 9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.
49
+
50
+ ## Generation (model `omni-flash`)
51
+
52
+ - **Inputs:** prompt only (t2v), prompt + ONE start frame (i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
53
+ - Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.
54
+ - **Name references inline** the standard Slates way ("Marcus (image 1) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
55
+ - **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language ("rain patters on the tin roof"). Negative direction as plain instructions ("Do not show text").
56
+ - Duration is an explicit 3–10s integer param; cost scales linearly per second.
57
+
58
+ ## Input conditioning (Slates handles this — know it exists)
59
+
60
+ Phone footage stores rotation as a metadata flag; models ignore it and edit the raw sideways pixels. Clips must be rotation-normalized (and oversized sources downscaled) before upload — receipt 2026-07-09: a portrait Pixel clip came back sideways until conditioned. If an edit output comes back rotated, the source wasn't normalized.
61
+
62
+ ## Content notes
63
+
64
+ - Google applies its own safety filters to input images/clips and output. Uploads containing recognizable real people are restricted by Google's policy — though own-footage editing of the uploader passed on our route 2026-07-09. See reference-content-policy.md.
65
+ - Output carries an invisible SynthID watermark (Google-side, programmatic detection only).
@@ -0,0 +1,362 @@
1
+ <!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
2
+
3
+ > Generated from the production Slates guide. Model-specific syntax and measured examples apply to the endpoints named below. For another generation tool, check its current schema and reference handling; its limits, billing and defaults may differ.
4
+
5
+ # Seedance 2.5 — prompting
6
+
7
+ ## Contents
8
+
9
+ - [The one fact that decides whether you use it at all](#the-one-fact-that-decides-whether-you-use-it-at-all)
10
+ - [🚨 Hazard 1 — the prompt-intent task classifier](#-hazard-1--the-prompt-intent-task-classifier)
11
+ - [🚨 Hazard 2 — resolution is not the price dial here. LENGTH is.](#-hazard-2--resolution-is-not-the-price-dial-here-length-is)
12
+ - [Timestamps — the one grammar change](#timestamps--the-one-grammar-change)
13
+ - [What the extra reference budget is actually for](#what-the-extra-reference-budget-is-actually-for)
14
+ - [Audio-only references — the genuinely new input](#audio-only-references--the-genuinely-new-input)
15
+ - [Video references](#video-references)
16
+ - [Mixing all three in one call](#mixing-all-three-in-one-call)
17
+ - [Sound: four bracket types, and they are the vendor's syntax](#sound-four-bracket-types-and-they-are-the-vendors-syntax)
18
+ - [Say what a reference is NOT for](#say-what-a-reference-is-not-for)
19
+ - [Seedance 2.5 Edit (model: 'seedance-2.5-edit')](#seedance-25-edit-model-seedance-25-edit)
20
+ - [Faces, and what does NOT change](#faces-and-what-does-not-change)
21
+
22
+ **Card — Seedance 2.5.** Shares 2.0's grammar exactly (subject binding, camera vocabulary, externalised emotion, inline constraints — read `reference-seedance.md` for those). Two things are different, and both matter.
23
+
24
+ **The five levers**
25
+ 1. **Timestamps work here**: integer seconds, and the model acts on them: `[0s-4s] she reads the letter. [4s-9s] she folds it and looks up.` 2.0 ignores exactly this syntax.
26
+ 2. **Length is the reason to be here**: takes up to 30 seconds, where 2.0 stops at 15. Write the beats as `[0s-6s]`, `[6s-12s]`, `[12s-18s]`; do not hope for them.
27
+ 3. **Up to 30 image references**, and a multi-view image can serve as ONE subject reference (up to 5 subjects). 2.0 cannot do either.
28
+ 4. **Audio-only references are accepted** without an image or video alongside — the only Seedance seat that takes one.
29
+ 5. **Keep the 2.0 discipline**: one camera move per beat (`slow track right`, `handheld follow`), physical action instead of stated emotion, and quality asked for in the image-quality slot vocabulary — `rich details`, `natural colors`, `cinematic texture`, `soft lighting`.
30
+
31
+ **Examples**
32
+ - `[0s-6s] Wide shot, <Subject_1>@<Image_1> crosses an empty car park toward a idling van, slow track right. [6s-12s] Medium, she stops as the driver's window comes down. [12s-18s] Close-up, she looks off past the lens and does not answer. Rich details, natural colors. Keep it subtitle-free.`
33
+ - `[0s-10s] A single continuous handheld follow behind a courier climbing a fire escape, rain. [10s-20s] She reaches the landing, turns, and the city opens behind her. Cinematic texture, soft lighting.`
34
+
35
+ **Hard constraint:** it is the default AND the dearer seat, and it has NO 4K — 480p/720p/1080p only, dearer than 2.0 at every resolution they share. Long takes multiply cost linearly: quote a 30-second take before you fire it.
36
+
37
+ **Never use** (with a reference video, 2.5 can reclassify the task and fail a fresh generation on these):
38
+ - `edit`, `extend`, `continue the video`, `same video but` — they make the provider read a fresh generation as an edit
39
+ - `f/1.4`, `Portra 400` and any other aperture or film-stock token, or a stacked list of gear — image-model vocabulary. The 2.5 guide's own example names one camera body and one 35 mm lens in a single style line, so a lone lens there is not on this list
40
+
41
+ **Read `reference-seedance.md` first.** The prompt GRAMMAR is the same model family: the
42
+ 8-slot advanced formula, subject binding by `<Subject_N>@<Image_N>`, camera vocabulary,
43
+ externalised emotion, inline constraint words, the anti-twin fix. None of it is restated here.
44
+ This file is only what 2.5 changes — and the biggest change is that **2.5 acts on timestamps
45
+ where 2.0 ignores them.**
46
+
47
+ ---
48
+
49
+ ## The one fact that decides whether you use it at all
50
+
51
+ **Seedance 2.5 is the EXPENSIVE seat, not the cheap one — and it has no 4K.**
52
+
53
+ It runs at 480p, 720p or 1080p (1080p landed on all three routes on 2026-08-24), and at every
54
+ resolution the two seats share it costs MORE than 2.0 — 720p $0.231/s against $0.15/s, **54% more**.
55
+ So 2.5 does not replace 2.0; it sits beside it, and you pay for what it buys:
56
+
57
+ | | Seedance 2.0 | Seedance 2.5 |
58
+ |---|---|---|
59
+ | Resolution | 480p / 720p / 1080p / **native 4K** | 480p / 720p / 1080p — **no 4K** |
60
+ | Price at 720p (faceless) | **$0.15/s** | $0.231/s |
61
+ | Length | 4–15s | **4–30s in one take** |
62
+ | Reference budget | 12 files total (9 image / 3 video / 3 audio caps) | **50 (30 image + 10 video + 10 audio)** |
63
+ | Combined reference video/audio | ≤15s | **≤30s** |
64
+ | Audio-only reference | ✗ (needs an image or video alongside) | **✓** |
65
+ | **Timestamps in the prompt** | **✗ — ignored; shot numbers only** | **✓ — integer seconds, acted on** |
66
+ | Multi-view image as ONE subject reference | ✗ (not recommended) | **✓ (up to 5 subjects)** |
67
+ | Video edit as its own task type | ✗ | **✓ (`seedance-2.5-edit`)** |
68
+ | Default video model | no | **yes** (since 2026-09-13) |
69
+
70
+ **2.5 is the default. Route to 2.0 for 4K delivery, or when the same resolution has to be
71
+ cheaper** — its 720p is $0.15/s against 2.5's $0.231/s.
72
+
73
+ ---
74
+
75
+ ## 🚨 Hazard 1 — the prompt-intent task classifier
76
+
77
+ This is the one that costs money and time, and it has no equivalent on 2.0.
78
+
79
+ **Seedance 2.5 sorts every request into one of five task types** — text-to-video,
80
+ reference-to-video, first/last-frame, **video edit**, **video extend** — from the reference roles
81
+ attached **plus the intent of your sentence**. Each type then has its own parameter constraints,
82
+ and a violation comes back **asynchronously**: the task queues, credits are reserved, and only then
83
+ does it fail.
84
+
85
+ The trigger words are ordinary English:
86
+
87
+ | Reclassified as | Words that do it (ByteDance's own list) |
88
+ |---|---|
89
+ | **video edit** | `edit video` · `add` · `insert` · `remove` · `delete` · `modify` · `replace` · `change to` |
90
+ | **video extend** | `extend forward` · `extend backward` · `continue` · `continue from` · `extend the story` |
91
+
92
+ So a perfectly legitimate prompt with a reference video, *"a wide shot of the workshop, **remove** the
93
+ tripod from frame"* — gets classified as an edit and fails on constraints it never set.
94
+
95
+ **What to do:**
96
+
97
+ 1. **If you mean to edit an existing clip, choose its dedicated video-edit endpoint.** The
98
+ task-typed endpoint removes the classifier's ambiguity.
99
+ 2. **If you mean a fresh shot, describe the finished frame rather than an instruction to change
100
+ one.** Not *"remove the tripod"* → *"the workshop bench, clear and uncluttered"*. Not
101
+ *"add rain"* → *"heavy rain falling through the streetlight"*. This is better prompting anyway:
102
+ the model renders what you describe, it does not take edits to an imagined draft.
103
+ 3. The trigger needs **a reference video plus edit or extend intent**. Image references alone do not trigger it. A plain text-to-video prompt is safe
104
+ however it is worded.
105
+
106
+ **Slates will warn you, and it will never rewrite your prompt.** When a 2.5 reference generation's
107
+ prompt contains one of these words, the composer shows a warning that NAMES the words and the agent
108
+ route returns the same string. Silently editing the user's sentence to dodge a provider classifier
109
+ is forbidden — the words that reach the model are always the words the user can see.
110
+
111
+ ---
112
+
113
+ ## 🚨 Hazard 2 — resolution is not the price dial here. LENGTH is.
114
+
115
+ Every other model in Slates trains the habit that lower resolution means lower cost. 2.5 breaks it,
116
+ because the thing that moves the bill is **length**, and 2.5's length ceiling is double 2.0's.
117
+
118
+ Worked, at the shipped rates:
119
+
120
+ | Generation | Credits |
121
+ |---|---|
122
+ | 2.5 · 480p · 5s · faceless | 26 |
123
+ | 2.5 · 720p · 5s · faceless | 58 |
124
+ | 2.5 · 1080p · 5s · faceless | 142 |
125
+ | 2.5 · 720p · 30s · faceless | 347 |
126
+ | 2.5 · 720p · 30s · AI-face route | **489** |
127
+ | 2.5 · 720p · 30s · consented real-face route | **710** |
128
+ | 2.5 · 1080p · 30s · faceless | **853** |
129
+ | 2.5 · 1080p · 30s · consented real-face route | **1,749** |
130
+ | *(for scale)* 2.0 · 1080p · 15s · AI-face route | 411 |
131
+
132
+ **A 30-second 720p clip can cost more than a 15-second 1080p one** — and a base licence starts
133
+ with 1,000 credits. Someone who reads "720p" as "cheap" and asks for a 30-second take on the
134
+ real-face route has spent 71% of their welcome grant on one clip; **on the real-face route a single
135
+ 30-second 1080p take is more than the whole grant.**
136
+
137
+ **Discipline:**
138
+
139
+ - **Find the shot at short LENGTH, not at low resolution.** Length is what moves the price, so cut
140
+ seconds while you are still exploring — 4–8s — and stay at the resolution you actually want.
141
+ **A 480p pass does not de-risk a 720p or 1080p render.** Generation is stochastic: the higher-
142
+ resolution run is a different take, not the same shot rendered better. So a 480p draft that looks
143
+ right buys you no guarantee, and one that looks wrong may have been fine at 720p — you paid 26
144
+ credits to learn nothing, when 58 would have bought a real candidate.
145
+ - **Length is a creative decision, not a default.** 30 seconds is available; it is rarely the right
146
+ answer for a single shot. Multi-shot storyboards inside one 30s generation are what the length is
147
+ actually for.
148
+
149
+ ---
150
+
151
+ ## Timestamps — the one grammar change
152
+
153
+ <!-- @inject:seedance-25-timestamps -->
154
+ **2.0 does not respond to timestamps and answers only to shot numbers. 2.5 responds to
155
+ integer-second timestamps.** That is ByteDance's own first line under "Differences from Seedance
156
+ 2.0", and it is why a 30-second take is usable at all: the length is only worth buying if you can
157
+ say *when* things happen inside it.
158
+
159
+ Both formats are valid on 2.5, and you can mix them — `Shot N` blocks for a storyboard whose
160
+ pacing you are happy to leave to the model, timestamps when a beat has to land at a moment.
161
+
162
+ **Three ways to control time, all first-party:**
163
+
164
+ | Form | Write it like |
165
+ |---|---|
166
+ | **Interval** | `0-3 seconds… 3-7 seconds… 7-15 seconds` or `[1s-4s]… [4s-8s]… [8s-12s]` |
167
+ | **Time point** | *"Quick left sideways transition at the 5-second mark."* |
168
+ | **Relative** | *"After 3 seconds, everyone around him shakes their head."* · *"The frame freezes for 1 second after he presses the shutter."* |
169
+
170
+ **The rules that come with them:**
171
+
172
+ - **One second is the smallest unit.** Integers only — no `2.5s`, no frames.
173
+ - **No gaps in the timeline.** `0-3s… 5-6s…` leaves 3-5s unspecified and the model fills it however
174
+ it likes. Intervals must abut: `0-3s`, `3-7s`, `7-15s`.
175
+ - **Budget the plot to the seconds.** Too little content in a range and the model improvises to
176
+ fill it; too much and you get extra cuts or dropped beats. This is the actual craft of a 30s take.
177
+ - **Never time-code a high-frequency action.** *"Shake your head three times per second"* is
178
+ explicitly called out as a misuse — timestamps schedule beats, they don't choreograph frames.
179
+ - **Transitions want both halves:** the moment AND the method — *"At the 5-second mark, the camera
180
+ transitions leftward with a left wipe into a natural dissolve."*
181
+ - **Timestamps work on an EDIT too**, and that is where they earn the most: they scope a change in
182
+ time as well as in content — *"Change the man's action from drinking coffee to mopping the floor
183
+ from 4-6 seconds in Video 1, and leave the rest of the content unchanged."* Without a range, a
184
+ whole-clip instruction is applied to the whole clip.
185
+
186
+ Do **not** carry this back to 2.0, and do not write `[00:00-00:02]` minute-second brackets (another
187
+ vendor's syntax) into either — 2.0 ignores time entirely, and the cross-model syntax swap is its own known failure.
188
+ <!-- @end:seedance-25-timestamps -->
189
+
190
+ ---
191
+
192
+ ## What the extra reference budget is actually for
193
+
194
+ 30 image references (up from 9) does **not** mean "attach 30 images". Every rule in
195
+ `reference-seedance.md` about references still holds — 2–4 strong references beat both
196
+ extremes, and one reference per role.
197
+
198
+ **Where 2.5 moves the ceiling, per ByteDance's own input recommendations:**
199
+
200
+ | | Stable | Works, but expect re-rolls |
201
+ |---|---|---|
202
+ | Subjects bound by IMAGE reference | 1–8 | 9–12 |
203
+ | Subjects bound by VIDEO or AUDIO reference | 1–5 | 6–10 |
204
+ | Reference clip length, per subject | 5–10s | longer drops stability |
205
+
206
+ **Multi-view images of one subject are supported on 2.5** — a turnaround sheet can be a single
207
+ reference image, where 2.0 wanted one authoritative rendering per subject. Past **5 subjects**,
208
+ go back to single-view images, one per view, rather than one image carrying several viewpoints.
209
+
210
+ The larger budget earns its keep in exactly two places:
211
+
212
+ - **A long multi-shot take** where different shots need different subjects and locations bound —
213
+ the budget is spread across the storyboard, not stacked on one frame.
214
+ - **Video and audio references alongside images**, which is where 2.5's 10 + 10 matters far more
215
+ than the image count.
216
+
217
+ ### Audio-only references — the genuinely new input
218
+
219
+ 2.0 required an image or video alongside any audio reference. **2.5 accepts audio on its own.**
220
+ That makes one recipe possible that was not before: drive a scene's timing, voice or ambience from
221
+ a recording with no visual anchor at all — a voice line, a music bed, a room tone — and let the
222
+ model build the picture to it. Cite it the same way as any other reference
223
+ (`Reference the timbre in <Audio_N> to generate…`), and remember that audio references carry **no
224
+ billing dimension** on any Seedance route: audio is included.
225
+
226
+ ### Video references
227
+
228
+ Up to 10 clips, ≤30s combined (2.0: 3 clips, ≤15s). A reference VIDEO switches the cost key to
229
+ `seedance-2.5*-vref-{res}-{T}s`, where **T = Σ input seconds + output seconds** on faceless and real-face routes;
230
+ on the AI-face route, **T = max(Σ input seconds, output seconds) + output seconds** on both 2.0 and 2.5. The sum is across
231
+ **every** clip attached, not just the longest. Three 6-second references on a 12-second output bills
232
+ 30 seconds, not 12 and not 18. Quote before confirming.
233
+
234
+ **The cap is a refusal, not a trim.** Attach an eleventh clip, or push past 30 combined seconds, and
235
+ the composition is rejected before anything uploads. That asymmetry is deliberate: reference images
236
+ warn-and-trim because dropping one doesn't change the price, and a dropped reference VIDEO would be
237
+ one you were quoted for and the model never saw.
238
+
239
+ ### Mixing all three in one call
240
+
241
+ 50 files total (30 image + 10 video + 10 audio) is a shared budget. Everything is cited positionally
242
+ by type — `image 1`, `video 2`, `audio 1` — in attachment order, so reordering the attachments
243
+ renumbers the citations. Write the prompt against those numbers:
244
+
245
+ ```
246
+ Marcus (image 1) performs the motion from video 1 in the workshop from image 2,
247
+ using the voice timbre from audio 1. Preserve his identity, appearance and outfit.
248
+ ```
249
+
250
+ 🚨 **SAY WHAT AN AUDIO REFERENCE IS FOR.** It can mean music, dialogue, voice, tone or timbre — five roles on one attachment — so an unroled clip falls back to **dialogue**: the model re-transcribes it and speaks ITS words. A real take came back as *"a map called Slates"* for *"an app called Slates"*. Name it as the voice timbre and the clip carries the voice while the prompt carries the words. ByteDance's own sentence: *"Image 1 depicts the protagonist John and uses the voice timbre from Audio 1."* Bind each speaker in a sentence, never by attachment order — position carries nothing.
251
+
252
+ Frames and reference media stay mutually exclusive, in every combination — the reference endpoint
253
+ has no first/last-frame parameters at all, so this is a shape mismatch rather than a preference.
254
+
255
+ ---
256
+
257
+ ## Sound: four bracket types, and they are the vendor's syntax
258
+
259
+ ByteDance's 2.5 API tutorial states this as a **prompt rule**, not a suggestion — verbatim: *"Use
260
+ special characters to distinguish sounds: `()` for music, `<>` for sound effects, `{}` for dialogue,
261
+ and `【】` for subtitles. For non-Chinese dialogue, it is recommended to specify the language before
262
+ the dialogue."*
263
+
264
+ ```
265
+ She sets the cup down {English: "We open in ten minutes."} <ceramic clink on wood>
266
+ (low piano, unhurried)
267
+ ```
268
+
269
+ - `()` **music** · `<>` **sound effects** · `{}` **dialogue** · `【】` **on-screen subtitles**
270
+ - **Name the language before non-Chinese dialogue.** `{English: "..."}`.
271
+ - Unbracketed sound description still works — this is a disambiguator, not a required wrapper. Reach
272
+ for it when one sentence carries more than one kind of sound and you need the model to tell them
273
+ apart, which is exactly where an unmarked prompt puts a line of dialogue into the score.
274
+
275
+ ⚠️ **These four are SEEDANCE 2.5's.** MiniMax H3 has its own three-layer scheme (body / soundscape /
276
+ score) and its angle brackets are documentation notation that must never be typed. Do not carry
277
+ either grammar onto the other model.
278
+
279
+ ## Say what a reference is NOT for
280
+
281
+ The same rule adds a half nobody uses: *"Specify what each asset provides, such as appearance,
282
+ action, or timbre, **and what should not be referenced**."* Negative scoping is a first-class part of
283
+ the citation, not a fallback — *"use her face and wardrobe from image 1, not its lighting or
284
+ background"* is a stronger instruction than naming the positive alone, because an unscoped reference
285
+ brings its whole frame with it.
286
+
287
+ **BytePlus documents disagree on reference sigils.** Its API tutorial says *"Use `@Image 1`,
288
+ `@Video 1`, and `@Audio 1`"*, while its 2.5 prompt guide uses the bare form
289
+ (`Image 1 / Video 1 / Audio 1`) in its normative sentence. Both are first-party sources; the
290
+ bare form is confirmed working on both Seedance models. Use the syntax accepted by the endpoint
291
+ you are calling, and preserve each asset's role and scope.
292
+
293
+ ## Seedance 2.5 Edit (`model: 'seedance-2.5-edit'`)
294
+
295
+ Its own picker row and its own op call, deliberately: the task type is **the model you chose**,
296
+ never something inferred from your sentence.
297
+
298
+ **Why route here at all:** it is the **only edit engine in Slates that accepts a clip longer than 15
299
+ seconds** (4–30s, versus Kling O3 Edit's 3–15s and Omni Flash Edit's 3–10s). For a clip inside the
300
+ others' range, choose on fidelity instead — Omni Flash Edit won the prompt-only head-to-head, and
301
+ Kling O3 Edit is the one that takes element and style reference images.
302
+
303
+ **How it behaves:**
304
+
305
+ - **Output length follows the SOURCE clip**, and the bill is the ceiled source length. The provider
306
+ requires an automatic duration on this task type, so there is no length knob — the clip you attach
307
+ is the quote. The returned clip can differ from the source by up to ~0.3s, which only compresses
308
+ transition frames; a clip that 2.5 itself generated comes back at exactly its input length.
309
+ - **Source clips under 20 seconds edit more reliably.** 4–30s is what the task type accepts;
310
+ ByteDance's own recommendation is to stay inside 20 for quality. A 28-second source is legal and
311
+ will need more attempts.
312
+ - **The aspect ratio follows the source clip too.** No ratio control; the frame is the clip's frame.
313
+ - **480p, 720p or 1080p output**, native audio — and an edit bills the video-reference tier ×2,
314
+ so quote the 1080p edit before confirming.
315
+ - **Prompt and source clip only** on this op. The MODEL takes reference images on an edit
316
+ (ByteDance recommends 1–5 — *"replace the man in dark clothing in @Video 1 with @Image 2"*);
317
+ **Slates has not wired that path**, so today an edit that must lock an identity from a photo
318
+ goes to Kling O3 Edit. Constraint of our build, not of the model — worth revisiting.
319
+ - **An edit costs about 1.2x a plain 2.5 generation of the same length**: every Seedance provider
320
+ charges an edit on input + output seconds, at the reduced video-reference rate. Read the confirm
321
+ gate's number; do not reason from the generation rate.
322
+ - **Set `seedanceFace: true` when a character's face is visible in the clip.** The faceless provider
323
+ blocks faces outright — this is not a price optimisation, it is whether the job runs at all.
324
+ - **There is no consented-real-face route for editing.** Real-person footage that the AI-face route
325
+ rejects has to go to Kling O3 Edit.
326
+
327
+ **Prompting an edit** — the same discipline as every other edit engine: **describe only what
328
+ changes.** The source already carries its composition, motion, timing and performance; re-describing
329
+ them fights the model. Use Seedance's own edit grammar from `reference-seedance.md`
330
+ (*"Strictly edit `<Video_1>`, and modify `<Original_Characteristic>` to `<New_Characteristic>`"*) and
331
+ **never** write *"reference video 1"* in an edit — the official guide is explicit that this phrasing
332
+ gets the request reclassified as a reference task, which is the same landmine as Hazard 1.
333
+
334
+ Two things sharpen an edit prompt, both first-party:
335
+
336
+ - **Say it as A → B, not as an outcome.** *"Change the man's action from drinking coffee to mopping
337
+ the floor"* beats *"the man mops the floor"* — naming what it currently is tells the model what
338
+ to overwrite.
339
+ - **Timestamp a partial edit** — the edit task type reads the same integer-second timestamps the
340
+ generation path does. Rules and forms are in § Timestamps above; this is the single most useful
341
+ thing they buy.
342
+
343
+ **Audio is editable too, and it is the least obvious use of this row.** The same op rewrites what
344
+ is heard while the picture stays put: change a spoken line, change the accent, translate the
345
+ dialogue and re-fit the lip movement, strip or replace the BGM or a sound effect. *"Only edit the
346
+ man's dialogue in Video 1: change it to 'Don't come over here,' in an American accent"* is an edit,
347
+ not a lip-sync job. Bill it like any other edit — on the source clip's length.
348
+
349
+ ---
350
+
351
+ ## Faces, and what does NOT change
352
+
353
+ Also unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:
354
+
355
+ - **One primary camera move per shot.**
356
+ - **No stacked lens / aperture / film-stock vocabulary.** One camera-and-lens style line is the
357
+ most the 2.5 guide itself uses; the full rule is in `reference-seedance.md` → "Don't
358
+ cross-pollinate image-model syntax".
359
+ - **No `negativePrompt` field** — constraints go inline, and 2.5 acts on negative phrasing in
360
+ exactly two dimensions: subtitles (*"no subtitles"*) and audio (*"no BGM; environmental and
361
+ action sounds only"*, *"no audio"*). Everywhere else, describe what you want, not what you don't.
362
+ - **Legible in-shot text still belongs in a baked start frame**, not in the video prompt.
@@ -1,10 +1,41 @@
1
1
  <!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
2
2
 
3
- > **This is the real thing.** Every rule below is the working doctrine Slates runs in production against this model — not a summary written for a handout. Slates automates it end to end; the doctrine works by hand too.
3
+ > Generated from the production Slates guide. Model-specific syntax and measured examples apply to the endpoints named below. For another generation tool, check its current schema and reference handling; its limits, billing and defaults may differ.
4
4
 
5
5
  # Seedance 2.0 — prompting
6
6
 
7
- <!-- @card:start -->
7
+ ## Contents
8
+
9
+ - [What Seedance actually is [official :1450-1452]](#what-seedance-actually-is-official-1450-1452)
10
+ - [The advanced formula — 8 slots [official :1455]](#the-advanced-formula--8-slots-official-1455)
11
+ - [Task-type sentence patterns [official :1389-1425]](#task-type-sentence-patterns-official-1389-1425)
12
+ - [⚠️ Edit / extend phrasing landmine [official :1431]](#-edit--extend-phrasing-landmine-official-1431)
13
+ - [Shot structure — "Shot 1 / Shot 2 / Shot 3" [official :1563-1598]](#shot-structure--shot-1--shot-2--shot-3-official-1563-1598)
14
+ - [Subject binding — names + image indexes [official :1488-1556]](#subject-binding--names--image-indexes-official-1488-1556)
15
+ - [Action description [official :1602-1621]](#action-description-official-1602-1621)
16
+ - [Externalize emotion [official :1623-1636]](#externalize-emotion-official-1623-1636)
17
+ - [Camera [official :1643-1648]](#camera-official-1643-1648)
18
+ - [Image quality, style, and constraints [official :1656-1679]](#image-quality-style-and-constraints-official-1656-1679)
19
+ - [🔴 Duplicated characters — the twin problem [official :1948-1994]](#-duplicated-characters--the-twin-problem-official-1948-1994)
20
+ - [Worked examples [official :1689-1745]](#worked-examples-official-1689-1745)
21
+ - [Other official notes](#other-official-notes)
22
+ - [Reference media — caps and transport](#reference-media--caps-and-transport)
23
+ - [All three modalities go in ONE call](#all-three-modalities-go-in-one-call)
24
+ - [Motion transfer & lip-sync recipes (reference video / audio)](#motion-transfer--lip-sync-recipes-reference-video--audio)
25
+ - [Reference rules (the verified ones)](#reference-rules-the-verified-ones)
26
+ - [For Seedance specifically](#for-seedance-specifically)
27
+ - [Length](#length)
28
+ - [Pin the subject in the first 20-30 words](#pin-the-subject-in-the-first-20-30-words)
29
+ - [Lighting is a top quality lever](#lighting-is-a-top-quality-lever)
30
+ - [Camera and subject motion — separate sentences](#camera-and-subject-motion--separate-sentences)
31
+ - [Slow-motion works; "fast" is a known bad token](#slow-motion-works-fast-is-a-known-bad-token)
32
+ - [Style block at the end](#style-block-at-the-end)
33
+ - [⚠️ Don't cross-pollinate image-model syntax](#-dont-cross-pollinate-image-model-syntax)
34
+ - [Negative prompting — inline only](#negative-prompting--inline-only)
35
+ - [Image-to-video / first-frame guidance](#image-to-video--first-frame-guidance)
36
+ - [Common failure modes + fixes](#common-failure-modes--fixes)
37
+ - [Sources](#sources)
38
+
8
39
  **Card — Seedance 2.0.** Not copywriting — an ENGINEERING instruction to a spatial layer and a temporal layer: who, in what scene, doing what, how the camera moves, and in what order. Multi-beat work is a `Shot 1 / Shot 2 / Shot 3` storyboard.
9
40
 
10
41
  **The five levers**
@@ -19,7 +50,6 @@
19
50
  - `Single continuous take. Wide shot of a fishing skiff crossing a grey swell, camera tracks from the starboard rail. Spray hits the lens once. Soft lighting, natural colors, film-grain texture. Avoid jitter and bent limbs.`
20
51
 
21
52
  **Hard constraint:** NO timestamps — 2.0 ignores them and answers only to shot numbers (2.5 acts on them). No lens, aperture, film stock or camera body: that is image-model vocabulary and a Seedance anti-pattern. There is no negativePrompt field.
22
- <!-- @card:end -->
23
53
 
24
54
  ByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K — 4K video is Pro-only, default 1080p), 4–15s, first+last frame, and up to 9 reference images / 3 videos / 3 audio clips.
25
55
 
@@ -235,7 +265,7 @@ using the voice timbre from audio 1. Preserve his identity, appearance and outfi
235
265
 
236
266
  ### Motion transfer & lip-sync recipes (reference video / audio)
237
267
 
238
- These aren't separate Seedance features — they're prompting strategies over reference media.
268
+ These aren't separate Seedance features; they're prompting strategies over reference media.
239
269
 
240
270
  - **Motion transfer:** subject image as a reference + the driving clip (2–15s) + `The character from image 1 performs the exact motion, choreography, and camera movement from video 1. Preserve the character's identity, appearance, and outfit.`
241
271
  - **Lip-sync / dialogue:** write the line in the prompt — `The person in video 1 says: "…"` — with audio generation on (always on in Slates). A **video** source's own voice is cloned natively; an **audio** reference (≤15s) drives speech from an existing recording: `…speaks the dialogue from audio 1 with accurate lip sync.`