@slatesvideo/shared 0.7.2 → 0.7.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/clients/cloud.d.ts +4 -0
- package/dist/clients/cloud.js +11 -3
- package/dist/index.d.ts +1 -0
- package/dist/index.js +1 -0
- package/dist/manual/content.d.ts +1 -1
- package/dist/manual/content.js +1 -1
- package/dist/operations/index.d.ts +12 -13
- package/dist/operations/index.js +158 -133
- package/dist/operations/surface.d.ts +6 -2
- package/dist/operations/surface.js +29 -5
- package/dist/prompts/agent-doctrine.d.ts +4 -4
- package/dist/prompts/agent-doctrine.js +17 -28
- package/dist/prompts/guide-discovery.d.ts +23 -0
- package/dist/prompts/guide-discovery.js +39 -0
- package/dist/prompts/guide-retrieval.js +1 -1
- package/dist/prompts/model-capabilities.d.ts +8 -9
- package/dist/prompts/model-capabilities.js +11 -51
- package/dist/prompts/model-facts.d.ts +2 -2
- package/dist/prompts/model-facts.js +15 -26
- package/dist/prompts/partials.generated.js +6 -3
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +21 -63
- package/dist/prompts/search-terms.d.ts +3 -0
- package/dist/prompts/search-terms.js +24 -0
- package/dist/skills/content.js +36 -37
- package/dist/skills/metadata.d.ts +7 -0
- package/dist/skills/metadata.js +29 -0
- package/exports/slates-chatgpt-images/generated/SKILL.md +7 -1
- package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
- package/exports/slates-prompt-builder/generated/SKILL.md +28 -16
- package/exports/slates-prompt-builder/generated/reference-character.md +12 -13
- package/exports/slates-prompt-builder/generated/reference-content-policy.md +2 -2
- package/exports/slates-prompt-builder/generated/reference-gpt-image-2-5.md +191 -0
- package/exports/slates-prompt-builder/generated/reference-kling.md +32 -11
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +24 -6
- package/exports/slates-prompt-builder/generated/reference-omni-flash.md +65 -0
- package/exports/slates-prompt-builder/generated/reference-seedance-2-5.md +362 -0
- package/exports/slates-prompt-builder/generated/reference-seedance.md +34 -4
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +77 -23
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +2 -1
- package/skills/_partials/blender-action-curves.md +24 -0
- package/skills/_partials/iteration-diagnosis.md +5 -0
- package/skills/_partials/model-routing.md +35 -0
- package/skills/_partials/seedance-25-timestamps.md +2 -2
- package/skills/_partials/still-gate.md +2 -2
- package/skills/_partials/thresholds.md +1 -1
- package/skills/slates-blocking-to-prompt.md +15 -13
- package/skills/slates-camera-language.md +45 -7
- package/skills/slates-character-identity.md +8 -6
- package/skills/slates-chatgpt-images.md +7 -1
- package/skills/slates-cinematic-look.md +1 -1
- package/skills/slates-content-policy.md +4 -6
- package/skills/slates-cost-discipline.md +18 -12
- package/skills/slates-dialogue-blocking.md +6 -6
- package/skills/slates-direct-response-ad.md +1 -1
- package/skills/slates-edit-and-iterate.md +12 -4
- package/skills/slates-model-selection.md +82 -90
- package/skills/slates-one-prompt-film.md +1 -1
- package/skills/slates-previs-blocking.md +44 -13
- package/skills/slates-project-organization.md +2 -2
- package/skills/slates-prompting-elevenlabs.md +4 -4
- package/skills/slates-prompting-flux-2-max.md +2 -3
- package/skills/slates-prompting-gpt-image-2-5.md +2 -2
- package/skills/slates-prompting-inworld-tts.md +174 -174
- package/skills/slates-prompting-kling-v3.md +11 -9
- package/skills/slates-prompting-lip-sync.md +15 -15
- package/skills/slates-prompting-ltx-2-5.md +5 -6
- package/skills/slates-prompting-minimax-h3.md +11 -11
- package/skills/slates-prompting-motion-transfer.md +8 -8
- package/skills/slates-prompting-nano-banana-2.md +8 -4
- package/skills/slates-prompting-omni-flash.md +9 -9
- package/skills/slates-prompting-seed-audio.md +24 -4
- package/skills/slates-prompting-seedance-2-5.md +40 -30
- package/skills/slates-prompting-seedance.md +4 -4
- package/skills/slates-prompting-seedream-5-lite.md +6 -6
- package/skills/slates-restyle-from-blocking.md +2 -2
- package/skills/slates-script-craft.md +1 -1
- package/skills/slates-shot-variety.md +1 -1
- package/skills/slates-storyboard-from-script.md +1 -1
- package/skills/slates-style-prompting.md +56 -54
- package/skills/slates-ugc-influencer-ad.md +1 -1
- package/skills/slates-vision-feedback-loop.md +118 -110
- package/skills/slates-prompting-veo-3.md +0 -224
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-kling-v3
|
|
3
|
-
description:
|
|
3
|
+
description: "Prompt Kling V3.0 video generation and Kling O3 video edits. Use with Kling models on slates_generate_video or slates_edit_video; covers subjects, dialogue, sound syntax, multi-shot direction and edit fidelity."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Kling V3.0 — prompting
|
|
@@ -15,7 +15,7 @@ description: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_gen
|
|
|
15
15
|
Keep it under 2,400 characters (the build fails above that) and keep the
|
|
16
16
|
rationale, the receipts and the worked examples in the body below. -->
|
|
17
17
|
<!-- /slates-only -->
|
|
18
|
-
**Card — Kling V3.0.**
|
|
18
|
+
**Card — Kling V3.0.** Define the core subjects clearly at the START and keep those descriptions identical across shots. Strong image-to-video identity hold; use the current capability surface for duration and multi-shot limits, and the model catalogue for routing.
|
|
19
19
|
|
|
20
20
|
**The five levers**
|
|
21
21
|
1. **Dialogue in quotes** — `Character says, "exact words here"`. On Omni, direct the voice with `Gender + Age + Voice quality + Speech rate + Emotional tone + Language`: `[Character A: Detective, mid-40s, raspy, slow cadence, weary]: "I've seen this before."`
|
|
@@ -44,7 +44,7 @@ description: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_gen
|
|
|
44
44
|
- `single continuous take` — Seedance's phrase, and it fights Kling's multi-shot
|
|
45
45
|
<!-- @banned:end -->
|
|
46
46
|
|
|
47
|
-
Kuaishou's video model. Three tiers: `kling-v3.0-std` (general use,
|
|
47
|
+
Kuaishou's video model. Three tiers: `kling-v3.0-std` (general use, sound supported), `kling-v3.0-pro` (higher visual quality, sound supported), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).
|
|
48
48
|
|
|
49
49
|
Up to 15s. Multi-shot supported (up to 6 cuts in 15s total). Strong on image-to-video — preserves identity, layout, and text from the input image well.
|
|
50
50
|
|
|
@@ -134,7 +134,9 @@ Miss conditions:
|
|
|
134
134
|
- Mixing camera moves within a shot ("pan then orbit then push in")
|
|
135
135
|
- Extreme wide → extreme close in adjacent shots without reference images
|
|
136
136
|
|
|
137
|
-
## Element references
|
|
137
|
+
## Element references
|
|
138
|
+
|
|
139
|
+
Standard and Pro take element references with a first frame; Omni also takes references without one. 4K refuses reference images.
|
|
138
140
|
|
|
139
141
|
Upload 2-4 multi-angle reference photos per character/object. Tag inline:
|
|
140
142
|
|
|
@@ -199,11 +201,11 @@ Layer scene-specific suppressions on top, and never suppress something the promp
|
|
|
199
201
|
|
|
200
202
|
## Tier choice
|
|
201
203
|
|
|
202
|
-
- **Standard**: general use,
|
|
203
|
-
- **Pro**: higher visual quality,
|
|
204
|
-
- **Omni**: multi-character dialogue, audio-visual co-gen, language codes,
|
|
204
|
+
- **Standard**: general use, sound supported
|
|
205
|
+
- **Pro**: higher visual quality, sound supported
|
|
206
|
+
- **Omni**: multi-character dialogue, audio-visual co-gen, language codes, references without a first frame
|
|
205
207
|
|
|
206
|
-
|
|
208
|
+
Every tier can generate dialogue and sound. Sound is on unless `sound: false` is passed; below 4K it bills the audio key, while 4K includes audio. Pick by visual quality and reference needs. Prices change; check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.
|
|
207
209
|
|
|
208
210
|
## Benchmark prompt structure
|
|
209
211
|
|
|
@@ -253,7 +255,7 @@ Rules:
|
|
|
253
255
|
- One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).
|
|
254
256
|
- Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.
|
|
255
257
|
- Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.
|
|
256
|
-
-
|
|
258
|
+
- Route by the required change: this edit seat supports element/style-reference control and original-audio retention. Read the current catalogue for defaults and competing seats<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.
|
|
257
259
|
|
|
258
260
|
## Sources
|
|
259
261
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-lip-sync
|
|
3
|
-
description:
|
|
3
|
+
description: "Prepare Kling lip-sync or avatar generation with slates_generate_lip_sync. Covers source selection, voice, framing, audio constraints and the separate Seedance video-reference alternative."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Lip-sync — setup guide
|
|
@@ -15,7 +15,7 @@ description: How to set up lip-sync — Kling-only (dedicated lip-sync and avata
|
|
|
15
15
|
Keep it under 2,400 characters (the build fails above that) and keep the
|
|
16
16
|
rationale, the receipts and the worked examples in the body below. -->
|
|
17
17
|
<!-- /slates-only -->
|
|
18
|
-
**Card
|
|
18
|
+
**Card: Lip-sync (Kling only).** Two different flows with different inputs and different prices; output follows the source clip for video or the voice track for a still, billed per 5s block.
|
|
19
19
|
|
|
20
20
|
**The five levers**
|
|
21
21
|
1. **Pick `sourceType` deliberately** — `video` re-dubs an existing talking head (cheapest); `image` animates a still portrait (avatar-standard, then avatar-pro only on the final selected take).
|
|
@@ -28,7 +28,7 @@ description: How to set up lip-sync — Kling-only (dedicated lip-sync and avata
|
|
|
28
28
|
- `Soft rim light, warm office, gentle confident smile between sentences.`
|
|
29
29
|
- `Cool blue evening light through a window, focused intent expression.` (Or `.` — an empty prompt is fine when you have nothing to add.)
|
|
30
30
|
|
|
31
|
-
**Hard constraint:** it is Kling-only and
|
|
31
|
+
**Hard constraint:** it is Kling-only; output follows the media and bills per 5s block. For a generated PERFORMANCE instead (head movement, gesture, delivery energy, with the dialogue as a native conditioning signal), that is a normal Seedance video generation with the clip attached as a video reference, not a mode of this tool. A real recording, or a cloned/cast voice rendered on `inworld-tts-2`, for production; this tool's built-in TTS is for scratch.
|
|
32
32
|
<!-- @card:end -->
|
|
33
33
|
|
|
34
34
|
<!-- @banned:start -->
|
|
@@ -43,13 +43,13 @@ description: How to set up lip-sync — Kling-only (dedicated lip-sync and avata
|
|
|
43
43
|
- `reader_en_m-v1` — listed in fal's docs, returns "Voice id not found" in production
|
|
44
44
|
<!-- @banned:end -->
|
|
45
45
|
|
|
46
|
-
**This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints;
|
|
46
|
+
**This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; output follows the source clip for video or the voice track for a still, billed per 5s block.
|
|
47
47
|
|
|
48
48
|
| Flow | Source | Model | Cost | Use case |
|
|
49
49
|
|------|--------|-------|-----------|----------|
|
|
50
|
-
| Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |
|
|
51
|
-
| Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |
|
|
52
|
-
| Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |
|
|
50
|
+
| Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s block | Replace dialogue on an existing talking head |
|
|
51
|
+
| Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s block with uploaded audio; typed text adds one flat voice block | Animate a portrait into a talking avatar |
|
|
52
|
+
| Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s block with uploaded audio; typed text adds one flat voice block | Higher facial fidelity for hero shots |
|
|
53
53
|
|
|
54
54
|
Pick `sourceType` deliberately — it decides the pricing tier and the underlying endpoint.
|
|
55
55
|
|
|
@@ -60,7 +60,7 @@ Seedance can generate the performance rather than bolting a mouth onto finished
|
|
|
60
60
|
That is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a "video 1" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.
|
|
61
61
|
|
|
62
62
|
- Driving clips must be 2–15s; output duration is whatever you set (4–15s).
|
|
63
|
-
- Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys)
|
|
63
|
+
- Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys); pass the clip duration and quote before confirming. On both Seedance 2.0 and 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
|
|
64
64
|
- Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.
|
|
65
65
|
|
|
66
66
|
Everything below is about the Kling tool.
|
|
@@ -81,7 +81,7 @@ Use **avatar** when:
|
|
|
81
81
|
|
|
82
82
|
### Video flow (`sourceType: 'video'`)
|
|
83
83
|
- Format: mp4 or mov
|
|
84
|
-
- Duration: 2–10s (
|
|
84
|
+
- Duration: 2–10s (output follows the source clip)
|
|
85
85
|
- Resolution: 720p or 1080p (480p will be rejected)
|
|
86
86
|
- Max file size: 100MB
|
|
87
87
|
- Face must be visible and roughly facing camera. Profile shots fail.
|
|
@@ -107,7 +107,7 @@ Two ways to drive the lips:
|
|
|
107
107
|
- Pass `audioFilePath` — absolute path to an audio file on the user's machine
|
|
108
108
|
- Format: mp3, wav, m4a, ogg, aac
|
|
109
109
|
- Max 5MB
|
|
110
|
-
- Duration: 2–60s (output
|
|
110
|
+
- Duration: 2–60s (avatar output follows the voice track)
|
|
111
111
|
- Single clean voice. Music underneath, multiple speakers, or noisy mics produce garbage lips.
|
|
112
112
|
|
|
113
113
|
Prefer upload for production-quality voice. TTS for fast iteration / placeholder dialogue.
|
|
@@ -181,15 +181,15 @@ Don't default to pro. The ~15-credit delta per take adds up across iteration.
|
|
|
181
181
|
|
|
182
182
|
## Cost discipline
|
|
183
183
|
|
|
184
|
-
- Video re-dub at ~4 credits is the cheapest dialogue iteration in the entire Slates stack
|
|
185
|
-
- Avatar standard at ~14 credits is fine for medium use
|
|
186
|
-
- Avatar pro at ~29 credits trips the confirm gate
|
|
187
|
-
-
|
|
184
|
+
- Video re-dub at ~4 credits per 5s block is the cheapest dialogue iteration in the entire Slates stack; use it for voice A/B testing
|
|
185
|
+
- Avatar standard at ~14 credits per 5s block is fine for medium use; typed text on a still adds one flat voice block
|
|
186
|
+
- Avatar pro at ~29 credits per 5s block trips the confirm gate; explicit user OK required every time
|
|
187
|
+
- Output follows the media; billing rounds up to whole 5s blocks.
|
|
188
188
|
|
|
189
189
|
## Workflow patterns
|
|
190
190
|
|
|
191
191
|
**Voice A/B test (cheap):**
|
|
192
|
-
1. Generate one base talking-head video clip with
|
|
192
|
+
1. Generate one base talking-head video clip with Seedance (~40 credits)
|
|
193
193
|
2. Run `slates_generate_lip_sync` with `sourceType: 'video'` against 3–5 different `ttsVoice` values
|
|
194
194
|
3. Total cost: ~40 + (5 × ~4) ≈ 60 credits to compare voices
|
|
195
195
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-ltx-2-5
|
|
3
|
-
description:
|
|
3
|
+
description: "Prompt LTX-2.5 or LTX-2.5 Pro (ltx-2-5, ltx-2-5-pro) with slates_generate_video. Covers sound-first direction, connected shots, frame inputs, variant differences and duration constraints."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# LTX-2.5 — prompting
|
|
@@ -125,11 +125,10 @@ The model renders actions. It does not render adjectives.
|
|
|
125
125
|
|
|
126
126
|
---
|
|
127
127
|
|
|
128
|
-
## 4. Multishot
|
|
128
|
+
## 4. Multishot
|
|
129
129
|
|
|
130
130
|
**One LTX generation can carry several connected shots**, holding character, environment, lighting,
|
|
131
|
-
voice and style across every cut.
|
|
132
|
-
you generate separate clips and stitch them, and identity drifts between them.
|
|
131
|
+
voice and style across every cut. It is one of the seats that carry several shots in one generation.
|
|
133
132
|
|
|
134
133
|
**Working range is two to four shots.** Three is the comfortable stopping point.
|
|
135
134
|
|
|
@@ -182,7 +181,7 @@ Choose the length the beat needs.
|
|
|
182
181
|
|
|
183
182
|
### Aspect ratios: 16:9 and 9:16, and nothing else
|
|
184
183
|
|
|
185
|
-
The narrowest set in the catalogue
|
|
184
|
+
The narrowest set in the catalogue. Square, 4:5 and 21:9 are not available on this
|
|
186
185
|
model at any resolution.
|
|
187
186
|
|
|
188
187
|
### Frames, not references
|
|
@@ -210,7 +209,7 @@ Native synchronised audio is **included at every resolution on both seats**, wit
|
|
|
210
209
|
no toggle that costs money — unlike Kling, where sound is a paid dimension. A 6-second 1080p LTX
|
|
211
210
|
clip **with sound** is 39 credits.
|
|
212
211
|
|
|
213
|
-
Combined with 1080p at $0.13/s
|
|
212
|
+
Combined with 1080p at $0.13/s, the cheapest native 1080p second with sound included in Slates, this makes LTX **the
|
|
214
213
|
coverage seat**: the one to reach for when the job is many takes rather than one hero shot, when a
|
|
215
214
|
sequence needs its own sound, or when the credit budget is the binding constraint.
|
|
216
215
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-minimax-h3
|
|
3
|
-
description:
|
|
3
|
+
description: "Prompt MiniMax H3, H3 Max or H3 Max Turbo with slates_generate_video. Covers separately authored dialogue, scene sound and score, declared reference relationships, frame inputs and variant-specific constraints."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# MiniMax H3 — prompting
|
|
@@ -15,18 +15,18 @@ description: How to prompt MiniMax H3, H3 Max and H3 Max Turbo. Read before call
|
|
|
15
15
|
Keep it under 2,400 characters (the build fails above that) and keep the
|
|
16
16
|
rationale, the receipts and the worked examples in the body below. -->
|
|
17
17
|
<!-- /slates-only -->
|
|
18
|
-
**Card
|
|
18
|
+
**Card: MiniMax H3.** The only seat with three separately authored audio layers: dialogue, scene sound and score are three separate sections of the prompt, generated in one pass, and putting a sound in the wrong section drops or doubles it.
|
|
19
19
|
|
|
20
20
|
**The five levers**
|
|
21
|
-
1. **Write the three audio layers separately
|
|
21
|
+
1. **Write the three audio layers separately**: `overall_soundscape:` for what is in the room, `non_diegetic_music:` for what only the audience hears, and the dialogue quoted inline. Section decides attribution.
|
|
22
22
|
2. **Quote dialogue and name the language** — `says in English`, `speaks in Spanish`. Eleven languages are stably supported; the language is part of the instruction, not an afterthought.
|
|
23
23
|
3. **Declare the reference RELATIONSHIP**, which no other seat has: `kept whole`, `partly kept`, `transferred`, or `a loose echo`. An undeclared reference is a guess.
|
|
24
24
|
4. **Give a beat of stillness before a line** — `sits still for a beat, then looks up`. The sync needs something to lock against; a character already mid-motion when the line starts drifts.
|
|
25
25
|
5. **Describe the beat structure** — `waits`, `then speaks`, `under the last three seconds`. H3 is a timeline, so write one.
|
|
26
26
|
|
|
27
27
|
**Examples**
|
|
28
|
-
- `A woman sits still at a kitchen table for a beat, then looks up. She says in English, "You said Tuesday."
|
|
29
|
-
- `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, "No es el alternador."
|
|
28
|
+
- `A woman sits still at a kitchen table for a beat, then looks up. She says in English, "You said Tuesday." overall_soundscape: a fridge hum, a spoon set down on formica. non_diegetic_music: N/A.`
|
|
29
|
+
- `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, "No es el alternador." overall_soundscape: a socket wrench, a radio two bays over. non_diegetic_music: a low sustained cello under the last three seconds, audience only.`
|
|
30
30
|
|
|
31
31
|
**Hard constraint:** the three seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p, takes the same 9+3+3 references, and costs MORE at the tier they share — a speed pick, never the cheap one; `minimax-h3-max-turbo` has Max's ladder at half its rate and takes frames only, NO references. Every tier above 768p is built from the native 768p render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.
|
|
32
32
|
<!-- @card:end -->
|
|
@@ -61,7 +61,7 @@ what the endpoint accepts:
|
|
|
61
61
|
| Price at 768p | **$0.060/s** | $0.080/s | $0.040/s |
|
|
62
62
|
| Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) | **price** — half Max's rate at every tier |
|
|
63
63
|
|
|
64
|
-
**Max is the premium seat, not the budget one.** It is 33% dearer at
|
|
64
|
+
**Max is the premium seat, not the budget one.** It is 33% dearer at 768p, equal at 480p, and it
|
|
65
65
|
tops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth
|
|
66
66
|
paying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.
|
|
67
67
|
|
|
@@ -91,15 +91,15 @@ reaching for 2K, which adds its own artifacting on top.
|
|
|
91
91
|
|
|
92
92
|
## The one thing that makes H3 different: audio is a THREE-LAYER instruction
|
|
93
93
|
|
|
94
|
-
|
|
94
|
+
Kling and Seedance have their own sound syntax. H3 splits audio into three separately authored layers, enforced by where
|
|
95
95
|
you write each thing. Get the section wrong and the sound is dropped, doubled, or attributed to the
|
|
96
96
|
wrong source.
|
|
97
97
|
|
|
98
98
|
| Layer | What belongs in it | Where it goes |
|
|
99
99
|
|---|---|---|
|
|
100
100
|
| **Synchronised events** | dialogue, singing, and any sound tied to a specific shot or action | the **body** of the prompt, on the beat it lands |
|
|
101
|
-
| **Scene sound** | ambience and physical sounds that run across the whole clip
|
|
102
|
-
| **Score** | music the characters cannot hear; audience-only | the **
|
|
101
|
+
| **Scene sound** | ambience and physical sounds that run across the whole clip, room tone, rain, traffic, a ventilation hum | the **overall_soundscape** section |
|
|
102
|
+
| **Score** | music the characters cannot hear; audience-only | the **non_diegetic_music** section |
|
|
103
103
|
|
|
104
104
|
**Three rules, all from MiniMax's own guide:**
|
|
105
105
|
|
|
@@ -125,10 +125,10 @@ the middle-aged baker with a calm, slightly raspy voice places a fresh loaf on t
|
|
|
125
125
|
says: "First batch of the morning." [Shot 2] At 00:05.000, the camera cuts to a close-up of
|
|
126
126
|
steam rising from the sliced bread while his final words carry over from the previous shot.
|
|
127
127
|
|
|
128
|
-
|
|
128
|
+
overall_soundscape: wooden shutters scrape open over a quiet street, trays clink softly inside, a
|
|
129
129
|
doorbell rings once, then light footsteps and the crisp sound of bread being sliced.
|
|
130
130
|
|
|
131
|
-
|
|
131
|
+
non_diegetic_music: a soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes,
|
|
132
132
|
gentle fade at the end.
|
|
133
133
|
```
|
|
134
134
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-motion-transfer
|
|
3
|
-
description:
|
|
3
|
+
description: "Prepare Kling Motion Control with slates_generate_motion_transfer: a character image plus a driving clip. Covers orientation, source selection, tiers and the separate Seedance video-reference alternative."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Motion transfer — setup guide
|
|
@@ -15,7 +15,7 @@ description: How to set up motion transfer — Kling Motion Control only (std an
|
|
|
15
15
|
Keep it under 2,400 characters (the build fails above that) and keep the
|
|
16
16
|
rationale, the receipts and the worked examples in the body below. -->
|
|
17
17
|
<!-- /slates-only -->
|
|
18
|
-
**Card
|
|
18
|
+
**Card: Motion transfer (Kling Motion Control only).** A target IMAGE (your character) plus a source VIDEO (the motion) produces your character performing that motion. Output follows the driving clip, up to 30s with video orientation or 10s with image orientation, billed per 5s block.
|
|
19
19
|
|
|
20
20
|
**The five levers**
|
|
21
21
|
1. **The target image must show body proportions clearly** and the character must occupy more than about 5% of the frame. A tiny figure in a wide shot has nothing to drive.
|
|
@@ -23,7 +23,7 @@ description: How to set up motion transfer — Kling Motion Control only (std an
|
|
|
23
23
|
3. **Choose `characterOrientation` on purpose** — `video` takes the source clip's framing, `image` preserves the portrait's. It is the most-missed choice here.
|
|
24
24
|
|
|
25
25
|
4. **The prompt is atmosphere only** — `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.` Motion verbs are ignored; the motion is already in the driving video.
|
|
26
|
-
5. **Pick the
|
|
26
|
+
5. **Pick the source section up front**, and write only atmosphere: `soft afternoon sunlight`, `vintage warm color grade`, `clean studio backdrop`. The output follows that section, up to the orientation's limit; longer clips cost more blocks.
|
|
27
27
|
|
|
28
28
|
**Examples**
|
|
29
29
|
- `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.`
|
|
@@ -63,8 +63,8 @@ movement from video 1. Preserve the character's identity, appearance, and outfit
|
|
|
63
63
|
|
|
64
64
|
That is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.
|
|
65
65
|
|
|
66
|
-
- **
|
|
67
|
-
- **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key
|
|
66
|
+
- **Reference videos must total 2–15s on 2.0, or 2–30s on 2.5.** Longer clips: trim first. Kling MC (`characterOrientation: 'video'`) takes up to 30s.
|
|
67
|
+
- **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key; quote via the confirm gate before spending. On both Seedance 2.0 and 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
|
|
68
68
|
- **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
|
|
69
69
|
- `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.
|
|
70
70
|
|
|
@@ -180,13 +180,13 @@ Leave it empty if you don't have a specific atmospheric note.
|
|
|
180
180
|
- Pro tier on first iteration — waste, switch to it once the motion + framing combo is locked
|
|
181
181
|
- Cartoon driving videos — guaranteed failure
|
|
182
182
|
- Cropped or partial target characters — identity will drift
|
|
183
|
-
-
|
|
183
|
+
- Driving videos longer than the needed motion; pick the source section upfront, within the orientation's limit
|
|
184
184
|
|
|
185
185
|
## Cost discipline
|
|
186
186
|
|
|
187
|
-
-
|
|
187
|
+
- Output follows the driving clip, up to 30s with video orientation or 10s with image orientation; billed per 5s block
|
|
188
188
|
- Both tiers trip the confirm gate — every call needs explicit user OK
|
|
189
|
-
- Iteration is expensive: 4 takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.
|
|
189
|
+
- Iteration is expensive: 4 five-second takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.
|
|
190
190
|
- Always run a single std take first to validate the motion + framing combo before committing to pro
|
|
191
191
|
|
|
192
192
|
## Confirm gate: cost + codes, no inline preview
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-nano-banana-2
|
|
3
|
-
description:
|
|
3
|
+
description: "Prompt Nano Banana 2, Lite or Pro image generation and edits. Use on these models; covers photographic craft, subject and reference binding, typography, prompt structure and family differences."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Nano Banana 2 — cinematic & photorealistic prompting
|
|
@@ -220,16 +220,20 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
|
|
|
220
220
|
|
|
221
221
|
✅ **Cinema:** "Extreme close on subject's mouth and nose, 135mm f/2.8, shallow depth of field. Breath pluming out, catching cold light from upper-left key. Lips slightly parted, peach fuzz visible. The breath holds. CineStill 800T halation around catchlights. Waiting."
|
|
222
222
|
|
|
223
|
-
|
|
223
|
+
<!-- @inject:iteration-diagnosis -->
|
|
224
|
+
## Diagnose repeated failures
|
|
224
225
|
|
|
225
|
-
|
|
226
|
+
After three failed attempts at the same requirement, pause unchanged re-rolls and diagnose the source reference, prompt structure, model fit and tool result. Three is a review checkpoint, not a universal limit or proof that the seed cannot matter. Preserve the attempts and name what each test changed.
|
|
227
|
+
|
|
228
|
+
Continue autonomously when the brief is clear, a specific correction is supported and the next request is already authorized. Hand control back when taste or intent cannot be inferred, the next request needs fresh consent, or the available tool cannot meet the requirement. A failed roll never authorizes an additional charge. Follow the existing batch and per-request cost policy.
|
|
229
|
+
<!-- @end:iteration-diagnosis -->
|
|
226
230
|
|
|
227
231
|
## Family variants — Lite and Pro
|
|
228
232
|
|
|
229
233
|
Everything in this skill applies to the whole Nano Banana family; two variants trade speed/ceiling around NB2 full:
|
|
230
234
|
|
|
231
235
|
- **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.
|
|
232
|
-
- **nano-banana-pro
|
|
236
|
+
- **nano-banana-pro**: the hero-frame/typography ceiling (2× NB2 at 1K, 1.33× at 2K, about 1.9× at 4K; 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs; it takes a full subject library in one call.
|
|
233
237
|
|
|
234
238
|
<!-- slates-only -->
|
|
235
239
|
Routing between them (and vs GPT Image 2.5 / FLUX / Seedream): `slates-model-selection`.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-omni-flash
|
|
3
|
-
description:
|
|
3
|
+
description: "Prompt Gemini Omni Flash video generation or Omni Flash edits (omni-flash, omni-flash-edit). Covers native sound, reference inputs and short change-only prompts that preserve edit fidelity."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Gemini Omni Flash — prompting
|
|
@@ -22,7 +22,7 @@ description: How to prompt Gemini Omni Flash (Google, via fal). Read before call
|
|
|
22
22
|
2. **Editing: always end with `Keep everything else the same.`** — the one documented preservation lever.
|
|
23
23
|
|
|
24
24
|
3. **Editing: describe the EFFECT, never a real object as a metaphor.** "Candle-like flame" rendered a literal candle in the subject's hand.
|
|
25
|
-
4. **Editing: no
|
|
25
|
+
4. **Editing: no chained stage directions.** Several beats cued to moments ("…flies onto his shoulder when he calls it, and perches as he walks…") hard-failed with `invalid_request`. One effect tied to an action already in the footage ("…when he snaps his fingers…") passed. Collapse to one continuous action; the model syncs to the footage's own motion.
|
|
26
26
|
5. **Generation: the opposite — describe fully.** Subject, action, setting, `camera tracking alongside`, `overcast flat light`, tone. Audio is prompt-driven with no parameters: dialogue in quotes, sound in plain language — `rain patters on the tin roof`, `spray from tyres`, `a horn somewhere behind`.
|
|
27
27
|
|
|
28
28
|
**Examples**
|
|
@@ -42,7 +42,7 @@ description: How to prompt Gemini Omni Flash (Google, via fal). Read before call
|
|
|
42
42
|
**Never use in an EDIT prompt** (each one has a receipt above):
|
|
43
43
|
- a long preservation preamble — it produces WORSE drift than `Keep everything else the same.`
|
|
44
44
|
- a real object as a metaphor: `candle-like`, `flame-like`, `laser-like`
|
|
45
|
-
-
|
|
45
|
+
- several staged beats cued to moments (`when he calls it, and perches as he walks`): these hard-fail, they do not merely drift. One effect tied to an action already in the clip (`when he snaps his fingers`) passed
|
|
46
46
|
- harm-to-person framing: `ignite`, `catch fire`, `on fire` applied to a person trips the safety filter
|
|
47
47
|
<!-- @banned:end -->
|
|
48
48
|
|
|
@@ -51,15 +51,15 @@ Google's fast video generation + editing model ("Nano Banana Pro for video" in c
|
|
|
51
51
|
## Where it routes
|
|
52
52
|
|
|
53
53
|
- **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.
|
|
54
|
-
- **
|
|
55
|
-
- **
|
|
54
|
+
- **Drafts and iteration volume with sound**: an audio-native video seat (~6.4 cr/s at 720p).
|
|
55
|
+
- **Hero-generation quality is still unproven.** Use the current model-routing guide for the final generation seat; the edit-fidelity receipt does not establish generation quality.
|
|
56
56
|
|
|
57
|
-
## Editing (
|
|
57
|
+
## Editing (<!-- slates-only -->`slates_edit_video`, <!-- /slates-only -->model `omni-flash-edit`) — THE RULES (receipts, not theory)
|
|
58
58
|
|
|
59
59
|
1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes."* Live receipt 2026-07-09: a long "keep every frame/word/movement identical…" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same."*
|
|
60
60
|
2. **Always end with "Keep everything else the same."** — the one documented preservation lever.
|
|
61
61
|
3. **Never name a real-world object as a metaphor.** "Candle-like flame" rendered a literal candle in his hand. Describe the effect itself ("small magical flames on his fingertips").
|
|
62
|
-
3b. **No
|
|
62
|
+
3b. **No chained stage directions: they HARD-FAIL, not drift.** Receipt 2026-07-09: "a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…" → deterministic `invalid_request` (2×, "could not generate with the given inputs"); collapsing to one continuous action — "A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke." — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video. The winning flames prompt above ("when he snaps his fingers") shows one effect tied to an action the footage already contains is fine; which part of the dragon prompt triggered the refusal is untested beyond that.
|
|
63
63
|
4. **Safety filter (Google's, strict about harm-to-person):** "fingertips ignite / catch fire" → `content_policy_violation`. Frame effects as magical/harmless VFX: "small magical flames appear on his fingertips" passed. See slates-content-policy §Gemini for the substitution patterns.
|
|
64
64
|
5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.
|
|
65
65
|
6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.
|
|
@@ -67,9 +67,9 @@ Google's fast video generation + editing model ("Nano Banana Pro for video" in c
|
|
|
67
67
|
8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).
|
|
68
68
|
9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.
|
|
69
69
|
|
|
70
|
-
## Generation (
|
|
70
|
+
## Generation (<!-- slates-only -->`slates_generate_video`, <!-- /slates-only -->model `omni-flash`)
|
|
71
71
|
|
|
72
|
-
- **Inputs:** prompt only (t2v), prompt + ONE start frame (
|
|
72
|
+
- **Inputs:** prompt only (t2v), prompt + ONE start frame (<!-- slates-only -->`firstFrameAssetId`, <!-- /slates-only -->i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
|
|
73
73
|
- Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.
|
|
74
74
|
- **Name references inline** the standard Slates way ("Marcus (image 1) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
|
|
75
75
|
- **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language ("rain patters on the tin roof"). Negative direction as plain instructions ("Do not show text").
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-seed-audio
|
|
3
|
-
description:
|
|
3
|
+
description: "Prompt Seed Audio 1.0 (seed-audio) with slates_generate_audio. Covers scene sentences, dialogue, crowd size, spatial sound, duration in prompt text and mutually exclusive audio or image references."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Seed Audio 1.0 — prompting
|
|
@@ -43,7 +43,29 @@ description: How to prompt Seed Audio 1.0 (ByteDance, via fal). Read before call
|
|
|
43
43
|
- shot language: `wide shot`, `slow push in`, `warm tungsten` — camera and lighting words are video-prompt words the model has to ignore
|
|
44
44
|
<!-- @banned:end -->
|
|
45
45
|
|
|
46
|
-
ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence.
|
|
46
|
+
ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence. Use the generated duration bounds below and the current catalogue for routing; it is useful for continuity beds.
|
|
47
|
+
|
|
48
|
+
<!-- @inject:thresholds -->
|
|
49
|
+
<!-- GENERATED from @slatesvideo/shared — do not edit between the markers.
|
|
50
|
+
Source: CONFIRM_CREDITS, DEVIATION_FACTOR and the audio bounds in
|
|
51
|
+
packages/shared/src/operations/index.ts. Every number here is REFUSED by an
|
|
52
|
+
op when a prompt gets it wrong, which is why none of them is typed by hand
|
|
53
|
+
any more: this block replaced four claims that contradicted the code. -->
|
|
54
|
+
|
|
55
|
+
**The thresholds, from the code that enforces them:**
|
|
56
|
+
|
|
57
|
+
- **Confirm gate:** above **17 credits** an op returns `requires_confirm` and will not
|
|
58
|
+
proceed until you re-call with `confirm: true`. This is a code gate, not permission to spend: every generation still needs the user-approved plan or quote.
|
|
59
|
+
- **Deviation pause:** the desktop Studio Agent stops and re-asks when projected generation spend
|
|
60
|
+
exceeds the approved plan by more than **20%**. You do not trigger this; the app does.
|
|
61
|
+
- **Seed Audio duration:** **3–120 seconds.** There is no duration
|
|
62
|
+
parameter on the model — the number you pass is written into the prompt AND is what the user is
|
|
63
|
+
billed. Outside that range the op refuses rather than clamping.
|
|
64
|
+
- **Sound Effects duration:** **1–22 seconds**, billed per second, never left for the
|
|
65
|
+
model to pick.
|
|
66
|
+
|
|
67
|
+
Never quote a credit figure from memory: `slates_estimate_generation_cost` returns the real one.
|
|
68
|
+
<!-- @end:thresholds -->
|
|
47
69
|
|
|
48
70
|
## Where it routes
|
|
49
71
|
|
|
@@ -134,8 +156,6 @@ Set `multilingual: true` for non-English or mixed-language lines.
|
|
|
134
156
|
| `volume` | 0.5–2.0 | Rarely — normalize on the timeline instead. |
|
|
135
157
|
| `pitch` | −12…+12 semitones | Ageing or shifting a voice. Small moves only; ±3 is already a lot. |
|
|
136
158
|
| `multilingual` | bool | Non-English or code-switched lines. |
|
|
137
|
-
| `sampleRate` | 8k–48k | Leave at 24000 unless you are matching an existing stem. |
|
|
138
|
-
| `outputFormat` | mp3 / wav / pcm / ogg_opus | wav when this is going into a mix; mp3 otherwise. |
|
|
139
159
|
|
|
140
160
|
## Iterating
|
|
141
161
|
|