@slatesvideo/shared 0.6.1 → 0.6.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/clients/blender.d.ts +50 -0
- package/dist/clients/blender.js +195 -0
- package/dist/index.d.ts +3 -1
- package/dist/index.js +26 -0
- package/dist/operations/index.d.ts +25 -1
- package/dist/operations/index.js +454 -24
- package/dist/prompts/agent-doctrine.d.ts +36 -0
- package/dist/prompts/agent-doctrine.js +194 -0
- package/dist/prompts/banned-tokens.d.ts +28 -0
- package/dist/prompts/banned-tokens.js +152 -0
- package/dist/prompts/model-capabilities.d.ts +13 -1
- package/dist/prompts/model-capabilities.js +97 -2
- package/dist/prompts/model-facts.d.ts +20 -0
- package/dist/prompts/model-facts.js +87 -23
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +65 -0
- package/dist/prompts/reference-composer.d.ts +57 -0
- package/dist/prompts/reference-composer.js +70 -1
- package/dist/skills/content.js +8 -2
- package/exports/slates-prompt-builder/generated/SKILL.md +3 -3
- package/exports/slates-prompt-builder/generated/reference-seedance.md +2 -1
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +8 -8
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +1 -1
- package/skills/slates-blocking-to-prompt.md +250 -0
- package/skills/slates-camera-language.md +196 -0
- package/skills/slates-dialogue-blocking.md +134 -0
- package/skills/slates-previs-blocking.md +153 -0
- package/skills/slates-prompting-ltx-2-5.md +180 -0
- package/skills/slates-prompting-nano-banana-2.md +10 -0
- package/skills/slates-prompting-seedance.md +10 -1
- package/skills/slates-restyle-from-blocking.md +121 -0
|
@@ -0,0 +1,134 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-dialogue-blocking
|
|
3
|
+
description: Keep multiple characters spatially consistent across cuts — seating, screen direction, eyelines, the 180-degree rule — by blocking the scene in 3D first. Use for any multi-character dialogue scene, conversations around a table or in a car, or when generated characters swap seats, change sides, or look the wrong way between shots.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Dialogue blocking — six people who stay where you put them
|
|
7
|
+
|
|
8
|
+
The hardest thing to generate, and the case where previs beats raw prompting by the widest margin.
|
|
9
|
+
|
|
10
|
+
## Why this is hard
|
|
11
|
+
|
|
12
|
+
Every cut is an independent guess unless something forces agreement. Prompt a six-person conversation four times and you get four different seating charts: characters swap places, the 180-degree line breaks, and nothing cuts together. The failure is not aesthetic — the shots are simply unusable as an edit, and you find out only after paying for all four.
|
|
13
|
+
|
|
14
|
+
**The blocking fixes it structurally.** Positions exist in 3D, so every camera sees the same arrangement, and consistency stops being something the model has to remember.
|
|
15
|
+
|
|
16
|
+
## Build
|
|
17
|
+
|
|
18
|
+
Follow `slates-previs-blocking` and add these.
|
|
19
|
+
|
|
20
|
+
### Seated proxies, colour-coded
|
|
21
|
+
|
|
22
|
+
Simple seated shapes. **Do not animate the heads** — a proxy head turning the wrong way is worse than one that never turns.
|
|
23
|
+
|
|
24
|
+
Give each character a distinct viewport colour and write the mapping down. This is the identity channel:
|
|
25
|
+
|
|
26
|
+
> red = the boss · green = the kid · blue = the driver · yellow = the fixer · purple = the cousin · cyan = the nephew
|
|
27
|
+
|
|
28
|
+
That mapping goes verbatim into the generation prompt. It is what lets the model bind a grey body to a character sheet across four cuts.
|
|
29
|
+
|
|
30
|
+
Give every proxy a material and set `mat.diffuse_color` to its identity colour — the blocking render pins Workbench to `MATERIAL` shading, so **the material's `diffuse_color` is what reaches the clip**. Set `object.color` to the same value too, so the user's viewport matches what renders. See `slates-previs-blocking` for the snippet.
|
|
31
|
+
|
|
32
|
+
### Fix the geography, then never move it
|
|
33
|
+
|
|
34
|
+
Place people once. Write down who sits where relative to the camera's opening position, in words, because that sentence is going into the prompt:
|
|
35
|
+
|
|
36
|
+
> Across the table, facing camera: yellow dead centre, purple far left, blue and green to the right.
|
|
37
|
+
|
|
38
|
+
### The camera plan
|
|
39
|
+
|
|
40
|
+
Per `slates-camera-language`, with two things specific to dialogue:
|
|
41
|
+
|
|
42
|
+
- **Below shoulder height, slow rail glides.** Eye-level-and-above reads as surveillance.
|
|
43
|
+
- **Decide who owns the near foreground in each cut and honour it.** A shoulder in frame is a spatial anchor; a different shoulder in the next cut relocates the whole room.
|
|
44
|
+
|
|
45
|
+
The move that earns the most: **a gaze handoff without a cut** — the camera keeps gliding while the target hands off across the table, face to face, slowing on each but never stopping. Build it by keyframing the Track To target's position between subjects.
|
|
46
|
+
|
|
47
|
+
### Crossing behind someone
|
|
48
|
+
|
|
49
|
+
A head wiping frame during a move is a strong depth cue. It is also a spatial claim, so pick who gets crossed and say so — *the camera crosses directly behind cyan's back mid-shot and his head wipes the frame once.*
|
|
50
|
+
|
|
51
|
+
## The prompt
|
|
52
|
+
|
|
53
|
+
Everything in `slates-blocking-to-prompt`, plus these blocks.
|
|
54
|
+
|
|
55
|
+
### Geography — restate it as a rule
|
|
56
|
+
|
|
57
|
+
> TABLE GEOGRAPHY — do not deviate: the camera is never parked behind red. Only in the opening seconds does his dark shoulder hang at the near frame RIGHT edge, and it slides out as the camera travels LEFT. The true near-foreground of this shot is cyan: the camera crosses directly behind him mid-shot. After the opening seconds red is gone from the foreground, and the camera never travels behind anyone except cyan.
|
|
58
|
+
|
|
59
|
+
### Screen direction, and the mirror that is not a swap
|
|
60
|
+
|
|
61
|
+
The 180-degree rule survives on its own in the blocking. What breaks is the model **"correcting" a legitimate mirror** — when the camera faces back through a scene, sides invert, and that inversion is correct. Say so explicitly or it gets flipped:
|
|
62
|
+
|
|
63
|
+
> The driver's seat is on the LEFT for the entire timeline; this layout never mirrors or flips. When a camera faces BACKWARD into the car, screen sides mirror naturally: the driver reads on the RIGHT of frame, the passenger on the LEFT — that is correct left-hand drive, not a swap. They never swap seats or roles anywhere in the timeline.
|
|
64
|
+
|
|
65
|
+
Then compress it into the HOLD block: *(backward camera mirrors them: he right of frame, she left)*.
|
|
66
|
+
|
|
67
|
+
### Presence
|
|
68
|
+
|
|
69
|
+
> A seated person stays drawn even when partially occluded — in every interior frame some part of each seat's owner is visible: a hand, an arm, a shoulder, a head above the bolster. Every occupied seat visibly holds its person.
|
|
70
|
+
|
|
71
|
+
### Keep everyone alive
|
|
72
|
+
|
|
73
|
+
Three orthogonal layers. Without them, whoever is not speaking freezes:
|
|
74
|
+
|
|
75
|
+
- **ONGOING BUSINESS** — a small continuous physical action per character, running whether or not they are speaking. *turns his glass a quarter every few seconds · thumbs a lighter without lighting it.*
|
|
76
|
+
- **BACKGROUND LIFE** — soft-focus, low contrast, never pulls attention, never crosses in front of a speaking face.
|
|
77
|
+
- **SCENE EVENT** — the unnamed thing everyone is playing but nobody says. One line, repeated verbatim in every character's direction: *keep tomorrow sounding like a fishing trip.*
|
|
78
|
+
|
|
79
|
+
### Acting, per character
|
|
80
|
+
|
|
81
|
+
Same six slots each. Terse:
|
|
82
|
+
|
|
83
|
+
```
|
|
84
|
+
ACTING TASK — <character>
|
|
85
|
+
SCENE DIRECTION (shared, unspoken): <the same line for everyone>
|
|
86
|
+
MOTIVE (his fuel): <what he wants underneath>
|
|
87
|
+
GOAL: <what he wants in this scene>
|
|
88
|
+
OBSTACLE: <what is in the way>
|
|
89
|
+
TACTIC: <how he goes about it>
|
|
90
|
+
Moment to moment: <2-3 beats keyed to timestamps>
|
|
91
|
+
(Safety: gaze always engaged in the task — never a frozen, glassy,
|
|
92
|
+
unfocused stare; natural blink cadence.)
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
That safety line is not filler. Dead eyes are the characteristic failure of generated faces in dialogue, and naming it is what prevents it.
|
|
96
|
+
|
|
97
|
+
### Split any strong emotion into phases
|
|
98
|
+
|
|
99
|
+
The other characteristic failure is a face that strikes one extreme expression and holds it for the whole shot — a mouth stuck open for three seconds. Give the beat two phases with a hinge, and name the failure you are excluding:
|
|
100
|
+
|
|
101
|
+
> PHASE 1 (17.3-18.6s) — she SCREAMS at him, mouth wide, eyes huge, hand clamped on the grab handle. PHASE 2 (18.6-19.9s) — the scream breaks off: she shuts her eyes tight and CLOSES her mouth, both hands now on the handle, head ducked, braced. Scream, then brace — never one frozen open mouth held through the whole shot.
|
|
102
|
+
|
|
103
|
+
The hinge timestamp is what makes it a performance instead of a pose.
|
|
104
|
+
|
|
105
|
+
### Dialogue must not restructure the edit
|
|
106
|
+
|
|
107
|
+
Both of these, verbatim, every time:
|
|
108
|
+
|
|
109
|
+
> DIALOGUE NEVER CREATES SHOTS: spoken lines happen inside the reference's takes exactly as blocked — no cutaways to a speaker, no reverse shots, no added close-ups. If a line plays while the camera is elsewhere, the line stays off-screen audio.
|
|
110
|
+
|
|
111
|
+
> OFF-SCREEN VOICES RULE: a line marked off-screen must STAY off-screen — never show the speaker, never move him into frame, never route the camera behind him because he spoke.
|
|
112
|
+
|
|
113
|
+
A sentence may cross a cut. Say so where it does: *the sentence does not pause for the edit.*
|
|
114
|
+
|
|
115
|
+
## Model routing
|
|
116
|
+
|
|
117
|
+
Dialogue directed as separate layers (voices, scene sound, score) is **minimax-h3**'s seat; it also takes declared reference relationships, which suits a colour-coded cast. Native synced audio is **Veo**'s niche. seedance-2.5 carries the reference-video capacity. Route per `slates-model-selection` and read the chosen model's prompting skill before writing the audio block.
|
|
118
|
+
|
|
119
|
+
## Checklist
|
|
120
|
+
|
|
121
|
+
- [ ] Colour→character mapping written down and pasted into the prompt
|
|
122
|
+
- [ ] Heads not animated in the blocking
|
|
123
|
+
- [ ] Seating stated as a geography rule
|
|
124
|
+
- [ ] Foreground owner named per cut
|
|
125
|
+
- [ ] Mirror-is-not-a-swap clause present if any camera faces back through the scene
|
|
126
|
+
- [ ] Presence rule present
|
|
127
|
+
- [ ] Ongoing business, background life and scene event all specified
|
|
128
|
+
- [ ] Acting task per character, safety line included
|
|
129
|
+
- [ ] Any strong emotion split into phases with a hinge timestamp
|
|
130
|
+
- [ ] Dialogue-never-creates-shots and off-screen-voices rules present
|
|
131
|
+
|
|
132
|
+
## Related
|
|
133
|
+
|
|
134
|
+
`slates-previs-blocking` · `slates-camera-language` · `slates-blocking-to-prompt` · `slates-character-identity` (the sheets) · `slates-prompting-minimax-h3`
|
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-previs-blocking
|
|
3
|
+
description: Build a 3D blocking pass in Blender, render it grey-box, and use it as a reference video so the generated shot follows a camera path you designed instead of one the model invented. Use when the user wants precise camera control, a multi-cut sequence, a one-take move, spatial consistency across shots, or says the camera keeps drifting / they keep burning credits re-rolling.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Previs blocking — design the shot, then generate it
|
|
7
|
+
|
|
8
|
+
The spine of the whole workflow. Read this first; the other four previs skills are branches off it.
|
|
9
|
+
|
|
10
|
+
## The mechanism (why this works at all)
|
|
11
|
+
|
|
12
|
+
A text prompt asks the model to *invent* camera motion, so it invents differently every roll. You cannot iterate on a variable you do not control, so you re-roll and pay again.
|
|
13
|
+
|
|
14
|
+
A **reference video** removes the invention. You build the shot in Blender as untextured proxies — a neutral grey set with colour-coded figures, free, instant, deterministic — render the camera's path to mp4, and hand the model that clip alongside the prompt. **Blender locks the motion; the model builds the world.** Iteration moves to the free half, and the paid half usually lands first try.
|
|
15
|
+
|
|
16
|
+
Two halves, and keeping them separate is the whole discipline:
|
|
17
|
+
|
|
18
|
+
| Half | Lives in | Changes when |
|
|
19
|
+
|---|---|---|
|
|
20
|
+
| **Structure** — cuts, camera, timing, who is where | the blocking clip | you re-block |
|
|
21
|
+
| **Style** — what any of it looks like | references + prompt text | you restyle (see `slates-restyle-from-blocking`) |
|
|
22
|
+
|
|
23
|
+
## Before you start
|
|
24
|
+
|
|
25
|
+
1. `slates_blender_status` — confirms the bridge is up and returns fps, frame range, existing camera. If it reports `connected: false`, relay its hint and stop; nothing else here works.
|
|
26
|
+
2. Settle **format first**, because the blocking render *is* the film's format: fps, aspect, duration. 24fps is the default and makes cut times land on clean frames. Duration ≤ 30s (seedance-2.5's reference-video ceiling; 15s on the others).
|
|
27
|
+
3. Know the shot count. "One take" and "19 cuts" are different builds.
|
|
28
|
+
|
|
29
|
+
## Build order
|
|
30
|
+
|
|
31
|
+
Do these in order. Each stage is verifiable on its own, and a camera built before the geometry has nothing to frame.
|
|
32
|
+
|
|
33
|
+
### 1. Set the format
|
|
34
|
+
|
|
35
|
+
```python
|
|
36
|
+
scene = bpy.context.scene
|
|
37
|
+
scene.render.fps = 24
|
|
38
|
+
scene.render.fps_base = 1.0
|
|
39
|
+
scene.render.resolution_x, scene.render.resolution_y = 1920, 1080
|
|
40
|
+
scene.frame_start, scene.frame_end = 1, 720 # 30s at 24fps
|
|
41
|
+
result = {"seconds": 720 / 24}
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Frame maths, stated once so you never redo it in your head: **frame = seconds × fps + 1**. A cut at 7.79s is frame 188.
|
|
45
|
+
|
|
46
|
+
### 2. Geometry and light — grey set, coded figures, named
|
|
47
|
+
|
|
48
|
+
Proxies only. A person is a box or a capsule with a sphere head. A car is a stretched cube. A can is a cylinder. **The SET is neutral grey — one light, a floor and enough wall that the space reads.** Colour is reserved for the figures, where it carries meaning (below); a grey set is what makes those few colours legible as notation rather than décor. Anything you spend on materials here you pay for twice, because the model repaints every surface anyway.
|
|
49
|
+
|
|
50
|
+
**Name every object for what it *is* in the story**, not `Cube.003`. The name is how you refer to it later, and it is how you keep your own timeline honest.
|
|
51
|
+
|
|
52
|
+
Two conventions that cost nothing now and save a re-roll later:
|
|
53
|
+
|
|
54
|
+
- **Colour is identity.** Give each character a distinct viewport colour and *write the mapping down* — `red = the boss, green = the kid, blue = the driver`. The generation prompt will restate that mapping so the model knows which grey body is which person across cuts. Without it, characters swap.
|
|
55
|
+
- **Encode facing on featureless proxies.** A box has no front. Mark one face red, the back black, the sides green, and say so in the prompt: `RED face = the direction he faces`. Otherwise the model guesses which way people are looking.
|
|
56
|
+
- **Checker a surface when SCALE or SPEED has to read.** Flat grey gives a model no parallax cue, so a fast move over a featureless floor reads as slow, and a big room reads as a small one. A black-and-white checker on the ground (or the wall a camera races past) gives it something to measure against. ⚠️ **Build it as GEOMETRY, never as a Checker Texture node.** The blocking render is Workbench, which draws one flat colour per material and never evaluates a shader node tree — a `TEX_CHECKER` comes out flat grey and you lose the cue without being told. Subdivide the plane and alternate `material_index` per face. Like every other colour here it is notation, so it goes in the translation list and gets dressed over.
|
|
57
|
+
|
|
58
|
+
```python
|
|
59
|
+
# Two materials, alternated per face. `TILE` is the square size in metres.
|
|
60
|
+
dark = bpy.data.materials.new("Checker_Dark")
|
|
61
|
+
dark.diffuse_color = (0.05, 0.05, 0.05, 1.0)
|
|
62
|
+
light = bpy.data.materials.new("Checker_Light")
|
|
63
|
+
light.diffuse_color = (0.80, 0.80, 0.80, 1.0)
|
|
64
|
+
floor.data.materials.append(dark) # material_index 0
|
|
65
|
+
floor.data.materials.append(light) # material_index 1
|
|
66
|
+
# Subdivide first (edit mode or a Subdivide modifier applied) so there ARE
|
|
67
|
+
# faces to alternate — a 2-triangle plane can only ever be one colour.
|
|
68
|
+
for face in floor.data.polygons:
|
|
69
|
+
cx, cy = face.center.x, face.center.y
|
|
70
|
+
face.material_index = (int(cx // TILE) + int(cy // TILE)) % 2
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
And the identity colour on each proxy:
|
|
74
|
+
|
|
75
|
+
```python
|
|
76
|
+
mat = bpy.data.materials.new("ID_Red")
|
|
77
|
+
mat.diffuse_color = (0.8, 0.1, 0.1, 1.0) # what the blocking render draws
|
|
78
|
+
obj.data.materials.append(mat)
|
|
79
|
+
obj.color = (0.8, 0.1, 0.1, 1.0) # same value, for viewport parity
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
The blocking render pins Workbench to `MATERIAL` shading, so **`mat.diffuse_color` is the value that reaches the clip** — and an object with no material at all falls back to a neutral grey, which is why an unpainted set still reads correctly. Set `obj.color` to the same value anyway: it costs one line, it makes the user's viewport match what renders, and keeping the two equal means you never have to remember which one is authoritative.
|
|
83
|
+
|
|
84
|
+
### 3. Camera
|
|
85
|
+
|
|
86
|
+
The whole of `slates-camera-language`. Build the rig, then keyframe it. Then **read back what you built** with `slates_blender_scene` — its `cutSeconds` is your cut list, and it is the number you will write timings against. That field is the authoritative one on EITHER rig — marker frames when cameras are bound to markers, the active camera's own keyframes when they are not. `camera.keyframeSeconds` is empty on a marker-bound edit, which is the rig `slates-camera-language` recommends for anything past a handful of cuts.
|
|
87
|
+
|
|
88
|
+
### 4. Handheld, last
|
|
89
|
+
|
|
90
|
+
Add it after the moves are right, never before — noise on top of a wrong path just hides the wrong path.
|
|
91
|
+
|
|
92
|
+
### 5. Verify the cuts
|
|
93
|
+
|
|
94
|
+
The one check that catches the most damage: on a multi-cut blocking, camera position, target and focal length must all change **exactly on the cut frame, with no transition frame between**. One interpolated frame reads as a whip-pan the model will faithfully reproduce.
|
|
95
|
+
|
|
96
|
+
```python
|
|
97
|
+
# Every camera f-curve keyframe on a cut frame must be CONSTANT out of the
|
|
98
|
+
# previous key, or the cut smears.
|
|
99
|
+
for fc in cam.animation_data.action.fcurves:
|
|
100
|
+
for kp in fc.keyframe_points:
|
|
101
|
+
if int(kp.co[0]) in CUT_FRAMES:
|
|
102
|
+
kp.interpolation = 'CONSTANT'
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
Also check nothing interpenetrates — proxies through floors, clones through the hero object, letters through each other. The model renders intersections as faithfully as it renders everything else.
|
|
106
|
+
|
|
107
|
+
### 6. Save a backup after every stage
|
|
108
|
+
|
|
109
|
+
Cheap, and blocking is iterative by nature.
|
|
110
|
+
|
|
111
|
+
```python
|
|
112
|
+
bpy.ops.wm.save_as_mainfile(filepath=path, copy=True)
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
## Render and generate
|
|
116
|
+
|
|
117
|
+
```
|
|
118
|
+
slates_blender_render_blocking { projectId, fps: 24 }
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Renders the **scene camera** through scene settings — never the user's viewport, so the result does not depend on where they left their mouse — imports the mp4 into the project, and returns `assetId` + `durationSeconds`.
|
|
122
|
+
|
|
123
|
+
Then:
|
|
124
|
+
|
|
125
|
+
```
|
|
126
|
+
slates_generate_video {
|
|
127
|
+
model: "seedance-2.5",
|
|
128
|
+
videoReferenceAssetIds: [<the blocking asset>],
|
|
129
|
+
videoReferenceSecondsEach: [<durationSeconds>],
|
|
130
|
+
characterAssetIds: [...], environmentAssetIds: [...], styleAssetIds: [...],
|
|
131
|
+
prompt: <written per slates-blocking-to-prompt>
|
|
132
|
+
}
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
**Four inputs, and that is the entire stack:** a character sheet each, one location/style reference, the blocking clip, and a prompt written against the blocking. Resist adding a fifth.
|
|
136
|
+
|
|
137
|
+
Model note: seedance-2.5 is the seat for this — 10 reference videos at up to 30s each. seedance-2 and minimax-h3 take 3 at 15s. Route per `slates-model-selection`.
|
|
138
|
+
|
|
139
|
+
## Leaving holes on purpose
|
|
140
|
+
|
|
141
|
+
Where the model outperforms any blockout you could build — liquid, smoke, fire, cloth — **block a black gap instead** and say so in the prompt: `CUT 7 (14.5-17.0, black gap in the reference)`. You are reserving a slot, not forgetting one.
|
|
142
|
+
|
|
143
|
+
## What not to do
|
|
144
|
+
|
|
145
|
+
- **Don't texture, light or material the blocking.** Grey is the specification. The reference supplies motion; the references supply look.
|
|
146
|
+
- **Don't animate what you don't need.** Heads especially — a proxy head turning wrong is worse than one that never turns.
|
|
147
|
+
- **Don't build the camera before the geometry.** It has nothing to aim at, and every value you set gets redone.
|
|
148
|
+
- **Don't skip reading the scene back.** Write timings from `slates_blender_scene`'s `cutSeconds`, never from what you intended to build.
|
|
149
|
+
- **Don't exceed the model's reference-video ceiling.** A 40s blocking against a 30s cap silently truncates.
|
|
150
|
+
|
|
151
|
+
## Related
|
|
152
|
+
|
|
153
|
+
`slates-camera-language` (rigs and moves) · `slates-blocking-to-prompt` (writing the prompt against the clip) · `slates-dialogue-blocking` (multi-character continuity) · `slates-restyle-from-blocking` (one blocking, many worlds) · `slates-model-selection` (routing)
|
|
@@ -0,0 +1,180 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-ltx-2-5
|
|
3
|
+
description: How to prompt LTX-2.5 and LTX-2.5 Pro. Read before calling slates_generate_video with model ltx-2-5 or ltx-2-5-pro. LTX scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE — Lightricks ranks the prompt sound, camera, character detail, shot type and scene, then scene dressing, all in one flowing paragraph. It is also the catalogue's native MULTISHOT seat: one generation carries two to four connected shots holding character, light and voice across the cuts. Base ltx-2-5 is the distilled build — 720p/1080p/1440p/4K, clips of 6 to 20 seconds in EVEN steps, and the cheapest native 1080p second in Slates; ltx-2-5-pro is the full diffusion build and is NOT a superset, reaching only 1080p and 10 seconds for about a third more money. Three hazards live here: durations are even numbers only from six (there is no 5s or 7s clip), the model has NO reference endpoint at all so identity references are unavailable, and any sound not anchored to something in frame gets invented for you.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# LTX-2.5 — prompting
|
|
7
|
+
|
|
8
|
+
LTX-2.5 generates picture and sound **in a single pass**, with a Gemma-4 12B text encoder reading
|
|
9
|
+
one flowing paragraph. That single fact drives everything below: the prompt is not a shot
|
|
10
|
+
description with audio bolted on, it is **a scene where the sound is load-bearing** — and
|
|
11
|
+
Lightricks' own priority order puts sound first, ahead of the camera.
|
|
12
|
+
|
|
13
|
+
Two seats, and the naming is a trap:
|
|
14
|
+
|
|
15
|
+
| | `ltx-2-5` (base) | `ltx-2-5-pro` |
|
|
16
|
+
|---|---|---|
|
|
17
|
+
| Build | Distilled, 8-step | Full diffusion ("Diffusion Fidelity Rendering") |
|
|
18
|
+
| Resolutions | 720p / 1080p / **1440p** / 4K | 720p / 1080p |
|
|
19
|
+
| Durations | 6–20s, even steps | 6 / 8 / 10s |
|
|
20
|
+
| Price | $0.09–$0.30 per second | $0.12–$0.17 per second |
|
|
21
|
+
| Reach for it when | iterating, long takes, 4K delivery, batch volume | one dense final render inside 1080p and 10s |
|
|
22
|
+
|
|
23
|
+
**Pro is not "base plus more."** It buys picture quality on a *narrower* envelope — it cannot make
|
|
24
|
+
a 1440p frame and it cannot make a 12-second clip. Reaching for it out of habit costs a third more
|
|
25
|
+
*and* takes away the reach.
|
|
26
|
+
|
|
27
|
+
---
|
|
28
|
+
|
|
29
|
+
## 1. The six parts, in priority order, in one paragraph
|
|
30
|
+
|
|
31
|
+
Lightricks ranks the elements of an LTX prompt like this. When a prompt sprawls, **cut from the
|
|
32
|
+
bottom.**
|
|
33
|
+
|
|
34
|
+
1. **Sound** — highest priority; the model scores the picture as it draws it.
|
|
35
|
+
2. **Camera** — framing decides visual weight and the feel of the shot.
|
|
36
|
+
3. **Character detail** — expressed as physical action.
|
|
37
|
+
4. **Shot type and scene** — the action itself.
|
|
38
|
+
5. **Scene dressing** — the first thing to trim.
|
|
39
|
+
|
|
40
|
+
Write it as **one flowing paragraph**, not a list of labelled sections. LTX is not Seedance (eight
|
|
41
|
+
engineering slots) and not H3 (three separate audio layers) — it wants continuous prose.
|
|
42
|
+
|
|
43
|
+
---
|
|
44
|
+
|
|
45
|
+
## 2. Sound: anchor it or it gets invented
|
|
46
|
+
|
|
47
|
+
**Write the audio line last, then go back and check every cue has a source you could point at.**
|
|
48
|
+
Anything unanchored, the model invents for you.
|
|
49
|
+
|
|
50
|
+
The test is **"visible, or at least locatable."** A distant whistle is fine *if* you have named the
|
|
51
|
+
marshal's post it comes from. A "distant whistle" with nothing to attach to is a coin flip.
|
|
52
|
+
|
|
53
|
+
> the rope creaks against the cleat as she leans back, gulls calling somewhere off the port bow,
|
|
54
|
+
> the hull knocking hollow against the fenders
|
|
55
|
+
|
|
56
|
+
**Never write mood adjectives as sound.** "Tense atmosphere", "a sense of dread" and "ominous
|
|
57
|
+
ambience" produce nothing usable. If a scene feels thin, the fix is **one more moving object in
|
|
58
|
+
frame with a sound attached to it** — never another adjective.
|
|
59
|
+
|
|
60
|
+
### Dialogue
|
|
61
|
+
|
|
62
|
+
Quote it, and name the language and accent:
|
|
63
|
+
|
|
64
|
+
> "We should not have come back," in English with a slight German accent.
|
|
65
|
+
|
|
66
|
+
Two rules that decide whether the lip sync lands:
|
|
67
|
+
|
|
68
|
+
- **Give the character a beat of stillness before they speak.** The sync needs something to lock
|
|
69
|
+
against; a character already mid-motion when the line starts drifts.
|
|
70
|
+
- **Describe the beat structure** — when they look, how long they wait, when they speak, where they
|
|
71
|
+
look afterwards.
|
|
72
|
+
|
|
73
|
+
Slates pins the frame rate at 25fps, which is also what Lightricks recommends for dialogue: at 50fps
|
|
74
|
+
the performance "pulls toward a video look."
|
|
75
|
+
|
|
76
|
+
---
|
|
77
|
+
|
|
78
|
+
## 3. Character emotion is physical
|
|
79
|
+
|
|
80
|
+
The model renders actions. It does not render adjectives.
|
|
81
|
+
|
|
82
|
+
| Instead of | Write |
|
|
83
|
+
|---|---|
|
|
84
|
+
| she looks anxious | her jaw sets, she turns the ring on her finger twice |
|
|
85
|
+
| he seems exhausted | he blinks slowly and lets his shoulder take the doorframe |
|
|
86
|
+
| a tense standoff | neither moves; his thumb finds the strap and stays there |
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## 4. Multishot — the thing this model is uniquely for
|
|
91
|
+
|
|
92
|
+
**One LTX generation can carry several connected shots**, holding character, environment, lighting,
|
|
93
|
+
voice and style across every cut. Nothing else in the catalogue does this natively; everywhere else
|
|
94
|
+
you generate separate clips and stitch them, and identity drifts between them.
|
|
95
|
+
|
|
96
|
+
**Working range is two to four shots.** Three is the comfortable stopping point.
|
|
97
|
+
|
|
98
|
+
At **every** transition you must supply four things:
|
|
99
|
+
|
|
100
|
+
1. **Name the edit in the prose** — "hard cut", "dissolve", "match cut".
|
|
101
|
+
2. **Re-establish the shot completely** — scale, angle, lens and light all reset at a cut. A cut is
|
|
102
|
+
not a continuation.
|
|
103
|
+
3. **Re-identify recurring characters by their original descriptor.** "The woman in the bronze
|
|
104
|
+
gown", never "she". Pronouns lose the character across a cut — this is the single most common
|
|
105
|
+
multishot failure.
|
|
106
|
+
4. **State what the sound does at the cut.** Silence is not assumed; if the room tone should drop
|
|
107
|
+
out, say so.
|
|
108
|
+
|
|
109
|
+
A shape that works:
|
|
110
|
+
|
|
111
|
+
> Wide establishing shot of the workshop, dust in the window light, a lathe turning somewhere off
|
|
112
|
+
> frame — hard cut — macro close-up of the brass fitting as it seats, the turning noise gone,
|
|
113
|
+
> replaced by a single dry click — match cut — medium shot of the woman in the bronze gown stepping
|
|
114
|
+
> back, the room tone returning underneath her.
|
|
115
|
+
|
|
116
|
+
---
|
|
117
|
+
|
|
118
|
+
## 5. Camera: write it, don't enumerate it
|
|
119
|
+
|
|
120
|
+
fal exposes a `camera_motion` enum (dolly in/out/left/right, jib up/down, static, focus shift).
|
|
121
|
+
**Slates does not surface it, deliberately** — and prose is the better instrument anyway:
|
|
122
|
+
|
|
123
|
+
- **A written move can be tied to a specific moment.** "A slow push-in that settles as she reaches
|
|
124
|
+
the door, then holds" is not expressible as an enum value.
|
|
125
|
+
- **For multishot it would be actively wrong** — one enum value would impose a single camera
|
|
126
|
+
behaviour on three shots that each want their own.
|
|
127
|
+
|
|
128
|
+
So name the lens, the framing, the move, and **the moment the move resolves**.
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
132
|
+
## 6. The hard constraints
|
|
133
|
+
|
|
134
|
+
### Durations are even numbers only, starting at six
|
|
135
|
+
|
|
136
|
+
**6, 8, 10, 12, 14, 16, 18, 20.** There is no 5-second LTX clip and no odd duration of any length.
|
|
137
|
+
Asking for 7s is not a rounding matter — that generation does not exist.
|
|
138
|
+
|
|
139
|
+
**And the long end is 1080p-and-below only.** At 1440p and 4K the ceiling drops to **6, 8 or 10**.
|
|
140
|
+
|
|
141
|
+
fal's own default is `auto`, which lets the model pick the length from the described action.
|
|
142
|
+
**Slates always sends an explicit length instead**, so what you choose is what you are billed for.
|
|
143
|
+
Choose the length the beat needs.
|
|
144
|
+
|
|
145
|
+
### Aspect ratios: 16:9 and 9:16, and nothing else
|
|
146
|
+
|
|
147
|
+
The narrowest set in the catalogue alongside Veo. Square, 4:5 and 21:9 are not available on this
|
|
148
|
+
model at any resolution.
|
|
149
|
+
|
|
150
|
+
### Frames, not references
|
|
151
|
+
|
|
152
|
+
LTX takes a **start frame** and an **optional end frame** (which generates a transition between the
|
|
153
|
+
two). It has **no reference-to-video endpoint at all** — no identity references, no style
|
|
154
|
+
references, no environment references, no reference video, no reference audio.
|
|
155
|
+
|
|
156
|
+
**For character consistency across separate shots, use MiniMax H3 or Kling.** Within a single LTX
|
|
157
|
+
generation, use multishot instead — that is precisely the gap it fills.
|
|
158
|
+
|
|
159
|
+
In image-to-video, **do not cut away from the opening frame too early.** You have paid for that
|
|
160
|
+
frame; let it play before the first move.
|
|
161
|
+
|
|
162
|
+
### Do not ask for text on screen
|
|
163
|
+
|
|
164
|
+
Neither the spelling nor its stability from frame to frame can be relied on. Signage, labels,
|
|
165
|
+
captions and lower-thirds belong in post.
|
|
166
|
+
|
|
167
|
+
---
|
|
168
|
+
|
|
169
|
+
## 7. Audio is free here, and that changes the routing
|
|
170
|
+
|
|
171
|
+
Native synchronised audio is **included at every resolution on both seats**, with no surcharge and
|
|
172
|
+
no toggle that costs money — unlike Kling, where sound is a paid dimension. A 6-second 1080p LTX
|
|
173
|
+
clip **with sound** is 39 credits.
|
|
174
|
+
|
|
175
|
+
Combined with 1080p at $0.13/s — the cheapest native 1080p second in Slates — this makes LTX **the
|
|
176
|
+
coverage seat**: the one to reach for when the job is many takes rather than one hero shot, when a
|
|
177
|
+
sequence needs its own sound, or when the credit budget is the binding constraint.
|
|
178
|
+
|
|
179
|
+
Route away from it when you need identity references (H3, Kling), a ratio other than 16:9 or 9:16
|
|
180
|
+
(Seedance, Kling), or authored multi-layer audio direction (H3).
|
|
@@ -78,6 +78,15 @@ Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and a
|
|
|
78
78
|
|
|
79
79
|
These are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.
|
|
80
80
|
|
|
81
|
+
<!-- @banned:start -->
|
|
82
|
+
<!-- slates-only -->
|
|
83
|
+
<!-- MACHINE-READ. Every `backticked` token between the @banned markers is extracted
|
|
84
|
+
by src/prompts/banned-tokens.ts, inlined verbatim into the slates_generate_image
|
|
85
|
+
op description (always in context on both surfaces), and matched against every
|
|
86
|
+
submitted prompt. Editing this list changes what the agent is told AND what it
|
|
87
|
+
is warned about — keep every entry backticked, and keep prose outside the
|
|
88
|
+
backticks. -->
|
|
89
|
+
<!-- /slates-only -->
|
|
81
90
|
**Never use:**
|
|
82
91
|
- `8k`, `4k` (as a quality token)
|
|
83
92
|
- `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone
|
|
@@ -86,6 +95,7 @@ These are Stable-Diffusion-era tag soup. The model treats them as low-signal noi
|
|
|
86
95
|
- `perfect skin`, `flawless`, `airbrushed`, `smooth skin`
|
|
87
96
|
- `cinematic` standing alone — always specify *which cinema* (director, lens, era, stock)
|
|
88
97
|
- `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue
|
|
98
|
+
<!-- @banned:end -->
|
|
89
99
|
|
|
90
100
|
## Negative prompting — there is no field
|
|
91
101
|
|
|
@@ -339,7 +339,16 @@ Speed ramps and slow-motion are supported in natural language, and `fast` is wid
|
|
|
339
339
|
the lid opens in slow-motion · the blade whips through the air
|
|
340
340
|
```
|
|
341
341
|
|
|
342
|
-
|
|
342
|
+
<!-- @banned:start -->
|
|
343
|
+
<!-- slates-only -->
|
|
344
|
+
<!-- MACHINE-READ — same contract as the anti-list in slates-prompting-nano-banana-2.
|
|
345
|
+
Extracted by src/prompts/banned-tokens.ts into the slates_generate_video op
|
|
346
|
+
description and matched against submitted prompts. The RECOMMENDED vocabulary
|
|
347
|
+
below sits OUTSIDE the markers on purpose — it is backticked too. -->
|
|
348
|
+
<!-- /slates-only -->
|
|
349
|
+
**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.
|
|
350
|
+
<!-- @banned:end -->
|
|
351
|
+
These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
|
|
343
352
|
|
|
344
353
|
## Style block at the end
|
|
345
354
|
|
|
@@ -0,0 +1,121 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-restyle-from-blocking
|
|
3
|
+
description: Render one blocking pass as several different visual worlds — live action, 2.5D painted, 2D ink, toybox — matching cut for cut. Use when a client needs style options, when someone wants to see the same edit in another look, or when an approved edit needs a new treatment without re-blocking.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Restyle — one edit, many worlds
|
|
7
|
+
|
|
8
|
+
The commercial payoff of the whole previs workflow, and the reason a blocking file is an asset rather than a step.
|
|
9
|
+
|
|
10
|
+
## The idea
|
|
11
|
+
|
|
12
|
+
Every prompt has two halves:
|
|
13
|
+
|
|
14
|
+
- **Structure** — cuts, camera, timing, who is where. Lives in the blocking clip. **Never changes.**
|
|
15
|
+
- **Style** — what any of it looks like. Lives in the references and the prompt text. **Changes freely.**
|
|
16
|
+
|
|
17
|
+
Hold the structure, swap the style, and the same edit comes back as live action, painted 2.5D, ink on paper or a toybox — **matching frame for frame across all of them.** Cuts land on the same frames, the car drifts at the same moment, the same head turns at the same beat.
|
|
18
|
+
|
|
19
|
+
For anyone pitching work: three visual worlds in a day, off one edit the client has already approved. The foundation is not up for renegotiation, so the conversation is only about look.
|
|
20
|
+
|
|
21
|
+
## Before you restyle
|
|
22
|
+
|
|
23
|
+
You need a blocking clip whose structure you are happy with, and a finished prompt for at least one style (per `slates-blocking-to-prompt`). The first style is the expensive one; every later style is an edit of its text.
|
|
24
|
+
|
|
25
|
+
## What stays fixed
|
|
26
|
+
|
|
27
|
+
Copy these across every style **verbatim**. Changing them is what desynchronises the outputs:
|
|
28
|
+
|
|
29
|
+
- The blocking reference's own contract — that it is the master for all movement, the placement-only clause, the tie-break clause, the disambiguation clause
|
|
30
|
+
- The shot count and every timestamp
|
|
31
|
+
- Every shot's camera position, angle, framing and cut point
|
|
32
|
+
- Screen direction and seating
|
|
33
|
+
- The `HOLD FOR THE FULL TIMELINE` block
|
|
34
|
+
- `videoReferenceAssetIds` and `videoReferenceSecondsEach`
|
|
35
|
+
|
|
36
|
+
Lead each style's prompt with a lock so the style layer cannot leak into the structure:
|
|
37
|
+
|
|
38
|
+
> VIDEO LOCK — the dominant rule of this prompt: the reference defines 100% of the motion, editing and object choreography. The text below defines only look, materials, locations and effects layered onto that motion. Wherever the text and the video could be read differently about motion, the video decides.
|
|
39
|
+
|
|
40
|
+
## What changes
|
|
41
|
+
|
|
42
|
+
| Layer | What you swap |
|
|
43
|
+
|---|---|
|
|
44
|
+
| Rendering style | photoreal · painted 2.5D · 2D ink · miniature/toybox |
|
|
45
|
+
| Characters | different sheets entirely — a couple, grandparents, a robot and a cat |
|
|
46
|
+
| Locations | the same four beats set in a different world |
|
|
47
|
+
| Time of day / weather | night after rain · golden hour · hard noon |
|
|
48
|
+
| Lighting and colour | per style |
|
|
49
|
+
| Audio | SFX-only, or scored, or lip-synced dialogue |
|
|
50
|
+
|
|
51
|
+
Characters can change species and still land, because the blocking only supplies where a body is and how it moves.
|
|
52
|
+
|
|
53
|
+
## Dummy mapping — the mechanism that makes it work
|
|
54
|
+
|
|
55
|
+
Each style needs its own explicit mapping from grey proxy to real object. The proxy is a slot; the style fills it:
|
|
56
|
+
|
|
57
|
+
> DUMMY MAPPING: the front-LEFT sphere-head dummy (with its grey arm at the shifter and grey leg at the pedals) is THE GRANDPA; the front-RIGHT sphere-head dummy is THE GRANDMA; a front-seat dummy together with its loose blocks is that ONE whole person. Blocks on the rear bench are the luggage. The low-poly flying model in SHOT 18 is THE HELICOPTER. The two vehicles behind the hero car in SHOT 19 are THE POLICE CARS.
|
|
58
|
+
|
|
59
|
+
Same clause per style, different right-hand side. And restate the placement-only rule in style terms:
|
|
60
|
+
|
|
61
|
+
> The source defines only placement and motion, never appearance: every placeholder becomes the real object its position implies — spheres are always people, cabin blocks are always cases and bags, fully drawn.
|
|
62
|
+
|
|
63
|
+
## Location continuity
|
|
64
|
+
|
|
65
|
+
If the piece travels, name the places and pin each shot to one. Reusing labels across styles keeps the four prompts diffable:
|
|
66
|
+
|
|
67
|
+
```
|
|
68
|
+
LOCATION CONTINUITY — one journey through four fixed places; each looks
|
|
69
|
+
identical in every shot where it appears:
|
|
70
|
+
LOC-A <opening> LOC-B <middle> LOC-C <turn> LOC-D <finale>
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
Then tag every beat: `SHOT 9 — 7.79-9.33s — LOCKED, LOC-B: <description>`.
|
|
74
|
+
|
|
75
|
+
**On a piece that visits many places, make the map absolute and countable** — otherwise the model reuses a room it liked and you get the same interior three times:
|
|
76
|
+
|
|
77
|
+
> The location map is absolute — SEVEN locations, each appearing EXACTLY ONCE, in this exact order: 1) yard 00:00-00:03.3 … 7) rooftop 00:20-00:30. No location ever appears twice, and the three interiors are three COMPLETELY DIFFERENT rooms — different walls, furniture, people and light — never the same room repeated.
|
|
78
|
+
|
|
79
|
+
## Style references
|
|
80
|
+
|
|
81
|
+
A style reference is **not a keyframe**, and saying so prevents the model reproducing its composition as a shot:
|
|
82
|
+
|
|
83
|
+
> STYLE MASTER — defines the painting and rendering style only: hand-painted look with visible brushstrokes, sculpted painterly volumes, textured matte surfaces, dramatic coloured rim light, deep moody shadows. NOT a keyframe, NOT a location to reproduce, NOT a frame that ever appears in the film. Its own subject, framing and composition are never seen in any shot.
|
|
84
|
+
|
|
85
|
+
A style can also be **text-only** — no reference image at all. Ink and toybox looks usually specify better in words than they match from a still.
|
|
86
|
+
|
|
87
|
+
## Keep performance inside the existing shots
|
|
88
|
+
|
|
89
|
+
Style changes tempt the model to earn new coverage. Refuse it:
|
|
90
|
+
|
|
91
|
+
> ACTING — inside the existing shots only: performance is visible only at the size and distance the reference already gives it, only where the source already shows a face; everywhere else it reads through posture and hands alone. The performance NEVER earns a new shot, a new angle or a closer framing.
|
|
92
|
+
|
|
93
|
+
## Text-free worlds
|
|
94
|
+
|
|
95
|
+
Stylised worlds are where invented signage and garbled lettering appear. One clause kills it:
|
|
96
|
+
|
|
97
|
+
> TEXT-FREE WORLD: every sign is a blank painted shape, every gauge face carries tick marks only, every licence plate is a blank plate.
|
|
98
|
+
|
|
99
|
+
## Running it
|
|
100
|
+
|
|
101
|
+
Generate each style as its own `slates_generate_video` call against the **same** `videoReferenceAssetIds`. Keep them in one project so they sit side by side; name assets by style so the comparison reads at a glance.
|
|
102
|
+
|
|
103
|
+
Quote the whole set before firing — `slates_estimate_generation_cost` per style — and confirm. Four styles is four generations, not one.
|
|
104
|
+
|
|
105
|
+
🚨 Never fire a batch of style variants without showing the user the prompts and the total cost first.
|
|
106
|
+
|
|
107
|
+
## Checklist per style
|
|
108
|
+
|
|
109
|
+
- [ ] Same blocking asset, same `videoReferenceSecondsEach`
|
|
110
|
+
- [ ] VIDEO LOCK leads the prompt
|
|
111
|
+
- [ ] Every timestamp and shot count identical to style 1
|
|
112
|
+
- [ ] Dummy mapping written for this style's cast
|
|
113
|
+
- [ ] Style reference declared as style-only, or none used
|
|
114
|
+
- [ ] Location labels reused; on a travelling piece the map is absolute and countable
|
|
115
|
+
- [ ] Acting-inside-existing-shots clause present
|
|
116
|
+
- [ ] HOLD block copied verbatim
|
|
117
|
+
- [ ] Cost quoted and confirmed
|
|
118
|
+
|
|
119
|
+
## Related
|
|
120
|
+
|
|
121
|
+
`slates-previs-blocking` · `slates-blocking-to-prompt` · `slates-style-prompting` (style vocabulary per model) · `slates-cost-discipline` (batch quoting)
|