@slatesvideo/shared 0.6.1 → 0.6.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/clients/blender.d.ts +50 -0
- package/dist/clients/blender.js +195 -0
- package/dist/operations/index.d.ts +24 -0
- package/dist/operations/index.js +244 -1
- package/dist/prompts/reference-composer.d.ts +57 -0
- package/dist/prompts/reference-composer.js +70 -1
- package/dist/skills/content.js +5 -0
- package/package.json +1 -1
- package/skills/slates-blocking-to-prompt.md +250 -0
- package/skills/slates-camera-language.md +196 -0
- package/skills/slates-dialogue-blocking.md +134 -0
- package/skills/slates-previs-blocking.md +153 -0
- package/skills/slates-restyle-from-blocking.md +121 -0
|
@@ -0,0 +1,250 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-blocking-to-prompt
|
|
3
|
+
description: Write the generation prompt that matches a blocking clip second by second, so the reference video and the text agree instead of fighting. Use after rendering a previs blocking pass, when a generated shot ignores the reference video, when timings drift, or when the model invents shots and camera angles that are not in the blocking.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Blocking → prompt
|
|
7
|
+
|
|
8
|
+
You have a blocking clip. This is how you write the prompt that goes with it.
|
|
9
|
+
|
|
10
|
+
## The one idea
|
|
11
|
+
|
|
12
|
+
The clip already contains the camera, the cuts and the timing. **The prompt's job is to say what everything looks like — and, where the grey boxes are ambiguous, to disambiguate them.** It is not a second, competing description of the motion.
|
|
13
|
+
|
|
14
|
+
State that contract inside the prompt, because the model needs it as much as you do:
|
|
15
|
+
|
|
16
|
+
> Where a timeline line below names a camera position or move, it is a restatement of what the reference video already does at that timestamp — a disambiguation, never a new instruction.
|
|
17
|
+
|
|
18
|
+
And give it a tie-break, because ambiguity is guaranteed:
|
|
19
|
+
|
|
20
|
+
> If any text in this prompt appears to disagree with the reference video about camera, framing, direction, motion, timing or object placement, the reference video wins.
|
|
21
|
+
|
|
22
|
+
Those two sentences do more work than any other part of the prompt.
|
|
23
|
+
|
|
24
|
+
## Get the real numbers first
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
slates_blender_scene
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
Write timings from `cutSeconds`, never from the shot list you intended to build. It resolves to the marker frames on a multi-camera edit and to the camera's own keyframes otherwise, so it is the one field that is never empty on a rig that has cuts. At 24fps cuts land on frame boundaries and the honest values are not round — `7.79s`, `9.33s`, `19.875s`. **Use the exact ones.** Rounding to `7.8` is a tenth of drift you are handing the model for free.
|
|
31
|
+
|
|
32
|
+
## Structure
|
|
33
|
+
|
|
34
|
+
Order matters — contract, then globals, then timeline, then the re-assertion.
|
|
35
|
+
|
|
36
|
+
```
|
|
37
|
+
TITLE — one line: what this is, how long, that it is video-to-video
|
|
38
|
+
|
|
39
|
+
LOGLINE — 2-4 sentences. The whole piece in plain language.
|
|
40
|
+
|
|
41
|
+
ACTIVE REFERENCES
|
|
42
|
+
<one entry per reference: what it defines, and what is NOT inherited>
|
|
43
|
+
|
|
44
|
+
TECHNICAL BLOCK (format, grade, lens, and the blanket negatives — see below)
|
|
45
|
+
|
|
46
|
+
STYLE / LOOK
|
|
47
|
+
LIGHTING
|
|
48
|
+
COLOR
|
|
49
|
+
CAMERA
|
|
50
|
+
PHYSICS (only if things move, collide or deform)
|
|
51
|
+
|
|
52
|
+
RULES (numbered — the invariants, see below)
|
|
53
|
+
|
|
54
|
+
ACTION TIMING (beat by beat, against the clip's real timestamps)
|
|
55
|
+
|
|
56
|
+
AUDIO
|
|
57
|
+
[Sound design] [Timed accents] [Dialogue] [Music]
|
|
58
|
+
|
|
59
|
+
ENDING LOCK (one line — where the film stops)
|
|
60
|
+
|
|
61
|
+
HOLD FOR THE FULL TIMELINE
|
|
62
|
+
<the 5-6 constraints most likely to drift, compressed>
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
## ACTIVE REFERENCES — every entry has an exclusion
|
|
66
|
+
|
|
67
|
+
The single highest-leverage format in this whole workflow. Each reference is a **positive claim plus an exclusion list**, because a reference the model over-reads is as damaging as one it ignores.
|
|
68
|
+
|
|
69
|
+
Label each by the badge code Slates echoes back (`IMG-A8`, `VID-C2`) or by an unmistakable role name, and use that same label everywhere below.
|
|
70
|
+
|
|
71
|
+
**The blocking clip:**
|
|
72
|
+
|
|
73
|
+
> VID-C2 = the blocking previz (30s, 720 frames, 24fps) — the MASTER for everything that moves and everything that stands. It defines the full edit one-to-one: every cut point, every camera position, angle, move and framing, all action timing, screen direction, and the geometry of the world. Its untextured grey surfaces, flat colours and viewport grid are NOT inherited — every grey proxy is dressed into a real object in the exact position the previz puts it. Proxies give position, angle, scale and motion only; never surface, shape detail or design.
|
|
74
|
+
|
|
75
|
+
That last sentence is the **placement-only clause** and it is not optional. Without it the model renders grey boxes.
|
|
76
|
+
|
|
77
|
+
**A character sheet:**
|
|
78
|
+
|
|
79
|
+
> IMG-A8 = the driver — defines his face, hair and wardrobe. Identity 100% consistent at every distance and through every motion blur. Background, lighting and pose are NOT inherited.
|
|
80
|
+
|
|
81
|
+
**A location/style reference:**
|
|
82
|
+
|
|
83
|
+
> IMG-B3 = the tunnel — defines location geometry, look and grade. Camera angle, framing and any people in it are NOT inherited; the camera comes exclusively from VID-C2.
|
|
84
|
+
|
|
85
|
+
**An atmosphere or style master — a reference that is never a shot:**
|
|
86
|
+
|
|
87
|
+
> IMG-D9 = ATMOSPHERE MASTER — NOT a keyframe, NOT a location to reproduce, NOT a frame that ever appears in the film: its own subject, framing and composition are never seen in any shot. It defines ONLY the weather, light, colour and grade: deep clean night just after rain, wet asphalt as a dark mirror, cool white-cyan lamps as the ambient key, teal-and-amber grade, deep clean blacks. Every shot is lit and graded in this regime for all 30 seconds.
|
|
88
|
+
|
|
89
|
+
Without those three NOTs the model reproduces the reference's composition as an actual shot — you get its street corner in your film. The same wording covers a rendering-style master; see `slates-restyle-from-blocking`.
|
|
90
|
+
|
|
91
|
+
**References can be scheduled.** If something is only true for part of the timeline, say so: *the hooded panel applies only to 0–7.0s and 27.5–30s*, or per-reference: *Active for 00:03.3–00:06.7 only.* On a piece that travels through several locations, every location still carries its own window and the model stops blending two sets into one shot.
|
|
92
|
+
|
|
93
|
+
## Translate the blocking's artifacts
|
|
94
|
+
|
|
95
|
+
Your grey-box render contains things that are *notation*, not content. Every one needs an explicit reinterpretation or it gets rendered literally:
|
|
96
|
+
|
|
97
|
+
| In the blocking | Say in the prompt |
|
|
98
|
+
|---|---|
|
|
99
|
+
| Colour-coded bodies | `red = the boss, green = the kid, blue = the driver` |
|
|
100
|
+
| A marked face on a proxy | `RED face = the direction he faces, BLACK = his back` |
|
|
101
|
+
| A checkered floor or wall | `the checkerboard is a scale reference, not a surface — it becomes <the real material>` |
|
|
102
|
+
| Flat black background | `a PLACEHOLDER — replace with the location assigned below` |
|
|
103
|
+
| A floor grid | `a motion-tracking aid — render as light on the surface, never as wireframe or tiles` |
|
|
104
|
+
| A deliberate black gap | `CUT 7 (14.5-17.0, black gap in the reference) — <what fills it>` |
|
|
105
|
+
| Frame goes dark mid-move | `the camera is passing through the ground — a doorway to the NEXT location, never back to a previous one` |
|
|
106
|
+
| The source hard-resets mid-move | `each reset begins a NEW, completely different room — never a replay of one already seen` |
|
|
107
|
+
| A proxy that is a PROP or VEHICLE | `the low-poly flying model in SHOT 18 is THE HELICOPTER · blocks on the rear bench are the luggage · the small dark block in his hand IS the pistol` |
|
|
108
|
+
| A blocky proxy limb in a tight insert | `the blocky low-poly leg is a stand-in and must NOT be replicated — generate complete human anatomy: a real boot, a real trouser leg, a correct ankle at this exact camera angle` |
|
|
109
|
+
| A stray object at the frame edge | `ignore it completely — never blend two sets into one shot` |
|
|
110
|
+
|
|
111
|
+
## TECHNICAL BLOCK — format, lens, and the blanket negatives
|
|
112
|
+
|
|
113
|
+
One paragraph, before the timeline. It carries the things that are true of every frame and that no beat should have to repeat:
|
|
114
|
+
|
|
115
|
+
> Cinematic, photoreal. 21:9. 30s. SFX only, no music. Kodak 500T film look, natural 35mm grain, organic colour, soft highlight roll-off, anamorphic lens character with oval bokeh and gentle barrel distortion at the edges, chromatic aberration creeping in at the frame edges, natural motion blur on every fast move, faint bloom on hot speculars. Every location well exposed — night interiors bright and readable, open shadows, no crushed blacks, no murk. NO CGI. NON-IP, no brand badges or logos anywhere, no text, no watermark.
|
|
116
|
+
|
|
117
|
+
Three parts worth naming:
|
|
118
|
+
|
|
119
|
+
- **Lens realism is a list, not an adjective.** Aberration, motion blur, depth of field, barrel distortion, bloom, grain. "Cinematic" buys you nothing; these buy you the look.
|
|
120
|
+
- **Exposure needs saying on dark work.** Models crush night scenes into murk. *Night interiors bright and readable, open shadows, no crushed blacks* is what keeps a scene legible.
|
|
121
|
+
- **The blanket negatives go here once** — `NON-IP`, no logos, no on-screen text, no subtitles, no watermark — rather than being scattered through the beats.
|
|
122
|
+
|
|
123
|
+
## RULES — the invariants
|
|
124
|
+
|
|
125
|
+
Numbered, short, absolute. These are the things that must hold in every frame, and they are where you put anything that has already gone wrong once.
|
|
126
|
+
|
|
127
|
+
Two patterns worth stealing outright:
|
|
128
|
+
|
|
129
|
+
**Countable state.** Give the model arithmetic it can check itself against:
|
|
130
|
+
|
|
131
|
+
> At every second: standing + fallen + on the lintel = 6. Never a seventh figure — no extras, no duplicates, no distant silhouettes, no half-bodies at frame edges.
|
|
132
|
+
|
|
133
|
+
> Bodies on the ground count exactly: 0 before 12s → 1 → 2 → 3 → 4 at 12/13/14/16s → 5 at 20s → 6 at 26s. Never more.
|
|
134
|
+
|
|
135
|
+
**Every mass is dressed, and nothing is invented.** The blocking is authority over what EXISTS, not just what moves — otherwise the model deletes the masses it finds boring and adds architecture you never blocked:
|
|
136
|
+
|
|
137
|
+
> Every lamppost, guardrail, road and terrain mass visible in the reference exists in the output in the same place, at the same scale, in the same position in frame — the opening blocks are dark-brick warehouse facades, the roadside masses are the waterfront skyline, the finale rocks are the city's tower walls. Nothing is deleted, and no structure is invented where the reference shows none.
|
|
138
|
+
|
|
139
|
+
**Anatomy is never inherited from a proxy.** Tight inserts on hands and feet are where blocking leaks straight into the render:
|
|
140
|
+
|
|
141
|
+
> The reference shows only WHERE hands and feet are. In the output they are always complete human anatomy — a five-fingered gloved hand with natural knuckles, a real leg in wool trousers, a real foot in a leather shoe — never the proxy's blocky shape.
|
|
142
|
+
|
|
143
|
+
**A ledger for anything that happens a countable number of times.** A ritual, a reload, a set of falls: state it as a linear sequence, each step exactly once, and close with the tally:
|
|
144
|
+
|
|
145
|
+
> Strictly linear, six steps in fixed order, each happening EXACTLY ONCE and never repeating; once a step is done it is done for good, and the sequence only ever moves FORWARD, never backward. Count of weapon events in the entire video: one draw, one magazine insertion, one slide rack, one shot.
|
|
146
|
+
|
|
147
|
+
Without the ledger the model loops the most cinematic beat — it will rack the slide four times because racking looks good.
|
|
148
|
+
|
|
149
|
+
**Named misreads.** When a generation gets something specifically wrong, do not rewrite the description — **name the wrong reading and kill it**:
|
|
150
|
+
|
|
151
|
+
> The lamp is a man-made steel structure — NOT an animal, NOT a snake, NOT any living or organic shape.
|
|
152
|
+
|
|
153
|
+
> The rear of the car and its tail lights are NOT visible in this shot.
|
|
154
|
+
|
|
155
|
+
This is the highest-value edit available after a failed roll, and it is why the prompt grows rather than changes between takes.
|
|
156
|
+
|
|
157
|
+
## ACTION TIMING — the beats
|
|
158
|
+
|
|
159
|
+
One block per shot or beat. Two notations; pick one and hold it.
|
|
160
|
+
|
|
161
|
+
**For a continuous take**, ranges with a camera note and a closing state audit:
|
|
162
|
+
|
|
163
|
+
```
|
|
164
|
+
8-12s — THE SWEEP (per VID-C2: elevated rear push, swinging to profile by 12s):
|
|
165
|
+
<what happens, in prose, with sub-beats on tenths and → chaining cause to effect>
|
|
166
|
+
END 12s: bodies 1 (behind him as he steps past) · standing — four ahead, holding.
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
**For a cut edit**, numbered shots ending on their cut:
|
|
170
|
+
|
|
171
|
+
```
|
|
172
|
+
9.33-10.33s — SHOT 10 — Interior over the centre console as in VID-C2: <what the
|
|
173
|
+
frame contains>. Hard cut at 10.33s.
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
Three habits that separate a beat that works from one that does not:
|
|
177
|
+
|
|
178
|
+
- **Declare the frame's contents as a closed set** when the shot is tight: *the frame holds exactly the console, the lever, his hand, and the edges of both seats.* An open description invites additions.
|
|
179
|
+
- **Chain cause to effect inside one sentence** with `→`. `he overcommits a lunge → the Hero drops low and sweeps his standing leg → he hits the earth at 12s`.
|
|
180
|
+
- **Put events on tenths.** `11.7s`, `19.5s`, `22.5s`. Vague beats generate vague timing.
|
|
181
|
+
|
|
182
|
+
Density: roughly 60–130 words per second of screen time is what these prompts actually run at. That is much denser than a normal video prompt, and it is the point.
|
|
183
|
+
|
|
184
|
+
## AUDIO
|
|
185
|
+
|
|
186
|
+
`[Timed accents]` uses the same timestamps as the beats:
|
|
187
|
+
|
|
188
|
+
> 3.1s tyres light up into the burnout squeal · 7.0s drift-entry screech · 10.1s hard mechanical shifter clack · 20.3s full-speed pass-by whoosh
|
|
189
|
+
|
|
190
|
+
`[Dialogue]` is a closed list — count the lines, give each a window, quote it verbatim, and forbid everything else:
|
|
191
|
+
|
|
192
|
+
> Exactly TWO vocal events in the entire 30 seconds, both screamed, in English, VERBATIM: 1. 17.3-18.6s "STOOOOOP!!" 2. 23.4-24.0s "You crazy!" Nothing else is ever spoken.
|
|
193
|
+
|
|
194
|
+
**Then forbid the lines it will invent anyway.** A closed list is a rule; an enumerated blacklist is enforcement, and the phrases to list are the clichés the scene invites:
|
|
195
|
+
|
|
196
|
+
> FORBIDDEN — she never says any of these and no one else says anything: "they're behind us", "cops", "go go go", "are you crazy", "you're insane", or ANY other invented phrase. All other human voice is wordless screaming or laughing.
|
|
197
|
+
|
|
198
|
+
**When lines are lip-synced, give each one a timestamp** in the same list, and say that they change nothing else:
|
|
199
|
+
|
|
200
|
+
> Timing: "You wind up for this one?" ~18.2s · "Three full turns." ~19.0s · "Company." ~24.3s. Every line lip-synced; the lines never change the camera.
|
|
201
|
+
|
|
202
|
+
Two rules that stop dialogue from breaking the edit:
|
|
203
|
+
|
|
204
|
+
> DIALOGUE NEVER CREATES SHOTS: spoken lines happen inside the reference's takes exactly as blocked — no cutaways to a speaker, no reverse shots, no added close-ups. If a line plays while the camera is elsewhere, the line stays off-screen audio.
|
|
205
|
+
|
|
206
|
+
> A line marked off-screen must STAY off-screen — never show the speaker, never move him into frame, never route the camera to him because he spoke.
|
|
207
|
+
|
|
208
|
+
Model note: dialogue direction as separate layers is minimax-h3's seat; native synced audio is Veo's. Route per `slates-model-selection` and read the model's own prompting skill before writing the audio block.
|
|
209
|
+
|
|
210
|
+
## ENDING LOCK
|
|
211
|
+
|
|
212
|
+
One line, and it is the cheapest fix in the document. Models drift at the end — they hold a frame too long, add a beat after the last one, or fade somewhere the reference does not:
|
|
213
|
+
|
|
214
|
+
> The film ends exactly where the reference ends: the final take runs unbroken to its last frame, and that source frame IS the final frame of the film. Nothing follows. The last five seconds follow the source exactly as strictly as the first five.
|
|
215
|
+
|
|
216
|
+
## HOLD FOR THE FULL TIMELINE
|
|
217
|
+
|
|
218
|
+
Close with a terminal re-assertion of only the constraints most prone to drift — five or six lines, compressed, no new information:
|
|
219
|
+
|
|
220
|
+
```
|
|
221
|
+
HOLD FOR THE FULL TIMELINE
|
|
222
|
+
- VID-C2 camera path 1:1 — any deviation = failure.
|
|
223
|
+
- Six and only six figures; the count above holds at every second.
|
|
224
|
+
- IMG-A8 identity constant at every distance and through motion blur.
|
|
225
|
+
- The IMG-B3 location in every frame; no subtitles, no watermarks.
|
|
226
|
+
```
|
|
227
|
+
|
|
228
|
+
Restating is not redundancy here. It is the last thing the model reads.
|
|
229
|
+
|
|
230
|
+
## Checklist before you generate
|
|
231
|
+
|
|
232
|
+
- [ ] Timings taken from `slates_blender_scene`'s `cutSeconds`, frame-exact, not rounded
|
|
233
|
+
- [ ] Every reference has an explicit "NOT inherited"
|
|
234
|
+
- [ ] The placement-only clause is present
|
|
235
|
+
- [ ] The tie-break clause is present
|
|
236
|
+
- [ ] The disambiguation clause is present
|
|
237
|
+
- [ ] Every blocking artifact is translated — colours, marked faces, checkers, grid, black background, gaps, dark dips, resets, prop proxies, proxy limbs, strays
|
|
238
|
+
- [ ] Any style/weather reference is declared NOT a keyframe and never a shot
|
|
239
|
+
- [ ] The every-mass-is-dressed / nothing-invented rule is present
|
|
240
|
+
- [ ] Tight inserts on hands or feet demand complete anatomy
|
|
241
|
+
- [ ] Counts are stated where anything is countable, and repeatable actions carry a ledger
|
|
242
|
+
- [ ] A TECHNICAL BLOCK carries format, lens realism, exposure and the blanket negatives
|
|
243
|
+
- [ ] `[Dialogue]` is a closed list with a FORBIDDEN blacklist
|
|
244
|
+
- [ ] An ENDING LOCK says where the film stops
|
|
245
|
+
- [ ] `videoReferenceSecondsEach` matches the clip's real duration
|
|
246
|
+
- [ ] A HOLD block closes it
|
|
247
|
+
|
|
248
|
+
## Related
|
|
249
|
+
|
|
250
|
+
`slates-previs-blocking` (producing the clip) · `slates-camera-language` (the moves being described) · `slates-dialogue-blocking` (multi-character continuity) · `slates-restyle-from-blocking` (reusing this prompt across styles) · `slates-prompting-seedance-2-5` / `slates-prompting-minimax-h3` (model-specific rules)
|
|
@@ -0,0 +1,196 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-camera-language
|
|
3
|
+
description: Turn director vocabulary into real Blender camera rigs — orbits, floor rises, robo-arm whips, handheld, speed ramps, over-the-shoulder cuts — as bpy code. Use when building or refining the camera on a previs blocking pass, when a move needs to accelerate/hold/snap, or when someone asks for a "cinematic" camera and you need to convert that into an actual shot list.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Camera language — from a shot list to a rig
|
|
7
|
+
|
|
8
|
+
Companion to `slates-previs-blocking`. That skill owns the workflow; this one owns the camera.
|
|
9
|
+
|
|
10
|
+
## The first rule
|
|
11
|
+
|
|
12
|
+
🚨 **Never build "a cinematic camera move." Brief the camera the way you would brief an operator:** rails, target, height, lens, and the frame each move starts and ends on. "Cinematic" is not a specification, and asking for one produces the drifting slop the whole blocking workflow exists to avoid.
|
|
13
|
+
|
|
14
|
+
Bad: *a dynamic cinematic orbit around the subject.*
|
|
15
|
+
Good: *a 3/4 orbit starting rear-left at 1.6m, ending front-right at 0.9m, frames 1–96, 35mm, easing out of the start and holding hard on the last 8 frames.*
|
|
16
|
+
|
|
17
|
+
If the user gives you the first, convert it to the second and say what you assumed.
|
|
18
|
+
|
|
19
|
+
## Look it up, don't recall it
|
|
20
|
+
|
|
21
|
+
Before any constraint or operator you are not certain of, call `slates_blender_docs` (e.g. `bpy.types.FollowPathConstraint`) or `slates_blender_search_docs`. Invented enum values are the most common failure here and they often fail *quietly* — the constraint gets added, the axis is wrong, and the camera points at nothing.
|
|
22
|
+
|
|
23
|
+
## The two primitives everything is built from
|
|
24
|
+
|
|
25
|
+
### Target-based aiming
|
|
26
|
+
|
|
27
|
+
Almost every move in this skill is *position on a path* plus *aim at a target*. Separating them is what lets framing vary while the subject stays in frame.
|
|
28
|
+
|
|
29
|
+
```python
|
|
30
|
+
import bpy
|
|
31
|
+
|
|
32
|
+
target = bpy.data.objects.new("CAM_TARGET", None) # an Empty
|
|
33
|
+
target.empty_display_type = 'PLAIN_AXES'
|
|
34
|
+
bpy.context.collection.objects.link(target)
|
|
35
|
+
|
|
36
|
+
con = cam.constraints.new('TRACK_TO')
|
|
37
|
+
con.target = target
|
|
38
|
+
con.track_axis = 'TRACK_NEGATIVE_Z' # cameras look down -Z
|
|
39
|
+
con.up_axis = 'UP_Y'
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
**Aim at a separate target, not at the subject's head.** A camera locked to the head produces dead, centred framing. Offset the target beside or ahead of the subject and the shot breathes — that small offset is most of what reads as "real operator."
|
|
43
|
+
|
|
44
|
+
For a subject that should notice the camera, keyframe the target's follow with a **few frames of lag** behind each camera move. Heads catch up; they don't teleport.
|
|
45
|
+
|
|
46
|
+
### Path-based movement
|
|
47
|
+
|
|
48
|
+
```python
|
|
49
|
+
curve = bpy.data.curves.new("CAM_PATH", 'CURVE')
|
|
50
|
+
curve.dimensions = '3D'
|
|
51
|
+
spline = curve.splines.new('BEZIER')
|
|
52
|
+
spline.bezier_points.add(len(points) - 1)
|
|
53
|
+
for bp, co in zip(spline.bezier_points, points):
|
|
54
|
+
bp.co = co
|
|
55
|
+
bp.handle_left_type = bp.handle_right_type = 'AUTO'
|
|
56
|
+
|
|
57
|
+
path = bpy.data.objects.new("CAM_PATH", curve)
|
|
58
|
+
bpy.context.collection.objects.link(path)
|
|
59
|
+
|
|
60
|
+
con = cam.constraints.new('FOLLOW_PATH')
|
|
61
|
+
con.target = path
|
|
62
|
+
con.use_curve_follow = False # aiming is the Track To constraint's job
|
|
63
|
+
|
|
64
|
+
# Animate progress explicitly rather than relying on the default path animation.
|
|
65
|
+
curve.path_duration = 96
|
|
66
|
+
curve.eval_time = 0
|
|
67
|
+
curve.keyframe_insert("eval_time", frame=1)
|
|
68
|
+
curve.eval_time = 96
|
|
69
|
+
curve.keyframe_insert("eval_time", frame=96)
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
**Why a path and not raw location keys:** the user can drag a control point to retime or reshape the move without you regenerating anything. That is the difference between "re-prompt and hope" and "nudge it."
|
|
73
|
+
|
|
74
|
+
## Speed — the part that reads as production value
|
|
75
|
+
|
|
76
|
+
Movement at one constant speed is the tell of a machine. Real moves accelerate, hold, and snap.
|
|
77
|
+
|
|
78
|
+
Speed lives in the **f-curve handles** of `eval_time` (or of location, if you keyed it directly):
|
|
79
|
+
|
|
80
|
+
- **Long, near-horizontal handle** at a key → slow near that key.
|
|
81
|
+
- **Short, steep handle** → fast.
|
|
82
|
+
- `interpolation = 'CONSTANT'` → no movement at all until the next key. This is how you get an absolute dead stop.
|
|
83
|
+
|
|
84
|
+
```python
|
|
85
|
+
fc = curve.animation_data.action.fcurves.find("eval_time")
|
|
86
|
+
for kp in fc.keyframe_points:
|
|
87
|
+
kp.interpolation = 'BEZIER'
|
|
88
|
+
kp.handle_left_type = kp.handle_right_type = 'FREE'
|
|
89
|
+
|
|
90
|
+
a, b, c = fc.keyframe_points # start, middle, end
|
|
91
|
+
# Speed ramp: fast in, sag in the middle, accelerate out.
|
|
92
|
+
a.handle_right = (a.co.x + 2, a.co.y + 18) # steep = launches fast
|
|
93
|
+
b.handle_left = (b.co.x - 14, b.co.y) # flat = holds
|
|
94
|
+
b.handle_right = (b.co.x + 14, b.co.y)
|
|
95
|
+
c.handle_left = (c.co.x - 2, c.co.y - 18) # steep = arrives fast
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
**Zero drift at a stop.** If a move is supposed to be locked off, it must be *actually* locked — a slow crawl at a "stop" reads as a mistake. Hold with `CONSTANT` interpolation, or duplicate the key so the segment is genuinely flat.
|
|
99
|
+
|
|
100
|
+
## The moves
|
|
101
|
+
|
|
102
|
+
### Orbit
|
|
103
|
+
|
|
104
|
+
Circle the subject on a path, target at subject height. Vary radius and height across the move so it does not read as a turntable. Half-orbits and 3/4 orbits look more intentional than full ones.
|
|
105
|
+
|
|
106
|
+
For a multi-scene continuous orbit, keep one unbroken `eval_time` curve and move the *world* under it — the camera never cuts, the set changes.
|
|
107
|
+
|
|
108
|
+
### Floor rise
|
|
109
|
+
|
|
110
|
+
Pure vertical translation, **no rotation**, smooth acceleration with a slow middle. Each "floor" is a different set stacked on Z at a fixed interval; the subject sits centre-frame at each pass.
|
|
111
|
+
|
|
112
|
+
```python
|
|
113
|
+
FLOOR_H = 4.0
|
|
114
|
+
for i in range(4):
|
|
115
|
+
cam.location = (0, -6, i * FLOOR_H)
|
|
116
|
+
cam.keyframe_insert("location", frame=1 + i * 48)
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
Rotation during a rise destroys the effect. Leave it out.
|
|
120
|
+
|
|
121
|
+
### Robo-arm
|
|
122
|
+
|
|
123
|
+
The whip-and-lock commercial move: a fast flight along a curved arc, an **absolute** dead stop at a completely different angle, repeat. Each relocation is roughly a third of a second; each stop is a distinct, readable frame.
|
|
124
|
+
|
|
125
|
+
Build it as a path with a control point per stop, then make the stops real:
|
|
126
|
+
|
|
127
|
+
```python
|
|
128
|
+
HOLD_FRAMES = 10
|
|
129
|
+
for kp in fc.keyframe_points:
|
|
130
|
+
kp.interpolation = 'CONSTANT' # hold dead still between flights
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
Keep the target separate and slightly offset per stop, so each lock-off is a different composition of the same subject rather than six centred portraits.
|
|
134
|
+
|
|
135
|
+
### Handheld
|
|
136
|
+
|
|
137
|
+
Applied **last**, on top of a finished move. Slow organic sway, not jitter: long waves plus a barely-perceptible tremor.
|
|
138
|
+
|
|
139
|
+
```python
|
|
140
|
+
for path in ("location", "rotation_euler"):
|
|
141
|
+
for i in range(3):
|
|
142
|
+
fc = cam.animation_data.action.fcurves.find(path, index=i)
|
|
143
|
+
if fc is None:
|
|
144
|
+
continue
|
|
145
|
+
n = fc.modifiers.new('NOISE')
|
|
146
|
+
n.scale = 120 # large scale = long lazy waves (4-6s at 24fps)
|
|
147
|
+
n.strength = 0.035 # small; raise for rotation, lower for location
|
|
148
|
+
n.phase = i * 7.3 # decorrelate the axes or it reads as a slide
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
**Never fast jitter, wobble or snap corrections.** Wrong-flavour handheld is more damaging than none.
|
|
152
|
+
|
|
153
|
+
### Over-the-shoulder cuts
|
|
154
|
+
|
|
155
|
+
Per cut: a camera position below shoulder height, a near-foreground body mass, and a target on the far face. Move barely — a slow sideways crawl. See `slates-dialogue-blocking` for who may occupy the foreground and why it matters.
|
|
156
|
+
|
|
157
|
+
### Lens
|
|
158
|
+
|
|
159
|
+
Animate focal length like any other channel; a slow lens breath under a move adds a lot for nothing.
|
|
160
|
+
|
|
161
|
+
```python
|
|
162
|
+
cam.data.lens = 35
|
|
163
|
+
cam.data.keyframe_insert("lens", frame=1)
|
|
164
|
+
cam.data.lens = 50
|
|
165
|
+
cam.data.keyframe_insert("lens", frame=96)
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
**But the lens must not drift across a cut.** Within a cut it can animate; on the cut frame it changes instantly with everything else.
|
|
169
|
+
|
|
170
|
+
## Multiple cameras and cuts
|
|
171
|
+
|
|
172
|
+
Two ways, and only one of them survives contact with a 19-shot edit:
|
|
173
|
+
|
|
174
|
+
- **One camera, jump-cut it.** Keyframe location/rotation/lens with `CONSTANT` interpolation on the cut frames. Fine up to a handful of cuts.
|
|
175
|
+
- **A camera per setup, bound to timeline markers.** Correct for anything bigger, because each setup stays independently editable.
|
|
176
|
+
|
|
177
|
+
```python
|
|
178
|
+
marker = bpy.context.scene.timeline_markers.new("SHOT_04", frame=188)
|
|
179
|
+
marker.camera = cam_shot_04
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
Either way the invariant is the same: **camera, target and lens all change on the cut frame, and no frame between two setups is interpolated.** One transition frame reads as a whip-pan, and the model will reproduce it faithfully.
|
|
183
|
+
|
|
184
|
+
## Verify before you render
|
|
185
|
+
|
|
186
|
+
```
|
|
187
|
+
slates_blender_scene
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
`cutSeconds` is your actual cut list in seconds. Check it against the shot list you were given — mismatches here become mismatched prompt timings, and the prompt is what you write next.
|
|
191
|
+
|
|
192
|
+
⚠️ **On a marker-bound rig, read `cutSeconds` or `markers`, never `camera.keyframeSeconds`.** The per-setup cameras are usually static (a Track To constraint does the aiming), so the active camera's action is empty and that field reads `[]` on a perfectly good three-cut edit. `cutSeconds` resolves to whichever the scene actually used.
|
|
193
|
+
|
|
194
|
+
## Related
|
|
195
|
+
|
|
196
|
+
`slates-previs-blocking` (the workflow) · `slates-blocking-to-prompt` (turning these moves into prompt text) · `slates-dialogue-blocking` (OTS and eyelines)
|
|
@@ -0,0 +1,134 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-dialogue-blocking
|
|
3
|
+
description: Keep multiple characters spatially consistent across cuts — seating, screen direction, eyelines, the 180-degree rule — by blocking the scene in 3D first. Use for any multi-character dialogue scene, conversations around a table or in a car, or when generated characters swap seats, change sides, or look the wrong way between shots.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Dialogue blocking — six people who stay where you put them
|
|
7
|
+
|
|
8
|
+
The hardest thing to generate, and the case where previs beats raw prompting by the widest margin.
|
|
9
|
+
|
|
10
|
+
## Why this is hard
|
|
11
|
+
|
|
12
|
+
Every cut is an independent guess unless something forces agreement. Prompt a six-person conversation four times and you get four different seating charts: characters swap places, the 180-degree line breaks, and nothing cuts together. The failure is not aesthetic — the shots are simply unusable as an edit, and you find out only after paying for all four.
|
|
13
|
+
|
|
14
|
+
**The blocking fixes it structurally.** Positions exist in 3D, so every camera sees the same arrangement, and consistency stops being something the model has to remember.
|
|
15
|
+
|
|
16
|
+
## Build
|
|
17
|
+
|
|
18
|
+
Follow `slates-previs-blocking` and add these.
|
|
19
|
+
|
|
20
|
+
### Seated proxies, colour-coded
|
|
21
|
+
|
|
22
|
+
Simple seated shapes. **Do not animate the heads** — a proxy head turning the wrong way is worse than one that never turns.
|
|
23
|
+
|
|
24
|
+
Give each character a distinct viewport colour and write the mapping down. This is the identity channel:
|
|
25
|
+
|
|
26
|
+
> red = the boss · green = the kid · blue = the driver · yellow = the fixer · purple = the cousin · cyan = the nephew
|
|
27
|
+
|
|
28
|
+
That mapping goes verbatim into the generation prompt. It is what lets the model bind a grey body to a character sheet across four cuts.
|
|
29
|
+
|
|
30
|
+
Give every proxy a material and set `mat.diffuse_color` to its identity colour — the blocking render pins Workbench to `MATERIAL` shading, so **the material's `diffuse_color` is what reaches the clip**. Set `object.color` to the same value too, so the user's viewport matches what renders. See `slates-previs-blocking` for the snippet.
|
|
31
|
+
|
|
32
|
+
### Fix the geography, then never move it
|
|
33
|
+
|
|
34
|
+
Place people once. Write down who sits where relative to the camera's opening position, in words, because that sentence is going into the prompt:
|
|
35
|
+
|
|
36
|
+
> Across the table, facing camera: yellow dead centre, purple far left, blue and green to the right.
|
|
37
|
+
|
|
38
|
+
### The camera plan
|
|
39
|
+
|
|
40
|
+
Per `slates-camera-language`, with two things specific to dialogue:
|
|
41
|
+
|
|
42
|
+
- **Below shoulder height, slow rail glides.** Eye-level-and-above reads as surveillance.
|
|
43
|
+
- **Decide who owns the near foreground in each cut and honour it.** A shoulder in frame is a spatial anchor; a different shoulder in the next cut relocates the whole room.
|
|
44
|
+
|
|
45
|
+
The move that earns the most: **a gaze handoff without a cut** — the camera keeps gliding while the target hands off across the table, face to face, slowing on each but never stopping. Build it by keyframing the Track To target's position between subjects.
|
|
46
|
+
|
|
47
|
+
### Crossing behind someone
|
|
48
|
+
|
|
49
|
+
A head wiping frame during a move is a strong depth cue. It is also a spatial claim, so pick who gets crossed and say so — *the camera crosses directly behind cyan's back mid-shot and his head wipes the frame once.*
|
|
50
|
+
|
|
51
|
+
## The prompt
|
|
52
|
+
|
|
53
|
+
Everything in `slates-blocking-to-prompt`, plus these blocks.
|
|
54
|
+
|
|
55
|
+
### Geography — restate it as a rule
|
|
56
|
+
|
|
57
|
+
> TABLE GEOGRAPHY — do not deviate: the camera is never parked behind red. Only in the opening seconds does his dark shoulder hang at the near frame RIGHT edge, and it slides out as the camera travels LEFT. The true near-foreground of this shot is cyan: the camera crosses directly behind him mid-shot. After the opening seconds red is gone from the foreground, and the camera never travels behind anyone except cyan.
|
|
58
|
+
|
|
59
|
+
### Screen direction, and the mirror that is not a swap
|
|
60
|
+
|
|
61
|
+
The 180-degree rule survives on its own in the blocking. What breaks is the model **"correcting" a legitimate mirror** — when the camera faces back through a scene, sides invert, and that inversion is correct. Say so explicitly or it gets flipped:
|
|
62
|
+
|
|
63
|
+
> The driver's seat is on the LEFT for the entire timeline; this layout never mirrors or flips. When a camera faces BACKWARD into the car, screen sides mirror naturally: the driver reads on the RIGHT of frame, the passenger on the LEFT — that is correct left-hand drive, not a swap. They never swap seats or roles anywhere in the timeline.
|
|
64
|
+
|
|
65
|
+
Then compress it into the HOLD block: *(backward camera mirrors them: he right of frame, she left)*.
|
|
66
|
+
|
|
67
|
+
### Presence
|
|
68
|
+
|
|
69
|
+
> A seated person stays drawn even when partially occluded — in every interior frame some part of each seat's owner is visible: a hand, an arm, a shoulder, a head above the bolster. Every occupied seat visibly holds its person.
|
|
70
|
+
|
|
71
|
+
### Keep everyone alive
|
|
72
|
+
|
|
73
|
+
Three orthogonal layers. Without them, whoever is not speaking freezes:
|
|
74
|
+
|
|
75
|
+
- **ONGOING BUSINESS** — a small continuous physical action per character, running whether or not they are speaking. *turns his glass a quarter every few seconds · thumbs a lighter without lighting it.*
|
|
76
|
+
- **BACKGROUND LIFE** — soft-focus, low contrast, never pulls attention, never crosses in front of a speaking face.
|
|
77
|
+
- **SCENE EVENT** — the unnamed thing everyone is playing but nobody says. One line, repeated verbatim in every character's direction: *keep tomorrow sounding like a fishing trip.*
|
|
78
|
+
|
|
79
|
+
### Acting, per character
|
|
80
|
+
|
|
81
|
+
Same six slots each. Terse:
|
|
82
|
+
|
|
83
|
+
```
|
|
84
|
+
ACTING TASK — <character>
|
|
85
|
+
SCENE DIRECTION (shared, unspoken): <the same line for everyone>
|
|
86
|
+
MOTIVE (his fuel): <what he wants underneath>
|
|
87
|
+
GOAL: <what he wants in this scene>
|
|
88
|
+
OBSTACLE: <what is in the way>
|
|
89
|
+
TACTIC: <how he goes about it>
|
|
90
|
+
Moment to moment: <2-3 beats keyed to timestamps>
|
|
91
|
+
(Safety: gaze always engaged in the task — never a frozen, glassy,
|
|
92
|
+
unfocused stare; natural blink cadence.)
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
That safety line is not filler. Dead eyes are the characteristic failure of generated faces in dialogue, and naming it is what prevents it.
|
|
96
|
+
|
|
97
|
+
### Split any strong emotion into phases
|
|
98
|
+
|
|
99
|
+
The other characteristic failure is a face that strikes one extreme expression and holds it for the whole shot — a mouth stuck open for three seconds. Give the beat two phases with a hinge, and name the failure you are excluding:
|
|
100
|
+
|
|
101
|
+
> PHASE 1 (17.3-18.6s) — she SCREAMS at him, mouth wide, eyes huge, hand clamped on the grab handle. PHASE 2 (18.6-19.9s) — the scream breaks off: she shuts her eyes tight and CLOSES her mouth, both hands now on the handle, head ducked, braced. Scream, then brace — never one frozen open mouth held through the whole shot.
|
|
102
|
+
|
|
103
|
+
The hinge timestamp is what makes it a performance instead of a pose.
|
|
104
|
+
|
|
105
|
+
### Dialogue must not restructure the edit
|
|
106
|
+
|
|
107
|
+
Both of these, verbatim, every time:
|
|
108
|
+
|
|
109
|
+
> DIALOGUE NEVER CREATES SHOTS: spoken lines happen inside the reference's takes exactly as blocked — no cutaways to a speaker, no reverse shots, no added close-ups. If a line plays while the camera is elsewhere, the line stays off-screen audio.
|
|
110
|
+
|
|
111
|
+
> OFF-SCREEN VOICES RULE: a line marked off-screen must STAY off-screen — never show the speaker, never move him into frame, never route the camera behind him because he spoke.
|
|
112
|
+
|
|
113
|
+
A sentence may cross a cut. Say so where it does: *the sentence does not pause for the edit.*
|
|
114
|
+
|
|
115
|
+
## Model routing
|
|
116
|
+
|
|
117
|
+
Dialogue directed as separate layers (voices, scene sound, score) is **minimax-h3**'s seat; it also takes declared reference relationships, which suits a colour-coded cast. Native synced audio is **Veo**'s niche. seedance-2.5 carries the reference-video capacity. Route per `slates-model-selection` and read the chosen model's prompting skill before writing the audio block.
|
|
118
|
+
|
|
119
|
+
## Checklist
|
|
120
|
+
|
|
121
|
+
- [ ] Colour→character mapping written down and pasted into the prompt
|
|
122
|
+
- [ ] Heads not animated in the blocking
|
|
123
|
+
- [ ] Seating stated as a geography rule
|
|
124
|
+
- [ ] Foreground owner named per cut
|
|
125
|
+
- [ ] Mirror-is-not-a-swap clause present if any camera faces back through the scene
|
|
126
|
+
- [ ] Presence rule present
|
|
127
|
+
- [ ] Ongoing business, background life and scene event all specified
|
|
128
|
+
- [ ] Acting task per character, safety line included
|
|
129
|
+
- [ ] Any strong emotion split into phases with a hinge timestamp
|
|
130
|
+
- [ ] Dialogue-never-creates-shots and off-screen-voices rules present
|
|
131
|
+
|
|
132
|
+
## Related
|
|
133
|
+
|
|
134
|
+
`slates-previs-blocking` · `slates-camera-language` · `slates-blocking-to-prompt` · `slates-character-identity` (the sheets) · `slates-prompting-minimax-h3`
|