@slatesvideo/shared 0.6.0 → 0.6.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -28,7 +28,7 @@
28
28
  },
29
29
  {
30
30
  "path": "src/prompts/model-facts.ts",
31
- "sha256": "98f77646a75ae2078c5015638c416d1eb84de4fe057079f705439f1a16e30631"
31
+ "sha256": "89a899444c021975df0a03a87b446217cdfc3b9086ef501e9e847936d7cfec2d"
32
32
  }
33
33
  ],
34
34
  "outputs": [
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@slatesvideo/shared",
3
- "version": "0.6.0",
3
+ "version": "0.6.2",
4
4
  "description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -0,0 +1,250 @@
1
+ ---
2
+ name: slates-blocking-to-prompt
3
+ description: Write the generation prompt that matches a blocking clip second by second, so the reference video and the text agree instead of fighting. Use after rendering a previs blocking pass, when a generated shot ignores the reference video, when timings drift, or when the model invents shots and camera angles that are not in the blocking.
4
+ ---
5
+
6
+ # Blocking → prompt
7
+
8
+ You have a blocking clip. This is how you write the prompt that goes with it.
9
+
10
+ ## The one idea
11
+
12
+ The clip already contains the camera, the cuts and the timing. **The prompt's job is to say what everything looks like — and, where the grey boxes are ambiguous, to disambiguate them.** It is not a second, competing description of the motion.
13
+
14
+ State that contract inside the prompt, because the model needs it as much as you do:
15
+
16
+ > Where a timeline line below names a camera position or move, it is a restatement of what the reference video already does at that timestamp — a disambiguation, never a new instruction.
17
+
18
+ And give it a tie-break, because ambiguity is guaranteed:
19
+
20
+ > If any text in this prompt appears to disagree with the reference video about camera, framing, direction, motion, timing or object placement, the reference video wins.
21
+
22
+ Those two sentences do more work than any other part of the prompt.
23
+
24
+ ## Get the real numbers first
25
+
26
+ ```
27
+ slates_blender_scene
28
+ ```
29
+
30
+ Write timings from `cutSeconds`, never from the shot list you intended to build. It resolves to the marker frames on a multi-camera edit and to the camera's own keyframes otherwise, so it is the one field that is never empty on a rig that has cuts. At 24fps cuts land on frame boundaries and the honest values are not round — `7.79s`, `9.33s`, `19.875s`. **Use the exact ones.** Rounding to `7.8` is a tenth of drift you are handing the model for free.
31
+
32
+ ## Structure
33
+
34
+ Order matters — contract, then globals, then timeline, then the re-assertion.
35
+
36
+ ```
37
+ TITLE — one line: what this is, how long, that it is video-to-video
38
+
39
+ LOGLINE — 2-4 sentences. The whole piece in plain language.
40
+
41
+ ACTIVE REFERENCES
42
+ <one entry per reference: what it defines, and what is NOT inherited>
43
+
44
+ TECHNICAL BLOCK (format, grade, lens, and the blanket negatives — see below)
45
+
46
+ STYLE / LOOK
47
+ LIGHTING
48
+ COLOR
49
+ CAMERA
50
+ PHYSICS (only if things move, collide or deform)
51
+
52
+ RULES (numbered — the invariants, see below)
53
+
54
+ ACTION TIMING (beat by beat, against the clip's real timestamps)
55
+
56
+ AUDIO
57
+ [Sound design] [Timed accents] [Dialogue] [Music]
58
+
59
+ ENDING LOCK (one line — where the film stops)
60
+
61
+ HOLD FOR THE FULL TIMELINE
62
+ <the 5-6 constraints most likely to drift, compressed>
63
+ ```
64
+
65
+ ## ACTIVE REFERENCES — every entry has an exclusion
66
+
67
+ The single highest-leverage format in this whole workflow. Each reference is a **positive claim plus an exclusion list**, because a reference the model over-reads is as damaging as one it ignores.
68
+
69
+ Label each by the badge code Slates echoes back (`IMG-A8`, `VID-C2`) or by an unmistakable role name, and use that same label everywhere below.
70
+
71
+ **The blocking clip:**
72
+
73
+ > VID-C2 = the blocking previz (30s, 720 frames, 24fps) — the MASTER for everything that moves and everything that stands. It defines the full edit one-to-one: every cut point, every camera position, angle, move and framing, all action timing, screen direction, and the geometry of the world. Its untextured grey surfaces, flat colours and viewport grid are NOT inherited — every grey proxy is dressed into a real object in the exact position the previz puts it. Proxies give position, angle, scale and motion only; never surface, shape detail or design.
74
+
75
+ That last sentence is the **placement-only clause** and it is not optional. Without it the model renders grey boxes.
76
+
77
+ **A character sheet:**
78
+
79
+ > IMG-A8 = the driver — defines his face, hair and wardrobe. Identity 100% consistent at every distance and through every motion blur. Background, lighting and pose are NOT inherited.
80
+
81
+ **A location/style reference:**
82
+
83
+ > IMG-B3 = the tunnel — defines location geometry, look and grade. Camera angle, framing and any people in it are NOT inherited; the camera comes exclusively from VID-C2.
84
+
85
+ **An atmosphere or style master — a reference that is never a shot:**
86
+
87
+ > IMG-D9 = ATMOSPHERE MASTER — NOT a keyframe, NOT a location to reproduce, NOT a frame that ever appears in the film: its own subject, framing and composition are never seen in any shot. It defines ONLY the weather, light, colour and grade: deep clean night just after rain, wet asphalt as a dark mirror, cool white-cyan lamps as the ambient key, teal-and-amber grade, deep clean blacks. Every shot is lit and graded in this regime for all 30 seconds.
88
+
89
+ Without those three NOTs the model reproduces the reference's composition as an actual shot — you get its street corner in your film. The same wording covers a rendering-style master; see `slates-restyle-from-blocking`.
90
+
91
+ **References can be scheduled.** If something is only true for part of the timeline, say so: *the hooded panel applies only to 0–7.0s and 27.5–30s*, or per-reference: *Active for 00:03.3–00:06.7 only.* On a piece that travels through several locations, every location still carries its own window and the model stops blending two sets into one shot.
92
+
93
+ ## Translate the blocking's artifacts
94
+
95
+ Your grey-box render contains things that are *notation*, not content. Every one needs an explicit reinterpretation or it gets rendered literally:
96
+
97
+ | In the blocking | Say in the prompt |
98
+ |---|---|
99
+ | Colour-coded bodies | `red = the boss, green = the kid, blue = the driver` |
100
+ | A marked face on a proxy | `RED face = the direction he faces, BLACK = his back` |
101
+ | A checkered floor or wall | `the checkerboard is a scale reference, not a surface — it becomes <the real material>` |
102
+ | Flat black background | `a PLACEHOLDER — replace with the location assigned below` |
103
+ | A floor grid | `a motion-tracking aid — render as light on the surface, never as wireframe or tiles` |
104
+ | A deliberate black gap | `CUT 7 (14.5-17.0, black gap in the reference) — <what fills it>` |
105
+ | Frame goes dark mid-move | `the camera is passing through the ground — a doorway to the NEXT location, never back to a previous one` |
106
+ | The source hard-resets mid-move | `each reset begins a NEW, completely different room — never a replay of one already seen` |
107
+ | A proxy that is a PROP or VEHICLE | `the low-poly flying model in SHOT 18 is THE HELICOPTER · blocks on the rear bench are the luggage · the small dark block in his hand IS the pistol` |
108
+ | A blocky proxy limb in a tight insert | `the blocky low-poly leg is a stand-in and must NOT be replicated — generate complete human anatomy: a real boot, a real trouser leg, a correct ankle at this exact camera angle` |
109
+ | A stray object at the frame edge | `ignore it completely — never blend two sets into one shot` |
110
+
111
+ ## TECHNICAL BLOCK — format, lens, and the blanket negatives
112
+
113
+ One paragraph, before the timeline. It carries the things that are true of every frame and that no beat should have to repeat:
114
+
115
+ > Cinematic, photoreal. 21:9. 30s. SFX only, no music. Kodak 500T film look, natural 35mm grain, organic colour, soft highlight roll-off, anamorphic lens character with oval bokeh and gentle barrel distortion at the edges, chromatic aberration creeping in at the frame edges, natural motion blur on every fast move, faint bloom on hot speculars. Every location well exposed — night interiors bright and readable, open shadows, no crushed blacks, no murk. NO CGI. NON-IP, no brand badges or logos anywhere, no text, no watermark.
116
+
117
+ Three parts worth naming:
118
+
119
+ - **Lens realism is a list, not an adjective.** Aberration, motion blur, depth of field, barrel distortion, bloom, grain. "Cinematic" buys you nothing; these buy you the look.
120
+ - **Exposure needs saying on dark work.** Models crush night scenes into murk. *Night interiors bright and readable, open shadows, no crushed blacks* is what keeps a scene legible.
121
+ - **The blanket negatives go here once** — `NON-IP`, no logos, no on-screen text, no subtitles, no watermark — rather than being scattered through the beats.
122
+
123
+ ## RULES — the invariants
124
+
125
+ Numbered, short, absolute. These are the things that must hold in every frame, and they are where you put anything that has already gone wrong once.
126
+
127
+ Two patterns worth stealing outright:
128
+
129
+ **Countable state.** Give the model arithmetic it can check itself against:
130
+
131
+ > At every second: standing + fallen + on the lintel = 6. Never a seventh figure — no extras, no duplicates, no distant silhouettes, no half-bodies at frame edges.
132
+
133
+ > Bodies on the ground count exactly: 0 before 12s → 1 → 2 → 3 → 4 at 12/13/14/16s → 5 at 20s → 6 at 26s. Never more.
134
+
135
+ **Every mass is dressed, and nothing is invented.** The blocking is authority over what EXISTS, not just what moves — otherwise the model deletes the masses it finds boring and adds architecture you never blocked:
136
+
137
+ > Every lamppost, guardrail, road and terrain mass visible in the reference exists in the output in the same place, at the same scale, in the same position in frame — the opening blocks are dark-brick warehouse facades, the roadside masses are the waterfront skyline, the finale rocks are the city's tower walls. Nothing is deleted, and no structure is invented where the reference shows none.
138
+
139
+ **Anatomy is never inherited from a proxy.** Tight inserts on hands and feet are where blocking leaks straight into the render:
140
+
141
+ > The reference shows only WHERE hands and feet are. In the output they are always complete human anatomy — a five-fingered gloved hand with natural knuckles, a real leg in wool trousers, a real foot in a leather shoe — never the proxy's blocky shape.
142
+
143
+ **A ledger for anything that happens a countable number of times.** A ritual, a reload, a set of falls: state it as a linear sequence, each step exactly once, and close with the tally:
144
+
145
+ > Strictly linear, six steps in fixed order, each happening EXACTLY ONCE and never repeating; once a step is done it is done for good, and the sequence only ever moves FORWARD, never backward. Count of weapon events in the entire video: one draw, one magazine insertion, one slide rack, one shot.
146
+
147
+ Without the ledger the model loops the most cinematic beat — it will rack the slide four times because racking looks good.
148
+
149
+ **Named misreads.** When a generation gets something specifically wrong, do not rewrite the description — **name the wrong reading and kill it**:
150
+
151
+ > The lamp is a man-made steel structure — NOT an animal, NOT a snake, NOT any living or organic shape.
152
+
153
+ > The rear of the car and its tail lights are NOT visible in this shot.
154
+
155
+ This is the highest-value edit available after a failed roll, and it is why the prompt grows rather than changes between takes.
156
+
157
+ ## ACTION TIMING — the beats
158
+
159
+ One block per shot or beat. Two notations; pick one and hold it.
160
+
161
+ **For a continuous take**, ranges with a camera note and a closing state audit:
162
+
163
+ ```
164
+ 8-12s — THE SWEEP (per VID-C2: elevated rear push, swinging to profile by 12s):
165
+ <what happens, in prose, with sub-beats on tenths and → chaining cause to effect>
166
+ END 12s: bodies 1 (behind him as he steps past) · standing — four ahead, holding.
167
+ ```
168
+
169
+ **For a cut edit**, numbered shots ending on their cut:
170
+
171
+ ```
172
+ 9.33-10.33s — SHOT 10 — Interior over the centre console as in VID-C2: <what the
173
+ frame contains>. Hard cut at 10.33s.
174
+ ```
175
+
176
+ Three habits that separate a beat that works from one that does not:
177
+
178
+ - **Declare the frame's contents as a closed set** when the shot is tight: *the frame holds exactly the console, the lever, his hand, and the edges of both seats.* An open description invites additions.
179
+ - **Chain cause to effect inside one sentence** with `→`. `he overcommits a lunge → the Hero drops low and sweeps his standing leg → he hits the earth at 12s`.
180
+ - **Put events on tenths.** `11.7s`, `19.5s`, `22.5s`. Vague beats generate vague timing.
181
+
182
+ Density: roughly 60–130 words per second of screen time is what these prompts actually run at. That is much denser than a normal video prompt, and it is the point.
183
+
184
+ ## AUDIO
185
+
186
+ `[Timed accents]` uses the same timestamps as the beats:
187
+
188
+ > 3.1s tyres light up into the burnout squeal · 7.0s drift-entry screech · 10.1s hard mechanical shifter clack · 20.3s full-speed pass-by whoosh
189
+
190
+ `[Dialogue]` is a closed list — count the lines, give each a window, quote it verbatim, and forbid everything else:
191
+
192
+ > Exactly TWO vocal events in the entire 30 seconds, both screamed, in English, VERBATIM: 1. 17.3-18.6s "STOOOOOP!!" 2. 23.4-24.0s "You crazy!" Nothing else is ever spoken.
193
+
194
+ **Then forbid the lines it will invent anyway.** A closed list is a rule; an enumerated blacklist is enforcement, and the phrases to list are the clichés the scene invites:
195
+
196
+ > FORBIDDEN — she never says any of these and no one else says anything: "they're behind us", "cops", "go go go", "are you crazy", "you're insane", or ANY other invented phrase. All other human voice is wordless screaming or laughing.
197
+
198
+ **When lines are lip-synced, give each one a timestamp** in the same list, and say that they change nothing else:
199
+
200
+ > Timing: "You wind up for this one?" ~18.2s · "Three full turns." ~19.0s · "Company." ~24.3s. Every line lip-synced; the lines never change the camera.
201
+
202
+ Two rules that stop dialogue from breaking the edit:
203
+
204
+ > DIALOGUE NEVER CREATES SHOTS: spoken lines happen inside the reference's takes exactly as blocked — no cutaways to a speaker, no reverse shots, no added close-ups. If a line plays while the camera is elsewhere, the line stays off-screen audio.
205
+
206
+ > A line marked off-screen must STAY off-screen — never show the speaker, never move him into frame, never route the camera to him because he spoke.
207
+
208
+ Model note: dialogue direction as separate layers is minimax-h3's seat; native synced audio is Veo's. Route per `slates-model-selection` and read the model's own prompting skill before writing the audio block.
209
+
210
+ ## ENDING LOCK
211
+
212
+ One line, and it is the cheapest fix in the document. Models drift at the end — they hold a frame too long, add a beat after the last one, or fade somewhere the reference does not:
213
+
214
+ > The film ends exactly where the reference ends: the final take runs unbroken to its last frame, and that source frame IS the final frame of the film. Nothing follows. The last five seconds follow the source exactly as strictly as the first five.
215
+
216
+ ## HOLD FOR THE FULL TIMELINE
217
+
218
+ Close with a terminal re-assertion of only the constraints most prone to drift — five or six lines, compressed, no new information:
219
+
220
+ ```
221
+ HOLD FOR THE FULL TIMELINE
222
+ - VID-C2 camera path 1:1 — any deviation = failure.
223
+ - Six and only six figures; the count above holds at every second.
224
+ - IMG-A8 identity constant at every distance and through motion blur.
225
+ - The IMG-B3 location in every frame; no subtitles, no watermarks.
226
+ ```
227
+
228
+ Restating is not redundancy here. It is the last thing the model reads.
229
+
230
+ ## Checklist before you generate
231
+
232
+ - [ ] Timings taken from `slates_blender_scene`'s `cutSeconds`, frame-exact, not rounded
233
+ - [ ] Every reference has an explicit "NOT inherited"
234
+ - [ ] The placement-only clause is present
235
+ - [ ] The tie-break clause is present
236
+ - [ ] The disambiguation clause is present
237
+ - [ ] Every blocking artifact is translated — colours, marked faces, checkers, grid, black background, gaps, dark dips, resets, prop proxies, proxy limbs, strays
238
+ - [ ] Any style/weather reference is declared NOT a keyframe and never a shot
239
+ - [ ] The every-mass-is-dressed / nothing-invented rule is present
240
+ - [ ] Tight inserts on hands or feet demand complete anatomy
241
+ - [ ] Counts are stated where anything is countable, and repeatable actions carry a ledger
242
+ - [ ] A TECHNICAL BLOCK carries format, lens realism, exposure and the blanket negatives
243
+ - [ ] `[Dialogue]` is a closed list with a FORBIDDEN blacklist
244
+ - [ ] An ENDING LOCK says where the film stops
245
+ - [ ] `videoReferenceSecondsEach` matches the clip's real duration
246
+ - [ ] A HOLD block closes it
247
+
248
+ ## Related
249
+
250
+ `slates-previs-blocking` (producing the clip) · `slates-camera-language` (the moves being described) · `slates-dialogue-blocking` (multi-character continuity) · `slates-restyle-from-blocking` (reusing this prompt across styles) · `slates-prompting-seedance-2-5` / `slates-prompting-minimax-h3` (model-specific rules)
@@ -0,0 +1,196 @@
1
+ ---
2
+ name: slates-camera-language
3
+ description: Turn director vocabulary into real Blender camera rigs — orbits, floor rises, robo-arm whips, handheld, speed ramps, over-the-shoulder cuts — as bpy code. Use when building or refining the camera on a previs blocking pass, when a move needs to accelerate/hold/snap, or when someone asks for a "cinematic" camera and you need to convert that into an actual shot list.
4
+ ---
5
+
6
+ # Camera language — from a shot list to a rig
7
+
8
+ Companion to `slates-previs-blocking`. That skill owns the workflow; this one owns the camera.
9
+
10
+ ## The first rule
11
+
12
+ 🚨 **Never build "a cinematic camera move." Brief the camera the way you would brief an operator:** rails, target, height, lens, and the frame each move starts and ends on. "Cinematic" is not a specification, and asking for one produces the drifting slop the whole blocking workflow exists to avoid.
13
+
14
+ Bad: *a dynamic cinematic orbit around the subject.*
15
+ Good: *a 3/4 orbit starting rear-left at 1.6m, ending front-right at 0.9m, frames 1–96, 35mm, easing out of the start and holding hard on the last 8 frames.*
16
+
17
+ If the user gives you the first, convert it to the second and say what you assumed.
18
+
19
+ ## Look it up, don't recall it
20
+
21
+ Before any constraint or operator you are not certain of, call `slates_blender_docs` (e.g. `bpy.types.FollowPathConstraint`) or `slates_blender_search_docs`. Invented enum values are the most common failure here and they often fail *quietly* — the constraint gets added, the axis is wrong, and the camera points at nothing.
22
+
23
+ ## The two primitives everything is built from
24
+
25
+ ### Target-based aiming
26
+
27
+ Almost every move in this skill is *position on a path* plus *aim at a target*. Separating them is what lets framing vary while the subject stays in frame.
28
+
29
+ ```python
30
+ import bpy
31
+
32
+ target = bpy.data.objects.new("CAM_TARGET", None) # an Empty
33
+ target.empty_display_type = 'PLAIN_AXES'
34
+ bpy.context.collection.objects.link(target)
35
+
36
+ con = cam.constraints.new('TRACK_TO')
37
+ con.target = target
38
+ con.track_axis = 'TRACK_NEGATIVE_Z' # cameras look down -Z
39
+ con.up_axis = 'UP_Y'
40
+ ```
41
+
42
+ **Aim at a separate target, not at the subject's head.** A camera locked to the head produces dead, centred framing. Offset the target beside or ahead of the subject and the shot breathes — that small offset is most of what reads as "real operator."
43
+
44
+ For a subject that should notice the camera, keyframe the target's follow with a **few frames of lag** behind each camera move. Heads catch up; they don't teleport.
45
+
46
+ ### Path-based movement
47
+
48
+ ```python
49
+ curve = bpy.data.curves.new("CAM_PATH", 'CURVE')
50
+ curve.dimensions = '3D'
51
+ spline = curve.splines.new('BEZIER')
52
+ spline.bezier_points.add(len(points) - 1)
53
+ for bp, co in zip(spline.bezier_points, points):
54
+ bp.co = co
55
+ bp.handle_left_type = bp.handle_right_type = 'AUTO'
56
+
57
+ path = bpy.data.objects.new("CAM_PATH", curve)
58
+ bpy.context.collection.objects.link(path)
59
+
60
+ con = cam.constraints.new('FOLLOW_PATH')
61
+ con.target = path
62
+ con.use_curve_follow = False # aiming is the Track To constraint's job
63
+
64
+ # Animate progress explicitly rather than relying on the default path animation.
65
+ curve.path_duration = 96
66
+ curve.eval_time = 0
67
+ curve.keyframe_insert("eval_time", frame=1)
68
+ curve.eval_time = 96
69
+ curve.keyframe_insert("eval_time", frame=96)
70
+ ```
71
+
72
+ **Why a path and not raw location keys:** the user can drag a control point to retime or reshape the move without you regenerating anything. That is the difference between "re-prompt and hope" and "nudge it."
73
+
74
+ ## Speed — the part that reads as production value
75
+
76
+ Movement at one constant speed is the tell of a machine. Real moves accelerate, hold, and snap.
77
+
78
+ Speed lives in the **f-curve handles** of `eval_time` (or of location, if you keyed it directly):
79
+
80
+ - **Long, near-horizontal handle** at a key → slow near that key.
81
+ - **Short, steep handle** → fast.
82
+ - `interpolation = 'CONSTANT'` → no movement at all until the next key. This is how you get an absolute dead stop.
83
+
84
+ ```python
85
+ fc = curve.animation_data.action.fcurves.find("eval_time")
86
+ for kp in fc.keyframe_points:
87
+ kp.interpolation = 'BEZIER'
88
+ kp.handle_left_type = kp.handle_right_type = 'FREE'
89
+
90
+ a, b, c = fc.keyframe_points # start, middle, end
91
+ # Speed ramp: fast in, sag in the middle, accelerate out.
92
+ a.handle_right = (a.co.x + 2, a.co.y + 18) # steep = launches fast
93
+ b.handle_left = (b.co.x - 14, b.co.y) # flat = holds
94
+ b.handle_right = (b.co.x + 14, b.co.y)
95
+ c.handle_left = (c.co.x - 2, c.co.y - 18) # steep = arrives fast
96
+ ```
97
+
98
+ **Zero drift at a stop.** If a move is supposed to be locked off, it must be *actually* locked — a slow crawl at a "stop" reads as a mistake. Hold with `CONSTANT` interpolation, or duplicate the key so the segment is genuinely flat.
99
+
100
+ ## The moves
101
+
102
+ ### Orbit
103
+
104
+ Circle the subject on a path, target at subject height. Vary radius and height across the move so it does not read as a turntable. Half-orbits and 3/4 orbits look more intentional than full ones.
105
+
106
+ For a multi-scene continuous orbit, keep one unbroken `eval_time` curve and move the *world* under it — the camera never cuts, the set changes.
107
+
108
+ ### Floor rise
109
+
110
+ Pure vertical translation, **no rotation**, smooth acceleration with a slow middle. Each "floor" is a different set stacked on Z at a fixed interval; the subject sits centre-frame at each pass.
111
+
112
+ ```python
113
+ FLOOR_H = 4.0
114
+ for i in range(4):
115
+ cam.location = (0, -6, i * FLOOR_H)
116
+ cam.keyframe_insert("location", frame=1 + i * 48)
117
+ ```
118
+
119
+ Rotation during a rise destroys the effect. Leave it out.
120
+
121
+ ### Robo-arm
122
+
123
+ The whip-and-lock commercial move: a fast flight along a curved arc, an **absolute** dead stop at a completely different angle, repeat. Each relocation is roughly a third of a second; each stop is a distinct, readable frame.
124
+
125
+ Build it as a path with a control point per stop, then make the stops real:
126
+
127
+ ```python
128
+ HOLD_FRAMES = 10
129
+ for kp in fc.keyframe_points:
130
+ kp.interpolation = 'CONSTANT' # hold dead still between flights
131
+ ```
132
+
133
+ Keep the target separate and slightly offset per stop, so each lock-off is a different composition of the same subject rather than six centred portraits.
134
+
135
+ ### Handheld
136
+
137
+ Applied **last**, on top of a finished move. Slow organic sway, not jitter: long waves plus a barely-perceptible tremor.
138
+
139
+ ```python
140
+ for path in ("location", "rotation_euler"):
141
+ for i in range(3):
142
+ fc = cam.animation_data.action.fcurves.find(path, index=i)
143
+ if fc is None:
144
+ continue
145
+ n = fc.modifiers.new('NOISE')
146
+ n.scale = 120 # large scale = long lazy waves (4-6s at 24fps)
147
+ n.strength = 0.035 # small; raise for rotation, lower for location
148
+ n.phase = i * 7.3 # decorrelate the axes or it reads as a slide
149
+ ```
150
+
151
+ **Never fast jitter, wobble or snap corrections.** Wrong-flavour handheld is more damaging than none.
152
+
153
+ ### Over-the-shoulder cuts
154
+
155
+ Per cut: a camera position below shoulder height, a near-foreground body mass, and a target on the far face. Move barely — a slow sideways crawl. See `slates-dialogue-blocking` for who may occupy the foreground and why it matters.
156
+
157
+ ### Lens
158
+
159
+ Animate focal length like any other channel; a slow lens breath under a move adds a lot for nothing.
160
+
161
+ ```python
162
+ cam.data.lens = 35
163
+ cam.data.keyframe_insert("lens", frame=1)
164
+ cam.data.lens = 50
165
+ cam.data.keyframe_insert("lens", frame=96)
166
+ ```
167
+
168
+ **But the lens must not drift across a cut.** Within a cut it can animate; on the cut frame it changes instantly with everything else.
169
+
170
+ ## Multiple cameras and cuts
171
+
172
+ Two ways, and only one of them survives contact with a 19-shot edit:
173
+
174
+ - **One camera, jump-cut it.** Keyframe location/rotation/lens with `CONSTANT` interpolation on the cut frames. Fine up to a handful of cuts.
175
+ - **A camera per setup, bound to timeline markers.** Correct for anything bigger, because each setup stays independently editable.
176
+
177
+ ```python
178
+ marker = bpy.context.scene.timeline_markers.new("SHOT_04", frame=188)
179
+ marker.camera = cam_shot_04
180
+ ```
181
+
182
+ Either way the invariant is the same: **camera, target and lens all change on the cut frame, and no frame between two setups is interpolated.** One transition frame reads as a whip-pan, and the model will reproduce it faithfully.
183
+
184
+ ## Verify before you render
185
+
186
+ ```
187
+ slates_blender_scene
188
+ ```
189
+
190
+ `cutSeconds` is your actual cut list in seconds. Check it against the shot list you were given — mismatches here become mismatched prompt timings, and the prompt is what you write next.
191
+
192
+ ⚠️ **On a marker-bound rig, read `cutSeconds` or `markers`, never `camera.keyframeSeconds`.** The per-setup cameras are usually static (a Track To constraint does the aiming), so the active camera's action is empty and that field reads `[]` on a perfectly good three-cut edit. `cutSeconds` resolves to whichever the scene actually used.
193
+
194
+ ## Related
195
+
196
+ `slates-previs-blocking` (the workflow) · `slates-blocking-to-prompt` (turning these moves into prompt text) · `slates-dialogue-blocking` (OTS and eyelines)
@@ -0,0 +1,134 @@
1
+ ---
2
+ name: slates-dialogue-blocking
3
+ description: Keep multiple characters spatially consistent across cuts — seating, screen direction, eyelines, the 180-degree rule — by blocking the scene in 3D first. Use for any multi-character dialogue scene, conversations around a table or in a car, or when generated characters swap seats, change sides, or look the wrong way between shots.
4
+ ---
5
+
6
+ # Dialogue blocking — six people who stay where you put them
7
+
8
+ The hardest thing to generate, and the case where previs beats raw prompting by the widest margin.
9
+
10
+ ## Why this is hard
11
+
12
+ Every cut is an independent guess unless something forces agreement. Prompt a six-person conversation four times and you get four different seating charts: characters swap places, the 180-degree line breaks, and nothing cuts together. The failure is not aesthetic — the shots are simply unusable as an edit, and you find out only after paying for all four.
13
+
14
+ **The blocking fixes it structurally.** Positions exist in 3D, so every camera sees the same arrangement, and consistency stops being something the model has to remember.
15
+
16
+ ## Build
17
+
18
+ Follow `slates-previs-blocking` and add these.
19
+
20
+ ### Seated proxies, colour-coded
21
+
22
+ Simple seated shapes. **Do not animate the heads** — a proxy head turning the wrong way is worse than one that never turns.
23
+
24
+ Give each character a distinct viewport colour and write the mapping down. This is the identity channel:
25
+
26
+ > red = the boss · green = the kid · blue = the driver · yellow = the fixer · purple = the cousin · cyan = the nephew
27
+
28
+ That mapping goes verbatim into the generation prompt. It is what lets the model bind a grey body to a character sheet across four cuts.
29
+
30
+ Give every proxy a material and set `mat.diffuse_color` to its identity colour — the blocking render pins Workbench to `MATERIAL` shading, so **the material's `diffuse_color` is what reaches the clip**. Set `object.color` to the same value too, so the user's viewport matches what renders. See `slates-previs-blocking` for the snippet.
31
+
32
+ ### Fix the geography, then never move it
33
+
34
+ Place people once. Write down who sits where relative to the camera's opening position, in words, because that sentence is going into the prompt:
35
+
36
+ > Across the table, facing camera: yellow dead centre, purple far left, blue and green to the right.
37
+
38
+ ### The camera plan
39
+
40
+ Per `slates-camera-language`, with two things specific to dialogue:
41
+
42
+ - **Below shoulder height, slow rail glides.** Eye-level-and-above reads as surveillance.
43
+ - **Decide who owns the near foreground in each cut and honour it.** A shoulder in frame is a spatial anchor; a different shoulder in the next cut relocates the whole room.
44
+
45
+ The move that earns the most: **a gaze handoff without a cut** — the camera keeps gliding while the target hands off across the table, face to face, slowing on each but never stopping. Build it by keyframing the Track To target's position between subjects.
46
+
47
+ ### Crossing behind someone
48
+
49
+ A head wiping frame during a move is a strong depth cue. It is also a spatial claim, so pick who gets crossed and say so — *the camera crosses directly behind cyan's back mid-shot and his head wipes the frame once.*
50
+
51
+ ## The prompt
52
+
53
+ Everything in `slates-blocking-to-prompt`, plus these blocks.
54
+
55
+ ### Geography — restate it as a rule
56
+
57
+ > TABLE GEOGRAPHY — do not deviate: the camera is never parked behind red. Only in the opening seconds does his dark shoulder hang at the near frame RIGHT edge, and it slides out as the camera travels LEFT. The true near-foreground of this shot is cyan: the camera crosses directly behind him mid-shot. After the opening seconds red is gone from the foreground, and the camera never travels behind anyone except cyan.
58
+
59
+ ### Screen direction, and the mirror that is not a swap
60
+
61
+ The 180-degree rule survives on its own in the blocking. What breaks is the model **"correcting" a legitimate mirror** — when the camera faces back through a scene, sides invert, and that inversion is correct. Say so explicitly or it gets flipped:
62
+
63
+ > The driver's seat is on the LEFT for the entire timeline; this layout never mirrors or flips. When a camera faces BACKWARD into the car, screen sides mirror naturally: the driver reads on the RIGHT of frame, the passenger on the LEFT — that is correct left-hand drive, not a swap. They never swap seats or roles anywhere in the timeline.
64
+
65
+ Then compress it into the HOLD block: *(backward camera mirrors them: he right of frame, she left)*.
66
+
67
+ ### Presence
68
+
69
+ > A seated person stays drawn even when partially occluded — in every interior frame some part of each seat's owner is visible: a hand, an arm, a shoulder, a head above the bolster. Every occupied seat visibly holds its person.
70
+
71
+ ### Keep everyone alive
72
+
73
+ Three orthogonal layers. Without them, whoever is not speaking freezes:
74
+
75
+ - **ONGOING BUSINESS** — a small continuous physical action per character, running whether or not they are speaking. *turns his glass a quarter every few seconds · thumbs a lighter without lighting it.*
76
+ - **BACKGROUND LIFE** — soft-focus, low contrast, never pulls attention, never crosses in front of a speaking face.
77
+ - **SCENE EVENT** — the unnamed thing everyone is playing but nobody says. One line, repeated verbatim in every character's direction: *keep tomorrow sounding like a fishing trip.*
78
+
79
+ ### Acting, per character
80
+
81
+ Same six slots each. Terse:
82
+
83
+ ```
84
+ ACTING TASK — <character>
85
+ SCENE DIRECTION (shared, unspoken): <the same line for everyone>
86
+ MOTIVE (his fuel): <what he wants underneath>
87
+ GOAL: <what he wants in this scene>
88
+ OBSTACLE: <what is in the way>
89
+ TACTIC: <how he goes about it>
90
+ Moment to moment: <2-3 beats keyed to timestamps>
91
+ (Safety: gaze always engaged in the task — never a frozen, glassy,
92
+ unfocused stare; natural blink cadence.)
93
+ ```
94
+
95
+ That safety line is not filler. Dead eyes are the characteristic failure of generated faces in dialogue, and naming it is what prevents it.
96
+
97
+ ### Split any strong emotion into phases
98
+
99
+ The other characteristic failure is a face that strikes one extreme expression and holds it for the whole shot — a mouth stuck open for three seconds. Give the beat two phases with a hinge, and name the failure you are excluding:
100
+
101
+ > PHASE 1 (17.3-18.6s) — she SCREAMS at him, mouth wide, eyes huge, hand clamped on the grab handle. PHASE 2 (18.6-19.9s) — the scream breaks off: she shuts her eyes tight and CLOSES her mouth, both hands now on the handle, head ducked, braced. Scream, then brace — never one frozen open mouth held through the whole shot.
102
+
103
+ The hinge timestamp is what makes it a performance instead of a pose.
104
+
105
+ ### Dialogue must not restructure the edit
106
+
107
+ Both of these, verbatim, every time:
108
+
109
+ > DIALOGUE NEVER CREATES SHOTS: spoken lines happen inside the reference's takes exactly as blocked — no cutaways to a speaker, no reverse shots, no added close-ups. If a line plays while the camera is elsewhere, the line stays off-screen audio.
110
+
111
+ > OFF-SCREEN VOICES RULE: a line marked off-screen must STAY off-screen — never show the speaker, never move him into frame, never route the camera behind him because he spoke.
112
+
113
+ A sentence may cross a cut. Say so where it does: *the sentence does not pause for the edit.*
114
+
115
+ ## Model routing
116
+
117
+ Dialogue directed as separate layers (voices, scene sound, score) is **minimax-h3**'s seat; it also takes declared reference relationships, which suits a colour-coded cast. Native synced audio is **Veo**'s niche. seedance-2.5 carries the reference-video capacity. Route per `slates-model-selection` and read the chosen model's prompting skill before writing the audio block.
118
+
119
+ ## Checklist
120
+
121
+ - [ ] Colour→character mapping written down and pasted into the prompt
122
+ - [ ] Heads not animated in the blocking
123
+ - [ ] Seating stated as a geography rule
124
+ - [ ] Foreground owner named per cut
125
+ - [ ] Mirror-is-not-a-swap clause present if any camera faces back through the scene
126
+ - [ ] Presence rule present
127
+ - [ ] Ongoing business, background life and scene event all specified
128
+ - [ ] Acting task per character, safety line included
129
+ - [ ] Any strong emotion split into phases with a hinge timestamp
130
+ - [ ] Dialogue-never-creates-shots and off-screen-voices rules present
131
+
132
+ ## Related
133
+
134
+ `slates-previs-blocking` · `slates-camera-language` · `slates-blocking-to-prompt` · `slates-character-identity` (the sheets) · `slates-prompting-minimax-h3`
@@ -31,7 +31,7 @@ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni F
31
31
  | **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs, and the only Seedance seat that **acts on timestamps** (rules in `slates-prompting-seedance-2-5` § Timestamps) — 480p / 720p / 1080p, **no 4K**, and **dearer than 2.0 at every shared resolution** (720p $0.231/s vs $0.15/s, +54%). If you want 4K, or the same resolution cheaper, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) LENGTH is the price dial, not resolution — a 30s 720p face gen is 489 credits and a 30s 1080p faceless gen is 614, against a 1,000-credit welcome grant. Quote before any take over ~10s. |
32
32
  | **The SOUND has to be directed, not just present** — a specific line delivered a specific way, scene sound that has to sit under it, and score that must stay out of the characters' world | **MiniMax H3** | The only seat where audio is authored in three separate layers in ONE pass (synchronised events in the body, ambience in a soundscape section, audience-only score in its own) rather than toggled on. 5–15s, 480p / 768p / 2K / 4K, 24fps, 32kHz stereo, 11 languages. Rules in `slates-prompting-minimax-h3`. |
33
33
  | **A reference has to keep a DECLARED amount of itself** — especially moving one subject's characteristic onto a *different* subject | **MiniMax H3** | The only seat that understands a stated retention relationship (kept whole / kept in part / transferred onto another subject / loose echo). 9 images + 3 video + 3 audio, 12 files total. 🚨 The first 5 reference images are free and every one after that costs 4 credits — pass `referenceImages` to `slates_estimate_generation_cost` before a reference-heavy job. |
34
- | **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p/768p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at 768p with NO references of any kind. Never the default; never reach for it to save money. |
34
+ | **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p/768p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at 768p. It still animates a start frame and an end frame — image-to-video is one of the two things it is for — but it has no reference-to-video endpoint, so the omni-reference set (9 images + video + audio) is base-H3 only. Never the default; never reach for it to save money. |
35
35
  | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | Narrow, and now narrower: if the sound needs DIRECTING rather than merely existing, MiniMax H3 is the better seat. |
36
36
 
37
37
  ### Named Seedance escalation triggers