@officexapp/vidfarm-devcli 0.21.52 → 0.21.54

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,1963 @@
1
+ ---
2
+ name: ugc-reaction-greenscreen
3
+ video_type: UGC reaction cutaways + a keyed greenscreen device carrying the customer's app demo — subtitled, music-led, silent or narrated (TikTok / Reels / Shorts)
4
+ checks:
5
+ duration_sec: 15-30
6
+ aspect: 9:16
7
+ first_frame_visual: required
8
+ first_frame_text: required
9
+ text_by_sec: 0.6
10
+ captions: required
11
+ audio: required # a bed in silent mode; a bed + voiceover in narrated mode
12
+ font_regime: required
13
+ safe_zone: required
14
+ # This counts EVERY visual layer, not beats — sticker images land in it too.
15
+ # The seven-beat structure is a review item, not a machine check: 7 beats plus
16
+ # up to ~4 stickers is the shape this range is sized for.
17
+ scenes: 4-12
18
+ max_scene_sec: 10
19
+ max_text_cards: 1 # every word is a timed subtitle cue, not a card
20
+ max_simultaneous_text: 1
21
+ max_words_per_cue: 7
22
+ # 1.5 is right for SILENT mode, where text is the only channel and a gap is a free exit.
23
+ # In NARRATED mode the payoff beat is SUPPOSED to have a stretch with no caption over it —
24
+ # the voice stops on purpose and the screen does the work. Tighten this back to 1.5 if you
25
+ # cut the voiceover.
26
+ max_dead_air_sec: 3.0
27
+ max_tail_sec: 1.2
28
+ forbid_text:
29
+ - link in bio
30
+ - sign up for a free trial
31
+ - get started today
32
+ - book a demo
33
+ - download now
34
+ ---
35
+
36
+ # UGC Reaction × Greenscreen × App Demo
37
+
38
+ > A public Vidfarm prompt. Read it end to end before you build. It assumes nothing about the
39
+ > product except that it has a screen you can film.
40
+
41
+ The format is three streams of footage cut against each other:
42
+
43
+ | Stream | What it is | Where it comes from | What it does |
44
+ |---|---|---|---|
45
+ | 🙂 **Reaction** | An ordinary person on camera, mid-expression | `vidfarm.cc/discover` → public raws, shelf `ugc-reaction` (180 clips) | Makes the video a *situation* instead of an ad. Carries the hook and the bait |
46
+ | 📱 **Greenscreen device** | A hand holding a phone whose screen is a flat green field | public raws, shelf `greenscreen`, `sourceType: Display Greenscreen` | The frame the product lives inside. Turns a screen recording into something someone is *holding* |
47
+ | 🖥 **App demo** | The customer's real product doing one real thing | the customer's own folder, or you drive their site in a browser | The proof. The only beat where the product exists |
48
+
49
+ **Only one voice is ever allowed: yours.** The reaction clips are muted, the customer's demo is
50
+ muted, and the cut runs either fully silent under subtitles or under a voiceover you wrote and a
51
+ ducked music bed. Two modes, no third — see *Audio DNA*.
52
+
53
+ ## The one test
54
+
55
+ > **Does the product arrive as something a person is holding, or as a screen recording someone pasted in?**
56
+
57
+ The greenscreen device beat is the entire reason this format exists. A screen recording dropped
58
+ full-bleed on a timeline is a demo video with reaction clips stapled to it. A screen recording
59
+ living inside a phone that a real hand is tilting is a person showing you something. The viewer
60
+ reads the difference in well under a second and never articulates it.
61
+
62
+ ## The second test — does the video SHOW what the product is about?
63
+
64
+ Ask it before you cut and again on the render: **name the noun the product is about, then find it in
65
+ the footage.** A dish-search app is about *food*. A rental app is about *rooms*. A running app is
66
+ about *running*.
67
+
68
+ This format is unusually good at failing that test, and the reason is structural: its three streams
69
+ are a **face**, a **phone**, and a **UI**. None of them is the subject. On the reference build the
70
+ product finds you food by dish — and the first cut contained no food at all. It was twenty-two
71
+ seconds of people reacting to a screen. Every individual beat was correct and the video still did
72
+ not read as being about dinner.
73
+
74
+ **The UI is not the subject.** A results list saying *Cucumber Salad · $9* is a row of text. It
75
+ proves the product works; it does not make anyone hungry, and hunger is the thing you are selling.
76
+
77
+ So: find the noun, and if the raws don't contain it, **put it in** — with a cutaway from the
78
+ `lifestyle` / `b-roll` shelves if one exists, or with a sticker if it doesn't. Getting the subject
79
+ on screen twice in 22 seconds is worth more than any amount of re-cutting.
80
+
81
+ ## Part 0 — who this is for
82
+
83
+ **Not you — the customer's buyer.** Before anything else, write one line naming them, and one line
84
+ naming the moment they are in. Everything below is decided by those two lines: which reaction face,
85
+ which craving/query/task you type into the demo, which limit you concede.
86
+
87
+ - **Register:** peer to peer. Someone who had the problem, found the thing, is telling you.
88
+ - **The product enters in beat 3, never beat 1.** Beat 1 is a situation.
89
+ - **Never a brand voice.** No "introducing", no "meet", no "the all-new".
90
+
91
+ ---
92
+
93
+ ## Structural DNA — the beats
94
+
95
+ Seven beats, ~22s. The rhythm is **reaction → device → reaction → device**: the video cuts *back*
96
+ to a face after the payoff and *back* to the app after the face. One unbroken block of app footage
97
+ is the single most common way this format collapses into an ad.
98
+
99
+ | # | Beat | ~t | Stream | Job |
100
+ |---|---|---|---|---|
101
+ | 1 | **Hook** | 0:00–0:02.5 | Reaction | The situation, on a real face, before any product exists |
102
+ | 2 | **The wrong way** | 0:02.5–0:05 | Reaction | What they tried instead, and why it failed. This is the concession |
103
+ | 3 | **The move** | 0:05–0:09 | Device | The one action inside the app. Typing, tapping, pasting — at real speed |
104
+ | 4 | **The result** | 0:09–0:14.5 | Device | What came back. Held, legible, uncut. **This is the payoff** |
105
+ | 5 | **The face** | 0:14.5–0:17 | Reaction | The reaction to the result — the emotional label for what was just shown |
106
+ | 6 | **The scope** | 0:17–0:19.5 | Device | The honest limit, said over the app still running |
107
+ | 7 | **The bait** | 0:19.5–0:22 | Reaction | One comment ask, and the same ask in the post caption |
108
+
109
+ **Cast ONE actor across every reaction beat.** The narration is first person singular — *"I didn't
110
+ want a restaurant… I just picked one"* — so a different face on each beat is not a stylistic choice,
111
+ it is an incoherence: one voice saying "I" over four different people. One face turns four cutaways
112
+ into an arc (exasperated → beaten → convinced → warm), and the arc is what makes the payoff land.
113
+
114
+ Use several faces only when the video genuinely is a vox-pop — nobody is "I", the narration is
115
+ third person, and the point is *lots of people have this problem*. That is a different video.
116
+
117
+ **Beat 6 is load-bearing and gets cut first by every agent that reads this.** The limit is what
118
+ buys the rest of the video. Put it over the app so it doesn't read as a disclaimer card.
119
+
120
+ **One clip per beat — seven beats, seven layers.** Beats 3 and 4 are the same recording, so the
121
+ lazy build puts them on one ~10s layer. Don't: against 2.5s reaction cuts a 10s hold is where the
122
+ retention curve falls off, and `vidfarm qa` will say so (`slow-scene`). Split at the frame the
123
+ result lands — which is a cut the *content* is already asking for — and no layer runs past ~6s.
124
+
125
+ ---
126
+
127
+ ## Sourcing — three streams, three methods
128
+
129
+ ### 🙂 Reaction raws — browse the shelf, don't search it
130
+
131
+ ```bash
132
+ vidfarm public-raws --categories # the shelves + live counts
133
+ vidfarm public-raws --category ugc-reaction --limit 100 # the whole shelf, read the descriptions
134
+ ```
135
+
136
+ The shelf's semantic search is thin; **the descriptions are the index**. Pull the shelf as JSON and
137
+ filter the `description` field on the emotion you want (`surpris`, `disgust`, `distress`, `smil`,
138
+ `laugh`, `disbelie`, `confus`). Pick by the description, then confirm by eye — build a contact
139
+ sheet of 4–5 stills per candidate and read them as one image before you commit.
140
+
141
+ #### Cast the actor, not the clip — the `actor_<uuid>` tag
142
+
143
+ **A shelf is dozens of clips of a much smaller number of creators**, and nothing in the taxonomy
144
+ says "same face". The `actor_<uuid>` tag does. Every tagged public raw carries one actor id in its
145
+ `summary` (and in `tags.actor`), so you read it off a card you already like and search for it as a
146
+ plain token:
147
+
148
+ ```bash
149
+ vidfarm public-raws --category ugc-reaction --limit 100 # 1. browse, pick a face
150
+ # 2. read the actor_<uuid> token off that card's summary — it is appended at the END
151
+ vidfarm public-raws --query actor_668363ea-… # 3. every other clip of that person
152
+ ```
153
+
154
+ The token lives at the tail of the card's `summary` string. In the feed response `tags` comes back
155
+ `null`, so **`summary` is the field to parse** — a reader that only looks at `tags.actor` finds
156
+ nothing and concludes the shelf is untagged.
157
+
158
+ REST twin `GET /api/v1/public-raws?q=actor_<uuid>`, and because it is an ordinary keyword search it
159
+ composes: `?category=ugc-reaction&q=actor_<uuid>`.
160
+
161
+ **Cast by ARC, not by clip.** You need four reaction beats — hook, the wrong way, the payoff face,
162
+ the bait — so pick the actor whose set can carry all four, not the single best-looking clip. On the
163
+ reference build one actor had five takes and four of them mapped straight onto the beats:
164
+
165
+ | Beat | Their take |
166
+ |---|---|
167
+ | 1 hook | distressed faces, hand at her mouth |
168
+ | 2 the wrong way | frowning at a **laptop** — which is what "Google gave me a list of fifty" is about |
169
+ | 5 the face | **thumbs up**, outdoors |
170
+ | 7 the bait | a warm smile to camera |
171
+
172
+ Practical notes, from the catalogue as it stands:
173
+
174
+ - **Measure the pool, don't take a doc's word for it.** Counted today: **180 `ugc-reaction` raws,
175
+ 53 tagged actors, 3 untagged — and 25 actors with 4+ takes.** That is a far bigger casting pool
176
+ than the shelf looks like from the outside, and it is the only part of it that can carry this
177
+ format. Re-count rather than trusting this line; the shelf is actively being tagged.
178
+ - ⚠️ **The feed caps `limit` at 100 and paginates on `cursor`, not `offset`.** `offset=100` is
179
+ silently ignored and returns page 1 again — a loop built on it looks like it works, reports a
180
+ plausible number, and has seen half the shelf. Follow `next_cursor` until it is null.
181
+ - ⚠️ **Read the id fresh, and re-read it if it stops matching.** The ids come from a tagging pass
182
+ that gets re-run: mid-build here, an id that had returned 5 clips started returning **0**, because
183
+ the shelf had been re-tagged and the same person's clips now carried a different id (with 7 clips
184
+ grouped under it instead of 5). So an actor id is a lookup key that is good for *this* build, not
185
+ a permanent name. If a saved id returns nothing, the actor has not gone — re-read the token off a
186
+ card. And never invent one; a made-up id matches nothing.
187
+ - **No tag means "not tagged yet", not "a different person".** Older raws and raws with nobody on
188
+ camera have none.
189
+ - **Wardrobe and location will NOT match between takes** — different top, different room, outdoors.
190
+ That is fine and it is not a continuity error: UGC creators film across days, and the audience
191
+ reads it as the same person at different moments. **Identity continuity is what matters; wardrobe
192
+ continuity is not.** What *would* break it is a different face.
193
+ - Same trick answers a director asking *"more of her"* — the actor id on the raw already in the
194
+ composition is the answer.
195
+
196
+ ⚠️ **`ffprobe` reports CODED dimensions, and phone video is rotated.** Almost every clip on this
197
+ shelf carries a 90° rotation matrix: `ffprobe` says `3840x2160` and ffmpeg **decodes it as
198
+ 2160x3840**, already 9:16. Hand-computing a crop from the probed width put the subject off-frame on
199
+ all four beats of the reference build — wall where the face should be. Never compute geometry from
200
+ the probe; let the filter work on the decoded frame:
201
+
202
+ ```bash
203
+ # right: operates on what was decoded, whatever the rotation
204
+ -vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1,fps=30"
205
+ ```
206
+
207
+ If you genuinely need an off-centre crop, read the decoded size first
208
+ (`ffprobe -show_entries side_data=rotation`, or decode one frame and measure it).
209
+
210
+ ### 📱 Greenscreen device raws — check `sourceType` before you key anything
211
+
212
+ The `greenscreen` shelf holds two populations that look identical to a category filter and are
213
+ completely different to a lawyer:
214
+
215
+ | `sourceType` | What it is | Use it for |
216
+ |---|---|---|
217
+ | **Display Greenscreen** | Device mockups — a hand holding a phone with a green screen, a phone on a notebook. No identifiable person, no character | ✅ Client work. This is the one you want |
218
+ | **MemeScreens** | Keyed celebrities, TV characters, pets, animated figures on a green field | ⚠️ Normal for a meme repost. A rights problem on a customer's paid ad running on owned accounts |
219
+
220
+ ```bash
221
+ vidfarm public-raws --category greenscreen --query "smartphone green screen" --limit 25
222
+ ```
223
+
224
+ There are only a handful of Display Greenscreen clips, so **look at all of them and pick on
225
+ motion, not on looks.** Ranked, best first:
226
+
227
+ 1. **Two hands, thumb on the glass, indoors, native 9:16.** Reads as a person using a phone. Drifts
228
+ the most — needs the tracked insert below, which is fine, because you are doing that anyway.
229
+ 2. **Phone flat on a desk or notebook, top-down.** Calm, minimal drift, but reads as a product shot
230
+ rather than a person.
231
+ 3. **Phone against a plain green background.** Avoid: the plate colour and the screen colour are the
232
+ same field, so a key gives you a floating phone body and a demo behind *everything*.
233
+
234
+ ### 🖥 App demo raws — the customer's folder first, the browser second
235
+
236
+ **Ask for a folder before you build one.** A customer who already records their own product gives
237
+ you footage with the right accounts, the right data and the right edge cases. Take it, and then:
238
+
239
+ > **Cut it. Do not narrate it, and do not keep their narration.**
240
+ >
241
+ > A narrated demo raw is *worse* input than a silent one. Their voice competes with your subtitles,
242
+ > the pacing is theirs and the tone is a webinar. Strip the audio, find the four seconds where the
243
+ > product actually does the thing, and throw the rest away. If they insist their narration ships,
244
+ > that is a different format — use `ugc-testimonial`.
245
+
246
+ **No folder? Drive their site yourself.** Free, local, repeatable, and it always has current data:
247
+
248
+ ```bash
249
+ vidfarm capture "https://<their site>" --out ./capture # tokens, copy, screenshots — read this first
250
+ vidfarm browser setup # then drive their real UI
251
+ ```
252
+
253
+ Record the interaction at the **device screen's aspect ratio**, not at desktop and not at a laptop
254
+ viewport. A 440×940 recording drops into a phone screen with no letterbox and no re-crop; a
255
+ 1440×900 one arrives as a stamp inside a phone and the whole beat is unreadable.
256
+
257
+ Four rules for the recording itself, all of them learned the expensive way:
258
+
259
+ 1. **Verify the query returns results before you film it.** Type it yourself first. A product's own
260
+ advertised example can be broken — in the worked example that produced this file, the exact phrase the
261
+ customer prints on their own homepage returned *No dishes found*, and the first take was unusable.
262
+ 2. **Park the interactive region at the top of the frame before you type.** Otherwise the results
263
+ render below the fold and the payoff never appears on screen.
264
+ 3. **Type at human speed** (~130ms/char) and hold ~2.5s after the result lands. Instant fills read
265
+ as a mockup; a real keystroke cadence is the cheapest authenticity you will ever buy.
266
+ 4. **Draw a soft synthetic cursor.** A UI that changes with nothing touching it looks like a video
267
+ of a website. A dot that moves to the field first looks like a person.
268
+
269
+ ---
270
+
271
+ ## Visual DNA
272
+
273
+ ### The tracked insert — the technique the whole format rests on
274
+
275
+ The device is handheld. It **drifts and it is tilted**, so a flat chroma key with the demo pinned
276
+ behind it fails in three named ways, all visible within two seconds:
277
+
278
+ | Defect | Cause | Fix |
279
+ |---|---|---|
280
+ | The demo **slides under the bezel** | The insert is static; the phone is not | Re-measure the green region *every frame* |
281
+ | A **green wedge** down one edge | The insert is axis-aligned; the phone is rotated a few degrees | Fit to the screen's own axes (PCA on the green mask), not to an upright bounding box |
282
+ | The demo **paints over the thumb** | The insert is pasted into a *box* | Paste by **mask**: every green pixel becomes demo, every non-green pixel (thumb, bezel, room) is left alone |
283
+
284
+ The mask paste is what sells it — the thumb ends up *on top of* the app, which is the single frame
285
+ that makes a viewer believe someone is holding it. Appendix A is a ~90-line script that does all
286
+ three. `vidfarm remove-greenscreen` is the right tool for a **full-frame** flat plate; it is the
287
+ wrong tool here, because it has no reason to know where the screen went.
288
+
289
+ ```bash
290
+ python3 track-insert.py phone-raw.mp4 demo.mp4 phone-demo.mp4 --demo-start 1.2
291
+ ```
292
+
293
+ `--demo-start` trims the demo's dead head (page load, blank white) so the insert opens on something.
294
+
295
+ ### Punch in on the device beats
296
+
297
+ A Display Greenscreen raw frames the phone small, with a lot of blurred room around it. Shipped as
298
+ is, the UI is unreadable at phone scale and the demo did not happen. Crop to the device **plus a
299
+ drift margin**, then scale back to canvas:
300
+
301
+ Measure the green region's **travel** before you choose the crop — a crop sized to one frame clips
302
+ the device on another:
303
+
304
+ ```python
305
+ # per-second bounding box of the green field, on the ORIGINAL device raw
306
+ import subprocess, numpy as np
307
+ W, H, f = 1080, 1920, "device-raw.mp4"
308
+ p = subprocess.run(["ffmpeg","-v","error","-i",f,"-vf","fps=1,scale=320:569",
309
+ "-pix_fmt","rgb24","-f","rawvideo","-"], capture_output=True)
310
+ b = np.frombuffer(p.stdout, np.uint8); n = len(b)//(320*569*3)
311
+ for i, fr in enumerate(b[:n*320*569*3].reshape(n,569,320,3).astype(np.int16)):
312
+ r, g, bl = fr[...,0], fr[...,1], fr[...,2]
313
+ m = (g > 90) & (g-r > 40) & (g-bl > 40)
314
+ ys, xs = np.nonzero(m)
315
+ print(f"t={i}s x[{xs.min()*W//320},{xs.max()*W//320}] y[{ys.min()*H//569},{ys.max()*H//569}]")
316
+ ```
317
+
318
+ Take the union of every box, add ~60px of margin on each side, and round the crop to the canvas
319
+ aspect:
320
+
321
+ ```bash
322
+ ffmpeg -i phone-demo.mp4 -vf "crop=880:1564:20:130,scale=1080:1920,setsar=1,fps=30" s3-move.mp4
323
+ ```
324
+
325
+ Typical punch-in is 1.2–1.4×. Use the **same crop for every device beat** in the video — a punch-in
326
+ that changes between beat 3 and beat 6 reads as two different videos.
327
+
328
+ ### Stickers — allowed, and only ever as annotation
329
+
330
+ Stickers earn their place in exactly two jobs, and both are about comprehension:
331
+
332
+ | Job | What it looks like | When you need it |
333
+ |---|---|---|
334
+ | **Annotate** | a ring around the row the narration just named, an arrow at the field being typed into, an underline under the number being read out | the device beat asks a viewer to read a dense UI in about two seconds |
335
+ | **Supply the subject** | the actual thing the product is about — the dish, the room, the shoe — laid over the shot | the raws are faces and a phone, and the product's own noun never appears (see *The second test*) |
336
+
337
+ **The test: does it make the product clearer?** If yes, use it. If it competes for attention with
338
+ what's already on screen, or a viewer has to work out why it's there, cut it — a confusing sticker
339
+ is worse than none, and this format is dense enough already.
340
+
341
+ **Never furniture.** A floating icon, a badge, a chip, a logo, a "NEW!" burst, an emoji dropped in
342
+ to fill space — none of those is either job, and `vidfarm qa` is right to call them slop.
343
+
344
+ - **One on screen at a time, and not on every beat.** Two or three across a 22s cut is plenty. A
345
+ sticker that sits for the whole video has stopped being punctuation.
346
+ - **It appears on the word that names it and leaves.** Same discipline as the captions: absolute
347
+ `animation-delay`, and it goes when the thought does.
348
+ - **An in-screen annotation is only valid while the UI under it is STATIC.** This is the one that
349
+ bites. The reference build's ring was placed over the picked row for 1.3s — and the results list
350
+ began scrolling 0.3s in, so the last third of the shot had a marker circling the wrong dish.
351
+ Step the beat frame by frame, find where the content moves, and end the sticker before it:
352
+ ```bash
353
+ ffmpeg -i beat.mp4 -vf "fps=2,scale=300:-1,drawgrid=w=iw/20:h=ih/20:t=1:c=red@0.5,tile=6x1" \
354
+ -frames:v 1 rows.png # read the row's box off the grid, and see when it moves
355
+ ```
356
+ - **Hand-drawn for annotation; photographic for the subject.** A crisp vector circle reads as a UI
357
+ element drawn over the video — ask for a rough felt-tip stroke instead. But a *subject* sticker
358
+ should look like the real thing: a photographic cutout of the actual dish sells the craving, an
359
+ illustration of one does not.
360
+ - **A tight face has no room for a sticker.** On a close portrait there is nowhere to put one that
361
+ isn't a forehead, a temple or an eye, and anything you place there reads as stuck to their head
362
+ rather than laid on the shot. Put subject stickers on the beats with air in them — the device
363
+ beats have defocused room on both sides, and a medium shot has wall. If the beat you *want* is a
364
+ tight face, move the sticker to the next beat rather than shrinking it into a corner.
365
+
366
+ **Sourcing, cheapest first.** `vidfarm iconscout --free` and `vidfarm media icon` are free but
367
+ almost all require an attribution credit — and an attribution line on screen is brand chrome you
368
+ already said you would not have. So for this format the practical route is one generated cutout:
369
+
370
+ ```bash
371
+ vidfarm cutout --generate "a single thick hand-drawn yellow felt-tip ELLIPSE OUTLINE, stroke
372
+ ~40px. The ENTIRE rest of the image, INCLUDING THE AREA INSIDE THE ELLIPSE, is solid magenta
373
+ #FF00FF. No white, no grey, no paper, no shadow, no text." \
374
+ --key-color "#FF00FF" --key-mode flat --tolerance 0.4 --out stickers/ring.png
375
+ ```
376
+
377
+ **Match the KEY MODE to the shape, and the PLATE COLOUR to the palette.** Getting this wrong is
378
+ silent — the file looks right in a thumbnail and is wrong in the video:
379
+
380
+ | Art | Key mode | Why |
381
+ |---|---|---|
382
+ | **Hollow** — a ring, a frame, an arrow outline | `--key-mode flat` | the interior never touches the frame edge, so the default connectivity keyer *keeps* it and you get a filled disc instead of a ring |
383
+ | **Solid** — a photographic dish, a product, a person | `--key-mode smart` (the default) | flat mode's soft alpha ramp eats a photograph. Measured on a first attempt: **65% of the cutout came back partially transparent**, so the background showed straight through the food. Smart mode returned the same art at **0.6% partial alpha** |
384
+
385
+ Pick the plate against the subject's own colours — green plate for red food, magenta for green food
386
+ — and check the result rather than trusting it:
387
+
388
+ ```python
389
+ # any cutout: how much of it is actually opaque?
390
+ a = alpha_channel(png)
391
+ print((a == 0).mean(), ((a > 0) & (a < 255)).mean(), (a == 255).mean())
392
+ # healthy: mostly 0 or 255, a percent or two in between. Tens of percent "partial" = a bad key.
393
+ ```
394
+
395
+ Glass and soft edges pick up the plate colour even after a good key. One numpy pass fixes it: where
396
+ the plate's channel is the dominant one, pull it down to the average of the other two.
397
+
398
+ Three things in the prompt are load-bearing, and each one cost a wasted image job to learn:
399
+
400
+ 1. **Pin the plate colour and say it twice.** Left to itself the model draws on **white**, the plate
401
+ never gets keyed, and you get a white box with art in it. Worse, if the art is also white (an
402
+ arrow, a highlight) you cannot key the background without destroying the subject.
403
+ 2. **Say "including the area inside the ellipse".** A ring's interior does not touch the frame
404
+ edge, so the default connectivity keyer *keeps* it — you get a filled disc, not a ring.
405
+ `--key-mode flat` removes the colour wherever it appears, which is what an open shape needs.
406
+ 3. **Ban white and grey explicitly.** Otherwise you get paper texture and a drop shadow, and the
407
+ shadow keys as a grey halo.
408
+
409
+ **Placement.** Drive it from a manifest so the set is reproducible and overlaps get caught —
410
+ Appendix I:
411
+
412
+ ```
413
+ # src | start | dur | left% | top% | width% | height% | label | rotate | fit
414
+ media/stickers/food-chilicrisp.png | 6.60 | 1.90 | 2 | 21 | 27 | 16 | the craving | -5 | contain
415
+ media/stickers/ring.png | 10.30| 1.40 | 18 | 38.5 | 62 | 14 | the pick | 0 | fill
416
+ media/stickers/food-cucumber.png | 14.70| 1.90 | 56 | 14 | 30 | 18 | what $9 buys | 6 | contain
417
+ ```
418
+
419
+ **`fit` is not cosmetic — the two sticker jobs need opposite values.** A subject sticker must keep
420
+ its aspect (`contain`); a photo of a dish squashed to a box is obviously wrong. An annotation
421
+ sticker must **squash to the shape it marks** (`fill`) — a near-square ring set to `contain` inside
422
+ a wide, short box collapses to the box's *height* and renders as a small circle floating beside the
423
+ row instead of around it. That regression is invisible in the markup and obvious in one frame, so
424
+ look at the frame.
425
+
426
+ The markup it writes — wrap the image in a `div.clip` exactly like a caption layer. A bare `<img>`
427
+ with `data-start` renders at t=0 and then never shows or hides, because the engine does not manage
428
+ it:
429
+
430
+ ```html
431
+ <div class="clip" id="ann-ring" data-hf-id="ann-ring" data-layer-mode="publish"
432
+ data-layer-kind="image" data-start="10.3" data-duration="1.4" data-end="11.7"
433
+ data-track-index="6" data-label="annotation: the dish that was picked"
434
+ style="position:absolute;inset:auto;left:18%;top:38.5%;width:62%;height:14%;z-index:6">
435
+ <img class="sticker" src="media/stickers/ring.png"
436
+ style="width:100%;height:100%;object-fit:fill;animation-delay:10.3s"></div>
437
+ ```
438
+
439
+ `inset:auto` first, because `.clip { inset: 0 }` would otherwise fight the geometry. `object-fit:
440
+ fill` is deliberate — a marker circle around a line of text *should* squash to the row's aspect.
441
+
442
+ ```css
443
+ .sticker{animation-name:vfStickIn;animation-duration:.42s;animation-fill-mode:both;
444
+ animation-timing-function:cubic-bezier(.2,1.1,.3,1);transform-origin:center}
445
+ @keyframes vfStickIn{
446
+ 0%{opacity:0;transform:scale(.55) rotate(-14deg)}
447
+ 60%{opacity:1;transform:scale(1.06) rotate(2deg)}
448
+ 100%{opacity:1;transform:scale(1) rotate(0deg)}}
449
+ ```
450
+
451
+ ### Everything else
452
+
453
+ - **9:16, 1080×1920, 30fps**, one canvas, no letterboxing anywhere.
454
+ - **Butt cuts.** No crossfades, no whips, no transitions of any kind. This format cuts on the
455
+ emotional beat; a transition softens exactly the moment you want hard.
456
+ - **The cut from device back to face lands on the result**, not two seconds after it.
457
+ - **No brand chrome.** No logo, no end card, no URL, no title card. The product's own UI is the
458
+ only branding, and it is already on screen.
459
+
460
+ ---
461
+
462
+ ## Audio DNA
463
+
464
+ **Two modes. Pick one before you cut, because the subtitles are built differently in each.**
465
+
466
+ | Mode | Track stack | Subtitles are… |
467
+ |---|---|---|
468
+ | **Silent** | music bed only | authored by hand, timed to the CUTS (`captions.srt` → Appendix C) |
469
+ | **Narrated** | voiceover over a ducked bed | transcribed from the VO, timed to the WORD (Appendix F) |
470
+
471
+ Narrated is the stronger default: it carries the beats where the picture is quiet, and the read-out
472
+ of a real number lands harder spoken than written. Reach for silent when the poster will drop
473
+ trending platform audio over the whole thing at upload — a voiceover fighting a trending track is
474
+ worse than either alone.
475
+
476
+ **What never changes, in either mode:**
477
+
478
+ - **Mute every reaction raw.** They came with real audio about something else entirely. A face
479
+ visibly saying words that don't match the subtitles is the loudest fake tell in the format, and it
480
+ survives every other quality pass because nobody watches the render with sound on.
481
+ - **Mute the demo.** Even if the customer's raw has narration. Their voice competes with yours,
482
+ the pacing is theirs, and the tone is a webinar.
483
+ - **The narration is YOURS, written to the beat table.** A voiceover is not permission to keep the
484
+ customer's — it is one script, one voice, one delivery, across the whole cut.
485
+
486
+ ### Narrated mode
487
+
488
+ - **One VO clip per beat, not one long take.** Each clip is placed at its own beat, so a re-cut of
489
+ one beat re-records one line instead of re-timing everything. It also keeps the VO from drifting
490
+ against the picture as the edit changes.
491
+ - **Pin the voice.** `vidfarm tts "<line>" --voice <name> --style "<one delivery note>"` — pass the
492
+ SAME `--voice` and the SAME `--style` string on every line. Generate them in separate calls
493
+ without pinning and the narrator subtly changes character between beats, which reads as
494
+ "assembled" long before anyone can say why.
495
+ - **Direct the pace, don't fix it later.** A "dry, considered" style prompt produced 4.6s of speech
496
+ for four words — nearly double its beat. Re-prompting for "brisk, punchy, quick clipped delivery,
497
+ no pauses between sentences" brought the same line to 2.2s and it sounds *more* native, because
498
+ short-form narration is fast. Prefer re-prompting over time-stretching.
499
+ - **Compress pauses; don't stretch speech.** When a line is still slightly long, strip the internal
500
+ silences rather than `atempo` the whole clip — it preserves the natural word rate and kills dead
501
+ air, which is what you wanted anyway:
502
+ ```bash
503
+ ffmpeg -i vo.wav -af "silenceremove=start_periods=1:start_silence=0.04:start_threshold=-45dB\
504
+ :detection=peak:stop_periods=-1:stop_duration=0.12:stop_threshold=-40dB,areverse,\
505
+ silenceremove=start_periods=1:start_silence=0.04:start_threshold=-45dB:detection=peak,areverse" tight.wav
506
+ ```
507
+ - **Leave the payoff alone.** Do not narrate over the whole runtime. On a 22s cut, ~15s of speech
508
+ and ~7s of silence is right — the result beat needs a stretch where nothing is being said and the
509
+ screen is just doing the thing.
510
+ - **A VO may cross a cut.** A line that starts on a face and finishes over the app binds the two
511
+ shots together. Don't clip narration to shot boundaries.
512
+ - **Duck the bed to `data-volume` ≈ 0.2** and normalise each VO clip to about -16 LUFS. Target
513
+ **12–15 dB of speech over bed**, and verify it by measuring the two STEMS, because `ebur128` on
514
+ the finished mix cannot separate them.
515
+
516
+ ### The bed, in both modes
517
+
518
+ - **Level the bed and measure it.** Target ~-16 LUFS integrated, peak < -1 dBFS
519
+ (`loudnorm=I=-16:TP=-1.5:LRA=11`). Fade in ≤0.6s, fade out over the last ~1.2s.
520
+ - **`data-volume` multiplies the already-normalised file.** A bed normalised to -16 LUFS sitting on
521
+ a layer at `data-volume="0.5"` ships at about **-22 LUFS** — quiet enough that a phone speaker in
522
+ a noisy room gets nothing. The 0.1–0.2 bed level everyone quotes is for ducking *under a
523
+ voiceover*, and there is no voiceover here. Normalise the file, then leave the layer at `1`, and
524
+ measure the RENDER (`ffmpeg -i out.mp4 -af ebur128=peak=true -f null -`) rather than the stem.
525
+ - **Source it free.** `vidfarm media bgm "<vibe>" --provider openverse` — prefer **CC0** so nothing
526
+ has to be credited on screen. A CC-BY track means an attribution line, and an attribution line is
527
+ brand chrome you just said you would not have.
528
+ ### Ship TWO cuts, from ONE render
529
+
530
+ Do not put the music bed in the composition. **The composition is the voiceover-only master**, and
531
+ the bed is applied at export — which gives you both deliverables for the price of one render:
532
+
533
+ ```bash
534
+ python3 export-versions.py renders/cut-vo.mp4 media/bgm.mp3 --bed-volume 0.22
535
+ # voiceover only cut-vo.mp4 publish this
536
+ # voiceover + music cut-music.mp4 review / fallback / watermarked
537
+ ```
538
+
539
+ | Cut | What it's for |
540
+ |---|---|
541
+ | **`-vo`** (voiceover, no music) | **The one you publish.** TikTok, Reels and Shorts all let the poster attach a track from the platform's own library at upload — licensed through the platform's deals with the labels, so the music is cleared where viewers actually hear it, and the post gets whatever is trending *this week* rather than whatever was trending when you rendered |
542
+ | **`-music`** (voiceover + bed) | The review copy you send a client, the fallback for anywhere without an in-app music library, and the one that carries a **watermark** if you need one |
543
+
544
+ **Derive, never re-render.** `-c:v copy` means both files carry the byte-identical video stream, so
545
+ approving one approves the other. Two separate renders are two different videos that merely look
546
+ alike, and the frame you approved is not the frame you shipped. A watermark is the one thing that
547
+ breaks this — it has to be burned in, so that version gets its own encode and stops being
548
+ frame-identical. Say so in the handoff.
549
+
550
+ **`normalize=0` on the `amix`.** Left at its default, `amix` pulls the voice down to make room for
551
+ the bed and quietly undoes the levelling you just measured.
552
+
553
+ - **Say which kind of bed it is in the handoff.** If the poster will attach trending audio at
554
+ upload, the bed in the `-music` cut is a *review* bed. If the `-music` cut ships as is, the bed is
555
+ the bed. These are different decisions and the customer has to make it.
556
+
557
+ ---
558
+
559
+ ## Subtitle DNA — TikTok-native, or don't bother
560
+
561
+ Subtitles are not accessibility here. They are the **only** delivery system for the hook, the loop,
562
+ the payoff, the limit and the bait. Get them wrong and the video has no words at all.
563
+
564
+ **Native means the captions are TIMED TO THE VOICE.** One short phrase on screen at a time, turning
565
+ over as the person speaks. Not a static two-line block, not a designed card, not a lower third.
566
+
567
+ A per-word *highlight* is one way to do that and it is not required — **plain white text with no
568
+ accent at all is TikTok's own default auto-caption look**, and it is the right choice more often
569
+ than the coloured mechanics: over busy footage, over a UI that already has its own colours, and on
570
+ any beat where the words should carry no emphasis of their own.
571
+
572
+ ### What is FIXED, and what VARIES
573
+
574
+ Native is a small set of invariants, not a house style. Ship the same colour, case, size and
575
+ mechanic on every video and the catalogue reads as a content farm — which is a worse failure than an
576
+ odd caption, because it is visible across the whole account at once.
577
+
578
+ | Fixed — this is what "native" means | Varies — pick per video |
579
+ |---|---|
580
+ | **TikTok Sans**, self-hosted, no fallback chain | **weight** 700 / 800 / 900 |
581
+ | captions are **timed to the voice** | **which mechanic** — including *none* |
582
+ | **no phrase plate**; text sits on the picture | **accent colour** |
583
+ | inside the **platform-safe core** | **size** within the standard's ~36–64px band, and **position** within the core |
584
+ | **verbatim**, ≤3–4 words a page, one cue at a time | **case** — uppercase, or sentence case like TikTok's own auto-captions |
585
+
586
+ **Hold one combination for the length of a single video; change it between videos.** Styling that
587
+ shifts mid-video reads as a bug; styling that never shifts across thirty videos reads as a factory.
588
+
589
+ ### The six mechanics — all native, none of them "the" one
590
+
591
+ Ordered quietest to loudest. The first two carry **no accent colour whatsoever**.
592
+
593
+ | `--highlight` | What it does | Reads as |
594
+ |---|---|---|
595
+ | `none` | plain white; the whole page appears at once and the page turns do the work | TikTok's stock auto-captions. The calmest, and the one that never fights the footage |
596
+ | `reveal` | plain white, but each word switches on at its own moment | still word-timed, still no colour — a middle setting when the beat wants pace but not emphasis |
597
+ | `colour` | the active word takes the accent, settles back to white | the quiet accented default; safe on busy footage |
598
+ | `pop` | accent **plus a vertical stretch** | more energy, still calm enough for a talking beat |
599
+ | `underline` | an accent bar wipes under the active word | "reading along"; good under a number being read out |
600
+ | `pill` | a filled accent box behind the active word, dark text inside | the loudest; excellent over a bright UI where white-on-white struggles |
601
+
602
+ **Every accented one is engineered to take no horizontal space.** That is not a stylistic choice —
603
+ an inline-block word that grows sideways eats the gap beside it and `a restaurant` renders as
604
+ `arestaurant` on the beat the word fires. `pop` therefore scales on **Y only**; `pill` and
605
+ `underline` live inside the word's own box.
606
+
607
+ **`reveal` deliberately does not re-centre.** A word that has not fired yet is transparent but still
608
+ occupies its slot, so the visible words sit where they will end up rather than sliding left as each
609
+ one arrives. Reflowing to keep the visible text centred makes every word jump, which is far worse
610
+ than the asymmetry.
611
+
612
+ ### Take the accent from the customer's own palette
613
+
614
+ `vidfarm capture` already wrote it. `capture/extracted/tokens.json` carries the site's CSS
615
+ variables, so the caption accent can come from the product rather than from your habits — different
616
+ per client, for free, and never brand chrome because it is one word at a time rather than a logo:
617
+
618
+ ```bash
619
+ python3 -c "import json;d=json.load(open('capture/extracted/tokens.json'));print(d['cssVariables'])"
620
+ # dishcover.io → --accent #2b3eff · --accent-orange #ff6b00 · --accent-yellow #ffd500 · --accent-red #e5432e
621
+ ```
622
+
623
+ Then **contrast-check it against the band you measured** — a mid-blue accent on a dark scene is a
624
+ word that disappears at the exact moment it matters. If the accent fails, use it as the `pill`
625
+ background with dark text instead of as the type colour; a filled box survives backgrounds that
626
+ coloured type does not.
627
+
628
+ **The highlight is a COLOUR SWAP, not a zoom.** A scaling active word is a CapCut habit, and it has
629
+ a concrete failure mode as well as a stylistic one: an inline-block word scales from its centre into
630
+ its neighbour, so `a restaurant` renders as `arestaurant` on the exact beat the word pops — and a
631
+ text stroke eats the remaining gap from both sides. Padding the words apart only trades collision
632
+ for letterspaced display type that stops reading as speech. Colour (plus, at most, a ~0.05em
633
+ vertical lift that costs no horizontal room) makes the collision **impossible** rather than tuned.
634
+
635
+ > ⚠️ **Verify the words actually animate, in the RENDER.** `--style word-pop` writes per-word
636
+ > `animation-delay` / `animation-duration` and leaves the keyframes to the render engine — and on
637
+ > some engine versions that binding silently does not happen, so the cue renders as **one static
638
+ > block with no active-word colour**. Nothing errors. `captions list` shows correct timings. The
639
+ > composition lints clean. You only find it by pulling ~7 consecutive frames from *inside a single
640
+ > cue* and looking:
641
+ >
642
+ > ```bash
643
+ > ffmpeg -i out.mp4 -ss <cue start> -t <cue length> -vf "fps=6,crop=1080:340:0:1180,tile=7x1" \
644
+ > -frames:v 1 cap-anim.png
645
+ > ```
646
+ >
647
+ > If all seven frames are identical, the format's whole subtitle design did not ship — **unless the
648
+ > mechanic is `none`, where identical frames are exactly right.** Check which one you chose before
649
+ > you go debugging. **Own the
650
+ > keyframes in the composition's own `<style>`** rather than depending on the engine — the emitter
651
+ > in Appendix C does this, and it is the one change that makes the animation a property of the file
652
+ > instead of a property of whatever version rendered it.
653
+ >
654
+ > Two traps inside the fix, both of which look like "the animation is still broken":
655
+ >
656
+ > 1. **`animation-delay` is measured from PAGE time, not from the layer's own start.** A cue at
657
+ > 9.1s with a relative delay has already run to completion by the time the layer is visible, so
658
+ > every word renders in its end state. Emit `animation-delay: <cue start + word offset>`.
659
+ > Symptom: cue 1 animates, every later cue is static — which is easy to misread as "it works".
660
+ > 2. **Don't fade the words in.** With a single translucent plate behind the phrase, a progressive
661
+ > reveal leaves an empty grey box on screen ahead of the words, and that reads as a loading
662
+ > state. Keep the whole phrase visible and pop only the ACTIVE word (scale + accent colour,
663
+ > settling back to white) — which is TikTok's own convention anyway. `animation-fill-mode:
664
+ > forwards`, base style = the settled state.
665
+
666
+ ```bash
667
+ # Appendix C, with the dials shown. There is no default combination to copy —
668
+ # choose one per video and write it into the handoff so the next one differs.
669
+ python3 srt-to-cues.py composition.html cues.json \
670
+ --font "TikTok Sans" --weight 900 --font-size 64 \
671
+ --highlight colour --active-color "#FFD500" \ # or: none | reveal | pop | pill | underline
672
+ --y 57 --x 14 --width 72 --plate none # add --no-uppercase for sentence case
673
+ ```
674
+
675
+ - **In SILENT mode, hand-author the SRT.** There is no narration to transcribe, and `--text` paging
676
+ spreads words evenly across a window, which desyncs from your cuts. Write the cue times against
677
+ the beat table.
678
+ - **In NARRATED mode, never hand-author them.** Captions must be verbatim and sit on the word, so
679
+ they are derived from the VO's own transcript — Appendix F does it. A karaoke highlight 200ms off
680
+ the voice is more distracting than no highlight at all.
681
+
682
+ > ⚠️ **`captions generate --srt` does not respect your cue boundaries.** It flattens the whole file
683
+ > into one word stream and re-pages it by `--max-words`, so words hop the cut: `the dish instead` +
684
+ > `chili crisp` renders as `dish instead chili crisp`, two beats late and over the wrong footage.
685
+ > That is invisible in `captions list` timings and obvious in a contact sheet, which is exactly why
686
+ > you build the sheet. **Always read the generated cue text back before you render.** When the cues
687
+ > are pinned to cuts rather than to speech — which in this format they always are — use the
688
+ > one-layer-per-cue emitter in **Appendix C** instead, and keep `captions generate` for the case
689
+ > where a real narration track exists.
690
+
691
+ - **A cue is legible at frame 0 — literally `data-start="0"`.** `first_frame_text` grades the
692
+ frame, not the intent, so a cue that opens at 0.15s fails it. Frame 0 is the thumbnail.
693
+ - **≤3 words per page, ≤7 words per cue.** Longer than that and the page turns faster than it reads.
694
+ - **One cue at a time, ever.** Two simultaneous text objects is this format's version of clutter.
695
+ ### The caption standard is NOT this harness's to define
696
+
697
+ The font regime, the size band, the safe zone and the legal caption backgrounds are **core Vidfarm
698
+ standards**. Read them there and follow them; do not take a restatement from a format harness:
699
+
700
+ | Where | What it covers |
701
+ |---|---|
702
+ | `vidfarm.cc/skill.md` → *Standards* → **Captions** | the one-paragraph rule: imported display font, size band, safe zone, emptiest part of the frame, 3–5-word cues, the four legal backgrounds |
703
+ | installed `SKILL.md` → **"the TikTok-native caption standard"** | the same, with the `set_captions` presets named |
704
+ | `references/editor-workflows.md` → **"Size → scaled to the line"** and **"Font → the composition regime"** | the numbers, the bundled font list, and why an un-imported font is the slop look |
705
+
706
+ The parts that bind here, quoted rather than paraphrased: **the bundled display fonts only**
707
+ (Montserrat 700–900 default, **TikTok Sans**, Abel, Source Code Pro, Yesteryear); **~36–64px on a
708
+ 1080-wide frame** — above ~64px is a *hook-word* size, one to three words on purpose; **the 8%–85%
709
+ safe zone**; and **exactly one of four backgrounds** — `outline`, `plain`, active-word
710
+ `spotlight`/`karaoke`, or a tight `highlight-solid` band.
711
+
712
+ > ⚠️ I had this format running at **76px**, outside the band. Measured against the same cut, 76px
713
+ > wrapped 2 of 18 pages and reached 85.5% of frame width; **64px wrapped none and reached 82%** —
714
+ > better on both counts. When a local harness and the core standard disagree, check the standard is
715
+ > not simply right first.
716
+
717
+ #### What this format ADDS to the standard
718
+
719
+ Three deltas, each earned, and each a *narrowing* of the standard rather than a departure from it:
720
+
721
+ 1. **Self-host the face; never `<link>` it, and set `font-display: block`.** The standard says use
722
+ an imported display font. It does not say how, and the obvious how is broken for rendering:
723
+ Google Fonts' own snippet ends in `&display=swap`, which draws a fallback first. A frame-capture
724
+ render screenshots inside that window — this shipped **Montserrat for all 666 frames** while the
725
+ stylesheet returned `200` and named TikTok Sans. Pull the `woff2`, `@font-face` it from disk,
726
+ `font-display: block`, and no fallback chain (a chain means a missing font becomes a *different*
727
+ font with nothing to tell you).
728
+ 2. **A tighter band than 8%–85%.** The safe zone says where text is *allowed*; the platform's own
729
+ UI covers part of it. For this format the working band is **y 15%–73%, x 14%–86%** — a subset,
730
+ not a contradiction. See *Subtitle placement* below for the chrome map.
731
+ 3. **Of the four legal backgrounds, this format uses `outline` and `plain`.** No phrase plate: the
732
+ `highlight-solid` band is legal everywhere but it is the third-party-editor tell on a
733
+ reaction-led cut. The active-word `pill` mechanic is the standard's `spotlight`/`karaoke`
734
+ highlight and is fine.
735
+
736
+ ⚠️ **One mechanic below is an extension, not a standard preset.** `underline` is not one of the four
737
+ legal backgrounds. It reads native and renders fine, but treat it as this harness's own and expect
738
+ `vidfarm qa` to have no opinion on it — if a director wants to stay strictly inside the core set,
739
+ use `none`, `reveal`, `colour`, `pop` or `pill`.
740
+
741
+ - **No plate. The text sits directly on the picture.** A translucent slab behind the phrase is the
742
+ loudest third-party-editor tell in the format — it is what makes a cut read as *made somewhere
743
+ else and uploaded*, and it gets worse the moment a page wraps to two lines, because the slab's
744
+ own edges become a second shape competing with the frame. TikTok's own captions have no plate.
745
+ Legibility comes from the letterform instead: **a dark stroke under the fill plus a soft diffuse
746
+ shadow.**
747
+ ```css
748
+ -webkit-text-stroke: 3px #000; paint-order: stroke fill; /* stroke OUTSIDE the glyph */
749
+ text-shadow: 0 3px 12px rgba(0,0,0,.42), 0 1px 3px rgba(0,0,0,.55);
750
+ ```
751
+ `paint-order: stroke fill` is the part people miss — without it the stroke paints *over* the
752
+ glyph and eats the face's counters, and TikTok Sans at weight 900 turns to mush.
753
+ - **Still measure the band — the measurement now sets the STROKE, not a plate.** Sample the
754
+ composited luma and variance in the caption band across every scene (Appendix B). A calm dark
755
+ band survives on shadow alone; a bright busy one (a white UI, a lit face) needs the full 3px.
756
+ One value for the whole video.
757
+ - **NOT the lower third. The lower third is under the platform's own UI.** Every platform pastes
758
+ chrome over your video that you never see in your own render:
759
+
760
+ | | blocked |
761
+ |---|---|
762
+ | top | 0–12% — status bar, "Following \| For You" tabs |
763
+ | bottom | 78–100% — username, post caption, music ticker, progress bar |
764
+ | right | 84–100% (roughly y 45–80%) — like / comment / share / profile rail |
765
+
766
+ So the **safe core is about y 15%–73%**, and the default every subtitle tool ships — a lower
767
+ third at y≈66 with a 14% box — puts its bottom half under the caption block. Put the caption box
768
+ at the **lowest position that still clears the chrome**: `--y 57` for a 14%-tall box.
769
+ - **Keep the box narrow enough that a long line can't reach the rail.** Centred text needs a
770
+ symmetric box, so `--x 14 --width 72` (14%–86%) is the widest that stays clear. Then size the
771
+ type so pages mostly fit on one line inside it — measure, don't guess:
772
+ ```js
773
+ // in the render browser: how many pages wrap, and how far right do they reach?
774
+ [...document.querySelectorAll('[data-layer-kind="caption"]')].map(c => {
775
+ const w = [...c.querySelectorAll('[data-cap-word]')]
776
+ return { lines: c.querySelector('[data-vf-text-inline]').getClientRects().length,
777
+ right: Math.max(...w.map(x => x.getBoundingClientRect().right)) / 10.8 }})
778
+ ```
779
+ Measured on the reference build: 88px wrapped 7 of 18 pages and reached 85.7%, 76px wrapped 2 and
780
+ reached 85.0%, **64px wrapped none and reached 82.0%.** Smaller type that fits beats bigger type
781
+ that wraps into the rail — and 64px is the top of the standard's band anyway.
782
+ - **Let the measurement rank the bands, but don't let it move the caption over a face.** Appendix G
783
+ scores every band in the safe core. It will usually report the TOP of the frame as calmest —
784
+ that reading means "forehead, hair and defocused background", and a caption there covers the
785
+ performance, which on this format is the content. Take the quiet band only when the subject
786
+ genuinely sits low in frame.
787
+ - **Never repeat what the app already says on screen.** If the subtitle and the UI render the same
788
+ words, drop the subtitle. The demo beat is allowed to be almost caption-free — that is correct,
789
+ not a bug.
790
+
791
+ ---
792
+
793
+ ## The rules
794
+
795
+ ### Rule 1 — the reaction and the app must be about the same moment
796
+
797
+ A surprised face cut against a beat where nothing surprising happened is worse than no cutaway. The
798
+ face is a *label* for the frame before it. Choose the reaction after you have cut the demo, never
799
+ before, and cut on the emotional beat rather than at a round timecode.
800
+
801
+ ### Rule 2 — align the demo CUT to the subtitle, not the subtitle to the demo
802
+
803
+ The subtitle is fixed by the beat table; the demo recording is a long take you own every frame of.
804
+ So when the subtitle names the query and the input is still empty, you do **not** move the cue —
805
+ you move the in-point. Lay one frame per second of the composited device clip into a strip, read
806
+ off the timecodes where the action actually happens (`input complete`, `result lands`,
807
+ `scroll begins`), and solve for the in-point:
808
+
809
+ ```bash
810
+ ffmpeg -i phone-demo.mp4 -vf "fps=1,scale=150:-1,tile=<N>x1" -frames:v 1 timeline.png
811
+ ```
812
+
813
+ ```
814
+ in_point = <device-clip time of the action> - (<composition time the cue fires> - <beat start>)
815
+ ```
816
+
817
+ Do this for the two anchors that matter — the moment the input is complete, and the moment the
818
+ result appears — and the whole beat locks. Getting this wrong is the defect the format produces
819
+ most often and the one an author is least likely to notice, because the subtitles read correctly
820
+ and the footage reads correctly; only the two together are wrong.
821
+
822
+ ### Rule 3 — the demo shows one action, at real speed, uncut
823
+
824
+ Speed-ramping dead time is fine and expected. Cutting away at the instant the product does its work
825
+ is not — that is the exact frame the viewer is deciding on. One feature per video. A video covering
826
+ four features teaches none.
827
+
828
+ ### Rule 4 — every number on screen was read off the demo footage
829
+
830
+ Prices, counts, distances, names: quote what the recording actually shows, exactly. This format
831
+ makes it easy to be honest, because the source of truth is playing in the same frame as the claim —
832
+ and easy to be caught, for the same reason. A number in the subtitle that contradicts the number in
833
+ the app is the defect this format produces most.
834
+
835
+ ### Rule 5 — concede one true, unflattering limit, in beat 6
836
+
837
+ Coverage gaps, geography, price, what it can't do yet. Take it from the product's own site so it is
838
+ defensible. Over the app, not on a card. It is not a disclaimer — it is the beat that makes the
839
+ other twenty seconds believable.
840
+
841
+ ### Rule 6 — claim the mechanism, never the outcome
842
+
843
+ Say what it does and what it costs. Don't promise the result. On money, health and appearance
844
+ topics it is also the claim that draws platform enforcement.
845
+
846
+ ### Rule 7 — client work: nothing on screen the customer doesn't already say
847
+
848
+ Every claim traceable to their own site or app. No real third party named in a negative light. Real
849
+ names, faces and emails visible in their screenshots were a deliberate decision or they don't ship.
850
+ If the customer's own example query is broken, tell them — don't quietly film a different one and
851
+ let them discover it in the comments.
852
+
853
+ ### Rule 8 — production floor
854
+
855
+ Butt cuts only · no logo, end card, URL or title card · no CTA button, pricing card, feature grid,
856
+ benefit chip row, frosted panel or caption plate · nothing on screen looks clickable · the ask is a subtitle line.
857
+
858
+ ---
859
+
860
+ ## Bulk-generation notes
861
+
862
+ The variant axis is **the query you type into the app**, paired with the reaction that matches it.
863
+ Same product, same beats, same crop, same bed — five different things a person was actually looking
864
+ for, and five faces that fit those five moments. That is five videos.
865
+
866
+ Varying the *reaction clip* alone gives you one video five times, and the algorithm treats
867
+ near-duplicates accordingly. Varying the *subtitle wording* alone is worse.
868
+
869
+ **Vary the caption treatment across the batch too** — mechanic, accent, weight, size. Not because
870
+ any one of them is better, but because thirty videos in identical captions read as one account
871
+ running a template, and that is exactly the read you are trying to avoid. It costs nothing: the
872
+ treatment is four flags.
873
+
874
+ Because beats 3–6 are one tracked composite, a variant costs one browser recording plus one
875
+ tracker pass — roughly two minutes of compute and $0. That is the reason this format is worth
876
+ building a harness for at all.
877
+
878
+ Use `vidfarm dedupe <file> --variants N` when the same cut posts to more than one account.
879
+
880
+ ---
881
+
882
+ ## Pre-flight checklist
883
+
884
+ **Sourcing**
885
+ - [ ] The reaction raws came off the `ugc-reaction` shelf and were chosen from a contact sheet, not from a description alone
886
+ - [ ] ONE actor carries every reaction beat, cast from their `actor_<uuid>` set — not four different faces under a first-person voiceover
887
+ - [ ] The actor was chosen for their ARC (a take per beat), not for one good-looking clip
888
+ - [ ] The actor id was read off a card in THIS build and confirmed to return that person's clips
889
+ - [ ] The shelf was paginated with `cursor`, not `offset` — you saw all of it, not page 1 twice
890
+ - [ ] Beats 1 and 2 are two different people
891
+ - [ ] The greenscreen raw's `sourceType` is `Display Greenscreen` — no identifiable celebrity or character is being keyed into a customer's ad
892
+ - [ ] The demo query was run by hand and confirmed to return results before filming
893
+ - [ ] The demo was recorded at the device screen's aspect ratio
894
+ - [ ] Every 16:9 reaction raw's 9:16 crop was chosen off a still, not centred by default
895
+
896
+ **The insert**
897
+ - [ ] The demo's in-point was solved against the cue times — the input completes while its own subtitle is up, and the result lands on the cue that names it
898
+ - [ ] The insert is tracked per frame — scrubbed at 3 points, the app does not slide under the bezel
899
+ - [ ] The insert is rotation-fitted — no green wedge down any edge, at any frame
900
+ - [ ] The thumb and the bezel are ON TOP of the app, not painted over
901
+ - [ ] The device beats are punched in, and the UI is readable at phone scale
902
+ - [ ] Every device beat uses the same crop
903
+
904
+ **Stickers** (skip if none — none is a valid answer)
905
+ - [ ] The product's own noun appears on screen at least twice — not just its UI
906
+ - [ ] Every sticker either annotates something on screen or supplies that missing subject; none is a badge, chip, icon or logo
907
+ - [ ] Any sticker that a viewer would have to puzzle over was cut — a confusing sticker is worse than none
908
+ - [ ] Cutouts were checked for partial alpha (hollow art → flat key, solid subject → smart key) and despilled
909
+ - [ ] Each sticker's `fit` matches its job — subject `contain`, annotation `fill` — checked on a frame, not in the markup
910
+ - [ ] One on screen at a time, and not on every beat
911
+ - [ ] Each sticker's window ends before the UI underneath it moves — checked frame by frame
912
+ - [ ] Hand-drawn, not geometric; wrapped in a `div.clip` so the engine shows and hides it
913
+
914
+ **Anatomy**
915
+ - [ ] Beat 1 is a situation, not a product claim; frame 0 works as a standalone thumbnail
916
+ - [ ] The product first appears in beat 3
917
+ - [ ] The result plays uncut and long enough to be believed
918
+ - [ ] The video cuts back to a face after the payoff, and back to the app after the face
919
+ - [ ] One honest limit, in beat 6, over the app
920
+ - [ ] One comment ask, in the final beat AND in the post caption
921
+ - [ ] Exactly one feature is covered
922
+
923
+ **Audio**
924
+ - [ ] The mode was chosen deliberately (silent, or narrated over a ducked bed)
925
+ - [ ] Every reaction raw is muted — no face is visibly saying words the subtitles don't
926
+ - [ ] The demo is muted; the customer's own narration does not ship
927
+ - [ ] Narrated: one pinned `--voice` and one `--style` string across every line
928
+ - [ ] Narrated: captions came from the VO transcript, and the DISPLAY text is the script, not the transcript
929
+ - [ ] Narrated: speech sits 12–15 dB over the bed, measured on the STEMS
930
+ - [ ] Narrated: the payoff beat has a stretch with no speech over it
931
+ - [ ] The bed is NOT in the composition — the render is the voiceover-only master
932
+ - [ ] Both cuts exported from one render; the `-vo` and `-music` video streams are identical (unless watermarked)
933
+ - [ ] The handoff names the `-vo` cut as the one to publish, with platform audio added at upload
934
+ - [ ] The bed is CC0 (or its attribution requirement was accepted deliberately)
935
+ - [ ] The bed was measured ON THE RENDER, not on the stem: ~-16 LUFS, peak < -1 dBFS (`data-volume` multiplies it)
936
+ - [ ] The handoff says whether the bed ships or gets replaced with platform audio at upload
937
+
938
+ **Subtitles**
939
+ - [ ] The caption box sits inside the platform-safe core (y 15–73%, x 14–86%) — NOT the lower third
940
+ - [ ] Page wrapping and right-edge reach were measured in the browser; nothing crosses into the action rail
941
+ - [ ] The core caption standard was followed (bundled font, ~36–64px, safe zone, one of the four backgrounds) — not a harness restatement of it
942
+ - [ ] The face is self-hosted with `font-display: block` and no fallback chain
943
+ - [ ] `document.fonts.check('900 88px "TikTok Sans"')` was asserted true in the render browser, not assumed
944
+ - [ ] A cue is legible at frame 0 (`data-start="0"`, not 0.15)
945
+ - [ ] ≤3–4 words per page, one cue on screen at a time, timed to the voice
946
+ - [ ] If the mechanic is accented, the words were confirmed to FIRE in the render — 7 consecutive frames from inside one cue, not identical (skip for `none`, where identical frames are correct)
947
+ - [ ] No plate, panel or translucent slab behind the captions — stroke + shadow only
948
+ - [ ] `paint-order: stroke fill` is set, so the stroke sits outside the glyph
949
+ - [ ] The active-word highlight is a colour swap, not a scale — no two words ever touch mid-pop
950
+ - [ ] The band was measured across every scene and ONE stroke weight was chosen for the whole video
951
+ - [ ] The caption treatment was CHOSEN for this video (mechanic, accent, weight, size, case) — not inherited from the last one
952
+ - [ ] The accent was contrast-checked against the measured band; if it failed, it moved to a `pill` background rather than staying as type colour
953
+ - [ ] In a batch: this video's caption treatment differs from its siblings
954
+ - [ ] No cue repeats words the app already has on screen in that frame
955
+ - [ ] Every number in a cue matches the number visible in the demo footage
956
+
957
+ **Whole-video review** — on the render, not the plan
958
+ - [ ] A contact sheet of ~12 stills was read as one image: one crop convention, one type scale, one accent colour
959
+ - [ ] The reaction beats and the device beats look like the same video — not two edits spliced
960
+ - [ ] Pacing is deliberate, not N identically-long beats; no join is jarring
961
+ - [ ] No frame rests empty >0.5s; nothing runs after the last word for more than ~1.2s
962
+ - [ ] Frames from two different scenes were compared (a frozen render passes duration and frame-count checks)
963
+ - [ ] Audio verified by measurement, not by "it sounds fine"
964
+ - [ ] The MP4's timestamp is newer than the last edit — you reviewed THIS cut, not the previous one
965
+
966
+ ---
967
+
968
+ ## Diagnosing a flop
969
+
970
+ | What the numbers say | Weak beat | Fix |
971
+ |---|---|---|
972
+ | Barely any views | 1 | The hook is a product claim, or frame 0 is a phone instead of a face |
973
+ | Views, mass exit at 3–6s | 3 | The app arrived before the situation landed, or the move is not legible |
974
+ | Watched through, no reaction | 4 | The result was described in a subtitle instead of shown playing |
975
+ | Good retention, dead comments | 7 | No ask, or the ask was "thoughts?" |
976
+ | Comments arguing it's fake | 5 / audio | A reaction that didn't match its frame, or an unmuted raw |
977
+
978
+ **Vary one beat at a time.** A batch where the query, the faces and the bed all changed at once
979
+ teaches you nothing.
980
+
981
+ ---
982
+
983
+ ## Appendix A — `track-insert.py`
984
+
985
+ Free, local, no account, numpy + ffmpeg only. Fits the demo to the device screen's own axes every
986
+ frame and pastes by mask.
987
+
988
+ ```python
989
+ #!/usr/bin/env python3
990
+ """
991
+ Composite an app-demo recording INTO the green screen of a device raw.
992
+
993
+ usage: track-insert.py <device.mp4> <demo.mp4> <out.mp4> [--demo-start SEC]
994
+ """
995
+ import subprocess, sys, numpy as np
996
+
997
+ device, demo, out = sys.argv[1], sys.argv[2], sys.argv[3]
998
+ demo_start = 0.0
999
+ if "--demo-start" in sys.argv:
1000
+ demo_start = float(sys.argv[sys.argv.index("--demo-start") + 1])
1001
+
1002
+ def probe(f, *keys):
1003
+ r = subprocess.run(["ffprobe", "-v", "error", "-select_streams", "v",
1004
+ "-show_entries", "stream=" + ",".join(keys),
1005
+ "-of", "default=nw=1:nk=1", f], capture_output=True, text=True)
1006
+ return r.stdout.split()
1007
+
1008
+ W, H, fr = probe(device, "width", "height", "r_frame_rate")
1009
+ W, H = int(W), int(H)
1010
+ num, den = fr.split("/")
1011
+ FPS = round(float(num) / float(den), 3)
1012
+
1013
+ PW, PH = 720, 1530 # demo pre-scaled ONCE with a good filter; per-frame we only
1014
+ # nearest-resample down a few percent, so there is no shimmer
1015
+ demo_pipe = subprocess.Popen(
1016
+ ["ffmpeg", "-v", "error", "-ss", str(demo_start), "-i", demo,
1017
+ "-vf", f"fps={FPS},scale={PW}:{PH}:flags=lanczos", "-pix_fmt", "rgb24",
1018
+ "-f", "rawvideo", "-"], stdout=subprocess.PIPE)
1019
+ dev_pipe = subprocess.Popen(
1020
+ ["ffmpeg", "-v", "error", "-i", device, "-pix_fmt", "rgb24", "-f", "rawvideo", "-"],
1021
+ stdout=subprocess.PIPE)
1022
+ enc = subprocess.Popen(
1023
+ ["ffmpeg", "-y", "-v", "error", "-f", "rawvideo", "-pix_fmt", "rgb24",
1024
+ "-s", f"{W}x{H}", "-r", str(FPS), "-i", "-", "-an",
1025
+ "-c:v", "libx264", "-pix_fmt", "yuv420p", "-crf", "18", out],
1026
+ stdin=subprocess.PIPE)
1027
+
1028
+ def dilate(m, n=2):
1029
+ for _ in range(n):
1030
+ m = (m | np.roll(m, 1, 0) | np.roll(m, -1, 0)
1031
+ | np.roll(m, 1, 1) | np.roll(m, -1, 1))
1032
+ return m
1033
+
1034
+ FSZ, DSZ = W * H * 3, PW * PH * 3
1035
+ fit = None
1036
+ last_demo = None
1037
+ n = 0
1038
+ while True:
1039
+ raw = dev_pipe.stdout.read(FSZ)
1040
+ if len(raw) < FSZ:
1041
+ break
1042
+ d = demo_pipe.stdout.read(DSZ)
1043
+ if len(d) < DSZ:
1044
+ d = last_demo # demo ran out — hold its last frame
1045
+ last_demo = d
1046
+ f = np.frombuffer(raw, np.uint8).reshape(H, W, 3).astype(np.float32)
1047
+ dm = np.frombuffer(d, np.uint8).reshape(PH, PW, 3).astype(np.float32)
1048
+
1049
+ r, g, b = f[..., 0], f[..., 1], f[..., 2]
1050
+ greenness = g - np.maximum(r, b)
1051
+ core = (greenness > 55) & (g > 80)
1052
+
1053
+ if core.sum() > 5000:
1054
+ ys, xs = np.nonzero(core)
1055
+ cx, cy = xs.mean(), ys.mean()
1056
+ cov = np.cov(np.stack([xs - cx, ys - cy])) # the screen's own axes
1057
+ w_, v_ = np.linalg.eigh(cov)
1058
+ major = v_[:, np.argmax(w_)]
1059
+ if major[1] < 0:
1060
+ major = -major
1061
+ sn, cs = major[0], major[1]
1062
+ px, py = xs - cx, ys - cy
1063
+ v = px * sn + py * cs # along the long axis
1064
+ u = px * cs - py * sn # across it
1065
+ lenV = (np.percentile(v, 99.5) - np.percentile(v, 0.5)) / 2
1066
+ lenU = (np.percentile(u, 99.5) - np.percentile(u, 0.5)) / 2
1067
+ expect = lenV * (PW / PH) # a thumb eats one side, so trust the LONG axis
1068
+ if lenU < expect * 0.92:
1069
+ lenU = expect
1070
+ fit = (cx, cy, cs, sn, lenU, lenV)
1071
+
1072
+ if fit is None:
1073
+ enc.stdin.write(f.astype(np.uint8).tobytes()); n += 1; continue
1074
+ cx, cy, cs, sn, lenU, lenV = fit
1075
+
1076
+ alpha = np.clip((greenness - 25) / 30.0, 0, 1)
1077
+ paint = dilate(core, 2) | (alpha > 0.15) # paste by MASK, not by box
1078
+ ys, xs = np.nonzero(paint)
1079
+ if len(xs) == 0:
1080
+ enc.stdin.write(f.astype(np.uint8).tobytes()); n += 1; continue
1081
+
1082
+ px, py = xs - cx, ys - cy
1083
+ vv = px * sn + py * cs
1084
+ uu = px * cs - py * sn
1085
+ sx = ((uu / lenU + 1) * 0.5 * (PW - 1)).astype(np.int32).clip(0, PW - 1)
1086
+ sy = ((vv / lenV + 1) * 0.5 * (PH - 1)).astype(np.int32).clip(0, PH - 1)
1087
+ f[ys, xs] = dm[sy, sx] # opaque: no green survives
1088
+
1089
+ enc.stdin.write(np.clip(f, 0, 255).astype(np.uint8).tobytes())
1090
+ n += 1
1091
+
1092
+ enc.stdin.close(); enc.wait()
1093
+ print(f"{out} {n} frames @ {FPS}fps ({n/FPS:.1f}s)")
1094
+ ```
1095
+
1096
+ **Tuning it.** If the device raw's plate is a darker or bluer green, drop `greenness > 55` toward
1097
+ 40. If the demo comes out mirrored or 90° out, the PCA picked the short axis — that only happens on
1098
+ a landscape device, and the fix is to swap `PW`/`PH`. If the screen is only *partly* green (a
1099
+ notch, a bezel reflection), raise the dilation from 2 to 3.
1100
+
1101
+ ## Appendix B — measuring the caption band
1102
+
1103
+ Run this before you choose a caption treatment, on every scene, and put the numbers in the handoff.
1104
+
1105
+ ```python
1106
+ import subprocess, sys, numpy as np
1107
+ for f in sys.argv[1:]:
1108
+ p = subprocess.run(["ffmpeg","-v","error","-i",f,"-vf","fps=1,scale=270:480",
1109
+ "-pix_fmt","rgb24","-f","rawvideo","-"], capture_output=True)
1110
+ b = np.frombuffer(p.stdout, np.uint8); n = len(b)//(270*480*3)
1111
+ fr = b[:n*270*480*3].reshape(n,480,270,3).astype(np.float32)
1112
+ luma = 0.2126*fr[...,0] + 0.7152*fr[...,1] + 0.0722*fr[...,2]
1113
+ band = luma[:, 298:374, :] # the 62-78% lower third
1114
+ print(f"{f} luma={band.mean():.1f} var={band.std():.1f}")
1115
+ ```
1116
+
1117
+ | Band | Treatment |
1118
+ |---|---|
1119
+ | luma < ~70, var < ~42 | White type, shadow only — the stroke can drop to 1–2px |
1120
+ | luma > ~160, var < ~42 | White type, full 3px stroke; the shadow does the separating |
1121
+ | anything else | White type, full 3px stroke + shadow — this format is nearly always this row |
1122
+
1123
+ **Never a plate.** The measurement chooses how much stroke, not whether to put a box behind the
1124
+ words. See *Subtitle DNA*.
1125
+
1126
+ ## Appendix C — `srt-to-cues.py` (one caption layer per SRT cue)
1127
+
1128
+ Use this instead of `vidfarm captions generate --srt` whenever the cue times are pinned to cuts.
1129
+ It emits the same word-pop markup the CLI emits, and fixes the two things that bite in this format:
1130
+ it **keeps every SRT cue as its own page** (no word crosses a cut), and it **carries its own
1131
+ `@keyframes`** so the word-by-word pop is a property of the file rather than of whichever render
1132
+ engine version picks it up.
1133
+
1134
+ ```python
1135
+ #!/usr/bin/env python3
1136
+ """
1137
+ Write word-pop caption layers into a HyperFrames composition, ONE LAYER PER SRT CUE.
1138
+
1139
+ `vidfarm captions generate --srt` flattens the SRT into a single word stream and
1140
+ re-pages it by --max-words, so words hop across cue boundaries ("the dish instead"
1141
+ + "chili crisp" becomes "dish instead chili crisp"). When your cue times are pinned
1142
+ to CUTS rather than to speech, that is fatal. This emits the same markup the CLI
1143
+ emits, but keeps each SRT cue as its own page.
1144
+
1145
+ usage: srt-to-cues.py <composition.html> <captions.srt|cues.json>
1146
+ [--y 57] [--x 14] [--width 72] [--font-size 64] [--weight 900]
1147
+ [--highlight none|reveal|colour|pop|pill|underline] [--active-color '#FFD500']
1148
+ [--font 'TikTok Sans'] [--plate none|phrase] [--stroke 3]
1149
+ [--no-uppercase] [--track 5]
1150
+
1151
+ WHAT IS FIXED vs WHAT VARIES. Native means the FONT is TikTok Sans, the captions
1152
+ sit on the picture inside the platform-safe core, and the highlight fires
1153
+ word-by-word. It does NOT mean one colour, one case, one size and one mechanic on
1154
+ every video a studio ships — that is a house style, and at volume it reads as a
1155
+ content farm. Vary --highlight / --active-color / --weight / --font-size / --y
1156
+ between videos; hold one combination for the length of a single video.
1157
+ """
1158
+ import re, sys, html, json
1159
+
1160
+ args = sys.argv[1:]
1161
+ comp, srt = args[0], args[1]
1162
+ def opt(name, default):
1163
+ return args[args.index(name) + 1] if name in args else default
1164
+ Y = opt("--y", "66")
1165
+ X = opt("--x", "8")
1166
+ WIDTH = opt("--width", "84")
1167
+ FONT_SIZE = opt("--font-size", "64") # core standard: ~36-64px on a 1080 frame
1168
+ ACTIVE = opt("--active-color", "#FFD500")
1169
+ TRACK = opt("--track", "5")
1170
+ FONT = opt("--font", "TikTok Sans")
1171
+ WEIGHT = opt("--weight", "900") # 700 | 800 | 900
1172
+ HILITE = opt("--highlight", "colour") # none | reveal | colour | pop | pill | underline
1173
+ PLATE_MODE = opt("--plate", "none") # none | phrase
1174
+ STROKE = opt("--stroke", "3") # px of dark outline, plate-free mode
1175
+
1176
+ # TikTok's own captions sit DIRECTLY on the picture. A translucent slab behind
1177
+ # the phrase is a third-party-editor tell — it is the thing that makes a cut read
1178
+ # as "made somewhere else and uploaded", and it gets worse the moment two lines
1179
+ # wrap, because the slab's own edges become a shape competing with the frame.
1180
+ # Legibility without a plate comes from the letterform itself: a dark stroke
1181
+ # under the fill (paint-order keeps it outside the glyph, so the face stays
1182
+ # crisp) plus a soft diffuse shadow to lift it off a busy background.
1183
+ if PLATE_MODE == "phrase":
1184
+ PLATE = ("padding:0.07em 0.46em 0.09em;border-radius:0.32em;"
1185
+ "background:rgba(0,0,0,0.34);")
1186
+ BG_STYLE = "highlight-translucent"
1187
+ else:
1188
+ PLATE = "padding:0;background:transparent;"
1189
+ BG_STYLE = "outline"
1190
+ # NOTE: no "sans-serif" / Montserrat in the stack on purpose. A fallback chain
1191
+ # means a missing font renders as a DIFFERENT font and nothing tells you — the
1192
+ # cut just quietly stops looking native. Self-host the face (see the @font-face
1193
+ # in the composition head) so there is nothing to fall back FROM.
1194
+ UPPER = "--no-uppercase" not in args
1195
+
1196
+ def tsec(t):
1197
+ h, m, rest = t.split(":")
1198
+ s, ms = rest.split(",")
1199
+ return int(h) * 3600 + int(m) * 60 + int(s) + int(ms) / 1000
1200
+
1201
+ # Two input shapes. An SRT carries page text only, so word timings are ESTIMATED
1202
+ # from word length — right for a silent, subtitle-only cut where the cues are
1203
+ # pinned to picture. A cues.json (from vo-align.py) carries a real time per word,
1204
+ # which is what you must use once a voiceover exists: a karaoke highlight that is
1205
+ # 200ms off the voice is more distracting than no highlight at all.
1206
+ cues = []
1207
+ if srt.endswith(".json"):
1208
+ for p in json.load(open(srt)):
1209
+ cues.append((p["start"], p["end"],
1210
+ " ".join(w["text"] for w in p["words"]), p["words"]))
1211
+ else:
1212
+ for block in re.split(r"\n\s*\n", open(srt).read().strip()):
1213
+ lines = [l for l in block.strip().splitlines() if l.strip()]
1214
+ if len(lines) < 2:
1215
+ continue
1216
+ tl = next((l for l in lines if "-->" in l), None)
1217
+ if not tl:
1218
+ continue
1219
+ a, b = [x.strip() for x in tl.split("-->")]
1220
+ text = " ".join(lines[lines.index(tl) + 1:]).strip()
1221
+ if text:
1222
+ cues.append((tsec(a), tsec(b), text, None))
1223
+ if cues:
1224
+ cues[0] = (0.0, cues[0][1], cues[0][2], cues[0][3]) # frame 0 must carry text
1225
+
1226
+ def layer(i, start, dur, text, timed=None):
1227
+ words = text.split()
1228
+ # spread the cue's time across its words by length, with a floor so short
1229
+ # words still get a readable beat
1230
+ weights = [max(len(w), 2) for w in words]
1231
+ total = sum(weights)
1232
+ # animation-delay is measured from PAGE time, not from the layer's own start.
1233
+ # A cue at 9.1s with a relative delay has already finished animating by the
1234
+ # time it becomes visible, and every word renders in its end state — which
1235
+ # looks exactly like "the animation is broken". So the delay is ABSOLUTE
1236
+ # (cue start + word offset). data-word-start/end stay relative, which is the
1237
+ # convention the CLI's own markup uses.
1238
+ #
1239
+ # The highlight is COLOUR plus a small vertical lift — deliberately no
1240
+ # horizontal scale. A scaling word grows from its centre into its neighbour:
1241
+ # "a restaurant" renders as "arestaurant" on the exact beat the word pops,
1242
+ # and a text stroke eats the remaining gap from both sides. Padding it apart
1243
+ # only trades one defect (collision) for another (letterspaced display type
1244
+ # that no longer reads as speech). TikTok's own word highlight is a colour
1245
+ # swap; keeping it that way makes the collision impossible instead of tuned.
1246
+ t, spans = 0.0, []
1247
+ for j, (w, wt) in enumerate(zip(words, weights)):
1248
+ if timed: # real word times from the voiceover
1249
+ a0 = timed[j]["start"] - start
1250
+ d = max(timed[j]["end"] - timed[j]["start"], 0.12)
1251
+ else: # estimated from word length
1252
+ a0, d = t, dur * wt / total
1253
+ # A highlight should last about as long as the word is spoken — except a
1254
+ # plain REVEAL, which is an on-switch, not a fade: stretched across a
1255
+ # 0.4s word it reads as the caption struggling to load.
1256
+ adur = 0.12 if HILITE == "reveal" else d
1257
+ spans.append(
1258
+ f'<span data-cap-word="true" data-word-start="{a0:.3f}" '
1259
+ f'data-word-end="{a0+d:.3f}" style="animation-delay:{start+a0:.3f}s;'
1260
+ f'animation-duration:{adur:.3f}s">{html.escape(w)}</span>')
1261
+ t += dur * wt / total
1262
+ inner = " ".join(spans)
1263
+ upper = "text-transform:uppercase;" if UPPER else ""
1264
+ return (
1265
+ f'<div class="clip" id="element_cap_srt_{i}" data-hf-id="element_cap_srt_{i}" '
1266
+ f'data-layer-mode="publish" data-layer-kind="caption" data-start="{start:.3f}" '
1267
+ f'data-duration="{dur:.3f}" data-track-index="{TRACK}" '
1268
+ f'data-label="{html.escape(text, quote=True)}" '
1269
+ f'data-text-background-style="{BG_STYLE}" data-text-background-color="#000000" '
1270
+ f'data-font-family="{FONT}" data-caption-animation="word-pop" '
1271
+ f'data-caption-uppercase="{1 if UPPER else 0}" '
1272
+ f'style="position:absolute;left:{X}%;top:{Y}%;width:{WIDTH}%;height:14%;z-index:{TRACK};'
1273
+ f'opacity:1;display:flex;align-items:center;justify-content:center;padding:3px;'
1274
+ f"font-family:'{FONT}', 'Noto Color Emoji';font-weight:{WEIGHT};"
1275
+ f'line-height:1.18;text-align:center;font-size:{FONT_SIZE}px;color:#ffffff;'
1276
+ f'background:transparent;--vf-cap-active:{ACTIVE};{upper}overflow:visible">'
1277
+ f'<span data-vf-text-lines="true" style="display:block;max-width:100%;text-align:inherit;'
1278
+ f'line-height:1.18;white-space:pre-wrap;overflow:visible">'
1279
+ f'<span data-vf-text-inline="true" style="display:inline;box-decoration-break:clone;'
1280
+ f'-webkit-box-decoration-break:clone;line-height:inherit;{PLATE}'
1281
+ f"font-family:'{FONT}', 'Noto Color Emoji';font-weight:{WEIGHT};color:#ffffff\">"
1282
+ f'{inner}</span></span></div>')
1283
+
1284
+ doc = open(comp).read()
1285
+ # The word-by-word pop itself. The render engine is SUPPOSED to drive this off
1286
+ # data-caption-animation + the per-word animation-delay/duration, and on some
1287
+ # versions it silently does not — the cue renders as one static block and the
1288
+ # video ships without the thing that makes it read as native. Owning the
1289
+ # keyframes here makes the animation a property of the file, not of the engine.
1290
+ #
1291
+ # The whole phrase is visible for the cue's whole life and the ACTIVE word pops
1292
+ # — that is TikTok's convention, and it is also the only version that survives a
1293
+ # single shared plate behind the phrase. A progressive fade-in leaves an empty
1294
+ # translucent box sitting on screen ahead of the words, which reads as a loading
1295
+ # state. So: fill-mode FORWARDS, base state is the settled state, and the
1296
+ # keyframes describe only the word's own moment.
1297
+ LEGIBILITY = ("" if PLATE_MODE == "phrase" else
1298
+ f"-webkit-text-stroke:{STROKE}px #000;paint-order:stroke fill;"
1299
+ "text-shadow:0 3px 12px rgba(0,0,0,.42),0 1px 3px rgba(0,0,0,.55);")
1300
+
1301
+ # ── The highlight mechanic ───────────────────────────────────────────────────
1302
+ # All four are TikTok-native; none of them is "the" house style. Vary this
1303
+ # ACROSS a batch and hold one for the length of a single video.
1304
+ #
1305
+ # Every one of them is engineered to take NO HORIZONTAL SPACE, because an
1306
+ # inline-block word that grows sideways eats the gap beside it and "a restaurant"
1307
+ # renders as "arestaurant" on the exact beat the word fires. `pop` scales on Y
1308
+ # only for that reason; `pill` and `underline` sit inside the word's own box.
1309
+ _HILITE = {
1310
+ "colour": """
1311
+ 0%{transform:translateY(-.05em);color:var(--vf-cap-active,#FFD500)}
1312
+ 45%{transform:translateY(0);color:var(--vf-cap-active,#FFD500)}
1313
+ 80%{transform:translateY(0);color:var(--vf-cap-active,#FFD500)}
1314
+ 100%{transform:translateY(0);color:#ffffff}""",
1315
+ "pop": """
1316
+ 0%{transform:scaleY(1.18) translateY(-.04em);color:var(--vf-cap-active,#FFD500)}
1317
+ 40%{transform:scaleY(1.05) translateY(0);color:var(--vf-cap-active,#FFD500)}
1318
+ 75%{transform:scaleY(1);color:var(--vf-cap-active,#FFD500)}
1319
+ 100%{transform:scaleY(1);color:#ffffff}""",
1320
+ "pill": """
1321
+ 0%{background:var(--vf-cap-active,#FFD500);color:#111;transform:translateY(-.04em)}
1322
+ 70%{background:var(--vf-cap-active,#FFD500);color:#111;transform:translateY(0)}
1323
+ 100%{background:transparent;color:#ffffff;transform:translateY(0)}""",
1324
+ "underline": """
1325
+ 0%{box-shadow:inset 0 -.10em 0 0 var(--vf-cap-active,#FFD500);color:#ffffff}
1326
+ 70%{box-shadow:inset 0 -.22em 0 0 var(--vf-cap-active,#FFD500);color:#ffffff}
1327
+ 100%{box-shadow:inset 0 -.22em 0 0 transparent;color:#ffffff}""",
1328
+ # PLAIN — no accent at all. This is TikTok's own auto-caption look, and it is
1329
+ # the right choice more often than the coloured ones: over busy footage, over a
1330
+ # UI that already has its own colours, and on any beat where the words should
1331
+ # carry no emphasis of their own. "reveal" is still word-timed (each word turns
1332
+ # on at its own moment); "none" turns the whole page on at once and lets the
1333
+ # page turns do the work.
1334
+ "reveal": """
1335
+ 0%{opacity:0}
1336
+ 100%{opacity:1}""",
1337
+ "none": None,
1338
+ }
1339
+ if HILITE not in _HILITE:
1340
+ sys.exit(f"--highlight must be one of {', '.join(_HILITE)}")
1341
+ # the pill needs its own breathing room and a radius; the others must not have it
1342
+ _PILL_BOX = ("padding:.02em .16em;border-radius:.18em;" if HILITE == "pill"
1343
+ else "padding:0 .02em;")
1344
+
1345
+ if _HILITE[HILITE] is None:
1346
+ # plain, page-timed: no per-word animation at all
1347
+ POP_CSS = f"""
1348
+ [data-cap-word]{{display:inline-block;{_PILL_BOX}{LEGIBILITY}}}
1349
+ """
1350
+ else:
1351
+ _EASE = "linear" if HILITE == "reveal" else "cubic-bezier(.2,.9,.3,1.1)"
1352
+ _FILL = "both" if HILITE == "reveal" else "forwards"
1353
+ POP_CSS = f"""
1354
+ [data-cap-word]{{display:inline-block;{_PILL_BOX}animation-name:vfWordPop;
1355
+ {LEGIBILITY}
1356
+ animation-fill-mode:{_FILL};animation-timing-function:{_EASE}}}
1357
+ @keyframes vfWordPop{{{_HILITE[HILITE]}
1358
+ }}
1359
+ """
1360
+ def strip_cues(h):
1361
+ """Remove existing caption layers by BALANCING <div> tags — a regex that
1362
+ guesses the closing sequence silently leaves the old cues in place, and you
1363
+ end up rendering two caption tracks stacked on each other."""
1364
+ out, i = [], 0
1365
+ while True:
1366
+ j = h.find('<div class="clip" id="element_cap_', i)
1367
+ if j < 0:
1368
+ out.append(h[i:]); break
1369
+ out.append(h[i:j])
1370
+ depth, k = 0, j
1371
+ while k < len(h):
1372
+ if h.startswith("<div", k):
1373
+ depth += 1; k += 4
1374
+ elif h.startswith("</div>", k):
1375
+ depth -= 1; k += 6
1376
+ if depth == 0:
1377
+ break
1378
+ else:
1379
+ k += 1
1380
+ i = k
1381
+ return "".join(out)
1382
+
1383
+ doc = strip_cues(doc)
1384
+ # strip any PREVIOUS pop CSS by shape, not by its exact text — otherwise tuning
1385
+ # the keyframes appends a second copy and the older rule wins the cascade
1386
+ doc = re.sub(r'\n?\[data-cap-word\]\{.*?@keyframes vfWordPop\{.*?\}\}\n?', "", doc, flags=re.S)
1387
+ if "</style>" in doc:
1388
+ doc = doc.replace("</style>", POP_CSS + "</style>", 1)
1389
+ else:
1390
+ doc = doc.replace("</head>", f"<style>{POP_CSS}</style></head>", 1)
1391
+ blocks = "\n".join(layer(i, a, b - a, t, wd) for i, (a, b, t, wd) in enumerate(cues))
1392
+ idx = doc.rfind("</div>")
1393
+ doc = doc[:idx] + blocks + "\n" + doc[idx:]
1394
+ open(comp, "w").write(doc)
1395
+ print(f"wrote {len(cues)} caption cues, one per SRT cue")
1396
+ for a, b, t, _ in cues:
1397
+ print(f" {a:6.2f} -> {b:6.2f} {t}")
1398
+ ```
1399
+
1400
+ ## Appendix D — rendering notes
1401
+
1402
+ - **Render with the standalone `hyperframes` CLI, not only `vidfarm render local`.** The copy of
1403
+ hyperframes bundled inside `vidfarm-devcli` can lag the released one, and a stale engine can abort
1404
+ the render at *Extracting video frames* with `captured 0 of expected N frames` on media that is
1405
+ perfectly valid. Check both versions (`hyperframes doctor`, `vidfarm doctor`) before you spend an
1406
+ hour re-encoding footage that was never the problem.
1407
+ - **Pass `--video-frame-format png` on this format.** Every device beat is a UI recording, and JPEG
1408
+ frame extraction softens small type exactly where the payoff lives.
1409
+ - **Hand-authored `<video class="clip">` layers render fine and grade as zero.** `vidfarm qa` counts
1410
+ scenes off `data-layer-kind`, so a composition you wrote by hand reports `scenes: 0` and
1411
+ `first_frame_visual: black` while rendering perfectly. Give every media layer the attributes
1412
+ `vidfarm place` writes — `data-layer-kind="video"|"audio"`, `data-layer-mode="publish"`,
1413
+ `data-hf-id`, `data-end`, `data-label`, `data-volume` — and the same file grades correctly. This
1414
+ is a reporting bug in your markup, not in the video, and it is easy to "fix" in the wrong place.
1415
+ - `hyperframes render` reads `index.html`; `vidfarm` writes and lints `composition.html`. Keep the
1416
+ two in sync (`cp composition.html index.html`) or you will render a stale cut.
1417
+ - **`[sub_timeline_readiness_timeout] … did not become ready within 45000ms` is benign here.** This
1418
+ format has no GSAP timeline — the motion is CSS keyframes and video playback — so there is no
1419
+ sub-timeline to become ready. The render completes and is correct. Don't chase it.
1420
+ - **Check the output file's timestamp, not the exit code.** A backgrounded render can exit 0 having
1421
+ written nothing. `ls -la` the MP4 and confirm it is newer than your last edit before you review
1422
+ it — otherwise you will grade the previous cut and conclude your fix didn't work.
1423
+
1424
+ ## Appendix E — a filled-in beat table
1425
+
1426
+ One 22.2s cut built end to end against this file, for a dish-search product with 140 restaurants
1427
+ indexed. It is here so the beat table above has a shape, not so you copy the words.
1428
+
1429
+ | # | Beat | t | Stream | Subtitle |
1430
+ |---|---|---|---|---|
1431
+ | 1 | Hook | 0.0–2.6 | Reaction — woman, hand on head, confused | "I didn't want" / "a restaurant" |
1432
+ | 2 | Wrong way | 2.6–4.8 | Reaction — man at a laptop, disbelief | "Google gave me" / "a list of fifty" |
1433
+ | 3 | Move | 4.8–9.1 | Device — search field, typing | "So I searched" / "the dish instead" / "chili crisp" |
1434
+ | 4 | Result | 9.1–14.6 | Device — results list | "It read the menus" / "not the reviews" / "I just picked one" |
1435
+ | 5 | Face | 14.6–16.8 | Reaction — hand over mouth, surprise | "Nine dollars" / "Moody Tongue" |
1436
+ | 6 | Scope | 16.8–19.4 | Device — results still scrolling | "NYC only" / "140 restaurants" |
1437
+ | 7 | Bait | 19.4–22.2 | Reaction — smiling | "Comment the dish" / "you can never find" |
1438
+
1439
+ Three things in that table are the harness working rather than the writer working:
1440
+
1441
+ - **Beat 4 says "I just picked one" instead of naming the dish and the price.** The app is showing
1442
+ the dish and the price in that same frame. Naming them would be the caption reading the screen
1443
+ aloud. The read-out is deferred to **beat 5**, where the footage is a face and there is nothing
1444
+ to duplicate — which is also where it lands hardest.
1445
+ - **Beat 6 is the customer's own published number, used against them.** "140 restaurants" is a
1446
+ small index and saying so out loud is what makes the previous fifteen seconds credible.
1447
+ - **Beat 2 characterises a search result, it does not name a competitor.** "A list of fifty" is a
1448
+ true thing about what happens; naming a named product doing something bad is Rule 7.
1449
+
1450
+ ## Appendix F — `vo-align.py` (captions from the voiceover, in narrated mode)
1451
+
1452
+ Once a voice exists, captions are no longer a design decision — they are a synchronisation problem.
1453
+ This transcribes each beat's VO clip locally (free, keyless whisper), places it at that beat, and
1454
+ emits `cues.json`: balanced pages of ≤3 words, each word carrying an absolute timestamp. Feed that
1455
+ straight to Appendix C in place of the SRT.
1456
+
1457
+ ```bash
1458
+ # beats.txt — one line per beat: <clip.wav>|<place-at-seconds>|<the script you fed the TTS>
1459
+ python3 vo-align.py vo/beats.txt cues.json --max-words 3
1460
+ python3 srt-to-cues.py composition.html cues.json --font "TikTok Sans" --font-size 88 --y 66
1461
+ ```
1462
+
1463
+ **Timing from the transcript, TEXT from the script.** This split is the whole design. Local STT
1464
+ mishears proper nouns precisely where this format needs them — a real run turned
1465
+ *"Nine dollars. Moody Tongue."* into *"$9.00 multi-tongue"* — and a caption that ships the
1466
+ mishearing is worse than no caption. So the words on screen are always the ones you wrote for the
1467
+ TTS; only their timings come from what was heard. Where the transcript's word count matches the
1468
+ script's, the two are zipped word-for-word and the sync is exact (6 of 7 beats on the reference
1469
+ build); where it doesn't, the script's words are distributed across the *measured* speech span,
1470
+ which is still far better than guessing. The script prints which path each beat took — read it.
1471
+
1472
+ **Write the script line the way you want it on screen.** "140 restaurants" and "a hundred forty
1473
+ restaurants" sound identical and page very differently; the first is two words and one clean page,
1474
+ the second breaks as "A HUNDRED AND / FORTY RESTAURANTS". You control this for free at the moment
1475
+ you write the TTS input, and not at all afterwards.
1476
+
1477
+ **Balance pages inside a sentence.** Fixed-size chunking leaves one-word orphans — a 7-word
1478
+ sentence at 3 words a page gives `3,3,1`, and a single word alone for 400ms reads as a flicker
1479
+ rather than as emphasis. Split into `ceil(n/max)` pages of as-equal-as-possible size instead
1480
+ (`3,2,2`), and never let a page run across a full stop.
1481
+
1482
+ ```python
1483
+ #!/usr/bin/env python3
1484
+ """
1485
+ Turn per-beat voiceover clips into caption pages with REAL word timings.
1486
+
1487
+ With a voiceover in the mix, captions stop being a design decision and become a
1488
+ synchronisation problem: they must be verbatim and they must sit on the word.
1489
+
1490
+ Two halves, and keeping them apart is the whole point:
1491
+
1492
+ * TIMING comes from the transcript. Where the transcript's word count matches
1493
+ the script's, the two are zipped word for word — exact. Where it doesn't,
1494
+ the script's words are distributed across the MEASURED speech span, which is
1495
+ still far better than guessing at a duration.
1496
+ * TEXT comes from the SCRIPT you fed the TTS, never from the transcript. Local
1497
+ STT mishears proper nouns exactly where this format needs them most — a real
1498
+ run turned "Nine dollars. Moody Tongue." into "$9.00 multi-tongue" — and a
1499
+ caption that ships the mishearing is worse than no caption at all.
1500
+
1501
+ Input: a beats file, one line per beat: <clip.wav>|<place-at-seconds>|<script>
1502
+ Output: cues.json — pages of <=MAXW words, each word carrying an absolute time.
1503
+
1504
+ usage: vo-align.py <beats.txt> <out.json> [--max-words 3]
1505
+ """
1506
+ import json, re, subprocess, sys, os
1507
+
1508
+ beats_file, out_path = sys.argv[1], sys.argv[2]
1509
+ MAXW = int(sys.argv[sys.argv.index("--max-words") + 1]) if "--max-words" in sys.argv else 3
1510
+
1511
+ def transcribe(wav):
1512
+ base = os.path.splitext(wav)[0] + ".stt"
1513
+ subprocess.run(["vidfarm", "stt", wav, "--engine", "whisper", "--out", base],
1514
+ capture_output=True)
1515
+ return json.load(open(base + ".json"))["words"]
1516
+
1517
+ pages, report = [], []
1518
+ for raw in open(beats_file):
1519
+ raw = raw.strip()
1520
+ if not raw or raw.startswith("#"):
1521
+ continue
1522
+ wav, at, script = [p.strip() for p in raw.split("|", 2)]
1523
+ at = float(at)
1524
+ heard = transcribe(wav)
1525
+ if not heard:
1526
+ continue
1527
+ span0, span1 = heard[0]["start"], heard[-1]["end"]
1528
+ script_words = script.split()
1529
+
1530
+ if len(heard) == len(script_words):
1531
+ timed = [{"text": s, "start": h["start"], "end": h["end"]}
1532
+ for s, h in zip(script_words, heard)]
1533
+ report.append(f"{os.path.basename(wav)}: exact ({len(heard)} words)")
1534
+ else:
1535
+ weights = [max(len(re.sub(r'\W', '', w)), 2) for w in script_words]
1536
+ total, t = sum(weights), span0
1537
+ timed = []
1538
+ for w, wt in zip(script_words, weights):
1539
+ d = (span1 - span0) * wt / total
1540
+ timed.append({"text": w, "start": t, "end": t + d})
1541
+ t += d
1542
+ report.append(f"{os.path.basename(wav)}: proportional "
1543
+ f"(heard {len(heard)}, script {len(script_words)})")
1544
+
1545
+ cur = []
1546
+ def flush():
1547
+ global cur
1548
+ if not cur:
1549
+ return
1550
+ pages.append({
1551
+ "start": round(at + cur[0]["start"], 3),
1552
+ "end": round(at + cur[-1]["end"], 3),
1553
+ "words": [{"text": re.sub(r'[,.!?;:]+$', "", w["text"]),
1554
+ "start": round(at + w["start"], 3),
1555
+ "end": round(at + w["end"], 3)} for w in cur],
1556
+ })
1557
+ cur = []
1558
+
1559
+ # Split into sentences first — a page that runs across a full stop reads as
1560
+ # one thought and isn't — then chunk each sentence into BALANCED pages.
1561
+ # Naive fixed-size chunking leaves one-word orphans ("...the reviews." ->
1562
+ # "It read the" / "menus not the" / "reviews"), and a single word alone on
1563
+ # screen for 400ms reads as a flicker, not as emphasis.
1564
+ # Break on COMMAS as well as full stops. A comma is the writer already telling
1565
+ # you where the thought divides, and honouring it is the difference between
1566
+ # "IT READ THE MENUS / NOT THE REVIEWS" and "IT READ THE / MENUS NOT / THE
1567
+ # REVIEWS" — same words, same timings, and only one of them is readable.
1568
+ fragments, s = [], []
1569
+ for w in timed:
1570
+ s.append(w)
1571
+ if re.search(r'[.!?,;:]$', w["text"]):
1572
+ fragments.append(s); s = []
1573
+ if s:
1574
+ fragments.append(s)
1575
+
1576
+ for frag in fragments:
1577
+ n = len(frag)
1578
+ # let a fragment run one word over the cap rather than split a clause
1579
+ # that fits: 4 words on one line beats 2 + 2 with a break mid-phrase
1580
+ k = 1 if n <= MAXW + 1 else -(-n // MAXW)
1581
+ base, extra = divmod(n, k) # spread the remainder over the first pages
1582
+ i = 0
1583
+ for p in range(k):
1584
+ size = base + (1 if p < extra else 0)
1585
+ cur = frag[i:i + size]
1586
+ i += size
1587
+ flush()
1588
+
1589
+ # a page whose last word is clipped short reads as a flicker; give every page a
1590
+ # floor, and never let one overlap the next
1591
+ for i, p in enumerate(pages):
1592
+ p["end"] = max(p["end"], p["start"] + 0.45)
1593
+ if i + 1 < len(pages):
1594
+ p["end"] = min(p["end"], pages[i + 1]["start"] - 0.02)
1595
+
1596
+ json.dump(pages, open(out_path, "w"), indent=1)
1597
+ for r in report:
1598
+ print(" " + r)
1599
+ print(f"{len(pages)} pages -> {out_path}")
1600
+ for p in pages:
1601
+ print(f" {p['start']:6.2f} -> {p['end']:6.2f} {' '.join(w['text'] for w in p['words'])}")
1602
+ ```
1603
+
1604
+ ## Appendix G — `caption-place.py` (where the captions go)
1605
+
1606
+ Scores every candidate band inside the platform-safe core and recommends one. Run it on the
1607
+ prepared beat clips, before you write a single cue.
1608
+
1609
+ ```bash
1610
+ python3 caption-place.py media/s*.mp4
1611
+ ```
1612
+
1613
+ ```python
1614
+ #!/usr/bin/env python3
1615
+ """
1616
+ Choose where the captions go: inside the platform-safe core, above the chrome.
1617
+
1618
+ THE CONSTRAINT THAT ACTUALLY DECIDES IT is not legibility — it is that TikTok,
1619
+ Reels and Shorts each paste their own UI over your video and you never see it in
1620
+ your own render:
1621
+
1622
+ top 0–12% status bar, "Following | For You" tabs
1623
+ bottom 78–100% username, post caption, music ticker, progress bar
1624
+ right 84–100% like / comment / share / profile rail (roughly y 45–80%)
1625
+
1626
+ So the safe core is about **y 15%–73%**, and the classic "lower third" — the
1627
+ default in every subtitle tool — is the single worst place to put text in
1628
+ vertical video. Half of it is under the caption block on at least one platform.
1629
+
1630
+ Inside the core, the default is the LOWEST band that still clears the chrome.
1631
+ That is where TikTok's own captions sit, and it is below the subject's face,
1632
+ which on this format is the content. Busyness is measured and reported, but it
1633
+ does not get to move the captions over someone's eyes: the top of the frame
1634
+ almost always scores "calmest" precisely because it is forehead, hair and
1635
+ defocused background, and a caption there covers the performance.
1636
+
1637
+ usage: caption-place.py <clip.mp4> [clip.mp4 ...] [--height 14] [--top 15] [--bottom 73]
1638
+ """
1639
+ import subprocess, sys, numpy as np
1640
+
1641
+ files = [a for a in sys.argv[1:] if not a.startswith("--")]
1642
+ def opt(n, d):
1643
+ return float(sys.argv[sys.argv.index(n) + 1]) if n in sys.argv else d
1644
+ H_PCT = opt("--height", 14)
1645
+ SAFE_TOP = opt("--top", 15)
1646
+ SAFE_BOTTOM = opt("--bottom", 73)
1647
+ AIR = 2.0 # don't sit flush against the chrome
1648
+
1649
+ W, H = 240, 426
1650
+ def frames(f):
1651
+ p = subprocess.run(["ffmpeg", "-v", "error", "-i", f, "-vf", f"fps=2,scale={W}:{H}",
1652
+ "-pix_fmt", "rgb24", "-f", "rawvideo", "-"], capture_output=True)
1653
+ b = np.frombuffer(p.stdout, np.uint8)
1654
+ n = len(b) // (W * H * 3)
1655
+ return b[:n * W * H * 3].reshape(n, H, W, 3).astype(np.float32)
1656
+
1657
+ F = np.concatenate([frames(f) for f in files])
1658
+ luma = 0.2126 * F[..., 0] + 0.7152 * F[..., 1] + 0.0722 * F[..., 2]
1659
+ # Busyness = local DETAIL behind the text, not global variance. A band split
1660
+ # between a bright wall and a dark jacket has huge variance and reads fine.
1661
+ gx = np.abs(np.diff(luma, axis=2))
1662
+ gy = np.abs(np.diff(luma, axis=1))
1663
+ edge = np.zeros_like(luma)
1664
+ edge[:, :, :-1] += gx
1665
+ edge[:, :-1, :] += gy
1666
+
1667
+ box_h = int(H_PCT / 100 * H)
1668
+ def score(top_pct):
1669
+ a = int(top_pct / 100 * H)
1670
+ band = edge[:, a:a + box_h, :]
1671
+ per_frame = band.mean(axis=(1, 2))
1672
+ # worst case matters more than the average: one scene where the caption is
1673
+ # unreadable is a defect however calm the other six are
1674
+ return per_frame.mean(), np.percentile(per_frame, 90)
1675
+
1676
+ cands = [(t, *score(t)) for t in np.arange(SAFE_TOP, SAFE_BOTTOM - H_PCT + 0.01, 1.0)]
1677
+ default_top = SAFE_BOTTOM - H_PCT - AIR
1678
+ d_avg, d_p90 = score(default_top)
1679
+ quietest = min(cands, key=lambda c: c[1] + 0.5 * c[2])
1680
+
1681
+ print(f"platform-safe core {SAFE_TOP:.0f}%–{SAFE_BOTTOM:.0f}% caption box {H_PCT:.0f}% tall")
1682
+ print(f" chrome assumed: top 0–12% bottom 78–100% right rail 84–100%")
1683
+ print()
1684
+ print(f" RECOMMENDED --y {default_top:.0f} busy {d_avg:5.2f} (p90 {d_p90:5.2f})"
1685
+ f" [lowest band clearing the chrome]")
1686
+ print(f" quietest --y {quietest[0]:.0f} busy {quietest[1]:5.2f} (p90 {quietest[2]:5.2f})")
1687
+ if quietest[1] < d_avg * 0.8:
1688
+ print()
1689
+ print(f" NOTE: the quietest band scores {(1-quietest[1]/d_avg)*100:.0f}% calmer. Look at it before")
1690
+ print( " taking it — near the top of the frame that reading usually means")
1691
+ print( " 'forehead, hair and defocused background', and a caption there covers")
1692
+ print( " the face. Only move up if the subject genuinely sits low in frame.")
1693
+ print()
1694
+ print(f" {'top%':>5} {'busy':>7} {'p90':>7}")
1695
+ for t, m, p in cands[::3]:
1696
+ mark = " <- recommended" if abs(t - default_top) < 0.5 else ""
1697
+ print(f" {t:5.0f} {m:7.2f} {p:7.2f}{mark}")
1698
+ ```
1699
+
1700
+ ## Appendix H — `export-versions.py` (the two cuts)
1701
+
1702
+ ```bash
1703
+ python3 export-versions.py renders/cut-vo.mp4 media/bgm.mp3 --bed-volume 0.22
1704
+ ```
1705
+
1706
+ ```python
1707
+ #!/usr/bin/env python3
1708
+ """
1709
+ Ship TWO cuts from ONE render: voiceover-only, and voiceover + music bed.
1710
+
1711
+ Why two, and why this order:
1712
+
1713
+ * The VOICEOVER-ONLY cut is the one you publish. TikTok, Reels and Shorts all
1714
+ let the poster attach a track from the platform's own library at upload,
1715
+ licensed through the platform's deals with the labels — so the music is
1716
+ cleared where viewers actually hear it, and the post gets whatever is
1717
+ trending that week instead of whatever was trending when you rendered.
1718
+ * The MUSIC cut is the review copy and the fallback: what you send a client to
1719
+ approve, what you post anywhere without an in-app music library, and the one
1720
+ that carries a watermark if you need one.
1721
+
1722
+ The music version is derived, never re-rendered. `-c:v copy` means both files
1723
+ carry the BYTE-IDENTICAL video stream, so approving one approves the other. Two
1724
+ separate renders would be two different videos that merely look alike, and the
1725
+ frame you approved would not be the frame you shipped. (A watermark breaks this
1726
+ on purpose — see --watermark.)
1727
+
1728
+ usage: export-versions.py <render.mp4> <bed.mp3> [--bed-volume 0.22]
1729
+ [--out-dir DIR] [--name NAME] [--watermark PNG]
1730
+ """
1731
+ import os, subprocess, sys
1732
+
1733
+ src, bed = sys.argv[1], sys.argv[2]
1734
+ def opt(n, d):
1735
+ return sys.argv[sys.argv.index(n) + 1] if n in sys.argv else d
1736
+ VOL = float(opt("--bed-volume", "0.22"))
1737
+ OUT = opt("--out-dir", os.path.dirname(src) or ".")
1738
+ NAME = opt("--name", os.path.splitext(os.path.basename(src))[0].removesuffix("-vo"))
1739
+ WM = opt("--watermark", None)
1740
+
1741
+ vo_path = os.path.join(OUT, f"{NAME}-vo.mp4")
1742
+ music_path = os.path.join(OUT, f"{NAME}-music.mp4")
1743
+
1744
+ def run(cmd):
1745
+ r = subprocess.run(cmd, capture_output=True, text=True)
1746
+ if r.returncode:
1747
+ sys.exit(f"failed: {' '.join(cmd)}\n{r.stderr[-1500:]}")
1748
+
1749
+ # 1. the publish master — untouched. If the render is already sitting at the
1750
+ # -vo path, leave it alone: ffmpeg refuses to read and write one file, and
1751
+ # "re-encode it to itself" would be the wrong fix anyway.
1752
+ if os.path.abspath(src) != os.path.abspath(vo_path):
1753
+ run(["ffmpeg", "-y", "-v", "error", "-i", src, "-c", "copy", vo_path])
1754
+
1755
+ # 2. the review / fallback cut. The bed is duration-matched to the video and
1756
+ # mixed UNDER the existing voiceover track; `normalize=0` on amix stops it
1757
+ # quietly pulling the voice down to make room for the music.
1758
+ afilter = (f"[1:a]volume={VOL},afade=t=in:st=0:d=0.6[bed];"
1759
+ f"[0:a][bed]amix=inputs=2:duration=first:normalize=0[a]")
1760
+ cmd = ["ffmpeg", "-y", "-v", "error", "-i", src, "-stream_loop", "-1", "-i", bed,
1761
+ "-filter_complex", afilter, "-map", "0:v", "-map", "[a]",
1762
+ "-c:a", "aac", "-b:a", "192k", "-shortest"]
1763
+ if WM:
1764
+ # a watermark has to be burned in, so this version gets its own video encode
1765
+ # and stops being frame-identical to the master. Say so in the handoff.
1766
+ cmd += ["-i", WM, "-filter_complex",
1767
+ afilter + ";[0:v][2:v]overlay=W-w-40:40:format=auto[v]",
1768
+ "-map", "[v]", "-c:v", "libx264", "-preset", "slow", "-crf", "18",
1769
+ "-pix_fmt", "yuv420p"]
1770
+ cmd = [c for c in cmd if c not in ("-map", "0:v")]
1771
+ else:
1772
+ cmd += ["-c:v", "copy"]
1773
+ run(cmd + [music_path])
1774
+
1775
+ def probe(f):
1776
+ r = subprocess.run(["ffmpeg", "-hide_banner", "-nostats", "-i", f,
1777
+ "-af", "ebur128=peak=true", "-f", "null", "-"],
1778
+ capture_output=True, text=True).stderr
1779
+ g = lambda k: next((l.split()[-2] for l in r.splitlines() if l.strip().startswith(k)), "?")
1780
+ return g("I:"), g("Peak:")
1781
+
1782
+ for label, f in (("voiceover only", vo_path), ("voiceover + music", music_path)):
1783
+ i, pk = probe(f)
1784
+ size = os.path.getsize(f) / 1e6
1785
+ print(f" {label:20s} {os.path.basename(f):32s} {size:6.1f} MB {i} LUFS peak {pk} dBFS")
1786
+ print()
1787
+ print(f" PUBLISH: {os.path.basename(vo_path)} — add music from the platform's own library at upload.")
1788
+ print(f" REVIEW : {os.path.basename(music_path)}"
1789
+ + (" (watermarked; video re-encoded, not frame-identical)" if WM else ""))
1790
+ ```
1791
+
1792
+ ## Appendix I — `place-stickers.py` (stickers from a manifest)
1793
+
1794
+ Writes the wrapper markup and the pop-in CSS, and warns when two stickers would be on screen at
1795
+ once — which is clutter in this format and the kind of thing you only notice on the render.
1796
+
1797
+ ```bash
1798
+ python3 place-stickers.py composition.html stickers.txt
1799
+ ```
1800
+
1801
+ ```python
1802
+ #!/usr/bin/env python3
1803
+ """
1804
+ Write sticker layers into a HyperFrames composition from a manifest.
1805
+
1806
+ Two things this exists to get right:
1807
+
1808
+ * A bare <img> with data-start renders at t=0 and then never shows or hides —
1809
+ the engine does not manage it. Each sticker has to be wrapped in a
1810
+ `div.clip` exactly like a caption layer.
1811
+ * `.clip { inset: 0 }` fights any geometry you set, so the wrapper has to
1812
+ reset `inset:auto` BEFORE left/top/width/height.
1813
+
1814
+ Manifest, one line per sticker (# comments and blank lines ignored):
1815
+
1816
+ <src>|<start>|<duration>|<left%>|<top%>|<width%>|<height%>|<label>[|<rotate>][|<fit>]
1817
+
1818
+ `fit` is `contain` (default) or `fill`, and the choice is not cosmetic:
1819
+
1820
+ * SUBJECT stickers — a dish, a product — must keep their aspect: `contain`.
1821
+ * ANNOTATION stickers — a ring, an underline — must SQUASH to the shape they
1822
+ are marking: `fill`. A near-square ring set to `contain` inside a wide, short
1823
+ box shrinks to the box's HEIGHT and ends up a small circle floating next to
1824
+ the row instead of around it.
1825
+
1826
+ usage: place-stickers.py <composition.html> <stickers.txt>
1827
+ """
1828
+ import os, re, sys
1829
+
1830
+ comp, manifest = sys.argv[1], sys.argv[2]
1831
+
1832
+ CSS = """
1833
+ .sticker{animation-name:vfStickIn;animation-duration:.42s;animation-fill-mode:both;
1834
+ animation-timing-function:cubic-bezier(.2,1.1,.3,1);transform-origin:center;
1835
+ filter:drop-shadow(0 10px 24px rgba(0,0,0,.35))}
1836
+ @keyframes vfStickIn{
1837
+ 0%{opacity:0;transform:scale(.55) rotate(-14deg)}
1838
+ 60%{opacity:1;transform:scale(1.06) rotate(2deg)}
1839
+ 100%{opacity:1;transform:scale(1) rotate(0deg)}}
1840
+ """
1841
+
1842
+ doc = open(comp).read()
1843
+ # re-runnable: drop any previous sticker layers and CSS
1844
+ doc = re.sub(r'<div class="clip" id="stk-[^"]*".*?</div>\s*', '', doc, flags=re.S)
1845
+ doc = re.sub(r'\n?\.sticker\{.*?@keyframes vfStickIn\{.*?\}\}\n?', '', doc, flags=re.S)
1846
+ doc = doc.replace("</style>", CSS + "</style>", 1)
1847
+
1848
+ rows, layers = [], []
1849
+ for i, raw in enumerate(open(manifest)):
1850
+ raw = raw.strip()
1851
+ if not raw or raw.startswith("#"):
1852
+ continue
1853
+ parts = [p.strip() for p in raw.split("|")]
1854
+ src, start, dur, left, top, w, h, label = parts[:8]
1855
+ rot = parts[8] if len(parts) > 8 else "0"
1856
+ fit = parts[9] if len(parts) > 9 else "contain"
1857
+ start, dur = float(start), float(dur)
1858
+ layers.append(
1859
+ f' <div class="clip" id="stk-{i}" data-hf-id="stk-{i}" data-layer-mode="publish" '
1860
+ f'data-layer-kind="image" data-start="{start}" data-duration="{dur}" '
1861
+ f'data-end="{round(start+dur,3)}" data-track-index="6" '
1862
+ f'data-label="{label}" '
1863
+ f'style="position:absolute;inset:auto;left:{left}%;top:{top}%;width:{w}%;height:{h}%;'
1864
+ f'z-index:6;transform:rotate({rot}deg)">'
1865
+ f'<img class="sticker" src="{src}" '
1866
+ f'style="width:100%;height:100%;object-fit:{fit};animation-delay:{start}s"></div>')
1867
+ rows.append((start, start + dur, label))
1868
+
1869
+ idx = doc.rfind("</div>")
1870
+ doc = doc[:idx] + "\n".join(layers) + "\n" + doc[idx:]
1871
+ open(comp, "w").write(doc)
1872
+
1873
+ rows.sort()
1874
+ print(f"{len(rows)} sticker(s) placed")
1875
+ for a, b, label in rows:
1876
+ print(f" {a:6.2f} -> {b:6.2f} {label}")
1877
+ # two stickers on screen at once is clutter in this format, and it is the kind of
1878
+ # thing you only notice on the render — so say it here instead
1879
+ for (a1, b1, l1), (a2, b2, l2) in zip(rows, rows[1:]):
1880
+ if a2 < b1:
1881
+ print(f" ⚠ OVERLAP: '{l1}' and '{l2}' are both on screen {a2:.2f}-{b1:.2f}")
1882
+ ```
1883
+
1884
+ ## Appendix J — `cast-actors.py` (group a shelf by actor)
1885
+
1886
+ Pulls the whole shelf (cursor-paginated), reads each card's `actor_<uuid>` out of its summary,
1887
+ and ranks actors by how many takes they have. Run it before you pick a single clip.
1888
+
1889
+ ```bash
1890
+ python3 cast-actors.py --category ugc-reaction --min 4
1891
+ # ugc-reaction: 180 raws · 53 tagged actors · 3 untagged
1892
+ # actor_668363ea-… — 7 takes …
1893
+ ```
1894
+
1895
+ ```python
1896
+ #!/usr/bin/env python3
1897
+ """
1898
+ Group a public-raws shelf by ACTOR, so you can cast a face that has enough takes.
1899
+
1900
+ A shelf is dozens of clips of a much smaller number of creators, and nothing in
1901
+ the taxonomy says "same face" — the `actor_<uuid>` tag does. This pulls the
1902
+ shelf, reads each card's actor id out of its summary/tags, and ranks the actors
1903
+ by how many takes they have.
1904
+
1905
+ This format needs FOUR reaction beats from one person, so the only actors worth
1906
+ casting from are the ones with 4+ takes. That list is short.
1907
+
1908
+ usage: cast-actors.py [--category ugc-reaction] [--min 2]
1909
+ env: VIDFARM_API_KEY
1910
+ """
1911
+ import collections, json, os, re, sys, urllib.parse, urllib.request
1912
+
1913
+ def opt(n, d):
1914
+ return sys.argv[sys.argv.index(n) + 1] if n in sys.argv else d
1915
+ CATEGORY = opt("--category", "ugc-reaction")
1916
+ MIN = int(opt("--min", "2"))
1917
+ HOST = os.environ.get("VIDFARM_HOST", "https://vidfarm.cc")
1918
+ KEY = os.environ["VIDFARM_API_KEY"]
1919
+
1920
+ # The feed caps `limit` at 100 and paginates on `cursor` — NOT `offset`, which is
1921
+ # silently ignored and hands you page 1 again. A loop built on offset looks like
1922
+ # it works, reports a plausible number, and has seen half the shelf.
1923
+ rows, cursor = {}, None
1924
+ while True:
1925
+ params = {"category": CATEGORY, "limit": 100}
1926
+ if cursor:
1927
+ params["cursor"] = cursor
1928
+ q = urllib.parse.urlencode(params)
1929
+ req = urllib.request.Request(f"{HOST}/api/v1/public-raws?{q}",
1930
+ headers={"Authorization": f"Bearer {KEY}"})
1931
+ page = json.load(urllib.request.urlopen(req))
1932
+ batch = page.get("raws", [])
1933
+ if not batch:
1934
+ break
1935
+ for r in batch:
1936
+ rows[r["rawId"]] = r
1937
+ cursor = page.get("next_cursor")
1938
+ if not cursor:
1939
+ break
1940
+
1941
+ ACTOR_RE = re.compile(r"actor_[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}", re.I)
1942
+ actors, untagged = collections.defaultdict(list), 0
1943
+ for r in rows.values():
1944
+ hay = (r.get("summary") or "") + " " + json.dumps(r.get("tags") or {})
1945
+ m = ACTOR_RE.search(hay)
1946
+ # "no tag" means NOT TAGGED YET — never treat it as "a different person"
1947
+ if m:
1948
+ actors[m.group(0)].append(r)
1949
+ else:
1950
+ untagged += 1
1951
+
1952
+ print(f"{CATEGORY}: {len(rows)} raws · {len(actors)} tagged actors · {untagged} untagged")
1953
+ ranked = sorted(actors.items(), key=lambda kv: -len(kv[1]))
1954
+ for actor_id, clips in ranked:
1955
+ if len(clips) < MIN:
1956
+ break
1957
+ print(f"\n{actor_id} — {len(clips)} takes")
1958
+ for c in sorted(clips, key=lambda c: -(c.get("durationSeconds") or 0)):
1959
+ print(f" {round(c.get('durationSeconds') or 0, 1):5}s {c['rawId']}")
1960
+ print(f" {(c.get('description') or '')[:96]}")
1961
+ print("\nCast from an actor with enough takes to carry all four reaction beats,")
1962
+ print("then pick each beat's take by EMOTION — see the harness's casting table.")
1963
+ ```