@officexapp/vidfarm-devcli 0.21.52 → 0.21.54
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/vidfarm/SKILL.md +6 -2
- package/.agents/skills/vidfarm/harnesses/hooks.HARNESS.md +10 -3
- package/.agents/skills/vidfarm/harnesses/product-explainer.HARNESS.md +1 -1
- package/.agents/skills/vidfarm/harnesses/short-form.HARNESS.md +2 -0
- package/.agents/skills/vidfarm/references/assets-and-sourcing.md +44 -4
- package/.agents/skills/vidfarm/references/automation-and-local-dev.md +1 -1
- package/.agents/skills/vidfarm/references/content-ideas.md +72 -1
- package/.agents/skills/vidfarm/references/core-workflows.md +17 -2
- package/.agents/skills/vidfarm/references/editor-workflows.md +39 -9
- package/.agents/skills/vidfarm/references/reviewing-renders.md +1 -0
- package/SKILL.director.md +180 -19
- package/SKILL.md +9 -2
- package/clipper.md +22 -0
- package/dist/src/cli.js +98 -5
- package/dist/src/devcli/composition-edit.js +36 -3
- package/dist/src/devcli/marketplace-gigs.js +2 -2
- package/dist/src/devcli/qa-check.js +7 -3
- package/dist/src/devcli/shared-folder.js +333 -44
- package/dist/src/devcli/skill-docs.js +2 -2
- package/experimental/meme-recaption.md +1756 -0
- package/experimental/sticker-slideshow-tips.md +1486 -0
- package/experimental/ugc-reaction-greenscreen.md +1963 -0
- package/experimental/wall-text-pov-ugc.md +2036 -0
- package/marketplace.md +82 -8
- package/package.json +6 -1
- package/public/assets/file-directory-app.js +33 -33
- package/public/assets/homepage-client-app.js +14 -14
|
@@ -0,0 +1,1963 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ugc-reaction-greenscreen
|
|
3
|
+
video_type: UGC reaction cutaways + a keyed greenscreen device carrying the customer's app demo — subtitled, music-led, silent or narrated (TikTok / Reels / Shorts)
|
|
4
|
+
checks:
|
|
5
|
+
duration_sec: 15-30
|
|
6
|
+
aspect: 9:16
|
|
7
|
+
first_frame_visual: required
|
|
8
|
+
first_frame_text: required
|
|
9
|
+
text_by_sec: 0.6
|
|
10
|
+
captions: required
|
|
11
|
+
audio: required # a bed in silent mode; a bed + voiceover in narrated mode
|
|
12
|
+
font_regime: required
|
|
13
|
+
safe_zone: required
|
|
14
|
+
# This counts EVERY visual layer, not beats — sticker images land in it too.
|
|
15
|
+
# The seven-beat structure is a review item, not a machine check: 7 beats plus
|
|
16
|
+
# up to ~4 stickers is the shape this range is sized for.
|
|
17
|
+
scenes: 4-12
|
|
18
|
+
max_scene_sec: 10
|
|
19
|
+
max_text_cards: 1 # every word is a timed subtitle cue, not a card
|
|
20
|
+
max_simultaneous_text: 1
|
|
21
|
+
max_words_per_cue: 7
|
|
22
|
+
# 1.5 is right for SILENT mode, where text is the only channel and a gap is a free exit.
|
|
23
|
+
# In NARRATED mode the payoff beat is SUPPOSED to have a stretch with no caption over it —
|
|
24
|
+
# the voice stops on purpose and the screen does the work. Tighten this back to 1.5 if you
|
|
25
|
+
# cut the voiceover.
|
|
26
|
+
max_dead_air_sec: 3.0
|
|
27
|
+
max_tail_sec: 1.2
|
|
28
|
+
forbid_text:
|
|
29
|
+
- link in bio
|
|
30
|
+
- sign up for a free trial
|
|
31
|
+
- get started today
|
|
32
|
+
- book a demo
|
|
33
|
+
- download now
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
# UGC Reaction × Greenscreen × App Demo
|
|
37
|
+
|
|
38
|
+
> A public Vidfarm prompt. Read it end to end before you build. It assumes nothing about the
|
|
39
|
+
> product except that it has a screen you can film.
|
|
40
|
+
|
|
41
|
+
The format is three streams of footage cut against each other:
|
|
42
|
+
|
|
43
|
+
| Stream | What it is | Where it comes from | What it does |
|
|
44
|
+
|---|---|---|---|
|
|
45
|
+
| 🙂 **Reaction** | An ordinary person on camera, mid-expression | `vidfarm.cc/discover` → public raws, shelf `ugc-reaction` (180 clips) | Makes the video a *situation* instead of an ad. Carries the hook and the bait |
|
|
46
|
+
| 📱 **Greenscreen device** | A hand holding a phone whose screen is a flat green field | public raws, shelf `greenscreen`, `sourceType: Display Greenscreen` | The frame the product lives inside. Turns a screen recording into something someone is *holding* |
|
|
47
|
+
| 🖥 **App demo** | The customer's real product doing one real thing | the customer's own folder, or you drive their site in a browser | The proof. The only beat where the product exists |
|
|
48
|
+
|
|
49
|
+
**Only one voice is ever allowed: yours.** The reaction clips are muted, the customer's demo is
|
|
50
|
+
muted, and the cut runs either fully silent under subtitles or under a voiceover you wrote and a
|
|
51
|
+
ducked music bed. Two modes, no third — see *Audio DNA*.
|
|
52
|
+
|
|
53
|
+
## The one test
|
|
54
|
+
|
|
55
|
+
> **Does the product arrive as something a person is holding, or as a screen recording someone pasted in?**
|
|
56
|
+
|
|
57
|
+
The greenscreen device beat is the entire reason this format exists. A screen recording dropped
|
|
58
|
+
full-bleed on a timeline is a demo video with reaction clips stapled to it. A screen recording
|
|
59
|
+
living inside a phone that a real hand is tilting is a person showing you something. The viewer
|
|
60
|
+
reads the difference in well under a second and never articulates it.
|
|
61
|
+
|
|
62
|
+
## The second test — does the video SHOW what the product is about?
|
|
63
|
+
|
|
64
|
+
Ask it before you cut and again on the render: **name the noun the product is about, then find it in
|
|
65
|
+
the footage.** A dish-search app is about *food*. A rental app is about *rooms*. A running app is
|
|
66
|
+
about *running*.
|
|
67
|
+
|
|
68
|
+
This format is unusually good at failing that test, and the reason is structural: its three streams
|
|
69
|
+
are a **face**, a **phone**, and a **UI**. None of them is the subject. On the reference build the
|
|
70
|
+
product finds you food by dish — and the first cut contained no food at all. It was twenty-two
|
|
71
|
+
seconds of people reacting to a screen. Every individual beat was correct and the video still did
|
|
72
|
+
not read as being about dinner.
|
|
73
|
+
|
|
74
|
+
**The UI is not the subject.** A results list saying *Cucumber Salad · $9* is a row of text. It
|
|
75
|
+
proves the product works; it does not make anyone hungry, and hunger is the thing you are selling.
|
|
76
|
+
|
|
77
|
+
So: find the noun, and if the raws don't contain it, **put it in** — with a cutaway from the
|
|
78
|
+
`lifestyle` / `b-roll` shelves if one exists, or with a sticker if it doesn't. Getting the subject
|
|
79
|
+
on screen twice in 22 seconds is worth more than any amount of re-cutting.
|
|
80
|
+
|
|
81
|
+
## Part 0 — who this is for
|
|
82
|
+
|
|
83
|
+
**Not you — the customer's buyer.** Before anything else, write one line naming them, and one line
|
|
84
|
+
naming the moment they are in. Everything below is decided by those two lines: which reaction face,
|
|
85
|
+
which craving/query/task you type into the demo, which limit you concede.
|
|
86
|
+
|
|
87
|
+
- **Register:** peer to peer. Someone who had the problem, found the thing, is telling you.
|
|
88
|
+
- **The product enters in beat 3, never beat 1.** Beat 1 is a situation.
|
|
89
|
+
- **Never a brand voice.** No "introducing", no "meet", no "the all-new".
|
|
90
|
+
|
|
91
|
+
---
|
|
92
|
+
|
|
93
|
+
## Structural DNA — the beats
|
|
94
|
+
|
|
95
|
+
Seven beats, ~22s. The rhythm is **reaction → device → reaction → device**: the video cuts *back*
|
|
96
|
+
to a face after the payoff and *back* to the app after the face. One unbroken block of app footage
|
|
97
|
+
is the single most common way this format collapses into an ad.
|
|
98
|
+
|
|
99
|
+
| # | Beat | ~t | Stream | Job |
|
|
100
|
+
|---|---|---|---|---|
|
|
101
|
+
| 1 | **Hook** | 0:00–0:02.5 | Reaction | The situation, on a real face, before any product exists |
|
|
102
|
+
| 2 | **The wrong way** | 0:02.5–0:05 | Reaction | What they tried instead, and why it failed. This is the concession |
|
|
103
|
+
| 3 | **The move** | 0:05–0:09 | Device | The one action inside the app. Typing, tapping, pasting — at real speed |
|
|
104
|
+
| 4 | **The result** | 0:09–0:14.5 | Device | What came back. Held, legible, uncut. **This is the payoff** |
|
|
105
|
+
| 5 | **The face** | 0:14.5–0:17 | Reaction | The reaction to the result — the emotional label for what was just shown |
|
|
106
|
+
| 6 | **The scope** | 0:17–0:19.5 | Device | The honest limit, said over the app still running |
|
|
107
|
+
| 7 | **The bait** | 0:19.5–0:22 | Reaction | One comment ask, and the same ask in the post caption |
|
|
108
|
+
|
|
109
|
+
**Cast ONE actor across every reaction beat.** The narration is first person singular — *"I didn't
|
|
110
|
+
want a restaurant… I just picked one"* — so a different face on each beat is not a stylistic choice,
|
|
111
|
+
it is an incoherence: one voice saying "I" over four different people. One face turns four cutaways
|
|
112
|
+
into an arc (exasperated → beaten → convinced → warm), and the arc is what makes the payoff land.
|
|
113
|
+
|
|
114
|
+
Use several faces only when the video genuinely is a vox-pop — nobody is "I", the narration is
|
|
115
|
+
third person, and the point is *lots of people have this problem*. That is a different video.
|
|
116
|
+
|
|
117
|
+
**Beat 6 is load-bearing and gets cut first by every agent that reads this.** The limit is what
|
|
118
|
+
buys the rest of the video. Put it over the app so it doesn't read as a disclaimer card.
|
|
119
|
+
|
|
120
|
+
**One clip per beat — seven beats, seven layers.** Beats 3 and 4 are the same recording, so the
|
|
121
|
+
lazy build puts them on one ~10s layer. Don't: against 2.5s reaction cuts a 10s hold is where the
|
|
122
|
+
retention curve falls off, and `vidfarm qa` will say so (`slow-scene`). Split at the frame the
|
|
123
|
+
result lands — which is a cut the *content* is already asking for — and no layer runs past ~6s.
|
|
124
|
+
|
|
125
|
+
---
|
|
126
|
+
|
|
127
|
+
## Sourcing — three streams, three methods
|
|
128
|
+
|
|
129
|
+
### 🙂 Reaction raws — browse the shelf, don't search it
|
|
130
|
+
|
|
131
|
+
```bash
|
|
132
|
+
vidfarm public-raws --categories # the shelves + live counts
|
|
133
|
+
vidfarm public-raws --category ugc-reaction --limit 100 # the whole shelf, read the descriptions
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
The shelf's semantic search is thin; **the descriptions are the index**. Pull the shelf as JSON and
|
|
137
|
+
filter the `description` field on the emotion you want (`surpris`, `disgust`, `distress`, `smil`,
|
|
138
|
+
`laugh`, `disbelie`, `confus`). Pick by the description, then confirm by eye — build a contact
|
|
139
|
+
sheet of 4–5 stills per candidate and read them as one image before you commit.
|
|
140
|
+
|
|
141
|
+
#### Cast the actor, not the clip — the `actor_<uuid>` tag
|
|
142
|
+
|
|
143
|
+
**A shelf is dozens of clips of a much smaller number of creators**, and nothing in the taxonomy
|
|
144
|
+
says "same face". The `actor_<uuid>` tag does. Every tagged public raw carries one actor id in its
|
|
145
|
+
`summary` (and in `tags.actor`), so you read it off a card you already like and search for it as a
|
|
146
|
+
plain token:
|
|
147
|
+
|
|
148
|
+
```bash
|
|
149
|
+
vidfarm public-raws --category ugc-reaction --limit 100 # 1. browse, pick a face
|
|
150
|
+
# 2. read the actor_<uuid> token off that card's summary — it is appended at the END
|
|
151
|
+
vidfarm public-raws --query actor_668363ea-… # 3. every other clip of that person
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
The token lives at the tail of the card's `summary` string. In the feed response `tags` comes back
|
|
155
|
+
`null`, so **`summary` is the field to parse** — a reader that only looks at `tags.actor` finds
|
|
156
|
+
nothing and concludes the shelf is untagged.
|
|
157
|
+
|
|
158
|
+
REST twin `GET /api/v1/public-raws?q=actor_<uuid>`, and because it is an ordinary keyword search it
|
|
159
|
+
composes: `?category=ugc-reaction&q=actor_<uuid>`.
|
|
160
|
+
|
|
161
|
+
**Cast by ARC, not by clip.** You need four reaction beats — hook, the wrong way, the payoff face,
|
|
162
|
+
the bait — so pick the actor whose set can carry all four, not the single best-looking clip. On the
|
|
163
|
+
reference build one actor had five takes and four of them mapped straight onto the beats:
|
|
164
|
+
|
|
165
|
+
| Beat | Their take |
|
|
166
|
+
|---|---|
|
|
167
|
+
| 1 hook | distressed faces, hand at her mouth |
|
|
168
|
+
| 2 the wrong way | frowning at a **laptop** — which is what "Google gave me a list of fifty" is about |
|
|
169
|
+
| 5 the face | **thumbs up**, outdoors |
|
|
170
|
+
| 7 the bait | a warm smile to camera |
|
|
171
|
+
|
|
172
|
+
Practical notes, from the catalogue as it stands:
|
|
173
|
+
|
|
174
|
+
- **Measure the pool, don't take a doc's word for it.** Counted today: **180 `ugc-reaction` raws,
|
|
175
|
+
53 tagged actors, 3 untagged — and 25 actors with 4+ takes.** That is a far bigger casting pool
|
|
176
|
+
than the shelf looks like from the outside, and it is the only part of it that can carry this
|
|
177
|
+
format. Re-count rather than trusting this line; the shelf is actively being tagged.
|
|
178
|
+
- ⚠️ **The feed caps `limit` at 100 and paginates on `cursor`, not `offset`.** `offset=100` is
|
|
179
|
+
silently ignored and returns page 1 again — a loop built on it looks like it works, reports a
|
|
180
|
+
plausible number, and has seen half the shelf. Follow `next_cursor` until it is null.
|
|
181
|
+
- ⚠️ **Read the id fresh, and re-read it if it stops matching.** The ids come from a tagging pass
|
|
182
|
+
that gets re-run: mid-build here, an id that had returned 5 clips started returning **0**, because
|
|
183
|
+
the shelf had been re-tagged and the same person's clips now carried a different id (with 7 clips
|
|
184
|
+
grouped under it instead of 5). So an actor id is a lookup key that is good for *this* build, not
|
|
185
|
+
a permanent name. If a saved id returns nothing, the actor has not gone — re-read the token off a
|
|
186
|
+
card. And never invent one; a made-up id matches nothing.
|
|
187
|
+
- **No tag means "not tagged yet", not "a different person".** Older raws and raws with nobody on
|
|
188
|
+
camera have none.
|
|
189
|
+
- **Wardrobe and location will NOT match between takes** — different top, different room, outdoors.
|
|
190
|
+
That is fine and it is not a continuity error: UGC creators film across days, and the audience
|
|
191
|
+
reads it as the same person at different moments. **Identity continuity is what matters; wardrobe
|
|
192
|
+
continuity is not.** What *would* break it is a different face.
|
|
193
|
+
- Same trick answers a director asking *"more of her"* — the actor id on the raw already in the
|
|
194
|
+
composition is the answer.
|
|
195
|
+
|
|
196
|
+
⚠️ **`ffprobe` reports CODED dimensions, and phone video is rotated.** Almost every clip on this
|
|
197
|
+
shelf carries a 90° rotation matrix: `ffprobe` says `3840x2160` and ffmpeg **decodes it as
|
|
198
|
+
2160x3840**, already 9:16. Hand-computing a crop from the probed width put the subject off-frame on
|
|
199
|
+
all four beats of the reference build — wall where the face should be. Never compute geometry from
|
|
200
|
+
the probe; let the filter work on the decoded frame:
|
|
201
|
+
|
|
202
|
+
```bash
|
|
203
|
+
# right: operates on what was decoded, whatever the rotation
|
|
204
|
+
-vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1,fps=30"
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
If you genuinely need an off-centre crop, read the decoded size first
|
|
208
|
+
(`ffprobe -show_entries side_data=rotation`, or decode one frame and measure it).
|
|
209
|
+
|
|
210
|
+
### 📱 Greenscreen device raws — check `sourceType` before you key anything
|
|
211
|
+
|
|
212
|
+
The `greenscreen` shelf holds two populations that look identical to a category filter and are
|
|
213
|
+
completely different to a lawyer:
|
|
214
|
+
|
|
215
|
+
| `sourceType` | What it is | Use it for |
|
|
216
|
+
|---|---|---|
|
|
217
|
+
| **Display Greenscreen** | Device mockups — a hand holding a phone with a green screen, a phone on a notebook. No identifiable person, no character | ✅ Client work. This is the one you want |
|
|
218
|
+
| **MemeScreens** | Keyed celebrities, TV characters, pets, animated figures on a green field | ⚠️ Normal for a meme repost. A rights problem on a customer's paid ad running on owned accounts |
|
|
219
|
+
|
|
220
|
+
```bash
|
|
221
|
+
vidfarm public-raws --category greenscreen --query "smartphone green screen" --limit 25
|
|
222
|
+
```
|
|
223
|
+
|
|
224
|
+
There are only a handful of Display Greenscreen clips, so **look at all of them and pick on
|
|
225
|
+
motion, not on looks.** Ranked, best first:
|
|
226
|
+
|
|
227
|
+
1. **Two hands, thumb on the glass, indoors, native 9:16.** Reads as a person using a phone. Drifts
|
|
228
|
+
the most — needs the tracked insert below, which is fine, because you are doing that anyway.
|
|
229
|
+
2. **Phone flat on a desk or notebook, top-down.** Calm, minimal drift, but reads as a product shot
|
|
230
|
+
rather than a person.
|
|
231
|
+
3. **Phone against a plain green background.** Avoid: the plate colour and the screen colour are the
|
|
232
|
+
same field, so a key gives you a floating phone body and a demo behind *everything*.
|
|
233
|
+
|
|
234
|
+
### 🖥 App demo raws — the customer's folder first, the browser second
|
|
235
|
+
|
|
236
|
+
**Ask for a folder before you build one.** A customer who already records their own product gives
|
|
237
|
+
you footage with the right accounts, the right data and the right edge cases. Take it, and then:
|
|
238
|
+
|
|
239
|
+
> **Cut it. Do not narrate it, and do not keep their narration.**
|
|
240
|
+
>
|
|
241
|
+
> A narrated demo raw is *worse* input than a silent one. Their voice competes with your subtitles,
|
|
242
|
+
> the pacing is theirs and the tone is a webinar. Strip the audio, find the four seconds where the
|
|
243
|
+
> product actually does the thing, and throw the rest away. If they insist their narration ships,
|
|
244
|
+
> that is a different format — use `ugc-testimonial`.
|
|
245
|
+
|
|
246
|
+
**No folder? Drive their site yourself.** Free, local, repeatable, and it always has current data:
|
|
247
|
+
|
|
248
|
+
```bash
|
|
249
|
+
vidfarm capture "https://<their site>" --out ./capture # tokens, copy, screenshots — read this first
|
|
250
|
+
vidfarm browser setup # then drive their real UI
|
|
251
|
+
```
|
|
252
|
+
|
|
253
|
+
Record the interaction at the **device screen's aspect ratio**, not at desktop and not at a laptop
|
|
254
|
+
viewport. A 440×940 recording drops into a phone screen with no letterbox and no re-crop; a
|
|
255
|
+
1440×900 one arrives as a stamp inside a phone and the whole beat is unreadable.
|
|
256
|
+
|
|
257
|
+
Four rules for the recording itself, all of them learned the expensive way:
|
|
258
|
+
|
|
259
|
+
1. **Verify the query returns results before you film it.** Type it yourself first. A product's own
|
|
260
|
+
advertised example can be broken — in the worked example that produced this file, the exact phrase the
|
|
261
|
+
customer prints on their own homepage returned *No dishes found*, and the first take was unusable.
|
|
262
|
+
2. **Park the interactive region at the top of the frame before you type.** Otherwise the results
|
|
263
|
+
render below the fold and the payoff never appears on screen.
|
|
264
|
+
3. **Type at human speed** (~130ms/char) and hold ~2.5s after the result lands. Instant fills read
|
|
265
|
+
as a mockup; a real keystroke cadence is the cheapest authenticity you will ever buy.
|
|
266
|
+
4. **Draw a soft synthetic cursor.** A UI that changes with nothing touching it looks like a video
|
|
267
|
+
of a website. A dot that moves to the field first looks like a person.
|
|
268
|
+
|
|
269
|
+
---
|
|
270
|
+
|
|
271
|
+
## Visual DNA
|
|
272
|
+
|
|
273
|
+
### The tracked insert — the technique the whole format rests on
|
|
274
|
+
|
|
275
|
+
The device is handheld. It **drifts and it is tilted**, so a flat chroma key with the demo pinned
|
|
276
|
+
behind it fails in three named ways, all visible within two seconds:
|
|
277
|
+
|
|
278
|
+
| Defect | Cause | Fix |
|
|
279
|
+
|---|---|---|
|
|
280
|
+
| The demo **slides under the bezel** | The insert is static; the phone is not | Re-measure the green region *every frame* |
|
|
281
|
+
| A **green wedge** down one edge | The insert is axis-aligned; the phone is rotated a few degrees | Fit to the screen's own axes (PCA on the green mask), not to an upright bounding box |
|
|
282
|
+
| The demo **paints over the thumb** | The insert is pasted into a *box* | Paste by **mask**: every green pixel becomes demo, every non-green pixel (thumb, bezel, room) is left alone |
|
|
283
|
+
|
|
284
|
+
The mask paste is what sells it — the thumb ends up *on top of* the app, which is the single frame
|
|
285
|
+
that makes a viewer believe someone is holding it. Appendix A is a ~90-line script that does all
|
|
286
|
+
three. `vidfarm remove-greenscreen` is the right tool for a **full-frame** flat plate; it is the
|
|
287
|
+
wrong tool here, because it has no reason to know where the screen went.
|
|
288
|
+
|
|
289
|
+
```bash
|
|
290
|
+
python3 track-insert.py phone-raw.mp4 demo.mp4 phone-demo.mp4 --demo-start 1.2
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
`--demo-start` trims the demo's dead head (page load, blank white) so the insert opens on something.
|
|
294
|
+
|
|
295
|
+
### Punch in on the device beats
|
|
296
|
+
|
|
297
|
+
A Display Greenscreen raw frames the phone small, with a lot of blurred room around it. Shipped as
|
|
298
|
+
is, the UI is unreadable at phone scale and the demo did not happen. Crop to the device **plus a
|
|
299
|
+
drift margin**, then scale back to canvas:
|
|
300
|
+
|
|
301
|
+
Measure the green region's **travel** before you choose the crop — a crop sized to one frame clips
|
|
302
|
+
the device on another:
|
|
303
|
+
|
|
304
|
+
```python
|
|
305
|
+
# per-second bounding box of the green field, on the ORIGINAL device raw
|
|
306
|
+
import subprocess, numpy as np
|
|
307
|
+
W, H, f = 1080, 1920, "device-raw.mp4"
|
|
308
|
+
p = subprocess.run(["ffmpeg","-v","error","-i",f,"-vf","fps=1,scale=320:569",
|
|
309
|
+
"-pix_fmt","rgb24","-f","rawvideo","-"], capture_output=True)
|
|
310
|
+
b = np.frombuffer(p.stdout, np.uint8); n = len(b)//(320*569*3)
|
|
311
|
+
for i, fr in enumerate(b[:n*320*569*3].reshape(n,569,320,3).astype(np.int16)):
|
|
312
|
+
r, g, bl = fr[...,0], fr[...,1], fr[...,2]
|
|
313
|
+
m = (g > 90) & (g-r > 40) & (g-bl > 40)
|
|
314
|
+
ys, xs = np.nonzero(m)
|
|
315
|
+
print(f"t={i}s x[{xs.min()*W//320},{xs.max()*W//320}] y[{ys.min()*H//569},{ys.max()*H//569}]")
|
|
316
|
+
```
|
|
317
|
+
|
|
318
|
+
Take the union of every box, add ~60px of margin on each side, and round the crop to the canvas
|
|
319
|
+
aspect:
|
|
320
|
+
|
|
321
|
+
```bash
|
|
322
|
+
ffmpeg -i phone-demo.mp4 -vf "crop=880:1564:20:130,scale=1080:1920,setsar=1,fps=30" s3-move.mp4
|
|
323
|
+
```
|
|
324
|
+
|
|
325
|
+
Typical punch-in is 1.2–1.4×. Use the **same crop for every device beat** in the video — a punch-in
|
|
326
|
+
that changes between beat 3 and beat 6 reads as two different videos.
|
|
327
|
+
|
|
328
|
+
### Stickers — allowed, and only ever as annotation
|
|
329
|
+
|
|
330
|
+
Stickers earn their place in exactly two jobs, and both are about comprehension:
|
|
331
|
+
|
|
332
|
+
| Job | What it looks like | When you need it |
|
|
333
|
+
|---|---|---|
|
|
334
|
+
| **Annotate** | a ring around the row the narration just named, an arrow at the field being typed into, an underline under the number being read out | the device beat asks a viewer to read a dense UI in about two seconds |
|
|
335
|
+
| **Supply the subject** | the actual thing the product is about — the dish, the room, the shoe — laid over the shot | the raws are faces and a phone, and the product's own noun never appears (see *The second test*) |
|
|
336
|
+
|
|
337
|
+
**The test: does it make the product clearer?** If yes, use it. If it competes for attention with
|
|
338
|
+
what's already on screen, or a viewer has to work out why it's there, cut it — a confusing sticker
|
|
339
|
+
is worse than none, and this format is dense enough already.
|
|
340
|
+
|
|
341
|
+
**Never furniture.** A floating icon, a badge, a chip, a logo, a "NEW!" burst, an emoji dropped in
|
|
342
|
+
to fill space — none of those is either job, and `vidfarm qa` is right to call them slop.
|
|
343
|
+
|
|
344
|
+
- **One on screen at a time, and not on every beat.** Two or three across a 22s cut is plenty. A
|
|
345
|
+
sticker that sits for the whole video has stopped being punctuation.
|
|
346
|
+
- **It appears on the word that names it and leaves.** Same discipline as the captions: absolute
|
|
347
|
+
`animation-delay`, and it goes when the thought does.
|
|
348
|
+
- **An in-screen annotation is only valid while the UI under it is STATIC.** This is the one that
|
|
349
|
+
bites. The reference build's ring was placed over the picked row for 1.3s — and the results list
|
|
350
|
+
began scrolling 0.3s in, so the last third of the shot had a marker circling the wrong dish.
|
|
351
|
+
Step the beat frame by frame, find where the content moves, and end the sticker before it:
|
|
352
|
+
```bash
|
|
353
|
+
ffmpeg -i beat.mp4 -vf "fps=2,scale=300:-1,drawgrid=w=iw/20:h=ih/20:t=1:c=red@0.5,tile=6x1" \
|
|
354
|
+
-frames:v 1 rows.png # read the row's box off the grid, and see when it moves
|
|
355
|
+
```
|
|
356
|
+
- **Hand-drawn for annotation; photographic for the subject.** A crisp vector circle reads as a UI
|
|
357
|
+
element drawn over the video — ask for a rough felt-tip stroke instead. But a *subject* sticker
|
|
358
|
+
should look like the real thing: a photographic cutout of the actual dish sells the craving, an
|
|
359
|
+
illustration of one does not.
|
|
360
|
+
- **A tight face has no room for a sticker.** On a close portrait there is nowhere to put one that
|
|
361
|
+
isn't a forehead, a temple or an eye, and anything you place there reads as stuck to their head
|
|
362
|
+
rather than laid on the shot. Put subject stickers on the beats with air in them — the device
|
|
363
|
+
beats have defocused room on both sides, and a medium shot has wall. If the beat you *want* is a
|
|
364
|
+
tight face, move the sticker to the next beat rather than shrinking it into a corner.
|
|
365
|
+
|
|
366
|
+
**Sourcing, cheapest first.** `vidfarm iconscout --free` and `vidfarm media icon` are free but
|
|
367
|
+
almost all require an attribution credit — and an attribution line on screen is brand chrome you
|
|
368
|
+
already said you would not have. So for this format the practical route is one generated cutout:
|
|
369
|
+
|
|
370
|
+
```bash
|
|
371
|
+
vidfarm cutout --generate "a single thick hand-drawn yellow felt-tip ELLIPSE OUTLINE, stroke
|
|
372
|
+
~40px. The ENTIRE rest of the image, INCLUDING THE AREA INSIDE THE ELLIPSE, is solid magenta
|
|
373
|
+
#FF00FF. No white, no grey, no paper, no shadow, no text." \
|
|
374
|
+
--key-color "#FF00FF" --key-mode flat --tolerance 0.4 --out stickers/ring.png
|
|
375
|
+
```
|
|
376
|
+
|
|
377
|
+
**Match the KEY MODE to the shape, and the PLATE COLOUR to the palette.** Getting this wrong is
|
|
378
|
+
silent — the file looks right in a thumbnail and is wrong in the video:
|
|
379
|
+
|
|
380
|
+
| Art | Key mode | Why |
|
|
381
|
+
|---|---|---|
|
|
382
|
+
| **Hollow** — a ring, a frame, an arrow outline | `--key-mode flat` | the interior never touches the frame edge, so the default connectivity keyer *keeps* it and you get a filled disc instead of a ring |
|
|
383
|
+
| **Solid** — a photographic dish, a product, a person | `--key-mode smart` (the default) | flat mode's soft alpha ramp eats a photograph. Measured on a first attempt: **65% of the cutout came back partially transparent**, so the background showed straight through the food. Smart mode returned the same art at **0.6% partial alpha** |
|
|
384
|
+
|
|
385
|
+
Pick the plate against the subject's own colours — green plate for red food, magenta for green food
|
|
386
|
+
— and check the result rather than trusting it:
|
|
387
|
+
|
|
388
|
+
```python
|
|
389
|
+
# any cutout: how much of it is actually opaque?
|
|
390
|
+
a = alpha_channel(png)
|
|
391
|
+
print((a == 0).mean(), ((a > 0) & (a < 255)).mean(), (a == 255).mean())
|
|
392
|
+
# healthy: mostly 0 or 255, a percent or two in between. Tens of percent "partial" = a bad key.
|
|
393
|
+
```
|
|
394
|
+
|
|
395
|
+
Glass and soft edges pick up the plate colour even after a good key. One numpy pass fixes it: where
|
|
396
|
+
the plate's channel is the dominant one, pull it down to the average of the other two.
|
|
397
|
+
|
|
398
|
+
Three things in the prompt are load-bearing, and each one cost a wasted image job to learn:
|
|
399
|
+
|
|
400
|
+
1. **Pin the plate colour and say it twice.** Left to itself the model draws on **white**, the plate
|
|
401
|
+
never gets keyed, and you get a white box with art in it. Worse, if the art is also white (an
|
|
402
|
+
arrow, a highlight) you cannot key the background without destroying the subject.
|
|
403
|
+
2. **Say "including the area inside the ellipse".** A ring's interior does not touch the frame
|
|
404
|
+
edge, so the default connectivity keyer *keeps* it — you get a filled disc, not a ring.
|
|
405
|
+
`--key-mode flat` removes the colour wherever it appears, which is what an open shape needs.
|
|
406
|
+
3. **Ban white and grey explicitly.** Otherwise you get paper texture and a drop shadow, and the
|
|
407
|
+
shadow keys as a grey halo.
|
|
408
|
+
|
|
409
|
+
**Placement.** Drive it from a manifest so the set is reproducible and overlaps get caught —
|
|
410
|
+
Appendix I:
|
|
411
|
+
|
|
412
|
+
```
|
|
413
|
+
# src | start | dur | left% | top% | width% | height% | label | rotate | fit
|
|
414
|
+
media/stickers/food-chilicrisp.png | 6.60 | 1.90 | 2 | 21 | 27 | 16 | the craving | -5 | contain
|
|
415
|
+
media/stickers/ring.png | 10.30| 1.40 | 18 | 38.5 | 62 | 14 | the pick | 0 | fill
|
|
416
|
+
media/stickers/food-cucumber.png | 14.70| 1.90 | 56 | 14 | 30 | 18 | what $9 buys | 6 | contain
|
|
417
|
+
```
|
|
418
|
+
|
|
419
|
+
**`fit` is not cosmetic — the two sticker jobs need opposite values.** A subject sticker must keep
|
|
420
|
+
its aspect (`contain`); a photo of a dish squashed to a box is obviously wrong. An annotation
|
|
421
|
+
sticker must **squash to the shape it marks** (`fill`) — a near-square ring set to `contain` inside
|
|
422
|
+
a wide, short box collapses to the box's *height* and renders as a small circle floating beside the
|
|
423
|
+
row instead of around it. That regression is invisible in the markup and obvious in one frame, so
|
|
424
|
+
look at the frame.
|
|
425
|
+
|
|
426
|
+
The markup it writes — wrap the image in a `div.clip` exactly like a caption layer. A bare `<img>`
|
|
427
|
+
with `data-start` renders at t=0 and then never shows or hides, because the engine does not manage
|
|
428
|
+
it:
|
|
429
|
+
|
|
430
|
+
```html
|
|
431
|
+
<div class="clip" id="ann-ring" data-hf-id="ann-ring" data-layer-mode="publish"
|
|
432
|
+
data-layer-kind="image" data-start="10.3" data-duration="1.4" data-end="11.7"
|
|
433
|
+
data-track-index="6" data-label="annotation: the dish that was picked"
|
|
434
|
+
style="position:absolute;inset:auto;left:18%;top:38.5%;width:62%;height:14%;z-index:6">
|
|
435
|
+
<img class="sticker" src="media/stickers/ring.png"
|
|
436
|
+
style="width:100%;height:100%;object-fit:fill;animation-delay:10.3s"></div>
|
|
437
|
+
```
|
|
438
|
+
|
|
439
|
+
`inset:auto` first, because `.clip { inset: 0 }` would otherwise fight the geometry. `object-fit:
|
|
440
|
+
fill` is deliberate — a marker circle around a line of text *should* squash to the row's aspect.
|
|
441
|
+
|
|
442
|
+
```css
|
|
443
|
+
.sticker{animation-name:vfStickIn;animation-duration:.42s;animation-fill-mode:both;
|
|
444
|
+
animation-timing-function:cubic-bezier(.2,1.1,.3,1);transform-origin:center}
|
|
445
|
+
@keyframes vfStickIn{
|
|
446
|
+
0%{opacity:0;transform:scale(.55) rotate(-14deg)}
|
|
447
|
+
60%{opacity:1;transform:scale(1.06) rotate(2deg)}
|
|
448
|
+
100%{opacity:1;transform:scale(1) rotate(0deg)}}
|
|
449
|
+
```
|
|
450
|
+
|
|
451
|
+
### Everything else
|
|
452
|
+
|
|
453
|
+
- **9:16, 1080×1920, 30fps**, one canvas, no letterboxing anywhere.
|
|
454
|
+
- **Butt cuts.** No crossfades, no whips, no transitions of any kind. This format cuts on the
|
|
455
|
+
emotional beat; a transition softens exactly the moment you want hard.
|
|
456
|
+
- **The cut from device back to face lands on the result**, not two seconds after it.
|
|
457
|
+
- **No brand chrome.** No logo, no end card, no URL, no title card. The product's own UI is the
|
|
458
|
+
only branding, and it is already on screen.
|
|
459
|
+
|
|
460
|
+
---
|
|
461
|
+
|
|
462
|
+
## Audio DNA
|
|
463
|
+
|
|
464
|
+
**Two modes. Pick one before you cut, because the subtitles are built differently in each.**
|
|
465
|
+
|
|
466
|
+
| Mode | Track stack | Subtitles are… |
|
|
467
|
+
|---|---|---|
|
|
468
|
+
| **Silent** | music bed only | authored by hand, timed to the CUTS (`captions.srt` → Appendix C) |
|
|
469
|
+
| **Narrated** | voiceover over a ducked bed | transcribed from the VO, timed to the WORD (Appendix F) |
|
|
470
|
+
|
|
471
|
+
Narrated is the stronger default: it carries the beats where the picture is quiet, and the read-out
|
|
472
|
+
of a real number lands harder spoken than written. Reach for silent when the poster will drop
|
|
473
|
+
trending platform audio over the whole thing at upload — a voiceover fighting a trending track is
|
|
474
|
+
worse than either alone.
|
|
475
|
+
|
|
476
|
+
**What never changes, in either mode:**
|
|
477
|
+
|
|
478
|
+
- **Mute every reaction raw.** They came with real audio about something else entirely. A face
|
|
479
|
+
visibly saying words that don't match the subtitles is the loudest fake tell in the format, and it
|
|
480
|
+
survives every other quality pass because nobody watches the render with sound on.
|
|
481
|
+
- **Mute the demo.** Even if the customer's raw has narration. Their voice competes with yours,
|
|
482
|
+
the pacing is theirs, and the tone is a webinar.
|
|
483
|
+
- **The narration is YOURS, written to the beat table.** A voiceover is not permission to keep the
|
|
484
|
+
customer's — it is one script, one voice, one delivery, across the whole cut.
|
|
485
|
+
|
|
486
|
+
### Narrated mode
|
|
487
|
+
|
|
488
|
+
- **One VO clip per beat, not one long take.** Each clip is placed at its own beat, so a re-cut of
|
|
489
|
+
one beat re-records one line instead of re-timing everything. It also keeps the VO from drifting
|
|
490
|
+
against the picture as the edit changes.
|
|
491
|
+
- **Pin the voice.** `vidfarm tts "<line>" --voice <name> --style "<one delivery note>"` — pass the
|
|
492
|
+
SAME `--voice` and the SAME `--style` string on every line. Generate them in separate calls
|
|
493
|
+
without pinning and the narrator subtly changes character between beats, which reads as
|
|
494
|
+
"assembled" long before anyone can say why.
|
|
495
|
+
- **Direct the pace, don't fix it later.** A "dry, considered" style prompt produced 4.6s of speech
|
|
496
|
+
for four words — nearly double its beat. Re-prompting for "brisk, punchy, quick clipped delivery,
|
|
497
|
+
no pauses between sentences" brought the same line to 2.2s and it sounds *more* native, because
|
|
498
|
+
short-form narration is fast. Prefer re-prompting over time-stretching.
|
|
499
|
+
- **Compress pauses; don't stretch speech.** When a line is still slightly long, strip the internal
|
|
500
|
+
silences rather than `atempo` the whole clip — it preserves the natural word rate and kills dead
|
|
501
|
+
air, which is what you wanted anyway:
|
|
502
|
+
```bash
|
|
503
|
+
ffmpeg -i vo.wav -af "silenceremove=start_periods=1:start_silence=0.04:start_threshold=-45dB\
|
|
504
|
+
:detection=peak:stop_periods=-1:stop_duration=0.12:stop_threshold=-40dB,areverse,\
|
|
505
|
+
silenceremove=start_periods=1:start_silence=0.04:start_threshold=-45dB:detection=peak,areverse" tight.wav
|
|
506
|
+
```
|
|
507
|
+
- **Leave the payoff alone.** Do not narrate over the whole runtime. On a 22s cut, ~15s of speech
|
|
508
|
+
and ~7s of silence is right — the result beat needs a stretch where nothing is being said and the
|
|
509
|
+
screen is just doing the thing.
|
|
510
|
+
- **A VO may cross a cut.** A line that starts on a face and finishes over the app binds the two
|
|
511
|
+
shots together. Don't clip narration to shot boundaries.
|
|
512
|
+
- **Duck the bed to `data-volume` ≈ 0.2** and normalise each VO clip to about -16 LUFS. Target
|
|
513
|
+
**12–15 dB of speech over bed**, and verify it by measuring the two STEMS, because `ebur128` on
|
|
514
|
+
the finished mix cannot separate them.
|
|
515
|
+
|
|
516
|
+
### The bed, in both modes
|
|
517
|
+
|
|
518
|
+
- **Level the bed and measure it.** Target ~-16 LUFS integrated, peak < -1 dBFS
|
|
519
|
+
(`loudnorm=I=-16:TP=-1.5:LRA=11`). Fade in ≤0.6s, fade out over the last ~1.2s.
|
|
520
|
+
- **`data-volume` multiplies the already-normalised file.** A bed normalised to -16 LUFS sitting on
|
|
521
|
+
a layer at `data-volume="0.5"` ships at about **-22 LUFS** — quiet enough that a phone speaker in
|
|
522
|
+
a noisy room gets nothing. The 0.1–0.2 bed level everyone quotes is for ducking *under a
|
|
523
|
+
voiceover*, and there is no voiceover here. Normalise the file, then leave the layer at `1`, and
|
|
524
|
+
measure the RENDER (`ffmpeg -i out.mp4 -af ebur128=peak=true -f null -`) rather than the stem.
|
|
525
|
+
- **Source it free.** `vidfarm media bgm "<vibe>" --provider openverse` — prefer **CC0** so nothing
|
|
526
|
+
has to be credited on screen. A CC-BY track means an attribution line, and an attribution line is
|
|
527
|
+
brand chrome you just said you would not have.
|
|
528
|
+
### Ship TWO cuts, from ONE render
|
|
529
|
+
|
|
530
|
+
Do not put the music bed in the composition. **The composition is the voiceover-only master**, and
|
|
531
|
+
the bed is applied at export — which gives you both deliverables for the price of one render:
|
|
532
|
+
|
|
533
|
+
```bash
|
|
534
|
+
python3 export-versions.py renders/cut-vo.mp4 media/bgm.mp3 --bed-volume 0.22
|
|
535
|
+
# voiceover only cut-vo.mp4 publish this
|
|
536
|
+
# voiceover + music cut-music.mp4 review / fallback / watermarked
|
|
537
|
+
```
|
|
538
|
+
|
|
539
|
+
| Cut | What it's for |
|
|
540
|
+
|---|---|
|
|
541
|
+
| **`-vo`** (voiceover, no music) | **The one you publish.** TikTok, Reels and Shorts all let the poster attach a track from the platform's own library at upload — licensed through the platform's deals with the labels, so the music is cleared where viewers actually hear it, and the post gets whatever is trending *this week* rather than whatever was trending when you rendered |
|
|
542
|
+
| **`-music`** (voiceover + bed) | The review copy you send a client, the fallback for anywhere without an in-app music library, and the one that carries a **watermark** if you need one |
|
|
543
|
+
|
|
544
|
+
**Derive, never re-render.** `-c:v copy` means both files carry the byte-identical video stream, so
|
|
545
|
+
approving one approves the other. Two separate renders are two different videos that merely look
|
|
546
|
+
alike, and the frame you approved is not the frame you shipped. A watermark is the one thing that
|
|
547
|
+
breaks this — it has to be burned in, so that version gets its own encode and stops being
|
|
548
|
+
frame-identical. Say so in the handoff.
|
|
549
|
+
|
|
550
|
+
**`normalize=0` on the `amix`.** Left at its default, `amix` pulls the voice down to make room for
|
|
551
|
+
the bed and quietly undoes the levelling you just measured.
|
|
552
|
+
|
|
553
|
+
- **Say which kind of bed it is in the handoff.** If the poster will attach trending audio at
|
|
554
|
+
upload, the bed in the `-music` cut is a *review* bed. If the `-music` cut ships as is, the bed is
|
|
555
|
+
the bed. These are different decisions and the customer has to make it.
|
|
556
|
+
|
|
557
|
+
---
|
|
558
|
+
|
|
559
|
+
## Subtitle DNA — TikTok-native, or don't bother
|
|
560
|
+
|
|
561
|
+
Subtitles are not accessibility here. They are the **only** delivery system for the hook, the loop,
|
|
562
|
+
the payoff, the limit and the bait. Get them wrong and the video has no words at all.
|
|
563
|
+
|
|
564
|
+
**Native means the captions are TIMED TO THE VOICE.** One short phrase on screen at a time, turning
|
|
565
|
+
over as the person speaks. Not a static two-line block, not a designed card, not a lower third.
|
|
566
|
+
|
|
567
|
+
A per-word *highlight* is one way to do that and it is not required — **plain white text with no
|
|
568
|
+
accent at all is TikTok's own default auto-caption look**, and it is the right choice more often
|
|
569
|
+
than the coloured mechanics: over busy footage, over a UI that already has its own colours, and on
|
|
570
|
+
any beat where the words should carry no emphasis of their own.
|
|
571
|
+
|
|
572
|
+
### What is FIXED, and what VARIES
|
|
573
|
+
|
|
574
|
+
Native is a small set of invariants, not a house style. Ship the same colour, case, size and
|
|
575
|
+
mechanic on every video and the catalogue reads as a content farm — which is a worse failure than an
|
|
576
|
+
odd caption, because it is visible across the whole account at once.
|
|
577
|
+
|
|
578
|
+
| Fixed — this is what "native" means | Varies — pick per video |
|
|
579
|
+
|---|---|
|
|
580
|
+
| **TikTok Sans**, self-hosted, no fallback chain | **weight** 700 / 800 / 900 |
|
|
581
|
+
| captions are **timed to the voice** | **which mechanic** — including *none* |
|
|
582
|
+
| **no phrase plate**; text sits on the picture | **accent colour** |
|
|
583
|
+
| inside the **platform-safe core** | **size** within the standard's ~36–64px band, and **position** within the core |
|
|
584
|
+
| **verbatim**, ≤3–4 words a page, one cue at a time | **case** — uppercase, or sentence case like TikTok's own auto-captions |
|
|
585
|
+
|
|
586
|
+
**Hold one combination for the length of a single video; change it between videos.** Styling that
|
|
587
|
+
shifts mid-video reads as a bug; styling that never shifts across thirty videos reads as a factory.
|
|
588
|
+
|
|
589
|
+
### The six mechanics — all native, none of them "the" one
|
|
590
|
+
|
|
591
|
+
Ordered quietest to loudest. The first two carry **no accent colour whatsoever**.
|
|
592
|
+
|
|
593
|
+
| `--highlight` | What it does | Reads as |
|
|
594
|
+
|---|---|---|
|
|
595
|
+
| `none` | plain white; the whole page appears at once and the page turns do the work | TikTok's stock auto-captions. The calmest, and the one that never fights the footage |
|
|
596
|
+
| `reveal` | plain white, but each word switches on at its own moment | still word-timed, still no colour — a middle setting when the beat wants pace but not emphasis |
|
|
597
|
+
| `colour` | the active word takes the accent, settles back to white | the quiet accented default; safe on busy footage |
|
|
598
|
+
| `pop` | accent **plus a vertical stretch** | more energy, still calm enough for a talking beat |
|
|
599
|
+
| `underline` | an accent bar wipes under the active word | "reading along"; good under a number being read out |
|
|
600
|
+
| `pill` | a filled accent box behind the active word, dark text inside | the loudest; excellent over a bright UI where white-on-white struggles |
|
|
601
|
+
|
|
602
|
+
**Every accented one is engineered to take no horizontal space.** That is not a stylistic choice —
|
|
603
|
+
an inline-block word that grows sideways eats the gap beside it and `a restaurant` renders as
|
|
604
|
+
`arestaurant` on the beat the word fires. `pop` therefore scales on **Y only**; `pill` and
|
|
605
|
+
`underline` live inside the word's own box.
|
|
606
|
+
|
|
607
|
+
**`reveal` deliberately does not re-centre.** A word that has not fired yet is transparent but still
|
|
608
|
+
occupies its slot, so the visible words sit where they will end up rather than sliding left as each
|
|
609
|
+
one arrives. Reflowing to keep the visible text centred makes every word jump, which is far worse
|
|
610
|
+
than the asymmetry.
|
|
611
|
+
|
|
612
|
+
### Take the accent from the customer's own palette
|
|
613
|
+
|
|
614
|
+
`vidfarm capture` already wrote it. `capture/extracted/tokens.json` carries the site's CSS
|
|
615
|
+
variables, so the caption accent can come from the product rather than from your habits — different
|
|
616
|
+
per client, for free, and never brand chrome because it is one word at a time rather than a logo:
|
|
617
|
+
|
|
618
|
+
```bash
|
|
619
|
+
python3 -c "import json;d=json.load(open('capture/extracted/tokens.json'));print(d['cssVariables'])"
|
|
620
|
+
# dishcover.io → --accent #2b3eff · --accent-orange #ff6b00 · --accent-yellow #ffd500 · --accent-red #e5432e
|
|
621
|
+
```
|
|
622
|
+
|
|
623
|
+
Then **contrast-check it against the band you measured** — a mid-blue accent on a dark scene is a
|
|
624
|
+
word that disappears at the exact moment it matters. If the accent fails, use it as the `pill`
|
|
625
|
+
background with dark text instead of as the type colour; a filled box survives backgrounds that
|
|
626
|
+
coloured type does not.
|
|
627
|
+
|
|
628
|
+
**The highlight is a COLOUR SWAP, not a zoom.** A scaling active word is a CapCut habit, and it has
|
|
629
|
+
a concrete failure mode as well as a stylistic one: an inline-block word scales from its centre into
|
|
630
|
+
its neighbour, so `a restaurant` renders as `arestaurant` on the exact beat the word pops — and a
|
|
631
|
+
text stroke eats the remaining gap from both sides. Padding the words apart only trades collision
|
|
632
|
+
for letterspaced display type that stops reading as speech. Colour (plus, at most, a ~0.05em
|
|
633
|
+
vertical lift that costs no horizontal room) makes the collision **impossible** rather than tuned.
|
|
634
|
+
|
|
635
|
+
> ⚠️ **Verify the words actually animate, in the RENDER.** `--style word-pop` writes per-word
|
|
636
|
+
> `animation-delay` / `animation-duration` and leaves the keyframes to the render engine — and on
|
|
637
|
+
> some engine versions that binding silently does not happen, so the cue renders as **one static
|
|
638
|
+
> block with no active-word colour**. Nothing errors. `captions list` shows correct timings. The
|
|
639
|
+
> composition lints clean. You only find it by pulling ~7 consecutive frames from *inside a single
|
|
640
|
+
> cue* and looking:
|
|
641
|
+
>
|
|
642
|
+
> ```bash
|
|
643
|
+
> ffmpeg -i out.mp4 -ss <cue start> -t <cue length> -vf "fps=6,crop=1080:340:0:1180,tile=7x1" \
|
|
644
|
+
> -frames:v 1 cap-anim.png
|
|
645
|
+
> ```
|
|
646
|
+
>
|
|
647
|
+
> If all seven frames are identical, the format's whole subtitle design did not ship — **unless the
|
|
648
|
+
> mechanic is `none`, where identical frames are exactly right.** Check which one you chose before
|
|
649
|
+
> you go debugging. **Own the
|
|
650
|
+
> keyframes in the composition's own `<style>`** rather than depending on the engine — the emitter
|
|
651
|
+
> in Appendix C does this, and it is the one change that makes the animation a property of the file
|
|
652
|
+
> instead of a property of whatever version rendered it.
|
|
653
|
+
>
|
|
654
|
+
> Two traps inside the fix, both of which look like "the animation is still broken":
|
|
655
|
+
>
|
|
656
|
+
> 1. **`animation-delay` is measured from PAGE time, not from the layer's own start.** A cue at
|
|
657
|
+
> 9.1s with a relative delay has already run to completion by the time the layer is visible, so
|
|
658
|
+
> every word renders in its end state. Emit `animation-delay: <cue start + word offset>`.
|
|
659
|
+
> Symptom: cue 1 animates, every later cue is static — which is easy to misread as "it works".
|
|
660
|
+
> 2. **Don't fade the words in.** With a single translucent plate behind the phrase, a progressive
|
|
661
|
+
> reveal leaves an empty grey box on screen ahead of the words, and that reads as a loading
|
|
662
|
+
> state. Keep the whole phrase visible and pop only the ACTIVE word (scale + accent colour,
|
|
663
|
+
> settling back to white) — which is TikTok's own convention anyway. `animation-fill-mode:
|
|
664
|
+
> forwards`, base style = the settled state.
|
|
665
|
+
|
|
666
|
+
```bash
|
|
667
|
+
# Appendix C, with the dials shown. There is no default combination to copy —
|
|
668
|
+
# choose one per video and write it into the handoff so the next one differs.
|
|
669
|
+
python3 srt-to-cues.py composition.html cues.json \
|
|
670
|
+
--font "TikTok Sans" --weight 900 --font-size 64 \
|
|
671
|
+
--highlight colour --active-color "#FFD500" \ # or: none | reveal | pop | pill | underline
|
|
672
|
+
--y 57 --x 14 --width 72 --plate none # add --no-uppercase for sentence case
|
|
673
|
+
```
|
|
674
|
+
|
|
675
|
+
- **In SILENT mode, hand-author the SRT.** There is no narration to transcribe, and `--text` paging
|
|
676
|
+
spreads words evenly across a window, which desyncs from your cuts. Write the cue times against
|
|
677
|
+
the beat table.
|
|
678
|
+
- **In NARRATED mode, never hand-author them.** Captions must be verbatim and sit on the word, so
|
|
679
|
+
they are derived from the VO's own transcript — Appendix F does it. A karaoke highlight 200ms off
|
|
680
|
+
the voice is more distracting than no highlight at all.
|
|
681
|
+
|
|
682
|
+
> ⚠️ **`captions generate --srt` does not respect your cue boundaries.** It flattens the whole file
|
|
683
|
+
> into one word stream and re-pages it by `--max-words`, so words hop the cut: `the dish instead` +
|
|
684
|
+
> `chili crisp` renders as `dish instead chili crisp`, two beats late and over the wrong footage.
|
|
685
|
+
> That is invisible in `captions list` timings and obvious in a contact sheet, which is exactly why
|
|
686
|
+
> you build the sheet. **Always read the generated cue text back before you render.** When the cues
|
|
687
|
+
> are pinned to cuts rather than to speech — which in this format they always are — use the
|
|
688
|
+
> one-layer-per-cue emitter in **Appendix C** instead, and keep `captions generate` for the case
|
|
689
|
+
> where a real narration track exists.
|
|
690
|
+
|
|
691
|
+
- **A cue is legible at frame 0 — literally `data-start="0"`.** `first_frame_text` grades the
|
|
692
|
+
frame, not the intent, so a cue that opens at 0.15s fails it. Frame 0 is the thumbnail.
|
|
693
|
+
- **≤3 words per page, ≤7 words per cue.** Longer than that and the page turns faster than it reads.
|
|
694
|
+
- **One cue at a time, ever.** Two simultaneous text objects is this format's version of clutter.
|
|
695
|
+
### The caption standard is NOT this harness's to define
|
|
696
|
+
|
|
697
|
+
The font regime, the size band, the safe zone and the legal caption backgrounds are **core Vidfarm
|
|
698
|
+
standards**. Read them there and follow them; do not take a restatement from a format harness:
|
|
699
|
+
|
|
700
|
+
| Where | What it covers |
|
|
701
|
+
|---|---|
|
|
702
|
+
| `vidfarm.cc/skill.md` → *Standards* → **Captions** | the one-paragraph rule: imported display font, size band, safe zone, emptiest part of the frame, 3–5-word cues, the four legal backgrounds |
|
|
703
|
+
| installed `SKILL.md` → **"the TikTok-native caption standard"** | the same, with the `set_captions` presets named |
|
|
704
|
+
| `references/editor-workflows.md` → **"Size → scaled to the line"** and **"Font → the composition regime"** | the numbers, the bundled font list, and why an un-imported font is the slop look |
|
|
705
|
+
|
|
706
|
+
The parts that bind here, quoted rather than paraphrased: **the bundled display fonts only**
|
|
707
|
+
(Montserrat 700–900 default, **TikTok Sans**, Abel, Source Code Pro, Yesteryear); **~36–64px on a
|
|
708
|
+
1080-wide frame** — above ~64px is a *hook-word* size, one to three words on purpose; **the 8%–85%
|
|
709
|
+
safe zone**; and **exactly one of four backgrounds** — `outline`, `plain`, active-word
|
|
710
|
+
`spotlight`/`karaoke`, or a tight `highlight-solid` band.
|
|
711
|
+
|
|
712
|
+
> ⚠️ I had this format running at **76px**, outside the band. Measured against the same cut, 76px
|
|
713
|
+
> wrapped 2 of 18 pages and reached 85.5% of frame width; **64px wrapped none and reached 82%** —
|
|
714
|
+
> better on both counts. When a local harness and the core standard disagree, check the standard is
|
|
715
|
+
> not simply right first.
|
|
716
|
+
|
|
717
|
+
#### What this format ADDS to the standard
|
|
718
|
+
|
|
719
|
+
Three deltas, each earned, and each a *narrowing* of the standard rather than a departure from it:
|
|
720
|
+
|
|
721
|
+
1. **Self-host the face; never `<link>` it, and set `font-display: block`.** The standard says use
|
|
722
|
+
an imported display font. It does not say how, and the obvious how is broken for rendering:
|
|
723
|
+
Google Fonts' own snippet ends in `&display=swap`, which draws a fallback first. A frame-capture
|
|
724
|
+
render screenshots inside that window — this shipped **Montserrat for all 666 frames** while the
|
|
725
|
+
stylesheet returned `200` and named TikTok Sans. Pull the `woff2`, `@font-face` it from disk,
|
|
726
|
+
`font-display: block`, and no fallback chain (a chain means a missing font becomes a *different*
|
|
727
|
+
font with nothing to tell you).
|
|
728
|
+
2. **A tighter band than 8%–85%.** The safe zone says where text is *allowed*; the platform's own
|
|
729
|
+
UI covers part of it. For this format the working band is **y 15%–73%, x 14%–86%** — a subset,
|
|
730
|
+
not a contradiction. See *Subtitle placement* below for the chrome map.
|
|
731
|
+
3. **Of the four legal backgrounds, this format uses `outline` and `plain`.** No phrase plate: the
|
|
732
|
+
`highlight-solid` band is legal everywhere but it is the third-party-editor tell on a
|
|
733
|
+
reaction-led cut. The active-word `pill` mechanic is the standard's `spotlight`/`karaoke`
|
|
734
|
+
highlight and is fine.
|
|
735
|
+
|
|
736
|
+
⚠️ **One mechanic below is an extension, not a standard preset.** `underline` is not one of the four
|
|
737
|
+
legal backgrounds. It reads native and renders fine, but treat it as this harness's own and expect
|
|
738
|
+
`vidfarm qa` to have no opinion on it — if a director wants to stay strictly inside the core set,
|
|
739
|
+
use `none`, `reveal`, `colour`, `pop` or `pill`.
|
|
740
|
+
|
|
741
|
+
- **No plate. The text sits directly on the picture.** A translucent slab behind the phrase is the
|
|
742
|
+
loudest third-party-editor tell in the format — it is what makes a cut read as *made somewhere
|
|
743
|
+
else and uploaded*, and it gets worse the moment a page wraps to two lines, because the slab's
|
|
744
|
+
own edges become a second shape competing with the frame. TikTok's own captions have no plate.
|
|
745
|
+
Legibility comes from the letterform instead: **a dark stroke under the fill plus a soft diffuse
|
|
746
|
+
shadow.**
|
|
747
|
+
```css
|
|
748
|
+
-webkit-text-stroke: 3px #000; paint-order: stroke fill; /* stroke OUTSIDE the glyph */
|
|
749
|
+
text-shadow: 0 3px 12px rgba(0,0,0,.42), 0 1px 3px rgba(0,0,0,.55);
|
|
750
|
+
```
|
|
751
|
+
`paint-order: stroke fill` is the part people miss — without it the stroke paints *over* the
|
|
752
|
+
glyph and eats the face's counters, and TikTok Sans at weight 900 turns to mush.
|
|
753
|
+
- **Still measure the band — the measurement now sets the STROKE, not a plate.** Sample the
|
|
754
|
+
composited luma and variance in the caption band across every scene (Appendix B). A calm dark
|
|
755
|
+
band survives on shadow alone; a bright busy one (a white UI, a lit face) needs the full 3px.
|
|
756
|
+
One value for the whole video.
|
|
757
|
+
- **NOT the lower third. The lower third is under the platform's own UI.** Every platform pastes
|
|
758
|
+
chrome over your video that you never see in your own render:
|
|
759
|
+
|
|
760
|
+
| | blocked |
|
|
761
|
+
|---|---|
|
|
762
|
+
| top | 0–12% — status bar, "Following \| For You" tabs |
|
|
763
|
+
| bottom | 78–100% — username, post caption, music ticker, progress bar |
|
|
764
|
+
| right | 84–100% (roughly y 45–80%) — like / comment / share / profile rail |
|
|
765
|
+
|
|
766
|
+
So the **safe core is about y 15%–73%**, and the default every subtitle tool ships — a lower
|
|
767
|
+
third at y≈66 with a 14% box — puts its bottom half under the caption block. Put the caption box
|
|
768
|
+
at the **lowest position that still clears the chrome**: `--y 57` for a 14%-tall box.
|
|
769
|
+
- **Keep the box narrow enough that a long line can't reach the rail.** Centred text needs a
|
|
770
|
+
symmetric box, so `--x 14 --width 72` (14%–86%) is the widest that stays clear. Then size the
|
|
771
|
+
type so pages mostly fit on one line inside it — measure, don't guess:
|
|
772
|
+
```js
|
|
773
|
+
// in the render browser: how many pages wrap, and how far right do they reach?
|
|
774
|
+
[...document.querySelectorAll('[data-layer-kind="caption"]')].map(c => {
|
|
775
|
+
const w = [...c.querySelectorAll('[data-cap-word]')]
|
|
776
|
+
return { lines: c.querySelector('[data-vf-text-inline]').getClientRects().length,
|
|
777
|
+
right: Math.max(...w.map(x => x.getBoundingClientRect().right)) / 10.8 }})
|
|
778
|
+
```
|
|
779
|
+
Measured on the reference build: 88px wrapped 7 of 18 pages and reached 85.7%, 76px wrapped 2 and
|
|
780
|
+
reached 85.0%, **64px wrapped none and reached 82.0%.** Smaller type that fits beats bigger type
|
|
781
|
+
that wraps into the rail — and 64px is the top of the standard's band anyway.
|
|
782
|
+
- **Let the measurement rank the bands, but don't let it move the caption over a face.** Appendix G
|
|
783
|
+
scores every band in the safe core. It will usually report the TOP of the frame as calmest —
|
|
784
|
+
that reading means "forehead, hair and defocused background", and a caption there covers the
|
|
785
|
+
performance, which on this format is the content. Take the quiet band only when the subject
|
|
786
|
+
genuinely sits low in frame.
|
|
787
|
+
- **Never repeat what the app already says on screen.** If the subtitle and the UI render the same
|
|
788
|
+
words, drop the subtitle. The demo beat is allowed to be almost caption-free — that is correct,
|
|
789
|
+
not a bug.
|
|
790
|
+
|
|
791
|
+
---
|
|
792
|
+
|
|
793
|
+
## The rules
|
|
794
|
+
|
|
795
|
+
### Rule 1 — the reaction and the app must be about the same moment
|
|
796
|
+
|
|
797
|
+
A surprised face cut against a beat where nothing surprising happened is worse than no cutaway. The
|
|
798
|
+
face is a *label* for the frame before it. Choose the reaction after you have cut the demo, never
|
|
799
|
+
before, and cut on the emotional beat rather than at a round timecode.
|
|
800
|
+
|
|
801
|
+
### Rule 2 — align the demo CUT to the subtitle, not the subtitle to the demo
|
|
802
|
+
|
|
803
|
+
The subtitle is fixed by the beat table; the demo recording is a long take you own every frame of.
|
|
804
|
+
So when the subtitle names the query and the input is still empty, you do **not** move the cue —
|
|
805
|
+
you move the in-point. Lay one frame per second of the composited device clip into a strip, read
|
|
806
|
+
off the timecodes where the action actually happens (`input complete`, `result lands`,
|
|
807
|
+
`scroll begins`), and solve for the in-point:
|
|
808
|
+
|
|
809
|
+
```bash
|
|
810
|
+
ffmpeg -i phone-demo.mp4 -vf "fps=1,scale=150:-1,tile=<N>x1" -frames:v 1 timeline.png
|
|
811
|
+
```
|
|
812
|
+
|
|
813
|
+
```
|
|
814
|
+
in_point = <device-clip time of the action> - (<composition time the cue fires> - <beat start>)
|
|
815
|
+
```
|
|
816
|
+
|
|
817
|
+
Do this for the two anchors that matter — the moment the input is complete, and the moment the
|
|
818
|
+
result appears — and the whole beat locks. Getting this wrong is the defect the format produces
|
|
819
|
+
most often and the one an author is least likely to notice, because the subtitles read correctly
|
|
820
|
+
and the footage reads correctly; only the two together are wrong.
|
|
821
|
+
|
|
822
|
+
### Rule 3 — the demo shows one action, at real speed, uncut
|
|
823
|
+
|
|
824
|
+
Speed-ramping dead time is fine and expected. Cutting away at the instant the product does its work
|
|
825
|
+
is not — that is the exact frame the viewer is deciding on. One feature per video. A video covering
|
|
826
|
+
four features teaches none.
|
|
827
|
+
|
|
828
|
+
### Rule 4 — every number on screen was read off the demo footage
|
|
829
|
+
|
|
830
|
+
Prices, counts, distances, names: quote what the recording actually shows, exactly. This format
|
|
831
|
+
makes it easy to be honest, because the source of truth is playing in the same frame as the claim —
|
|
832
|
+
and easy to be caught, for the same reason. A number in the subtitle that contradicts the number in
|
|
833
|
+
the app is the defect this format produces most.
|
|
834
|
+
|
|
835
|
+
### Rule 5 — concede one true, unflattering limit, in beat 6
|
|
836
|
+
|
|
837
|
+
Coverage gaps, geography, price, what it can't do yet. Take it from the product's own site so it is
|
|
838
|
+
defensible. Over the app, not on a card. It is not a disclaimer — it is the beat that makes the
|
|
839
|
+
other twenty seconds believable.
|
|
840
|
+
|
|
841
|
+
### Rule 6 — claim the mechanism, never the outcome
|
|
842
|
+
|
|
843
|
+
Say what it does and what it costs. Don't promise the result. On money, health and appearance
|
|
844
|
+
topics it is also the claim that draws platform enforcement.
|
|
845
|
+
|
|
846
|
+
### Rule 7 — client work: nothing on screen the customer doesn't already say
|
|
847
|
+
|
|
848
|
+
Every claim traceable to their own site or app. No real third party named in a negative light. Real
|
|
849
|
+
names, faces and emails visible in their screenshots were a deliberate decision or they don't ship.
|
|
850
|
+
If the customer's own example query is broken, tell them — don't quietly film a different one and
|
|
851
|
+
let them discover it in the comments.
|
|
852
|
+
|
|
853
|
+
### Rule 8 — production floor
|
|
854
|
+
|
|
855
|
+
Butt cuts only · no logo, end card, URL or title card · no CTA button, pricing card, feature grid,
|
|
856
|
+
benefit chip row, frosted panel or caption plate · nothing on screen looks clickable · the ask is a subtitle line.
|
|
857
|
+
|
|
858
|
+
---
|
|
859
|
+
|
|
860
|
+
## Bulk-generation notes
|
|
861
|
+
|
|
862
|
+
The variant axis is **the query you type into the app**, paired with the reaction that matches it.
|
|
863
|
+
Same product, same beats, same crop, same bed — five different things a person was actually looking
|
|
864
|
+
for, and five faces that fit those five moments. That is five videos.
|
|
865
|
+
|
|
866
|
+
Varying the *reaction clip* alone gives you one video five times, and the algorithm treats
|
|
867
|
+
near-duplicates accordingly. Varying the *subtitle wording* alone is worse.
|
|
868
|
+
|
|
869
|
+
**Vary the caption treatment across the batch too** — mechanic, accent, weight, size. Not because
|
|
870
|
+
any one of them is better, but because thirty videos in identical captions read as one account
|
|
871
|
+
running a template, and that is exactly the read you are trying to avoid. It costs nothing: the
|
|
872
|
+
treatment is four flags.
|
|
873
|
+
|
|
874
|
+
Because beats 3–6 are one tracked composite, a variant costs one browser recording plus one
|
|
875
|
+
tracker pass — roughly two minutes of compute and $0. That is the reason this format is worth
|
|
876
|
+
building a harness for at all.
|
|
877
|
+
|
|
878
|
+
Use `vidfarm dedupe <file> --variants N` when the same cut posts to more than one account.
|
|
879
|
+
|
|
880
|
+
---
|
|
881
|
+
|
|
882
|
+
## Pre-flight checklist
|
|
883
|
+
|
|
884
|
+
**Sourcing**
|
|
885
|
+
- [ ] The reaction raws came off the `ugc-reaction` shelf and were chosen from a contact sheet, not from a description alone
|
|
886
|
+
- [ ] ONE actor carries every reaction beat, cast from their `actor_<uuid>` set — not four different faces under a first-person voiceover
|
|
887
|
+
- [ ] The actor was chosen for their ARC (a take per beat), not for one good-looking clip
|
|
888
|
+
- [ ] The actor id was read off a card in THIS build and confirmed to return that person's clips
|
|
889
|
+
- [ ] The shelf was paginated with `cursor`, not `offset` — you saw all of it, not page 1 twice
|
|
890
|
+
- [ ] Beats 1 and 2 are two different people
|
|
891
|
+
- [ ] The greenscreen raw's `sourceType` is `Display Greenscreen` — no identifiable celebrity or character is being keyed into a customer's ad
|
|
892
|
+
- [ ] The demo query was run by hand and confirmed to return results before filming
|
|
893
|
+
- [ ] The demo was recorded at the device screen's aspect ratio
|
|
894
|
+
- [ ] Every 16:9 reaction raw's 9:16 crop was chosen off a still, not centred by default
|
|
895
|
+
|
|
896
|
+
**The insert**
|
|
897
|
+
- [ ] The demo's in-point was solved against the cue times — the input completes while its own subtitle is up, and the result lands on the cue that names it
|
|
898
|
+
- [ ] The insert is tracked per frame — scrubbed at 3 points, the app does not slide under the bezel
|
|
899
|
+
- [ ] The insert is rotation-fitted — no green wedge down any edge, at any frame
|
|
900
|
+
- [ ] The thumb and the bezel are ON TOP of the app, not painted over
|
|
901
|
+
- [ ] The device beats are punched in, and the UI is readable at phone scale
|
|
902
|
+
- [ ] Every device beat uses the same crop
|
|
903
|
+
|
|
904
|
+
**Stickers** (skip if none — none is a valid answer)
|
|
905
|
+
- [ ] The product's own noun appears on screen at least twice — not just its UI
|
|
906
|
+
- [ ] Every sticker either annotates something on screen or supplies that missing subject; none is a badge, chip, icon or logo
|
|
907
|
+
- [ ] Any sticker that a viewer would have to puzzle over was cut — a confusing sticker is worse than none
|
|
908
|
+
- [ ] Cutouts were checked for partial alpha (hollow art → flat key, solid subject → smart key) and despilled
|
|
909
|
+
- [ ] Each sticker's `fit` matches its job — subject `contain`, annotation `fill` — checked on a frame, not in the markup
|
|
910
|
+
- [ ] One on screen at a time, and not on every beat
|
|
911
|
+
- [ ] Each sticker's window ends before the UI underneath it moves — checked frame by frame
|
|
912
|
+
- [ ] Hand-drawn, not geometric; wrapped in a `div.clip` so the engine shows and hides it
|
|
913
|
+
|
|
914
|
+
**Anatomy**
|
|
915
|
+
- [ ] Beat 1 is a situation, not a product claim; frame 0 works as a standalone thumbnail
|
|
916
|
+
- [ ] The product first appears in beat 3
|
|
917
|
+
- [ ] The result plays uncut and long enough to be believed
|
|
918
|
+
- [ ] The video cuts back to a face after the payoff, and back to the app after the face
|
|
919
|
+
- [ ] One honest limit, in beat 6, over the app
|
|
920
|
+
- [ ] One comment ask, in the final beat AND in the post caption
|
|
921
|
+
- [ ] Exactly one feature is covered
|
|
922
|
+
|
|
923
|
+
**Audio**
|
|
924
|
+
- [ ] The mode was chosen deliberately (silent, or narrated over a ducked bed)
|
|
925
|
+
- [ ] Every reaction raw is muted — no face is visibly saying words the subtitles don't
|
|
926
|
+
- [ ] The demo is muted; the customer's own narration does not ship
|
|
927
|
+
- [ ] Narrated: one pinned `--voice` and one `--style` string across every line
|
|
928
|
+
- [ ] Narrated: captions came from the VO transcript, and the DISPLAY text is the script, not the transcript
|
|
929
|
+
- [ ] Narrated: speech sits 12–15 dB over the bed, measured on the STEMS
|
|
930
|
+
- [ ] Narrated: the payoff beat has a stretch with no speech over it
|
|
931
|
+
- [ ] The bed is NOT in the composition — the render is the voiceover-only master
|
|
932
|
+
- [ ] Both cuts exported from one render; the `-vo` and `-music` video streams are identical (unless watermarked)
|
|
933
|
+
- [ ] The handoff names the `-vo` cut as the one to publish, with platform audio added at upload
|
|
934
|
+
- [ ] The bed is CC0 (or its attribution requirement was accepted deliberately)
|
|
935
|
+
- [ ] The bed was measured ON THE RENDER, not on the stem: ~-16 LUFS, peak < -1 dBFS (`data-volume` multiplies it)
|
|
936
|
+
- [ ] The handoff says whether the bed ships or gets replaced with platform audio at upload
|
|
937
|
+
|
|
938
|
+
**Subtitles**
|
|
939
|
+
- [ ] The caption box sits inside the platform-safe core (y 15–73%, x 14–86%) — NOT the lower third
|
|
940
|
+
- [ ] Page wrapping and right-edge reach were measured in the browser; nothing crosses into the action rail
|
|
941
|
+
- [ ] The core caption standard was followed (bundled font, ~36–64px, safe zone, one of the four backgrounds) — not a harness restatement of it
|
|
942
|
+
- [ ] The face is self-hosted with `font-display: block` and no fallback chain
|
|
943
|
+
- [ ] `document.fonts.check('900 88px "TikTok Sans"')` was asserted true in the render browser, not assumed
|
|
944
|
+
- [ ] A cue is legible at frame 0 (`data-start="0"`, not 0.15)
|
|
945
|
+
- [ ] ≤3–4 words per page, one cue on screen at a time, timed to the voice
|
|
946
|
+
- [ ] If the mechanic is accented, the words were confirmed to FIRE in the render — 7 consecutive frames from inside one cue, not identical (skip for `none`, where identical frames are correct)
|
|
947
|
+
- [ ] No plate, panel or translucent slab behind the captions — stroke + shadow only
|
|
948
|
+
- [ ] `paint-order: stroke fill` is set, so the stroke sits outside the glyph
|
|
949
|
+
- [ ] The active-word highlight is a colour swap, not a scale — no two words ever touch mid-pop
|
|
950
|
+
- [ ] The band was measured across every scene and ONE stroke weight was chosen for the whole video
|
|
951
|
+
- [ ] The caption treatment was CHOSEN for this video (mechanic, accent, weight, size, case) — not inherited from the last one
|
|
952
|
+
- [ ] The accent was contrast-checked against the measured band; if it failed, it moved to a `pill` background rather than staying as type colour
|
|
953
|
+
- [ ] In a batch: this video's caption treatment differs from its siblings
|
|
954
|
+
- [ ] No cue repeats words the app already has on screen in that frame
|
|
955
|
+
- [ ] Every number in a cue matches the number visible in the demo footage
|
|
956
|
+
|
|
957
|
+
**Whole-video review** — on the render, not the plan
|
|
958
|
+
- [ ] A contact sheet of ~12 stills was read as one image: one crop convention, one type scale, one accent colour
|
|
959
|
+
- [ ] The reaction beats and the device beats look like the same video — not two edits spliced
|
|
960
|
+
- [ ] Pacing is deliberate, not N identically-long beats; no join is jarring
|
|
961
|
+
- [ ] No frame rests empty >0.5s; nothing runs after the last word for more than ~1.2s
|
|
962
|
+
- [ ] Frames from two different scenes were compared (a frozen render passes duration and frame-count checks)
|
|
963
|
+
- [ ] Audio verified by measurement, not by "it sounds fine"
|
|
964
|
+
- [ ] The MP4's timestamp is newer than the last edit — you reviewed THIS cut, not the previous one
|
|
965
|
+
|
|
966
|
+
---
|
|
967
|
+
|
|
968
|
+
## Diagnosing a flop
|
|
969
|
+
|
|
970
|
+
| What the numbers say | Weak beat | Fix |
|
|
971
|
+
|---|---|---|
|
|
972
|
+
| Barely any views | 1 | The hook is a product claim, or frame 0 is a phone instead of a face |
|
|
973
|
+
| Views, mass exit at 3–6s | 3 | The app arrived before the situation landed, or the move is not legible |
|
|
974
|
+
| Watched through, no reaction | 4 | The result was described in a subtitle instead of shown playing |
|
|
975
|
+
| Good retention, dead comments | 7 | No ask, or the ask was "thoughts?" |
|
|
976
|
+
| Comments arguing it's fake | 5 / audio | A reaction that didn't match its frame, or an unmuted raw |
|
|
977
|
+
|
|
978
|
+
**Vary one beat at a time.** A batch where the query, the faces and the bed all changed at once
|
|
979
|
+
teaches you nothing.
|
|
980
|
+
|
|
981
|
+
---
|
|
982
|
+
|
|
983
|
+
## Appendix A — `track-insert.py`
|
|
984
|
+
|
|
985
|
+
Free, local, no account, numpy + ffmpeg only. Fits the demo to the device screen's own axes every
|
|
986
|
+
frame and pastes by mask.
|
|
987
|
+
|
|
988
|
+
```python
|
|
989
|
+
#!/usr/bin/env python3
|
|
990
|
+
"""
|
|
991
|
+
Composite an app-demo recording INTO the green screen of a device raw.
|
|
992
|
+
|
|
993
|
+
usage: track-insert.py <device.mp4> <demo.mp4> <out.mp4> [--demo-start SEC]
|
|
994
|
+
"""
|
|
995
|
+
import subprocess, sys, numpy as np
|
|
996
|
+
|
|
997
|
+
device, demo, out = sys.argv[1], sys.argv[2], sys.argv[3]
|
|
998
|
+
demo_start = 0.0
|
|
999
|
+
if "--demo-start" in sys.argv:
|
|
1000
|
+
demo_start = float(sys.argv[sys.argv.index("--demo-start") + 1])
|
|
1001
|
+
|
|
1002
|
+
def probe(f, *keys):
|
|
1003
|
+
r = subprocess.run(["ffprobe", "-v", "error", "-select_streams", "v",
|
|
1004
|
+
"-show_entries", "stream=" + ",".join(keys),
|
|
1005
|
+
"-of", "default=nw=1:nk=1", f], capture_output=True, text=True)
|
|
1006
|
+
return r.stdout.split()
|
|
1007
|
+
|
|
1008
|
+
W, H, fr = probe(device, "width", "height", "r_frame_rate")
|
|
1009
|
+
W, H = int(W), int(H)
|
|
1010
|
+
num, den = fr.split("/")
|
|
1011
|
+
FPS = round(float(num) / float(den), 3)
|
|
1012
|
+
|
|
1013
|
+
PW, PH = 720, 1530 # demo pre-scaled ONCE with a good filter; per-frame we only
|
|
1014
|
+
# nearest-resample down a few percent, so there is no shimmer
|
|
1015
|
+
demo_pipe = subprocess.Popen(
|
|
1016
|
+
["ffmpeg", "-v", "error", "-ss", str(demo_start), "-i", demo,
|
|
1017
|
+
"-vf", f"fps={FPS},scale={PW}:{PH}:flags=lanczos", "-pix_fmt", "rgb24",
|
|
1018
|
+
"-f", "rawvideo", "-"], stdout=subprocess.PIPE)
|
|
1019
|
+
dev_pipe = subprocess.Popen(
|
|
1020
|
+
["ffmpeg", "-v", "error", "-i", device, "-pix_fmt", "rgb24", "-f", "rawvideo", "-"],
|
|
1021
|
+
stdout=subprocess.PIPE)
|
|
1022
|
+
enc = subprocess.Popen(
|
|
1023
|
+
["ffmpeg", "-y", "-v", "error", "-f", "rawvideo", "-pix_fmt", "rgb24",
|
|
1024
|
+
"-s", f"{W}x{H}", "-r", str(FPS), "-i", "-", "-an",
|
|
1025
|
+
"-c:v", "libx264", "-pix_fmt", "yuv420p", "-crf", "18", out],
|
|
1026
|
+
stdin=subprocess.PIPE)
|
|
1027
|
+
|
|
1028
|
+
def dilate(m, n=2):
|
|
1029
|
+
for _ in range(n):
|
|
1030
|
+
m = (m | np.roll(m, 1, 0) | np.roll(m, -1, 0)
|
|
1031
|
+
| np.roll(m, 1, 1) | np.roll(m, -1, 1))
|
|
1032
|
+
return m
|
|
1033
|
+
|
|
1034
|
+
FSZ, DSZ = W * H * 3, PW * PH * 3
|
|
1035
|
+
fit = None
|
|
1036
|
+
last_demo = None
|
|
1037
|
+
n = 0
|
|
1038
|
+
while True:
|
|
1039
|
+
raw = dev_pipe.stdout.read(FSZ)
|
|
1040
|
+
if len(raw) < FSZ:
|
|
1041
|
+
break
|
|
1042
|
+
d = demo_pipe.stdout.read(DSZ)
|
|
1043
|
+
if len(d) < DSZ:
|
|
1044
|
+
d = last_demo # demo ran out — hold its last frame
|
|
1045
|
+
last_demo = d
|
|
1046
|
+
f = np.frombuffer(raw, np.uint8).reshape(H, W, 3).astype(np.float32)
|
|
1047
|
+
dm = np.frombuffer(d, np.uint8).reshape(PH, PW, 3).astype(np.float32)
|
|
1048
|
+
|
|
1049
|
+
r, g, b = f[..., 0], f[..., 1], f[..., 2]
|
|
1050
|
+
greenness = g - np.maximum(r, b)
|
|
1051
|
+
core = (greenness > 55) & (g > 80)
|
|
1052
|
+
|
|
1053
|
+
if core.sum() > 5000:
|
|
1054
|
+
ys, xs = np.nonzero(core)
|
|
1055
|
+
cx, cy = xs.mean(), ys.mean()
|
|
1056
|
+
cov = np.cov(np.stack([xs - cx, ys - cy])) # the screen's own axes
|
|
1057
|
+
w_, v_ = np.linalg.eigh(cov)
|
|
1058
|
+
major = v_[:, np.argmax(w_)]
|
|
1059
|
+
if major[1] < 0:
|
|
1060
|
+
major = -major
|
|
1061
|
+
sn, cs = major[0], major[1]
|
|
1062
|
+
px, py = xs - cx, ys - cy
|
|
1063
|
+
v = px * sn + py * cs # along the long axis
|
|
1064
|
+
u = px * cs - py * sn # across it
|
|
1065
|
+
lenV = (np.percentile(v, 99.5) - np.percentile(v, 0.5)) / 2
|
|
1066
|
+
lenU = (np.percentile(u, 99.5) - np.percentile(u, 0.5)) / 2
|
|
1067
|
+
expect = lenV * (PW / PH) # a thumb eats one side, so trust the LONG axis
|
|
1068
|
+
if lenU < expect * 0.92:
|
|
1069
|
+
lenU = expect
|
|
1070
|
+
fit = (cx, cy, cs, sn, lenU, lenV)
|
|
1071
|
+
|
|
1072
|
+
if fit is None:
|
|
1073
|
+
enc.stdin.write(f.astype(np.uint8).tobytes()); n += 1; continue
|
|
1074
|
+
cx, cy, cs, sn, lenU, lenV = fit
|
|
1075
|
+
|
|
1076
|
+
alpha = np.clip((greenness - 25) / 30.0, 0, 1)
|
|
1077
|
+
paint = dilate(core, 2) | (alpha > 0.15) # paste by MASK, not by box
|
|
1078
|
+
ys, xs = np.nonzero(paint)
|
|
1079
|
+
if len(xs) == 0:
|
|
1080
|
+
enc.stdin.write(f.astype(np.uint8).tobytes()); n += 1; continue
|
|
1081
|
+
|
|
1082
|
+
px, py = xs - cx, ys - cy
|
|
1083
|
+
vv = px * sn + py * cs
|
|
1084
|
+
uu = px * cs - py * sn
|
|
1085
|
+
sx = ((uu / lenU + 1) * 0.5 * (PW - 1)).astype(np.int32).clip(0, PW - 1)
|
|
1086
|
+
sy = ((vv / lenV + 1) * 0.5 * (PH - 1)).astype(np.int32).clip(0, PH - 1)
|
|
1087
|
+
f[ys, xs] = dm[sy, sx] # opaque: no green survives
|
|
1088
|
+
|
|
1089
|
+
enc.stdin.write(np.clip(f, 0, 255).astype(np.uint8).tobytes())
|
|
1090
|
+
n += 1
|
|
1091
|
+
|
|
1092
|
+
enc.stdin.close(); enc.wait()
|
|
1093
|
+
print(f"{out} {n} frames @ {FPS}fps ({n/FPS:.1f}s)")
|
|
1094
|
+
```
|
|
1095
|
+
|
|
1096
|
+
**Tuning it.** If the device raw's plate is a darker or bluer green, drop `greenness > 55` toward
|
|
1097
|
+
40. If the demo comes out mirrored or 90° out, the PCA picked the short axis — that only happens on
|
|
1098
|
+
a landscape device, and the fix is to swap `PW`/`PH`. If the screen is only *partly* green (a
|
|
1099
|
+
notch, a bezel reflection), raise the dilation from 2 to 3.
|
|
1100
|
+
|
|
1101
|
+
## Appendix B — measuring the caption band
|
|
1102
|
+
|
|
1103
|
+
Run this before you choose a caption treatment, on every scene, and put the numbers in the handoff.
|
|
1104
|
+
|
|
1105
|
+
```python
|
|
1106
|
+
import subprocess, sys, numpy as np
|
|
1107
|
+
for f in sys.argv[1:]:
|
|
1108
|
+
p = subprocess.run(["ffmpeg","-v","error","-i",f,"-vf","fps=1,scale=270:480",
|
|
1109
|
+
"-pix_fmt","rgb24","-f","rawvideo","-"], capture_output=True)
|
|
1110
|
+
b = np.frombuffer(p.stdout, np.uint8); n = len(b)//(270*480*3)
|
|
1111
|
+
fr = b[:n*270*480*3].reshape(n,480,270,3).astype(np.float32)
|
|
1112
|
+
luma = 0.2126*fr[...,0] + 0.7152*fr[...,1] + 0.0722*fr[...,2]
|
|
1113
|
+
band = luma[:, 298:374, :] # the 62-78% lower third
|
|
1114
|
+
print(f"{f} luma={band.mean():.1f} var={band.std():.1f}")
|
|
1115
|
+
```
|
|
1116
|
+
|
|
1117
|
+
| Band | Treatment |
|
|
1118
|
+
|---|---|
|
|
1119
|
+
| luma < ~70, var < ~42 | White type, shadow only — the stroke can drop to 1–2px |
|
|
1120
|
+
| luma > ~160, var < ~42 | White type, full 3px stroke; the shadow does the separating |
|
|
1121
|
+
| anything else | White type, full 3px stroke + shadow — this format is nearly always this row |
|
|
1122
|
+
|
|
1123
|
+
**Never a plate.** The measurement chooses how much stroke, not whether to put a box behind the
|
|
1124
|
+
words. See *Subtitle DNA*.
|
|
1125
|
+
|
|
1126
|
+
## Appendix C — `srt-to-cues.py` (one caption layer per SRT cue)
|
|
1127
|
+
|
|
1128
|
+
Use this instead of `vidfarm captions generate --srt` whenever the cue times are pinned to cuts.
|
|
1129
|
+
It emits the same word-pop markup the CLI emits, and fixes the two things that bite in this format:
|
|
1130
|
+
it **keeps every SRT cue as its own page** (no word crosses a cut), and it **carries its own
|
|
1131
|
+
`@keyframes`** so the word-by-word pop is a property of the file rather than of whichever render
|
|
1132
|
+
engine version picks it up.
|
|
1133
|
+
|
|
1134
|
+
```python
|
|
1135
|
+
#!/usr/bin/env python3
|
|
1136
|
+
"""
|
|
1137
|
+
Write word-pop caption layers into a HyperFrames composition, ONE LAYER PER SRT CUE.
|
|
1138
|
+
|
|
1139
|
+
`vidfarm captions generate --srt` flattens the SRT into a single word stream and
|
|
1140
|
+
re-pages it by --max-words, so words hop across cue boundaries ("the dish instead"
|
|
1141
|
+
+ "chili crisp" becomes "dish instead chili crisp"). When your cue times are pinned
|
|
1142
|
+
to CUTS rather than to speech, that is fatal. This emits the same markup the CLI
|
|
1143
|
+
emits, but keeps each SRT cue as its own page.
|
|
1144
|
+
|
|
1145
|
+
usage: srt-to-cues.py <composition.html> <captions.srt|cues.json>
|
|
1146
|
+
[--y 57] [--x 14] [--width 72] [--font-size 64] [--weight 900]
|
|
1147
|
+
[--highlight none|reveal|colour|pop|pill|underline] [--active-color '#FFD500']
|
|
1148
|
+
[--font 'TikTok Sans'] [--plate none|phrase] [--stroke 3]
|
|
1149
|
+
[--no-uppercase] [--track 5]
|
|
1150
|
+
|
|
1151
|
+
WHAT IS FIXED vs WHAT VARIES. Native means the FONT is TikTok Sans, the captions
|
|
1152
|
+
sit on the picture inside the platform-safe core, and the highlight fires
|
|
1153
|
+
word-by-word. It does NOT mean one colour, one case, one size and one mechanic on
|
|
1154
|
+
every video a studio ships — that is a house style, and at volume it reads as a
|
|
1155
|
+
content farm. Vary --highlight / --active-color / --weight / --font-size / --y
|
|
1156
|
+
between videos; hold one combination for the length of a single video.
|
|
1157
|
+
"""
|
|
1158
|
+
import re, sys, html, json
|
|
1159
|
+
|
|
1160
|
+
args = sys.argv[1:]
|
|
1161
|
+
comp, srt = args[0], args[1]
|
|
1162
|
+
def opt(name, default):
|
|
1163
|
+
return args[args.index(name) + 1] if name in args else default
|
|
1164
|
+
Y = opt("--y", "66")
|
|
1165
|
+
X = opt("--x", "8")
|
|
1166
|
+
WIDTH = opt("--width", "84")
|
|
1167
|
+
FONT_SIZE = opt("--font-size", "64") # core standard: ~36-64px on a 1080 frame
|
|
1168
|
+
ACTIVE = opt("--active-color", "#FFD500")
|
|
1169
|
+
TRACK = opt("--track", "5")
|
|
1170
|
+
FONT = opt("--font", "TikTok Sans")
|
|
1171
|
+
WEIGHT = opt("--weight", "900") # 700 | 800 | 900
|
|
1172
|
+
HILITE = opt("--highlight", "colour") # none | reveal | colour | pop | pill | underline
|
|
1173
|
+
PLATE_MODE = opt("--plate", "none") # none | phrase
|
|
1174
|
+
STROKE = opt("--stroke", "3") # px of dark outline, plate-free mode
|
|
1175
|
+
|
|
1176
|
+
# TikTok's own captions sit DIRECTLY on the picture. A translucent slab behind
|
|
1177
|
+
# the phrase is a third-party-editor tell — it is the thing that makes a cut read
|
|
1178
|
+
# as "made somewhere else and uploaded", and it gets worse the moment two lines
|
|
1179
|
+
# wrap, because the slab's own edges become a shape competing with the frame.
|
|
1180
|
+
# Legibility without a plate comes from the letterform itself: a dark stroke
|
|
1181
|
+
# under the fill (paint-order keeps it outside the glyph, so the face stays
|
|
1182
|
+
# crisp) plus a soft diffuse shadow to lift it off a busy background.
|
|
1183
|
+
if PLATE_MODE == "phrase":
|
|
1184
|
+
PLATE = ("padding:0.07em 0.46em 0.09em;border-radius:0.32em;"
|
|
1185
|
+
"background:rgba(0,0,0,0.34);")
|
|
1186
|
+
BG_STYLE = "highlight-translucent"
|
|
1187
|
+
else:
|
|
1188
|
+
PLATE = "padding:0;background:transparent;"
|
|
1189
|
+
BG_STYLE = "outline"
|
|
1190
|
+
# NOTE: no "sans-serif" / Montserrat in the stack on purpose. A fallback chain
|
|
1191
|
+
# means a missing font renders as a DIFFERENT font and nothing tells you — the
|
|
1192
|
+
# cut just quietly stops looking native. Self-host the face (see the @font-face
|
|
1193
|
+
# in the composition head) so there is nothing to fall back FROM.
|
|
1194
|
+
UPPER = "--no-uppercase" not in args
|
|
1195
|
+
|
|
1196
|
+
def tsec(t):
|
|
1197
|
+
h, m, rest = t.split(":")
|
|
1198
|
+
s, ms = rest.split(",")
|
|
1199
|
+
return int(h) * 3600 + int(m) * 60 + int(s) + int(ms) / 1000
|
|
1200
|
+
|
|
1201
|
+
# Two input shapes. An SRT carries page text only, so word timings are ESTIMATED
|
|
1202
|
+
# from word length — right for a silent, subtitle-only cut where the cues are
|
|
1203
|
+
# pinned to picture. A cues.json (from vo-align.py) carries a real time per word,
|
|
1204
|
+
# which is what you must use once a voiceover exists: a karaoke highlight that is
|
|
1205
|
+
# 200ms off the voice is more distracting than no highlight at all.
|
|
1206
|
+
cues = []
|
|
1207
|
+
if srt.endswith(".json"):
|
|
1208
|
+
for p in json.load(open(srt)):
|
|
1209
|
+
cues.append((p["start"], p["end"],
|
|
1210
|
+
" ".join(w["text"] for w in p["words"]), p["words"]))
|
|
1211
|
+
else:
|
|
1212
|
+
for block in re.split(r"\n\s*\n", open(srt).read().strip()):
|
|
1213
|
+
lines = [l for l in block.strip().splitlines() if l.strip()]
|
|
1214
|
+
if len(lines) < 2:
|
|
1215
|
+
continue
|
|
1216
|
+
tl = next((l for l in lines if "-->" in l), None)
|
|
1217
|
+
if not tl:
|
|
1218
|
+
continue
|
|
1219
|
+
a, b = [x.strip() for x in tl.split("-->")]
|
|
1220
|
+
text = " ".join(lines[lines.index(tl) + 1:]).strip()
|
|
1221
|
+
if text:
|
|
1222
|
+
cues.append((tsec(a), tsec(b), text, None))
|
|
1223
|
+
if cues:
|
|
1224
|
+
cues[0] = (0.0, cues[0][1], cues[0][2], cues[0][3]) # frame 0 must carry text
|
|
1225
|
+
|
|
1226
|
+
def layer(i, start, dur, text, timed=None):
|
|
1227
|
+
words = text.split()
|
|
1228
|
+
# spread the cue's time across its words by length, with a floor so short
|
|
1229
|
+
# words still get a readable beat
|
|
1230
|
+
weights = [max(len(w), 2) for w in words]
|
|
1231
|
+
total = sum(weights)
|
|
1232
|
+
# animation-delay is measured from PAGE time, not from the layer's own start.
|
|
1233
|
+
# A cue at 9.1s with a relative delay has already finished animating by the
|
|
1234
|
+
# time it becomes visible, and every word renders in its end state — which
|
|
1235
|
+
# looks exactly like "the animation is broken". So the delay is ABSOLUTE
|
|
1236
|
+
# (cue start + word offset). data-word-start/end stay relative, which is the
|
|
1237
|
+
# convention the CLI's own markup uses.
|
|
1238
|
+
#
|
|
1239
|
+
# The highlight is COLOUR plus a small vertical lift — deliberately no
|
|
1240
|
+
# horizontal scale. A scaling word grows from its centre into its neighbour:
|
|
1241
|
+
# "a restaurant" renders as "arestaurant" on the exact beat the word pops,
|
|
1242
|
+
# and a text stroke eats the remaining gap from both sides. Padding it apart
|
|
1243
|
+
# only trades one defect (collision) for another (letterspaced display type
|
|
1244
|
+
# that no longer reads as speech). TikTok's own word highlight is a colour
|
|
1245
|
+
# swap; keeping it that way makes the collision impossible instead of tuned.
|
|
1246
|
+
t, spans = 0.0, []
|
|
1247
|
+
for j, (w, wt) in enumerate(zip(words, weights)):
|
|
1248
|
+
if timed: # real word times from the voiceover
|
|
1249
|
+
a0 = timed[j]["start"] - start
|
|
1250
|
+
d = max(timed[j]["end"] - timed[j]["start"], 0.12)
|
|
1251
|
+
else: # estimated from word length
|
|
1252
|
+
a0, d = t, dur * wt / total
|
|
1253
|
+
# A highlight should last about as long as the word is spoken — except a
|
|
1254
|
+
# plain REVEAL, which is an on-switch, not a fade: stretched across a
|
|
1255
|
+
# 0.4s word it reads as the caption struggling to load.
|
|
1256
|
+
adur = 0.12 if HILITE == "reveal" else d
|
|
1257
|
+
spans.append(
|
|
1258
|
+
f'<span data-cap-word="true" data-word-start="{a0:.3f}" '
|
|
1259
|
+
f'data-word-end="{a0+d:.3f}" style="animation-delay:{start+a0:.3f}s;'
|
|
1260
|
+
f'animation-duration:{adur:.3f}s">{html.escape(w)}</span>')
|
|
1261
|
+
t += dur * wt / total
|
|
1262
|
+
inner = " ".join(spans)
|
|
1263
|
+
upper = "text-transform:uppercase;" if UPPER else ""
|
|
1264
|
+
return (
|
|
1265
|
+
f'<div class="clip" id="element_cap_srt_{i}" data-hf-id="element_cap_srt_{i}" '
|
|
1266
|
+
f'data-layer-mode="publish" data-layer-kind="caption" data-start="{start:.3f}" '
|
|
1267
|
+
f'data-duration="{dur:.3f}" data-track-index="{TRACK}" '
|
|
1268
|
+
f'data-label="{html.escape(text, quote=True)}" '
|
|
1269
|
+
f'data-text-background-style="{BG_STYLE}" data-text-background-color="#000000" '
|
|
1270
|
+
f'data-font-family="{FONT}" data-caption-animation="word-pop" '
|
|
1271
|
+
f'data-caption-uppercase="{1 if UPPER else 0}" '
|
|
1272
|
+
f'style="position:absolute;left:{X}%;top:{Y}%;width:{WIDTH}%;height:14%;z-index:{TRACK};'
|
|
1273
|
+
f'opacity:1;display:flex;align-items:center;justify-content:center;padding:3px;'
|
|
1274
|
+
f"font-family:'{FONT}', 'Noto Color Emoji';font-weight:{WEIGHT};"
|
|
1275
|
+
f'line-height:1.18;text-align:center;font-size:{FONT_SIZE}px;color:#ffffff;'
|
|
1276
|
+
f'background:transparent;--vf-cap-active:{ACTIVE};{upper}overflow:visible">'
|
|
1277
|
+
f'<span data-vf-text-lines="true" style="display:block;max-width:100%;text-align:inherit;'
|
|
1278
|
+
f'line-height:1.18;white-space:pre-wrap;overflow:visible">'
|
|
1279
|
+
f'<span data-vf-text-inline="true" style="display:inline;box-decoration-break:clone;'
|
|
1280
|
+
f'-webkit-box-decoration-break:clone;line-height:inherit;{PLATE}'
|
|
1281
|
+
f"font-family:'{FONT}', 'Noto Color Emoji';font-weight:{WEIGHT};color:#ffffff\">"
|
|
1282
|
+
f'{inner}</span></span></div>')
|
|
1283
|
+
|
|
1284
|
+
doc = open(comp).read()
|
|
1285
|
+
# The word-by-word pop itself. The render engine is SUPPOSED to drive this off
|
|
1286
|
+
# data-caption-animation + the per-word animation-delay/duration, and on some
|
|
1287
|
+
# versions it silently does not — the cue renders as one static block and the
|
|
1288
|
+
# video ships without the thing that makes it read as native. Owning the
|
|
1289
|
+
# keyframes here makes the animation a property of the file, not of the engine.
|
|
1290
|
+
#
|
|
1291
|
+
# The whole phrase is visible for the cue's whole life and the ACTIVE word pops
|
|
1292
|
+
# — that is TikTok's convention, and it is also the only version that survives a
|
|
1293
|
+
# single shared plate behind the phrase. A progressive fade-in leaves an empty
|
|
1294
|
+
# translucent box sitting on screen ahead of the words, which reads as a loading
|
|
1295
|
+
# state. So: fill-mode FORWARDS, base state is the settled state, and the
|
|
1296
|
+
# keyframes describe only the word's own moment.
|
|
1297
|
+
LEGIBILITY = ("" if PLATE_MODE == "phrase" else
|
|
1298
|
+
f"-webkit-text-stroke:{STROKE}px #000;paint-order:stroke fill;"
|
|
1299
|
+
"text-shadow:0 3px 12px rgba(0,0,0,.42),0 1px 3px rgba(0,0,0,.55);")
|
|
1300
|
+
|
|
1301
|
+
# ── The highlight mechanic ───────────────────────────────────────────────────
|
|
1302
|
+
# All four are TikTok-native; none of them is "the" house style. Vary this
|
|
1303
|
+
# ACROSS a batch and hold one for the length of a single video.
|
|
1304
|
+
#
|
|
1305
|
+
# Every one of them is engineered to take NO HORIZONTAL SPACE, because an
|
|
1306
|
+
# inline-block word that grows sideways eats the gap beside it and "a restaurant"
|
|
1307
|
+
# renders as "arestaurant" on the exact beat the word fires. `pop` scales on Y
|
|
1308
|
+
# only for that reason; `pill` and `underline` sit inside the word's own box.
|
|
1309
|
+
_HILITE = {
|
|
1310
|
+
"colour": """
|
|
1311
|
+
0%{transform:translateY(-.05em);color:var(--vf-cap-active,#FFD500)}
|
|
1312
|
+
45%{transform:translateY(0);color:var(--vf-cap-active,#FFD500)}
|
|
1313
|
+
80%{transform:translateY(0);color:var(--vf-cap-active,#FFD500)}
|
|
1314
|
+
100%{transform:translateY(0);color:#ffffff}""",
|
|
1315
|
+
"pop": """
|
|
1316
|
+
0%{transform:scaleY(1.18) translateY(-.04em);color:var(--vf-cap-active,#FFD500)}
|
|
1317
|
+
40%{transform:scaleY(1.05) translateY(0);color:var(--vf-cap-active,#FFD500)}
|
|
1318
|
+
75%{transform:scaleY(1);color:var(--vf-cap-active,#FFD500)}
|
|
1319
|
+
100%{transform:scaleY(1);color:#ffffff}""",
|
|
1320
|
+
"pill": """
|
|
1321
|
+
0%{background:var(--vf-cap-active,#FFD500);color:#111;transform:translateY(-.04em)}
|
|
1322
|
+
70%{background:var(--vf-cap-active,#FFD500);color:#111;transform:translateY(0)}
|
|
1323
|
+
100%{background:transparent;color:#ffffff;transform:translateY(0)}""",
|
|
1324
|
+
"underline": """
|
|
1325
|
+
0%{box-shadow:inset 0 -.10em 0 0 var(--vf-cap-active,#FFD500);color:#ffffff}
|
|
1326
|
+
70%{box-shadow:inset 0 -.22em 0 0 var(--vf-cap-active,#FFD500);color:#ffffff}
|
|
1327
|
+
100%{box-shadow:inset 0 -.22em 0 0 transparent;color:#ffffff}""",
|
|
1328
|
+
# PLAIN — no accent at all. This is TikTok's own auto-caption look, and it is
|
|
1329
|
+
# the right choice more often than the coloured ones: over busy footage, over a
|
|
1330
|
+
# UI that already has its own colours, and on any beat where the words should
|
|
1331
|
+
# carry no emphasis of their own. "reveal" is still word-timed (each word turns
|
|
1332
|
+
# on at its own moment); "none" turns the whole page on at once and lets the
|
|
1333
|
+
# page turns do the work.
|
|
1334
|
+
"reveal": """
|
|
1335
|
+
0%{opacity:0}
|
|
1336
|
+
100%{opacity:1}""",
|
|
1337
|
+
"none": None,
|
|
1338
|
+
}
|
|
1339
|
+
if HILITE not in _HILITE:
|
|
1340
|
+
sys.exit(f"--highlight must be one of {', '.join(_HILITE)}")
|
|
1341
|
+
# the pill needs its own breathing room and a radius; the others must not have it
|
|
1342
|
+
_PILL_BOX = ("padding:.02em .16em;border-radius:.18em;" if HILITE == "pill"
|
|
1343
|
+
else "padding:0 .02em;")
|
|
1344
|
+
|
|
1345
|
+
if _HILITE[HILITE] is None:
|
|
1346
|
+
# plain, page-timed: no per-word animation at all
|
|
1347
|
+
POP_CSS = f"""
|
|
1348
|
+
[data-cap-word]{{display:inline-block;{_PILL_BOX}{LEGIBILITY}}}
|
|
1349
|
+
"""
|
|
1350
|
+
else:
|
|
1351
|
+
_EASE = "linear" if HILITE == "reveal" else "cubic-bezier(.2,.9,.3,1.1)"
|
|
1352
|
+
_FILL = "both" if HILITE == "reveal" else "forwards"
|
|
1353
|
+
POP_CSS = f"""
|
|
1354
|
+
[data-cap-word]{{display:inline-block;{_PILL_BOX}animation-name:vfWordPop;
|
|
1355
|
+
{LEGIBILITY}
|
|
1356
|
+
animation-fill-mode:{_FILL};animation-timing-function:{_EASE}}}
|
|
1357
|
+
@keyframes vfWordPop{{{_HILITE[HILITE]}
|
|
1358
|
+
}}
|
|
1359
|
+
"""
|
|
1360
|
+
def strip_cues(h):
|
|
1361
|
+
"""Remove existing caption layers by BALANCING <div> tags — a regex that
|
|
1362
|
+
guesses the closing sequence silently leaves the old cues in place, and you
|
|
1363
|
+
end up rendering two caption tracks stacked on each other."""
|
|
1364
|
+
out, i = [], 0
|
|
1365
|
+
while True:
|
|
1366
|
+
j = h.find('<div class="clip" id="element_cap_', i)
|
|
1367
|
+
if j < 0:
|
|
1368
|
+
out.append(h[i:]); break
|
|
1369
|
+
out.append(h[i:j])
|
|
1370
|
+
depth, k = 0, j
|
|
1371
|
+
while k < len(h):
|
|
1372
|
+
if h.startswith("<div", k):
|
|
1373
|
+
depth += 1; k += 4
|
|
1374
|
+
elif h.startswith("</div>", k):
|
|
1375
|
+
depth -= 1; k += 6
|
|
1376
|
+
if depth == 0:
|
|
1377
|
+
break
|
|
1378
|
+
else:
|
|
1379
|
+
k += 1
|
|
1380
|
+
i = k
|
|
1381
|
+
return "".join(out)
|
|
1382
|
+
|
|
1383
|
+
doc = strip_cues(doc)
|
|
1384
|
+
# strip any PREVIOUS pop CSS by shape, not by its exact text — otherwise tuning
|
|
1385
|
+
# the keyframes appends a second copy and the older rule wins the cascade
|
|
1386
|
+
doc = re.sub(r'\n?\[data-cap-word\]\{.*?@keyframes vfWordPop\{.*?\}\}\n?', "", doc, flags=re.S)
|
|
1387
|
+
if "</style>" in doc:
|
|
1388
|
+
doc = doc.replace("</style>", POP_CSS + "</style>", 1)
|
|
1389
|
+
else:
|
|
1390
|
+
doc = doc.replace("</head>", f"<style>{POP_CSS}</style></head>", 1)
|
|
1391
|
+
blocks = "\n".join(layer(i, a, b - a, t, wd) for i, (a, b, t, wd) in enumerate(cues))
|
|
1392
|
+
idx = doc.rfind("</div>")
|
|
1393
|
+
doc = doc[:idx] + blocks + "\n" + doc[idx:]
|
|
1394
|
+
open(comp, "w").write(doc)
|
|
1395
|
+
print(f"wrote {len(cues)} caption cues, one per SRT cue")
|
|
1396
|
+
for a, b, t, _ in cues:
|
|
1397
|
+
print(f" {a:6.2f} -> {b:6.2f} {t}")
|
|
1398
|
+
```
|
|
1399
|
+
|
|
1400
|
+
## Appendix D — rendering notes
|
|
1401
|
+
|
|
1402
|
+
- **Render with the standalone `hyperframes` CLI, not only `vidfarm render local`.** The copy of
|
|
1403
|
+
hyperframes bundled inside `vidfarm-devcli` can lag the released one, and a stale engine can abort
|
|
1404
|
+
the render at *Extracting video frames* with `captured 0 of expected N frames` on media that is
|
|
1405
|
+
perfectly valid. Check both versions (`hyperframes doctor`, `vidfarm doctor`) before you spend an
|
|
1406
|
+
hour re-encoding footage that was never the problem.
|
|
1407
|
+
- **Pass `--video-frame-format png` on this format.** Every device beat is a UI recording, and JPEG
|
|
1408
|
+
frame extraction softens small type exactly where the payoff lives.
|
|
1409
|
+
- **Hand-authored `<video class="clip">` layers render fine and grade as zero.** `vidfarm qa` counts
|
|
1410
|
+
scenes off `data-layer-kind`, so a composition you wrote by hand reports `scenes: 0` and
|
|
1411
|
+
`first_frame_visual: black` while rendering perfectly. Give every media layer the attributes
|
|
1412
|
+
`vidfarm place` writes — `data-layer-kind="video"|"audio"`, `data-layer-mode="publish"`,
|
|
1413
|
+
`data-hf-id`, `data-end`, `data-label`, `data-volume` — and the same file grades correctly. This
|
|
1414
|
+
is a reporting bug in your markup, not in the video, and it is easy to "fix" in the wrong place.
|
|
1415
|
+
- `hyperframes render` reads `index.html`; `vidfarm` writes and lints `composition.html`. Keep the
|
|
1416
|
+
two in sync (`cp composition.html index.html`) or you will render a stale cut.
|
|
1417
|
+
- **`[sub_timeline_readiness_timeout] … did not become ready within 45000ms` is benign here.** This
|
|
1418
|
+
format has no GSAP timeline — the motion is CSS keyframes and video playback — so there is no
|
|
1419
|
+
sub-timeline to become ready. The render completes and is correct. Don't chase it.
|
|
1420
|
+
- **Check the output file's timestamp, not the exit code.** A backgrounded render can exit 0 having
|
|
1421
|
+
written nothing. `ls -la` the MP4 and confirm it is newer than your last edit before you review
|
|
1422
|
+
it — otherwise you will grade the previous cut and conclude your fix didn't work.
|
|
1423
|
+
|
|
1424
|
+
## Appendix E — a filled-in beat table
|
|
1425
|
+
|
|
1426
|
+
One 22.2s cut built end to end against this file, for a dish-search product with 140 restaurants
|
|
1427
|
+
indexed. It is here so the beat table above has a shape, not so you copy the words.
|
|
1428
|
+
|
|
1429
|
+
| # | Beat | t | Stream | Subtitle |
|
|
1430
|
+
|---|---|---|---|---|
|
|
1431
|
+
| 1 | Hook | 0.0–2.6 | Reaction — woman, hand on head, confused | "I didn't want" / "a restaurant" |
|
|
1432
|
+
| 2 | Wrong way | 2.6–4.8 | Reaction — man at a laptop, disbelief | "Google gave me" / "a list of fifty" |
|
|
1433
|
+
| 3 | Move | 4.8–9.1 | Device — search field, typing | "So I searched" / "the dish instead" / "chili crisp" |
|
|
1434
|
+
| 4 | Result | 9.1–14.6 | Device — results list | "It read the menus" / "not the reviews" / "I just picked one" |
|
|
1435
|
+
| 5 | Face | 14.6–16.8 | Reaction — hand over mouth, surprise | "Nine dollars" / "Moody Tongue" |
|
|
1436
|
+
| 6 | Scope | 16.8–19.4 | Device — results still scrolling | "NYC only" / "140 restaurants" |
|
|
1437
|
+
| 7 | Bait | 19.4–22.2 | Reaction — smiling | "Comment the dish" / "you can never find" |
|
|
1438
|
+
|
|
1439
|
+
Three things in that table are the harness working rather than the writer working:
|
|
1440
|
+
|
|
1441
|
+
- **Beat 4 says "I just picked one" instead of naming the dish and the price.** The app is showing
|
|
1442
|
+
the dish and the price in that same frame. Naming them would be the caption reading the screen
|
|
1443
|
+
aloud. The read-out is deferred to **beat 5**, where the footage is a face and there is nothing
|
|
1444
|
+
to duplicate — which is also where it lands hardest.
|
|
1445
|
+
- **Beat 6 is the customer's own published number, used against them.** "140 restaurants" is a
|
|
1446
|
+
small index and saying so out loud is what makes the previous fifteen seconds credible.
|
|
1447
|
+
- **Beat 2 characterises a search result, it does not name a competitor.** "A list of fifty" is a
|
|
1448
|
+
true thing about what happens; naming a named product doing something bad is Rule 7.
|
|
1449
|
+
|
|
1450
|
+
## Appendix F — `vo-align.py` (captions from the voiceover, in narrated mode)
|
|
1451
|
+
|
|
1452
|
+
Once a voice exists, captions are no longer a design decision — they are a synchronisation problem.
|
|
1453
|
+
This transcribes each beat's VO clip locally (free, keyless whisper), places it at that beat, and
|
|
1454
|
+
emits `cues.json`: balanced pages of ≤3 words, each word carrying an absolute timestamp. Feed that
|
|
1455
|
+
straight to Appendix C in place of the SRT.
|
|
1456
|
+
|
|
1457
|
+
```bash
|
|
1458
|
+
# beats.txt — one line per beat: <clip.wav>|<place-at-seconds>|<the script you fed the TTS>
|
|
1459
|
+
python3 vo-align.py vo/beats.txt cues.json --max-words 3
|
|
1460
|
+
python3 srt-to-cues.py composition.html cues.json --font "TikTok Sans" --font-size 88 --y 66
|
|
1461
|
+
```
|
|
1462
|
+
|
|
1463
|
+
**Timing from the transcript, TEXT from the script.** This split is the whole design. Local STT
|
|
1464
|
+
mishears proper nouns precisely where this format needs them — a real run turned
|
|
1465
|
+
*"Nine dollars. Moody Tongue."* into *"$9.00 multi-tongue"* — and a caption that ships the
|
|
1466
|
+
mishearing is worse than no caption. So the words on screen are always the ones you wrote for the
|
|
1467
|
+
TTS; only their timings come from what was heard. Where the transcript's word count matches the
|
|
1468
|
+
script's, the two are zipped word-for-word and the sync is exact (6 of 7 beats on the reference
|
|
1469
|
+
build); where it doesn't, the script's words are distributed across the *measured* speech span,
|
|
1470
|
+
which is still far better than guessing. The script prints which path each beat took — read it.
|
|
1471
|
+
|
|
1472
|
+
**Write the script line the way you want it on screen.** "140 restaurants" and "a hundred forty
|
|
1473
|
+
restaurants" sound identical and page very differently; the first is two words and one clean page,
|
|
1474
|
+
the second breaks as "A HUNDRED AND / FORTY RESTAURANTS". You control this for free at the moment
|
|
1475
|
+
you write the TTS input, and not at all afterwards.
|
|
1476
|
+
|
|
1477
|
+
**Balance pages inside a sentence.** Fixed-size chunking leaves one-word orphans — a 7-word
|
|
1478
|
+
sentence at 3 words a page gives `3,3,1`, and a single word alone for 400ms reads as a flicker
|
|
1479
|
+
rather than as emphasis. Split into `ceil(n/max)` pages of as-equal-as-possible size instead
|
|
1480
|
+
(`3,2,2`), and never let a page run across a full stop.
|
|
1481
|
+
|
|
1482
|
+
```python
|
|
1483
|
+
#!/usr/bin/env python3
|
|
1484
|
+
"""
|
|
1485
|
+
Turn per-beat voiceover clips into caption pages with REAL word timings.
|
|
1486
|
+
|
|
1487
|
+
With a voiceover in the mix, captions stop being a design decision and become a
|
|
1488
|
+
synchronisation problem: they must be verbatim and they must sit on the word.
|
|
1489
|
+
|
|
1490
|
+
Two halves, and keeping them apart is the whole point:
|
|
1491
|
+
|
|
1492
|
+
* TIMING comes from the transcript. Where the transcript's word count matches
|
|
1493
|
+
the script's, the two are zipped word for word — exact. Where it doesn't,
|
|
1494
|
+
the script's words are distributed across the MEASURED speech span, which is
|
|
1495
|
+
still far better than guessing at a duration.
|
|
1496
|
+
* TEXT comes from the SCRIPT you fed the TTS, never from the transcript. Local
|
|
1497
|
+
STT mishears proper nouns exactly where this format needs them most — a real
|
|
1498
|
+
run turned "Nine dollars. Moody Tongue." into "$9.00 multi-tongue" — and a
|
|
1499
|
+
caption that ships the mishearing is worse than no caption at all.
|
|
1500
|
+
|
|
1501
|
+
Input: a beats file, one line per beat: <clip.wav>|<place-at-seconds>|<script>
|
|
1502
|
+
Output: cues.json — pages of <=MAXW words, each word carrying an absolute time.
|
|
1503
|
+
|
|
1504
|
+
usage: vo-align.py <beats.txt> <out.json> [--max-words 3]
|
|
1505
|
+
"""
|
|
1506
|
+
import json, re, subprocess, sys, os
|
|
1507
|
+
|
|
1508
|
+
beats_file, out_path = sys.argv[1], sys.argv[2]
|
|
1509
|
+
MAXW = int(sys.argv[sys.argv.index("--max-words") + 1]) if "--max-words" in sys.argv else 3
|
|
1510
|
+
|
|
1511
|
+
def transcribe(wav):
|
|
1512
|
+
base = os.path.splitext(wav)[0] + ".stt"
|
|
1513
|
+
subprocess.run(["vidfarm", "stt", wav, "--engine", "whisper", "--out", base],
|
|
1514
|
+
capture_output=True)
|
|
1515
|
+
return json.load(open(base + ".json"))["words"]
|
|
1516
|
+
|
|
1517
|
+
pages, report = [], []
|
|
1518
|
+
for raw in open(beats_file):
|
|
1519
|
+
raw = raw.strip()
|
|
1520
|
+
if not raw or raw.startswith("#"):
|
|
1521
|
+
continue
|
|
1522
|
+
wav, at, script = [p.strip() for p in raw.split("|", 2)]
|
|
1523
|
+
at = float(at)
|
|
1524
|
+
heard = transcribe(wav)
|
|
1525
|
+
if not heard:
|
|
1526
|
+
continue
|
|
1527
|
+
span0, span1 = heard[0]["start"], heard[-1]["end"]
|
|
1528
|
+
script_words = script.split()
|
|
1529
|
+
|
|
1530
|
+
if len(heard) == len(script_words):
|
|
1531
|
+
timed = [{"text": s, "start": h["start"], "end": h["end"]}
|
|
1532
|
+
for s, h in zip(script_words, heard)]
|
|
1533
|
+
report.append(f"{os.path.basename(wav)}: exact ({len(heard)} words)")
|
|
1534
|
+
else:
|
|
1535
|
+
weights = [max(len(re.sub(r'\W', '', w)), 2) for w in script_words]
|
|
1536
|
+
total, t = sum(weights), span0
|
|
1537
|
+
timed = []
|
|
1538
|
+
for w, wt in zip(script_words, weights):
|
|
1539
|
+
d = (span1 - span0) * wt / total
|
|
1540
|
+
timed.append({"text": w, "start": t, "end": t + d})
|
|
1541
|
+
t += d
|
|
1542
|
+
report.append(f"{os.path.basename(wav)}: proportional "
|
|
1543
|
+
f"(heard {len(heard)}, script {len(script_words)})")
|
|
1544
|
+
|
|
1545
|
+
cur = []
|
|
1546
|
+
def flush():
|
|
1547
|
+
global cur
|
|
1548
|
+
if not cur:
|
|
1549
|
+
return
|
|
1550
|
+
pages.append({
|
|
1551
|
+
"start": round(at + cur[0]["start"], 3),
|
|
1552
|
+
"end": round(at + cur[-1]["end"], 3),
|
|
1553
|
+
"words": [{"text": re.sub(r'[,.!?;:]+$', "", w["text"]),
|
|
1554
|
+
"start": round(at + w["start"], 3),
|
|
1555
|
+
"end": round(at + w["end"], 3)} for w in cur],
|
|
1556
|
+
})
|
|
1557
|
+
cur = []
|
|
1558
|
+
|
|
1559
|
+
# Split into sentences first — a page that runs across a full stop reads as
|
|
1560
|
+
# one thought and isn't — then chunk each sentence into BALANCED pages.
|
|
1561
|
+
# Naive fixed-size chunking leaves one-word orphans ("...the reviews." ->
|
|
1562
|
+
# "It read the" / "menus not the" / "reviews"), and a single word alone on
|
|
1563
|
+
# screen for 400ms reads as a flicker, not as emphasis.
|
|
1564
|
+
# Break on COMMAS as well as full stops. A comma is the writer already telling
|
|
1565
|
+
# you where the thought divides, and honouring it is the difference between
|
|
1566
|
+
# "IT READ THE MENUS / NOT THE REVIEWS" and "IT READ THE / MENUS NOT / THE
|
|
1567
|
+
# REVIEWS" — same words, same timings, and only one of them is readable.
|
|
1568
|
+
fragments, s = [], []
|
|
1569
|
+
for w in timed:
|
|
1570
|
+
s.append(w)
|
|
1571
|
+
if re.search(r'[.!?,;:]$', w["text"]):
|
|
1572
|
+
fragments.append(s); s = []
|
|
1573
|
+
if s:
|
|
1574
|
+
fragments.append(s)
|
|
1575
|
+
|
|
1576
|
+
for frag in fragments:
|
|
1577
|
+
n = len(frag)
|
|
1578
|
+
# let a fragment run one word over the cap rather than split a clause
|
|
1579
|
+
# that fits: 4 words on one line beats 2 + 2 with a break mid-phrase
|
|
1580
|
+
k = 1 if n <= MAXW + 1 else -(-n // MAXW)
|
|
1581
|
+
base, extra = divmod(n, k) # spread the remainder over the first pages
|
|
1582
|
+
i = 0
|
|
1583
|
+
for p in range(k):
|
|
1584
|
+
size = base + (1 if p < extra else 0)
|
|
1585
|
+
cur = frag[i:i + size]
|
|
1586
|
+
i += size
|
|
1587
|
+
flush()
|
|
1588
|
+
|
|
1589
|
+
# a page whose last word is clipped short reads as a flicker; give every page a
|
|
1590
|
+
# floor, and never let one overlap the next
|
|
1591
|
+
for i, p in enumerate(pages):
|
|
1592
|
+
p["end"] = max(p["end"], p["start"] + 0.45)
|
|
1593
|
+
if i + 1 < len(pages):
|
|
1594
|
+
p["end"] = min(p["end"], pages[i + 1]["start"] - 0.02)
|
|
1595
|
+
|
|
1596
|
+
json.dump(pages, open(out_path, "w"), indent=1)
|
|
1597
|
+
for r in report:
|
|
1598
|
+
print(" " + r)
|
|
1599
|
+
print(f"{len(pages)} pages -> {out_path}")
|
|
1600
|
+
for p in pages:
|
|
1601
|
+
print(f" {p['start']:6.2f} -> {p['end']:6.2f} {' '.join(w['text'] for w in p['words'])}")
|
|
1602
|
+
```
|
|
1603
|
+
|
|
1604
|
+
## Appendix G — `caption-place.py` (where the captions go)
|
|
1605
|
+
|
|
1606
|
+
Scores every candidate band inside the platform-safe core and recommends one. Run it on the
|
|
1607
|
+
prepared beat clips, before you write a single cue.
|
|
1608
|
+
|
|
1609
|
+
```bash
|
|
1610
|
+
python3 caption-place.py media/s*.mp4
|
|
1611
|
+
```
|
|
1612
|
+
|
|
1613
|
+
```python
|
|
1614
|
+
#!/usr/bin/env python3
|
|
1615
|
+
"""
|
|
1616
|
+
Choose where the captions go: inside the platform-safe core, above the chrome.
|
|
1617
|
+
|
|
1618
|
+
THE CONSTRAINT THAT ACTUALLY DECIDES IT is not legibility — it is that TikTok,
|
|
1619
|
+
Reels and Shorts each paste their own UI over your video and you never see it in
|
|
1620
|
+
your own render:
|
|
1621
|
+
|
|
1622
|
+
top 0–12% status bar, "Following | For You" tabs
|
|
1623
|
+
bottom 78–100% username, post caption, music ticker, progress bar
|
|
1624
|
+
right 84–100% like / comment / share / profile rail (roughly y 45–80%)
|
|
1625
|
+
|
|
1626
|
+
So the safe core is about **y 15%–73%**, and the classic "lower third" — the
|
|
1627
|
+
default in every subtitle tool — is the single worst place to put text in
|
|
1628
|
+
vertical video. Half of it is under the caption block on at least one platform.
|
|
1629
|
+
|
|
1630
|
+
Inside the core, the default is the LOWEST band that still clears the chrome.
|
|
1631
|
+
That is where TikTok's own captions sit, and it is below the subject's face,
|
|
1632
|
+
which on this format is the content. Busyness is measured and reported, but it
|
|
1633
|
+
does not get to move the captions over someone's eyes: the top of the frame
|
|
1634
|
+
almost always scores "calmest" precisely because it is forehead, hair and
|
|
1635
|
+
defocused background, and a caption there covers the performance.
|
|
1636
|
+
|
|
1637
|
+
usage: caption-place.py <clip.mp4> [clip.mp4 ...] [--height 14] [--top 15] [--bottom 73]
|
|
1638
|
+
"""
|
|
1639
|
+
import subprocess, sys, numpy as np
|
|
1640
|
+
|
|
1641
|
+
files = [a for a in sys.argv[1:] if not a.startswith("--")]
|
|
1642
|
+
def opt(n, d):
|
|
1643
|
+
return float(sys.argv[sys.argv.index(n) + 1]) if n in sys.argv else d
|
|
1644
|
+
H_PCT = opt("--height", 14)
|
|
1645
|
+
SAFE_TOP = opt("--top", 15)
|
|
1646
|
+
SAFE_BOTTOM = opt("--bottom", 73)
|
|
1647
|
+
AIR = 2.0 # don't sit flush against the chrome
|
|
1648
|
+
|
|
1649
|
+
W, H = 240, 426
|
|
1650
|
+
def frames(f):
|
|
1651
|
+
p = subprocess.run(["ffmpeg", "-v", "error", "-i", f, "-vf", f"fps=2,scale={W}:{H}",
|
|
1652
|
+
"-pix_fmt", "rgb24", "-f", "rawvideo", "-"], capture_output=True)
|
|
1653
|
+
b = np.frombuffer(p.stdout, np.uint8)
|
|
1654
|
+
n = len(b) // (W * H * 3)
|
|
1655
|
+
return b[:n * W * H * 3].reshape(n, H, W, 3).astype(np.float32)
|
|
1656
|
+
|
|
1657
|
+
F = np.concatenate([frames(f) for f in files])
|
|
1658
|
+
luma = 0.2126 * F[..., 0] + 0.7152 * F[..., 1] + 0.0722 * F[..., 2]
|
|
1659
|
+
# Busyness = local DETAIL behind the text, not global variance. A band split
|
|
1660
|
+
# between a bright wall and a dark jacket has huge variance and reads fine.
|
|
1661
|
+
gx = np.abs(np.diff(luma, axis=2))
|
|
1662
|
+
gy = np.abs(np.diff(luma, axis=1))
|
|
1663
|
+
edge = np.zeros_like(luma)
|
|
1664
|
+
edge[:, :, :-1] += gx
|
|
1665
|
+
edge[:, :-1, :] += gy
|
|
1666
|
+
|
|
1667
|
+
box_h = int(H_PCT / 100 * H)
|
|
1668
|
+
def score(top_pct):
|
|
1669
|
+
a = int(top_pct / 100 * H)
|
|
1670
|
+
band = edge[:, a:a + box_h, :]
|
|
1671
|
+
per_frame = band.mean(axis=(1, 2))
|
|
1672
|
+
# worst case matters more than the average: one scene where the caption is
|
|
1673
|
+
# unreadable is a defect however calm the other six are
|
|
1674
|
+
return per_frame.mean(), np.percentile(per_frame, 90)
|
|
1675
|
+
|
|
1676
|
+
cands = [(t, *score(t)) for t in np.arange(SAFE_TOP, SAFE_BOTTOM - H_PCT + 0.01, 1.0)]
|
|
1677
|
+
default_top = SAFE_BOTTOM - H_PCT - AIR
|
|
1678
|
+
d_avg, d_p90 = score(default_top)
|
|
1679
|
+
quietest = min(cands, key=lambda c: c[1] + 0.5 * c[2])
|
|
1680
|
+
|
|
1681
|
+
print(f"platform-safe core {SAFE_TOP:.0f}%–{SAFE_BOTTOM:.0f}% caption box {H_PCT:.0f}% tall")
|
|
1682
|
+
print(f" chrome assumed: top 0–12% bottom 78–100% right rail 84–100%")
|
|
1683
|
+
print()
|
|
1684
|
+
print(f" RECOMMENDED --y {default_top:.0f} busy {d_avg:5.2f} (p90 {d_p90:5.2f})"
|
|
1685
|
+
f" [lowest band clearing the chrome]")
|
|
1686
|
+
print(f" quietest --y {quietest[0]:.0f} busy {quietest[1]:5.2f} (p90 {quietest[2]:5.2f})")
|
|
1687
|
+
if quietest[1] < d_avg * 0.8:
|
|
1688
|
+
print()
|
|
1689
|
+
print(f" NOTE: the quietest band scores {(1-quietest[1]/d_avg)*100:.0f}% calmer. Look at it before")
|
|
1690
|
+
print( " taking it — near the top of the frame that reading usually means")
|
|
1691
|
+
print( " 'forehead, hair and defocused background', and a caption there covers")
|
|
1692
|
+
print( " the face. Only move up if the subject genuinely sits low in frame.")
|
|
1693
|
+
print()
|
|
1694
|
+
print(f" {'top%':>5} {'busy':>7} {'p90':>7}")
|
|
1695
|
+
for t, m, p in cands[::3]:
|
|
1696
|
+
mark = " <- recommended" if abs(t - default_top) < 0.5 else ""
|
|
1697
|
+
print(f" {t:5.0f} {m:7.2f} {p:7.2f}{mark}")
|
|
1698
|
+
```
|
|
1699
|
+
|
|
1700
|
+
## Appendix H — `export-versions.py` (the two cuts)
|
|
1701
|
+
|
|
1702
|
+
```bash
|
|
1703
|
+
python3 export-versions.py renders/cut-vo.mp4 media/bgm.mp3 --bed-volume 0.22
|
|
1704
|
+
```
|
|
1705
|
+
|
|
1706
|
+
```python
|
|
1707
|
+
#!/usr/bin/env python3
|
|
1708
|
+
"""
|
|
1709
|
+
Ship TWO cuts from ONE render: voiceover-only, and voiceover + music bed.
|
|
1710
|
+
|
|
1711
|
+
Why two, and why this order:
|
|
1712
|
+
|
|
1713
|
+
* The VOICEOVER-ONLY cut is the one you publish. TikTok, Reels and Shorts all
|
|
1714
|
+
let the poster attach a track from the platform's own library at upload,
|
|
1715
|
+
licensed through the platform's deals with the labels — so the music is
|
|
1716
|
+
cleared where viewers actually hear it, and the post gets whatever is
|
|
1717
|
+
trending that week instead of whatever was trending when you rendered.
|
|
1718
|
+
* The MUSIC cut is the review copy and the fallback: what you send a client to
|
|
1719
|
+
approve, what you post anywhere without an in-app music library, and the one
|
|
1720
|
+
that carries a watermark if you need one.
|
|
1721
|
+
|
|
1722
|
+
The music version is derived, never re-rendered. `-c:v copy` means both files
|
|
1723
|
+
carry the BYTE-IDENTICAL video stream, so approving one approves the other. Two
|
|
1724
|
+
separate renders would be two different videos that merely look alike, and the
|
|
1725
|
+
frame you approved would not be the frame you shipped. (A watermark breaks this
|
|
1726
|
+
on purpose — see --watermark.)
|
|
1727
|
+
|
|
1728
|
+
usage: export-versions.py <render.mp4> <bed.mp3> [--bed-volume 0.22]
|
|
1729
|
+
[--out-dir DIR] [--name NAME] [--watermark PNG]
|
|
1730
|
+
"""
|
|
1731
|
+
import os, subprocess, sys
|
|
1732
|
+
|
|
1733
|
+
src, bed = sys.argv[1], sys.argv[2]
|
|
1734
|
+
def opt(n, d):
|
|
1735
|
+
return sys.argv[sys.argv.index(n) + 1] if n in sys.argv else d
|
|
1736
|
+
VOL = float(opt("--bed-volume", "0.22"))
|
|
1737
|
+
OUT = opt("--out-dir", os.path.dirname(src) or ".")
|
|
1738
|
+
NAME = opt("--name", os.path.splitext(os.path.basename(src))[0].removesuffix("-vo"))
|
|
1739
|
+
WM = opt("--watermark", None)
|
|
1740
|
+
|
|
1741
|
+
vo_path = os.path.join(OUT, f"{NAME}-vo.mp4")
|
|
1742
|
+
music_path = os.path.join(OUT, f"{NAME}-music.mp4")
|
|
1743
|
+
|
|
1744
|
+
def run(cmd):
|
|
1745
|
+
r = subprocess.run(cmd, capture_output=True, text=True)
|
|
1746
|
+
if r.returncode:
|
|
1747
|
+
sys.exit(f"failed: {' '.join(cmd)}\n{r.stderr[-1500:]}")
|
|
1748
|
+
|
|
1749
|
+
# 1. the publish master — untouched. If the render is already sitting at the
|
|
1750
|
+
# -vo path, leave it alone: ffmpeg refuses to read and write one file, and
|
|
1751
|
+
# "re-encode it to itself" would be the wrong fix anyway.
|
|
1752
|
+
if os.path.abspath(src) != os.path.abspath(vo_path):
|
|
1753
|
+
run(["ffmpeg", "-y", "-v", "error", "-i", src, "-c", "copy", vo_path])
|
|
1754
|
+
|
|
1755
|
+
# 2. the review / fallback cut. The bed is duration-matched to the video and
|
|
1756
|
+
# mixed UNDER the existing voiceover track; `normalize=0` on amix stops it
|
|
1757
|
+
# quietly pulling the voice down to make room for the music.
|
|
1758
|
+
afilter = (f"[1:a]volume={VOL},afade=t=in:st=0:d=0.6[bed];"
|
|
1759
|
+
f"[0:a][bed]amix=inputs=2:duration=first:normalize=0[a]")
|
|
1760
|
+
cmd = ["ffmpeg", "-y", "-v", "error", "-i", src, "-stream_loop", "-1", "-i", bed,
|
|
1761
|
+
"-filter_complex", afilter, "-map", "0:v", "-map", "[a]",
|
|
1762
|
+
"-c:a", "aac", "-b:a", "192k", "-shortest"]
|
|
1763
|
+
if WM:
|
|
1764
|
+
# a watermark has to be burned in, so this version gets its own video encode
|
|
1765
|
+
# and stops being frame-identical to the master. Say so in the handoff.
|
|
1766
|
+
cmd += ["-i", WM, "-filter_complex",
|
|
1767
|
+
afilter + ";[0:v][2:v]overlay=W-w-40:40:format=auto[v]",
|
|
1768
|
+
"-map", "[v]", "-c:v", "libx264", "-preset", "slow", "-crf", "18",
|
|
1769
|
+
"-pix_fmt", "yuv420p"]
|
|
1770
|
+
cmd = [c for c in cmd if c not in ("-map", "0:v")]
|
|
1771
|
+
else:
|
|
1772
|
+
cmd += ["-c:v", "copy"]
|
|
1773
|
+
run(cmd + [music_path])
|
|
1774
|
+
|
|
1775
|
+
def probe(f):
|
|
1776
|
+
r = subprocess.run(["ffmpeg", "-hide_banner", "-nostats", "-i", f,
|
|
1777
|
+
"-af", "ebur128=peak=true", "-f", "null", "-"],
|
|
1778
|
+
capture_output=True, text=True).stderr
|
|
1779
|
+
g = lambda k: next((l.split()[-2] for l in r.splitlines() if l.strip().startswith(k)), "?")
|
|
1780
|
+
return g("I:"), g("Peak:")
|
|
1781
|
+
|
|
1782
|
+
for label, f in (("voiceover only", vo_path), ("voiceover + music", music_path)):
|
|
1783
|
+
i, pk = probe(f)
|
|
1784
|
+
size = os.path.getsize(f) / 1e6
|
|
1785
|
+
print(f" {label:20s} {os.path.basename(f):32s} {size:6.1f} MB {i} LUFS peak {pk} dBFS")
|
|
1786
|
+
print()
|
|
1787
|
+
print(f" PUBLISH: {os.path.basename(vo_path)} — add music from the platform's own library at upload.")
|
|
1788
|
+
print(f" REVIEW : {os.path.basename(music_path)}"
|
|
1789
|
+
+ (" (watermarked; video re-encoded, not frame-identical)" if WM else ""))
|
|
1790
|
+
```
|
|
1791
|
+
|
|
1792
|
+
## Appendix I — `place-stickers.py` (stickers from a manifest)
|
|
1793
|
+
|
|
1794
|
+
Writes the wrapper markup and the pop-in CSS, and warns when two stickers would be on screen at
|
|
1795
|
+
once — which is clutter in this format and the kind of thing you only notice on the render.
|
|
1796
|
+
|
|
1797
|
+
```bash
|
|
1798
|
+
python3 place-stickers.py composition.html stickers.txt
|
|
1799
|
+
```
|
|
1800
|
+
|
|
1801
|
+
```python
|
|
1802
|
+
#!/usr/bin/env python3
|
|
1803
|
+
"""
|
|
1804
|
+
Write sticker layers into a HyperFrames composition from a manifest.
|
|
1805
|
+
|
|
1806
|
+
Two things this exists to get right:
|
|
1807
|
+
|
|
1808
|
+
* A bare <img> with data-start renders at t=0 and then never shows or hides —
|
|
1809
|
+
the engine does not manage it. Each sticker has to be wrapped in a
|
|
1810
|
+
`div.clip` exactly like a caption layer.
|
|
1811
|
+
* `.clip { inset: 0 }` fights any geometry you set, so the wrapper has to
|
|
1812
|
+
reset `inset:auto` BEFORE left/top/width/height.
|
|
1813
|
+
|
|
1814
|
+
Manifest, one line per sticker (# comments and blank lines ignored):
|
|
1815
|
+
|
|
1816
|
+
<src>|<start>|<duration>|<left%>|<top%>|<width%>|<height%>|<label>[|<rotate>][|<fit>]
|
|
1817
|
+
|
|
1818
|
+
`fit` is `contain` (default) or `fill`, and the choice is not cosmetic:
|
|
1819
|
+
|
|
1820
|
+
* SUBJECT stickers — a dish, a product — must keep their aspect: `contain`.
|
|
1821
|
+
* ANNOTATION stickers — a ring, an underline — must SQUASH to the shape they
|
|
1822
|
+
are marking: `fill`. A near-square ring set to `contain` inside a wide, short
|
|
1823
|
+
box shrinks to the box's HEIGHT and ends up a small circle floating next to
|
|
1824
|
+
the row instead of around it.
|
|
1825
|
+
|
|
1826
|
+
usage: place-stickers.py <composition.html> <stickers.txt>
|
|
1827
|
+
"""
|
|
1828
|
+
import os, re, sys
|
|
1829
|
+
|
|
1830
|
+
comp, manifest = sys.argv[1], sys.argv[2]
|
|
1831
|
+
|
|
1832
|
+
CSS = """
|
|
1833
|
+
.sticker{animation-name:vfStickIn;animation-duration:.42s;animation-fill-mode:both;
|
|
1834
|
+
animation-timing-function:cubic-bezier(.2,1.1,.3,1);transform-origin:center;
|
|
1835
|
+
filter:drop-shadow(0 10px 24px rgba(0,0,0,.35))}
|
|
1836
|
+
@keyframes vfStickIn{
|
|
1837
|
+
0%{opacity:0;transform:scale(.55) rotate(-14deg)}
|
|
1838
|
+
60%{opacity:1;transform:scale(1.06) rotate(2deg)}
|
|
1839
|
+
100%{opacity:1;transform:scale(1) rotate(0deg)}}
|
|
1840
|
+
"""
|
|
1841
|
+
|
|
1842
|
+
doc = open(comp).read()
|
|
1843
|
+
# re-runnable: drop any previous sticker layers and CSS
|
|
1844
|
+
doc = re.sub(r'<div class="clip" id="stk-[^"]*".*?</div>\s*', '', doc, flags=re.S)
|
|
1845
|
+
doc = re.sub(r'\n?\.sticker\{.*?@keyframes vfStickIn\{.*?\}\}\n?', '', doc, flags=re.S)
|
|
1846
|
+
doc = doc.replace("</style>", CSS + "</style>", 1)
|
|
1847
|
+
|
|
1848
|
+
rows, layers = [], []
|
|
1849
|
+
for i, raw in enumerate(open(manifest)):
|
|
1850
|
+
raw = raw.strip()
|
|
1851
|
+
if not raw or raw.startswith("#"):
|
|
1852
|
+
continue
|
|
1853
|
+
parts = [p.strip() for p in raw.split("|")]
|
|
1854
|
+
src, start, dur, left, top, w, h, label = parts[:8]
|
|
1855
|
+
rot = parts[8] if len(parts) > 8 else "0"
|
|
1856
|
+
fit = parts[9] if len(parts) > 9 else "contain"
|
|
1857
|
+
start, dur = float(start), float(dur)
|
|
1858
|
+
layers.append(
|
|
1859
|
+
f' <div class="clip" id="stk-{i}" data-hf-id="stk-{i}" data-layer-mode="publish" '
|
|
1860
|
+
f'data-layer-kind="image" data-start="{start}" data-duration="{dur}" '
|
|
1861
|
+
f'data-end="{round(start+dur,3)}" data-track-index="6" '
|
|
1862
|
+
f'data-label="{label}" '
|
|
1863
|
+
f'style="position:absolute;inset:auto;left:{left}%;top:{top}%;width:{w}%;height:{h}%;'
|
|
1864
|
+
f'z-index:6;transform:rotate({rot}deg)">'
|
|
1865
|
+
f'<img class="sticker" src="{src}" '
|
|
1866
|
+
f'style="width:100%;height:100%;object-fit:{fit};animation-delay:{start}s"></div>')
|
|
1867
|
+
rows.append((start, start + dur, label))
|
|
1868
|
+
|
|
1869
|
+
idx = doc.rfind("</div>")
|
|
1870
|
+
doc = doc[:idx] + "\n".join(layers) + "\n" + doc[idx:]
|
|
1871
|
+
open(comp, "w").write(doc)
|
|
1872
|
+
|
|
1873
|
+
rows.sort()
|
|
1874
|
+
print(f"{len(rows)} sticker(s) placed")
|
|
1875
|
+
for a, b, label in rows:
|
|
1876
|
+
print(f" {a:6.2f} -> {b:6.2f} {label}")
|
|
1877
|
+
# two stickers on screen at once is clutter in this format, and it is the kind of
|
|
1878
|
+
# thing you only notice on the render — so say it here instead
|
|
1879
|
+
for (a1, b1, l1), (a2, b2, l2) in zip(rows, rows[1:]):
|
|
1880
|
+
if a2 < b1:
|
|
1881
|
+
print(f" ⚠ OVERLAP: '{l1}' and '{l2}' are both on screen {a2:.2f}-{b1:.2f}")
|
|
1882
|
+
```
|
|
1883
|
+
|
|
1884
|
+
## Appendix J — `cast-actors.py` (group a shelf by actor)
|
|
1885
|
+
|
|
1886
|
+
Pulls the whole shelf (cursor-paginated), reads each card's `actor_<uuid>` out of its summary,
|
|
1887
|
+
and ranks actors by how many takes they have. Run it before you pick a single clip.
|
|
1888
|
+
|
|
1889
|
+
```bash
|
|
1890
|
+
python3 cast-actors.py --category ugc-reaction --min 4
|
|
1891
|
+
# ugc-reaction: 180 raws · 53 tagged actors · 3 untagged
|
|
1892
|
+
# actor_668363ea-… — 7 takes …
|
|
1893
|
+
```
|
|
1894
|
+
|
|
1895
|
+
```python
|
|
1896
|
+
#!/usr/bin/env python3
|
|
1897
|
+
"""
|
|
1898
|
+
Group a public-raws shelf by ACTOR, so you can cast a face that has enough takes.
|
|
1899
|
+
|
|
1900
|
+
A shelf is dozens of clips of a much smaller number of creators, and nothing in
|
|
1901
|
+
the taxonomy says "same face" — the `actor_<uuid>` tag does. This pulls the
|
|
1902
|
+
shelf, reads each card's actor id out of its summary/tags, and ranks the actors
|
|
1903
|
+
by how many takes they have.
|
|
1904
|
+
|
|
1905
|
+
This format needs FOUR reaction beats from one person, so the only actors worth
|
|
1906
|
+
casting from are the ones with 4+ takes. That list is short.
|
|
1907
|
+
|
|
1908
|
+
usage: cast-actors.py [--category ugc-reaction] [--min 2]
|
|
1909
|
+
env: VIDFARM_API_KEY
|
|
1910
|
+
"""
|
|
1911
|
+
import collections, json, os, re, sys, urllib.parse, urllib.request
|
|
1912
|
+
|
|
1913
|
+
def opt(n, d):
|
|
1914
|
+
return sys.argv[sys.argv.index(n) + 1] if n in sys.argv else d
|
|
1915
|
+
CATEGORY = opt("--category", "ugc-reaction")
|
|
1916
|
+
MIN = int(opt("--min", "2"))
|
|
1917
|
+
HOST = os.environ.get("VIDFARM_HOST", "https://vidfarm.cc")
|
|
1918
|
+
KEY = os.environ["VIDFARM_API_KEY"]
|
|
1919
|
+
|
|
1920
|
+
# The feed caps `limit` at 100 and paginates on `cursor` — NOT `offset`, which is
|
|
1921
|
+
# silently ignored and hands you page 1 again. A loop built on offset looks like
|
|
1922
|
+
# it works, reports a plausible number, and has seen half the shelf.
|
|
1923
|
+
rows, cursor = {}, None
|
|
1924
|
+
while True:
|
|
1925
|
+
params = {"category": CATEGORY, "limit": 100}
|
|
1926
|
+
if cursor:
|
|
1927
|
+
params["cursor"] = cursor
|
|
1928
|
+
q = urllib.parse.urlencode(params)
|
|
1929
|
+
req = urllib.request.Request(f"{HOST}/api/v1/public-raws?{q}",
|
|
1930
|
+
headers={"Authorization": f"Bearer {KEY}"})
|
|
1931
|
+
page = json.load(urllib.request.urlopen(req))
|
|
1932
|
+
batch = page.get("raws", [])
|
|
1933
|
+
if not batch:
|
|
1934
|
+
break
|
|
1935
|
+
for r in batch:
|
|
1936
|
+
rows[r["rawId"]] = r
|
|
1937
|
+
cursor = page.get("next_cursor")
|
|
1938
|
+
if not cursor:
|
|
1939
|
+
break
|
|
1940
|
+
|
|
1941
|
+
ACTOR_RE = re.compile(r"actor_[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}", re.I)
|
|
1942
|
+
actors, untagged = collections.defaultdict(list), 0
|
|
1943
|
+
for r in rows.values():
|
|
1944
|
+
hay = (r.get("summary") or "") + " " + json.dumps(r.get("tags") or {})
|
|
1945
|
+
m = ACTOR_RE.search(hay)
|
|
1946
|
+
# "no tag" means NOT TAGGED YET — never treat it as "a different person"
|
|
1947
|
+
if m:
|
|
1948
|
+
actors[m.group(0)].append(r)
|
|
1949
|
+
else:
|
|
1950
|
+
untagged += 1
|
|
1951
|
+
|
|
1952
|
+
print(f"{CATEGORY}: {len(rows)} raws · {len(actors)} tagged actors · {untagged} untagged")
|
|
1953
|
+
ranked = sorted(actors.items(), key=lambda kv: -len(kv[1]))
|
|
1954
|
+
for actor_id, clips in ranked:
|
|
1955
|
+
if len(clips) < MIN:
|
|
1956
|
+
break
|
|
1957
|
+
print(f"\n{actor_id} — {len(clips)} takes")
|
|
1958
|
+
for c in sorted(clips, key=lambda c: -(c.get("durationSeconds") or 0)):
|
|
1959
|
+
print(f" {round(c.get('durationSeconds') or 0, 1):5}s {c['rawId']}")
|
|
1960
|
+
print(f" {(c.get('description') or '')[:96]}")
|
|
1961
|
+
print("\nCast from an actor with enough takes to carry all four reaction beats,")
|
|
1962
|
+
print("then pick each beat's take by EMOTION — see the harness's casting table.")
|
|
1963
|
+
```
|