@officexapp/vidfarm-devcli 0.21.42 → 0.21.43

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,855 @@
1
+ # Unique Product Explainer Videos — reusable prompt & method
2
+
3
+ Turns **N customer URLs → N product-introduction videos that do not look like each other**,
4
+ for **$0.00**. Proven on 32 videos (2026-08-03, `vidfarm-testing/bespoke5/`).
5
+
6
+ The hard part is not making one good video. It's making thirty good videos that don't read
7
+ as thirty runs of the same template. Everything below exists to prevent convergence.
8
+
9
+ > ## ⚠️ Read this first
10
+ >
11
+ > **This is a GENERAL FLOW, not a fixed script.** It was captured from one specific run on
12
+ > one machine. Expect it to differ depending on your **machine** (paths, Chrome location,
13
+ > core count, ffmpeg build), your **harness** (how subagents are spawned, whether they can be
14
+ > resumed, background-task behaviour, context limits), and the **AI model** driving it
15
+ > (stronger models need less prescriptive briefs and catch more of their own defects; weaker
16
+ > ones need the checklists spelled out and more revision rounds).
17
+ >
18
+ > Treat the _structure_ as the durable part — split deterministic from judgement work, assign
19
+ > differentiation as an input, review by looking at frames, verify by measurement. Treat the
20
+ > exact commands, flags and file paths as things to re-derive locally. If a step fails, the
21
+ > reasoning behind it is written down so you can adapt rather than guess.
22
+ > use vidfarm.cc/skill.md in cost saving mode with kokoro tts and if image gen is needed, give me the prompt and i will manually generate the images for you to use.
23
+ >
24
+ > **In cost-saving mode, stickers come from IconScout, not from an image model.** `vidfarm
25
+ > iconscout "<thing>" --free` is a $0 designer catalog that needs no key and returns a finished
26
+ > transparent SVG — so the sticker requirement below costs nothing and needs no manual
27
+ > image-gen round trip. **Search with no `--style` filter**: `--style sticker` is nearly empty
28
+ > on the free tier and will fool you into thinking the catalog is bare (§8a has the numbers).
29
+ > Illustrated PEOPLE are the exception — those want `--asset illustration --premium` at ~$0.02.
30
+ > Reach for `sticker-pack --generate` only when the catalog genuinely has nothing that fits.
31
+ > Details and the free-tier credit caveat: §8a.
32
+
33
+ ---
34
+
35
+ ## ⭐ THE GOLDEN RULE — first 3 seconds & first frame
36
+
37
+ If you remember nothing else from this document, remember this. It is the single biggest
38
+ driver of view-through, and it is the rule that most often has to be enforced against an
39
+ agent's instincts.
40
+
41
+ 1. **The first 3 seconds must be visual eye candy.** Something is **moving and visually
42
+ interesting from frame 1** — not motion that begins at 1.5s. Lead with the most arresting
43
+ visual the brand has, not with exposition.
44
+ 2. **Never open on a wall of static text** — not a paragraph, not a dense document, not an
45
+ inert title card, not a table sitting still.
46
+ 3. **Frame 0 must work as a thumbnail**, because it will be used as one. That means a
47
+ **complete headline** — never a half sentence, never a fragment mid-animation.
48
+ 4. **These two are compatible:** a striking, already-moving visual _with_ a short complete
49
+ line over it. That is the target for every single video.
50
+ 5. **Say what the product IS by t=5s.** On a *product explainer* the hook is not enough. A
51
+ viewer who has never heard of the brand must be able to state what it does, in plain
52
+ English, within the first five seconds — said in the VO and shown on screen. An abstract
53
+ mood piece that only pays off at 15s fails, however beautiful the motion is. This is
54
+ compatible with rule 1: lead with the arresting visual **and** land the plain what-it-is
55
+ under it. (Explainers only — a sizzle reel or a brand film is not bound by this.)
56
+ 6. **If the concept genuinely needs a dense or static surface, earn it** — arrive at it
57
+ through motion instead of cutting to it cold.
58
+ 7. **Test it every time:** extract t=0, 1, 2, 3, 4, 5 and look at them. If a viewer would
59
+ swipe past any of them, **or still could not say what the product is at t=5**, redo the open.
60
+
61
+ ---
62
+
63
+ ## ⭐⭐⭐ THE OPEN MUST BE SIMPLE, AND IT SHOULD CARRY A STICKER (customer directive, 2026-08-14)
64
+
65
+ The customer's words: *"anytime we make explainer videos the start looks simple and no wall of
66
+ text, and ideally uses relevant stickers if budget allows."*
67
+
68
+ Rules 1–2 above already say "no wall of static text", and videos kept failing anyway — because
69
+ agents read "wall of text" as "a paragraph" and then open on a **dense interface**, a **six-field
70
+ form**, a **detailed table**, or **three stacked cards each carrying a sentence**. That is a wall of
71
+ text with a layout. This section is the enforceable version.
72
+
73
+ ### 1. Count the text runs in frame 0. Three or fewer.
74
+
75
+ A **text run** is any contiguous piece of copy the eye has to read: a headline, a label, a field
76
+ value, a caption, a footer URL. Frame 0 gets **at most three**, and one of them is the
77
+ plain-English line. If your open has a headline, a sub-line, four form labels and a footer, it has
78
+ seven, and it will be sent back.
79
+
80
+ Practical consequences, all learned the hard way:
81
+
82
+ - **One subject, read large.** One voice letter at 80 px beats four at 40 px. One control lit in a
83
+ quiet panel beats a form of a dozen fields. When something is removed, **make what remains
84
+ bigger** — a simplification that leaves the type the same size just creates dead space.
85
+ - **Never open on a form, a settings panel or a table.** If the product's landing state is one,
86
+ open on the single row that matters and let the rest arrive later, dimmed.
87
+ - **Supporting surfaces carry no readable copy in the open.** Cards behind the subject are shapes,
88
+ colour fields or one word — not sentences competing with the line.
89
+ - **The footer/URL band does not count as content.** If it is the only thing in the lower third,
90
+ the frame is empty, not simple.
91
+
92
+ ### 2. A relevant sticker belongs in the first beat
93
+
94
+ **Put a sticker on screen in the opening beat, ideally in frame 0**, and make it the thing the
95
+ first spoken line is about. It does three jobs at once: it is a moving object for rule 1, it says
96
+ what kind of product this is before any copy is read, and it gives the thumbnail a focal point that
97
+ type alone never gives it.
98
+
99
+ - **Relevant, not decorative.** The sticker is the subject of the sentence being spoken — an
100
+ envelope for "send them a link", a football for "fixtures", a receipt for "photograph a receipt".
101
+ A generic sparkle or checkmark is decoration and does not count.
102
+ - **If the subject is a person or an audience, show them as an illustrated die-cut figure.** The
103
+ Families Echo cut opens on a grandmother in frame 0, the grandfather joining at 1.2 s and the
104
+ grandchild at 2.7 s; the family then recurs through the film. That single change did more for
105
+ that video than any motion work.
106
+ - **Budget: the opening sticker is now nearly always FREE — there is no excuse for not having one.**
107
+ Look it up first: `vidfarm iconscout "<the thing>" --free` costs $0 and needs no key (credit
108
+ line required — §8a; search with NO `--style` filter, and use `--asset illustration --premium`
109
+ at ~$0.02 for an illustrated person). If the catalog has nothing that fits, a generated sheet is
110
+ ~$0.01–0.05 against a $0.10 cap, so buy it. Authored CSS/SVG is the last fallback, for when the
111
+ pack comes back wrong or the sticker is really a word-mark.
112
+ - **Recur, do not decorate once.** If a motif is worth opening on, bring it back through the film.
113
+ Aim for the subject to be present for most of the runtime with no gap longer than ~4 s.
114
+
115
+ ### 3. One sticker on screen at a time — and mind the handover
116
+
117
+ Two stickers in one place read as a rendering fault, not a transition. Clamp every sticker exit to
118
+ `min(out, nextIn − 0.18)` and **ease `power2.out`, never `power2.in`**. A slow-leaving ease is what
119
+ causes the overlap: with a 0.18 s gap and a cubic ease-in, the outgoing pill is still at ~0.42
120
+ opacity when the next one pops. Sample the sticker band at 0.1 s intervals across every handover and
121
+ confirm exactly one is legible.
122
+
123
+ ### 4. The check
124
+
125
+ Extract t=0 and count the text runs out loud. More than three, or no sticker in the first beat, is a
126
+ redo — before anything else about the open is discussed.
127
+
128
+ ---
129
+
130
+ ## ⭐⭐ THE PLAIN-ENGLISH LINE — enforce rule 5 as a written input (customer directive, 2026-08-12)
131
+
132
+ Rule 5 above was in this document already and **two of four videos still failed it** in batch 17.
133
+ The customer's words: *"they need to be more clear on what it is, early asap."* Rule 5 as prose is
134
+ not enough — agents read it, agree with it, and then open on their concept anyway, because the
135
+ concept is the interesting part and the plain line feels like a downgrade. Make it a **deliverable
136
+ you write for them, not a principle you ask them to honour.**
137
+
138
+ ### 1. You write the line. In the brief. Verbatim.
139
+
140
+ Before you launch the agent, write the sentence yourself into the brief:
141
+
142
+ > **PLAIN-ENGLISH LINE (must be understood by t=5s):** "Paste a YouTube link, get an MP3 file."
143
+
144
+ Not "explain what the product is". The actual sentence. One clause, no metaphor, no brand voice, a
145
+ noun and a verb. If you cannot write it in one clause, you do not understand the product yet and
146
+ neither will the agent.
147
+
148
+ ### 2. It must land in BOTH channels, and it must be the FIRST thing said
149
+
150
+ - **In the VO**, as the **first or second spoken line** — not the third, not after the hook.
151
+ - **On screen**, as type, in the same window.
152
+ - Captions alone do not count as "on screen" if they are just transcribing the VO — that is one
153
+ channel wearing two hats. There must be an anchor in the art.
154
+
155
+ ### 3. The metaphor is the failure mode — name it in the brief
156
+
157
+ Every failure looked the same: a strong, ownable **concept** occupying the first five seconds while
158
+ the product went unnamed. Both batch-17 failures were metaphors that pay off beautifully at t=10
159
+ and say nothing at t=3:
160
+
161
+ | Brand | What the open actually said | What a first-time viewer could NOT say |
162
+ | ----- | --------------------------- | --------------------------------------- |
163
+ | AudioFetcher | "One link in… audio come out", over abstract machine hardware | that it converts a YouTube video into an MP3 file, free, in the browser |
164
+ | IdeaBoxd | "Finish the idea without losing the idea" · "one sentence, typed badly, at a red light" | that it is a place to capture a rough thought and build it into a plan |
165
+
166
+ Both are *good lines*. Neither is an answer to "what is this?".
167
+
168
+ **The concept and the plain line are not in competition** — the concept is what makes it worth
169
+ watching, the plain line is what makes it worth installing. Run them together: the metaphor is the
170
+ picture, the plain line is the type over it. Say so in the brief, in those words, or the agent will
171
+ assume you want one or the other.
172
+
173
+ ### 4. Check it on the frames, not on the script
174
+
175
+ The script always looks like it explains the product, because you already know what the product is.
176
+ Extract t=0–5, read them as images, and **ask the question as a stranger**: from these six frames
177
+ and nothing else, what does this thing do? If the honest answer is "something about audio" or
178
+ "something about ideas", the open fails, whatever the VO says.
179
+
180
+ A useful harsher version: cover the frames and read only the caption text for the first 5 seconds.
181
+ If that text alone does not identify the product, it is not landing.
182
+
183
+ ### 5. Fixing it later costs a re-record
184
+
185
+ Unlike the visual defects in §6, this one cannot be fixed by editing the generator — the plain line
186
+ has to be **spoken**, so the VO is re-recorded and every caption timing moves. That is still cheap
187
+ ($0, Kokoro + whisper), but it is a full pass rather than a re-render. **Cheaper to write the line
188
+ into the brief than to add it in revision.** This is the single highest-value line in the brief.
189
+
190
+ ---
191
+
192
+ ## 🚀 Slim prompt (start here)
193
+
194
+ For a quick run, or to hand to another agent, this is the whole method compressed. Expand
195
+ into the full document when you need the commands or hit a bug.
196
+
197
+ ```text
198
+ Build N bespoke product-introduction videos, one per customer URL, for $0.
199
+ They must NOT look like each other — no shared template.
200
+
201
+ MAIN LOOP (you, deterministically — costs no agent capacity):
202
+ 1. For each URL: headless DPR-2 capture of 1440x900 scroll slices, pull the raw HTML,
203
+ build a 3x3 contact sheet, and sample the palette from the RENDERED hero
204
+ (luminance < 110 = dark-themed brand → build the video dark; else light).
205
+ 2. Write a DIFFERENTIATION list: every concept + motion device already used, plus the
206
+ generic shapes they occupy (rail through stages, filling a container, tessellating
207
+ lattice, text morphing, funnel, travelling scan line, cards flying apart, pinned
208
+ element with swapping background, rings aligning).
209
+
210
+ PER BRAND (one subagent each, in parallel):
211
+ Give each an ASSIGNED format (9:16 / 16:9 / 4:5 / 1:1), theme (light/dark + temperature),
212
+ energy, type system and VO voice (**default `af_heart`** on Kokoro — assign a different one
213
+ only when the brand gives you a reason) — but make it DERIVE ITS OWN CONCEPT from the brand's own
214
+ words and assets, checking the differentiation list first. Name the predictable answer it
215
+ must avoid. Tell it the palette you sampled is a hypothesis to verify against real CSS.
216
+
217
+ HARD RULES for every video:
218
+ - First 3 seconds = visual eye candy, motion from frame 1, never static body text.
219
+ - Frame 0 = complete headline, works as a thumbnail.
220
+ - THE OPEN IS SIMPLE: at most THREE text runs in frame 0, one of them the plain-English line.
221
+ One subject read LARGE — never a form, a settings panel, a table, or three cards each
222
+ carrying a sentence. Removing an element means enlarging what remains, not leaving a hole.
223
+ - A RELEVANT STICKER IN THE FIRST BEAT, ideally in frame 0 — the subject of the first spoken
224
+ line, not decoration. If the subject is a person or an audience, use an illustrated die-cut
225
+ figure. LOOK IT UP FIRST: `vidfarm iconscout "<thing>" --free` is a $0 designer catalog, no
226
+ key, finished transparent SVG, no keying work (credit line required). Search with NO --style
227
+ filter (--style sticker is nearly empty on the free tier); for an illustrated PERSON use
228
+ --asset illustration --premium (~$0.02). Only generate a sheet (~$0.01-0.05 against a $0.10
229
+ cap) when the catalog has nothing that fits. Then recur the motif.
230
+ - One sticker on screen at a time: clamp exits to min(out, nextIn-0.18), ease power2.OUT
231
+ (power2.in leaves the outgoing pill at ~0.42 opacity when the next pops — two stickers, one
232
+ spot, reads as a render fault).
233
+ - By t=5s the viewer must be able to say WHAT THE PRODUCT IS, in plain English.
234
+ YOU write that sentence into the brief verbatim ("Paste a YouTube link, get an MP3 file"),
235
+ it is the FIRST or second spoken line, and it appears as type in the art, not only in the
236
+ captions. Say explicitly that the concept and the plain line run TOGETHER — metaphor as the
237
+ picture, plain line as the type over it — or the agent ships the metaphor alone.
238
+ - No large flat dead regions; compose for the whole frame.
239
+ - Never let two contradictory numbers share a frame.
240
+ - CTA settled >= 2s before the end.
241
+ - Claim only what the site claims. No real third party named in a negative context.
242
+ - Budget: <= $0.10 per video, and most videos should now land at $0.00. Local Kokoro TTS,
243
+ local whisper word timings, local render, CC0 music beds, the customer's own graphics AND
244
+ `vidfarm iconscout --free` stickers are all free. The only sanctioned spend is 1-2
245
+ `vidfarm sticker-pack --generate` sheets (~$0.01-0.05 each) for stickers the catalog lacks,
246
+ or `iconscout --premium` (~$0.02 each) when a customer deliverable can't carry the free
247
+ tier's credit line. Never `vidfarm generate image`.
248
+ - USE STICKERS GENEROUSLY: 5-8 across the video, one per beat, each carrying the meaning of
249
+ the line being spoken. SEARCH BEFORE YOU GENERATE — `vidfarm iconscout "<thing>" --free`
250
+ (no --style filter) already has the envelope, receipt, football, calendar and person, in the
251
+ hundreds-to-thousands, for $0 and with no keying/despill/sheet-splitting work at all. A
252
+ generated die-cut pack is for objects only this brand has; CSS/SVG stickers are the last
253
+ fallback, and best for word-marks.
254
+ - Deliver work/<slug>/<slug>.mp4 + NOTES.md. No watermark (applied centrally).
255
+
256
+ REVIEW (you, not the agent — agents pass their own broken work):
257
+ Extract ~6 frames per video into a contact sheet and LOOK at it. Send back anything with
258
+ dead space, a placeholder that reads as a missing asset, contradictory numbers, superimposed
259
+ headlines, or an unlanded CTA. Then read t=0-5 AS A STRANGER: from those frames alone, what
260
+ does this thing do? "Something about audio" = the open failed, however good the VO reads. Verify audio by measurement (12-15 dB separation, peak < 0).
261
+
262
+ WATERMARK + DELIVER:
263
+ Tiled diagonal mark, embossed two-tone, two variants (dark pass stronger on light videos,
264
+ light pass weaker on dark videos). Apply with ffmpeg overlay using `-loop 1` on the PNG and
265
+ `-c:a copy`. Then verify each watermarked frame against its OWN clean frame (tiny diff)
266
+ while consecutive frames still differ a lot (motion preserved) — a missing `-loop 1`
267
+ silently freezes the whole video while duration, frame count and audio hash all still pass.
268
+ Keep the un-watermarked masters. Present the cumulative delivery list every time.
269
+ ```
270
+
271
+ ---
272
+
273
+ ## 0. The core idea
274
+
275
+ > **Differentiation is an INPUT, not a hope.**
276
+ > Agents given "make it good" converge. Agents given an assigned format, theme, energy,
277
+ > type system and voice — plus an explicit list of concepts already used — do not.
278
+
279
+ Split the work:
280
+
281
+ | Layer | Who does it | Why |
282
+ | ---------------------------------------------------------------- | --------------------------------------- | ---------------------------------------------------------------------------------------------------------------- |
283
+ | **Deterministic** — capture, palettes, contact sheets, HTML pull | **Main loop** | Costs no agent capacity, is identical for every brand, and lets agents spend their whole budget on creative work |
284
+ | **Judgement** — concept, layout, motion, script | **One subagent per brand, in parallel** | This is the part that must differ |
285
+ | **Review** — frame sheets, defect calls | **Main loop** | Agents self-report "verified, looks good" on defects that are obvious in a frame |
286
+ | **Watermark + delivery** | **Main loop** | Central, uniform, verifiable |
287
+
288
+ ---
289
+
290
+ ## 1. Prerequisites
291
+
292
+ ```bash
293
+ export HYPERFRAMES_PYTHON="<repo>/vidfarm-testing/.venv-kokoro/bin/python"
294
+ export HYPERFRAMES_SKIP_SKILLS=1 HYPERFRAMES_NO_TELEMETRY=1
295
+ ```
296
+
297
+ - `vidfarm` + `hyperframes` on PATH; `.venv-kokoro` holding `kokoro-onnx` + `soundfile`
298
+ - `fonts/*.woff2` (self-host; a Google Fonts `<link>` trips lint)
299
+ - `bgm/hp_c*.wav` — CC0 beds, already high-passed 200 Hz ×2
300
+ - Copy in: `capture.mjs`, `audio.sh`, `watermark.py`, `PLAYBOOK.md`, `DIFFERENTIATION.md`
301
+
302
+ Everything else is local and keyless: Kokoro TTS, whisper timings, local render, CC0 music.
303
+ **IconScout sticker search is also free and needs no key** (`vidfarm iconscout "<thing>"
304
+ --free`, no `--style` filter) — it is the first place to look for any generic object, and it is
305
+ what keeps a `$0.00` run at `$0.00`. The remaining sanctioned spend is the generated sticker
306
+ pack (§8a) at **<= $0.10 per video**, for the stickers the catalog does not have.
307
+ **Never** call `vidfarm generate image` — that is a different, unbudgeted job. If a full AI
308
+ image would genuinely help, write the prompt to `AI_PROMPTS-<slug>.md` and ship without it.
309
+
310
+ ---
311
+
312
+ ## 2. Phase 1 — recon (main loop, ~10 min for 12 sites)
313
+
314
+ Run in **waves of 4–5**; more than that and headless Chrome starts timing out.
315
+
316
+ ```bash
317
+ # a) headless DPR-2 capture — 1440x900 scroll slices (already 16:10, so one slice = one
318
+ # "browser card" with zero crop tuning)
319
+ LOCALE="en-US,en;q=0.9" node capture.mjs <slug> <url> > cap-<slug>.log 2>&1
320
+
321
+ # b) raw HTML for every site (server-rendered copy, CSS tokens, asset URLs)
322
+ curl -sL -A "<desktop UA>" -H "Accept-Language: en-US,en;q=0.9" --max-time 45 "$url" \
323
+ -o "recon/<slug>.html"
324
+
325
+ # c) contact sheet per site — ONE image read instead of 8
326
+ ffmpeg -y -pattern_type glob -i "shots/<slug>/home-s*.png" \
327
+ -vf "scale=440:-1,tile=3x3:margin=6:padding=6:color=0x888888" -frames:v 1 "sheets/<slug>.png"
328
+
329
+ # d) palette sampled from the RENDERED hero (not the CSS) — this is what reveals dark vs light
330
+ ```
331
+
332
+ Palette sampler: most-common colours of `home-s00.png` downsampled 8×; luminance < 110 = DARK.
333
+
334
+ **Treat the sampled palette as a hypothesis, not fact.** Across 32 videos agents overturned
335
+ it perhaps a third of the time and were right every time — it had been sampled off a hero
336
+ _photograph_, a seasonal login page, or Google Play's own chrome. Always tell the agent to
337
+ confirm against the compiled CSS custom properties and to say so if it overturns.
338
+
339
+ **Special cases seen in the wild:**
340
+
341
+ - **Blank/tiny capture** → it's an app whose landing state is a form. The agent must _drive_
342
+ it headlessly (adapt `capture.mjs` into `work/<slug>/drive.mjs`; never edit the shared one).
343
+ - **Tiny HTML shell** → SPA. Harvest the JS bundle for asset paths + verbatim copy, and the
344
+ compiled CSS for tokens. **Check `manifest.json` → `screenshots[]` FIRST** — it repeatedly
345
+ held ten real store screenshots on apps I'd written off as unreachable.
346
+ - **No website at all** (Play/App Store listing only) → the listing IS the asset base. Pull
347
+ every screenshot at max resolution. **Do not use the store's chrome colour as the brand.**
348
+ - **Animated WebP/GIF on the site** → free motion footage. Extract frames with Pillow → build
349
+ a sprite sheet → drive it off the paused timeline. Unlike `<video>` it can be base64-inlined
350
+ and stays seek-safe.
351
+ - **A real screen recording in the bundle** → that's your A-roll. Best asset find of the run.
352
+
353
+ ---
354
+
355
+ ## 3. Phase 2 — the differentiation matrix
356
+
357
+ Maintain `DIFFERENTIATION.md` across the whole engagement. Two sections:
358
+
359
+ **(a) Concepts already used** — a table of brand → concept → motion device, appended after
360
+ every batch. Agents must read it and must not reuse or reskin any row.
361
+
362
+ **(b) Forbidden generic shapes** — the abstractions those concepts occupy. After 20 videos
363
+ this list was:
364
+
365
+ > a progress rail advancing through N labelled stages · items dropping/filling a container ·
366
+ > a grid/lattice tessellating as a transition · text morphing into other text · a funnel
367
+ > narrowing many to one · a scan/sweep line travelling across a surface and changing it ·
368
+ > cards flying apart and snapping back together · one element pinned while the background
369
+ > swaps · concentric rings resolving into alignment
370
+
371
+ Also steer each agent off **its own most predictable answer** by name: no waveform for a
372
+ music app, no node-graph for something called Nodebase, no word-flip-to-translation for a
373
+ language app, no stage-rail for a project tool.
374
+
375
+ **Where good concepts come from:** every strong one came from something the brand already
376
+ said or showed — umadum's own "No more FOMO", SkillDiscs' "two front doors", Halvy's "split
377
+ costs right down the middle", Sync.camera's "reveal live, or at the end". Tell the agent to
378
+ find the one idea that is _theirs_ and build the motion around it.
379
+
380
+ ---
381
+
382
+ ## 4. Phase 3 — the agent brief (THE REUSABLE PROMPT)
383
+
384
+ One subagent per brand, all launched in parallel. Template:
385
+
386
+ ---
387
+
388
+ > Build ONE bespoke product-introduction video for **<BRAND>** (<URL>). Slug: `<slug>`.
389
+ >
390
+ > Read FIRST, in this order: `PLAYBOOK.md` (the $0 pipeline, the **first-3-seconds hard
391
+ > requirement**, and the renderer-bug section) then `DIFFERENTIATION.md` (concepts already
392
+ > used — do not reuse; you derive your own concept from the site).
393
+ >
394
+ > Set up: `export HYPERFRAMES_PYTHON="$PWD/../.venv-kokoro/bin/python" HYPERFRAMES_SKIP_SKILLS=1 HYPERFRAMES_NO_TELEMETRY=1`
395
+ >
396
+ > ## Recon done for you
397
+ >
398
+ > - `sheets/<slug>.png` — contact sheet of all N slices. **Read this first** (one image read),
399
+ > then read only the 2–4 slices you actually want from `shots/<slug>/`.
400
+ > - `recon/<slug>.html` — raw HTML. Mine for verbatim copy, asset URLs, compiled CSS tokens.
401
+ > - Palette sampled from the rendered hero: stage `<hex>`, panels `<hex>`. **<DARK|LIGHT>.**
402
+ > Verify against the real CSS and overturn it if it's wrong — say so if you do.
403
+ > - <any special-case note: SPA / drive-the-app / store-listing-only / animated WebP>
404
+ >
405
+ > ## Assigned direction — the other agents in this batch have very different briefs; do not converge
406
+ >
407
+ > - **Format: <1080x1920 | 1920x1080 | 1080x1350 | 1080x1080>.**
408
+ > - **Theme: <LIGHT|DARK>** <+ the specific temperature: warm charcoal vs cold slate vs pure black>
409
+ > - **Energy: <one line>.** <Name which other videos in the batch are calm/punchy so this one
410
+ > leans the right way.>
411
+ > - **Type:** <2–3 candidate faces from `../fonts/`>, matched to the site's real faces.
412
+ > - **VO: `af_heart`** (the default — see §4a) _or_ `<bf_emma|am_adam|am_michael|bm_george>`
413
+ > when the brand calls for it — <why this brand is an exception, if it is>.
414
+ > - **Concept: derive it yourself** from the site's own strongest idea. Check
415
+ > `DIFFERENTIATION.md` first. <Name the predictable answer to avoid.>
416
+ >
417
+ > ## Accuracy (non-negotiable)
418
+ >
419
+ > Claim only what the site claims — platform, price, ratings, compliance. Don't upgrade
420
+ > "built on X guidelines" into an endorsement. Never let two contradictory numbers share a
421
+ > frame. <Any brand-specific care: no fear-selling, no medical authority, no named third
422
+ > party in a negative context, keep humour off real identifiable businesses.>
423
+ >
424
+ > Aim <20–30>s. Deliver `work/<slug>/<slug>.mp4` + `NOTES.md` (concept in one sentence, VO
425
+ > script, assets and where they came from, bed id, customer flags). **Do not apply a
426
+ > watermark** — that happens centrally. Do not publish or approve.
427
+ >
428
+ > Report back: path, duration, what the product actually is, your concept in one sentence,
429
+ > and anything you'd flag.
430
+
431
+ ---
432
+
433
+ ## 4a. Voice — `af_heart` is the default
434
+
435
+ **Use `af_heart` on local Kokoro (`--engine local`) unless there is a specific reason not
436
+ to.** It is warm, even, and unobtrusive; it sits under a wide range of brands without
437
+ fighting them, and it does not editorialise the copy. Assigning it as the default also
438
+ removes one more axis on which agents can make an arbitrary choice.
439
+
440
+ **Deviate when the brand gives you a reason**, not for variety's sake. Real reasons:
441
+
442
+ | Voice | Reach for it when |
443
+ | ------------ | -------------------------------------------------------------------------- |
444
+ | `af_heart` | **default** — warm, even, unobtrusive; correct for most brands |
445
+ | `bf_emma` | bright and brisk — a utility product sold on speed, a punchy high-key piece |
446
+ | `bm_george` | dry, measured, understated — editorial restraint, confidence without a pitch |
447
+ | `am_michael` | low and unexcited — severe, clinical, instrument-like reads |
448
+ | `am_adam` | plain and neutral — when the copy must carry everything itself |
449
+
450
+ The test is whether you can state the reason in one clause in the brief. "The read should
451
+ move because the product is sold on speed" is a reason. "For variety across the batch" is
452
+ not — differentiation is carried by format, theme, energy, type and concept, which is far
453
+ more visible than voice. A batch where three of four videos use `af_heart` is fine.
454
+
455
+ **Whatever you pick, these still apply:**
456
+
457
+ - **Say the reason in the brief.** The agent should know why it has that voice.
458
+ - **Check the brand name's pronunciation** and spell it phonetically in the TTS input if
459
+ needed — Kokoro said "Jenki" for *Genki* until the input was respelled "Ghenki". Confirm
460
+ with a whisper round-trip, since you are also running whisper for word timings anyway.
461
+ - **For a calm, unhurried read, render line-by-line** — separate Kokoro takes concatenated
462
+ with measured silences — so the pauses are real instead of TTS filler. One continuous
463
+ pass reads rushed however slow the copy is.
464
+ - **Non-English needs `--model large-v3 --language <code>` on whisper.** The default model
465
+ is English-only and will hallucinate fluent English over another language (see §10).
466
+
467
+ ---
468
+
469
+ **Sandbox discipline:** each agent owns `work/<slug>/`, `assets/<slug>/`, `shots/<slug>/`,
470
+ `recon/<slug>.*`. Shared read-only: `bgm/`, `../fonts/`, `capture.mjs`, `audio.sh`, the two
471
+ markdown files. Tell them explicitly never to edit another slug's files or the shared scripts.
472
+
473
+ **Scale:** 5 agents at a time is comfortable; 10–12 works. Recon in waves regardless.
474
+
475
+ ---
476
+
477
+ ## 5. THE FIRST 3 SECONDS — hard requirement (enforcement detail)
478
+
479
+ This restates the ⭐ Golden Rule at the top of this document, in the form that goes into
480
+ `PLAYBOOK.md` for the agents. It is repeated deliberately — it is the rule most often lost.
481
+
482
+ - **Never open on a wall of static text** — not a paragraph, not a dense document, not a
483
+ full-width headline sitting still, not a table.
484
+ - **Something must be MOVING from frame 1**, not motion that starts at 1.5s.
485
+ - **Lead with the most arresting visual you have, not with exposition.**
486
+ - **Frame 0 must still work as a thumbnail** — a complete headline, never a half sentence.
487
+ These two requirements are compatible: a striking visual _with_ a short complete line.
488
+ - If the concept genuinely needs a dense/static surface, **earn it** — arrive at it through
489
+ motion. (The "jargon → plain English" video flies the radiology page in at 62° with motion
490
+ blur under a hook line, then sweeps a scan line across it. Concept intact, hook fixed.)
491
+ - **Test it:** extract t=0, 1, 2, 3. If a viewer would swipe past any of them, redo the open.
492
+
493
+ ---
494
+
495
+ ## 6. Phase 4 — the review loop (this is what makes them good)
496
+
497
+ **Every single first-pass video in batch 1 had a real defect that its own agent reported as
498
+ "verified, looks good."** Do not skip this.
499
+
500
+ ```bash
501
+ for t in 0 5 10 15 20 25; do ffmpeg -y -ss $t -i <slug>.mp4 -frames:v 1 qa/f$t.png; done
502
+ ffmpeg -y -pattern_type glob -i "qa/f*.png" \
503
+ -vf "scale=360:-1,tile=3x2:margin=6:padding=6:color=0x999999" -frames:v 1 qa/sheet.png
504
+ # then READ qa/sheet.png as an image
505
+ ```
506
+
507
+ What this catches, in observed frequency order:
508
+
509
+ 0. **The product is never plainly named in the first 5 seconds** — the open is a metaphor.
510
+ Two of four in batch 17. See the PLAIN-ENGLISH LINE section; this is the one defect that
511
+ costs a VO re-record, so catch it in the brief instead.
512
+ 1. **Large flat dead regions** — content top-anchored with an empty band below. Agents do
513
+ this constantly and never notice.
514
+ 2. **A placeholder empty state that reads as a missing asset** — a big empty dashed rectangle
515
+ held for 2s looks like the render failed.
516
+ 3. **Two contradictory numbers in one frame** (~200px apart, on a data product).
517
+ 4. **A CTA still building when the video ends** — it must be settled ≥2s before the last frame.
518
+ 5. **Two headlines superimposed at a scene handoff** — see the clamp rule below.
519
+ 6. **Type colliding with a busy background layer** (pins, particles) exactly as it's spoken.
520
+
521
+ If a frame looks empty, **check whether it rests there**: sample at 0.2–0.25s intervals
522
+ through the transition. A transient near-empty wipe frame is fine; >0.5s is not.
523
+
524
+ **Verify audio by measurement, not by ear** (you can't hear it):
525
+ speech RMS over whisper word spans, minus bed-alone RMS. Target 12–15 dB, peak < 0 dBFS.
526
+ Note `audio.sh`'s printed "delivered separation" measures the **music-only tail**, not
527
+ in-speech separation — it over-reports. Don't "fix" a mix based on it.
528
+
529
+ ---
530
+
531
+ ## 7. Phase 5 — revision agents
532
+
533
+ **A completed subagent could not be resumed** (`SendMessage` failed by both id and name).
534
+ Spawn a _fresh_ fix-agent instead. This works well because every agent leaves a `NOTES.md`.
535
+
536
+ Fix-brief template:
537
+
538
+ > A previous agent built a finished video for <BRAND>. It's good — you are making <N>
539
+ > targeted fixes, nothing else. Read `PLAYBOOK.md` and `work/<slug>/NOTES.md`. Edit the
540
+ > **generator**, not the generated `composition.html` (check one exists — some are hand-authored).
541
+ > **Back up to `work/<slug>/<slug>-v1.mp4` before overwriting.**
542
+ > Keep VO, script, captions, bed and duration identical — re-running the audio step reuses
543
+ > the existing `vo.wav`/`words.json` so audio stays bit-identical; no re-record, no retiming.
544
+ > <the defects, each with the timestamp and what's wrong>
545
+ > Then sweep the whole video for the same class of problem and report what else you found.
546
+ > Verify by extracting the same timestamps and **reading them as images**.
547
+
548
+ Give the agent the _problem_, not just your proposed solution — three times an agent found a
549
+ better fix than the one I specified (using an app's own collapsed UI state instead of
550
+ redacting; a type safe-zone solver instead of a scrim; making a document's _arrival_ the
551
+ spectacle instead of cutting the document).
552
+
553
+ ---
554
+
555
+ ## 8. Phase 6 — watermark (two variants)
556
+
557
+ `watermark.py <W> <H> <autolight|autodark> <out.png> [opacity] [fs_frac] [gapx_frac] [gapy_frac]`
558
+
559
+ Tiled diagonal text at 30°, offset every other row into a diamond lattice, scaled to the
560
+ frame's **short edge** so it reads the same at 9:16, 16:9, 4:5 and 1:1.
561
+
562
+ **Embossed two-tone** — a dark pass offset one pixel under a light pass — because every one of
563
+ these videos mixes light and dark surfaces. A single-tone mark vanishes over content of its
564
+ own polarity (black-on-white disappeared over a dark screenshot sitting in a light frame).
565
+
566
+ - `autolight` — dark pass ×1.7, light pass ×1.0. Black needs more alpha to read on white.
567
+ - `autodark` — light pass ×0.55, dark pass ×1.2. White-on-near-black is the highest-contrast
568
+ pairing there is and _glares_ at equal alpha.
569
+
570
+ Shipped settings: `0.10 opacity, 0.048 size, 0.20/0.105 gaps` (~52px type on a 1080 frame).
571
+
572
+ ```bash
573
+ cp work/$s/$s.mp4 work/$s/$s-clean.mp4 # preserve the un-watermarked master
574
+ ffmpeg -y -i work/$s/$s-clean.mp4 -loop 1 -i wm/$s.png \
575
+ -filter_complex "[0:v][1:v]overlay=0:0:shortest=1:format=auto[v]" \
576
+ -map "[v]" -map 0:a -c:v libx264 -crf 17 -preset slow -pix_fmt yuv420p \
577
+ -c:a copy -movflags +faststart work/$s/$s.mp4
578
+ ```
579
+
580
+ > ### The `-loop 1` is not optional — I shipped 5 broken videos without it
581
+ >
582
+ > A single-frame PNG as a plain second input made the frame-sync **collapse every video onto
583
+ > one frozen frame**. Duration, frame count and audio hash all survived, so a naive check
584
+ > passed. **Always verify by comparing each watermarked frame against its OWN clean frame**
585
+ > (difference should be tiny — the mark only) **while consecutive frames still differ a lot**
586
+ > (motion preserved). Never verify a video by one frame.
587
+
588
+ ---
589
+
590
+ ## 8a. Stickers — use them GENEROUSLY (customer directive, 2026-08-12)
591
+
592
+ > *"use stickers to help explain where possible"* — and the follow-up: **be more generous with
593
+ > them.** The budget was raised from $0.00 to **$0.10 per video** specifically to pay for this.
594
+
595
+ ### How many
596
+
597
+ **5–8 stickers across a 20–30 s video — roughly one per beat.** One or two per video is the old,
598
+ too-timid setting; it reads as decoration rather than as a way of explaining. The ceiling is one
599
+ sticker on screen at a time, not one per video.
600
+
601
+ A sticker **carries the meaning of the line being spoken at that instant**. It is not garnish. If
602
+ the VO says "spam recovery", the sticker is the thing that says *spam recovery* — and then the
603
+ caption tile and the sticker between them are the whole frame, with no headline needed.
604
+
605
+ ### Where they come from
606
+
607
+ **In cost-saving mode, look them up on IconScout before you generate anything.** A designer
608
+ already drew the envelope, the receipt, the football and the grandmother. `vidfarm iconscout`
609
+ searches that catalog and hands back a finished transparent SVG or PNG on the first try — no
610
+ prompt loop, no keying, no despill, no off-by-one item naming, and none of the sheet-splitting
611
+ work in the next section. **Search costs nothing and needs no key**; vidfarm's own IconScout
612
+ account serves it.
613
+
614
+ ```bash
615
+ vidfarm iconscout "envelope" --free # search — $0, no wallet cost
616
+ vidfarm iconscout get <uuid> --format svg --out stickers/envelope.svg
617
+ ```
618
+
619
+ ### Do NOT lead with `--style sticker` on the free tier — it is nearly empty
620
+
621
+ This is the one trap, and it is counter-intuitive because `sticker` is the style this document
622
+ wants. Measured against the live catalog (2026-08-16), free `icon` results:
623
+
624
+ | query | free, no style | free `sticker` | free `flat` | free `colored-outline` | **premium** `sticker` |
625
+ | --- | --- | --- | --- | --- | --- |
626
+ | envelope | **1629** | 1 | 214 | 226 | 539 |
627
+ | receipt | **727** | 3 | 109 | 116 | 314 |
628
+ | calendar | **3462** | 12 | 445 | 494 | 1090 |
629
+ | person | **8259** | 54 | 1103 | 1166 | 3934 |
630
+ | grandmother | **77** | 0 | 46 | 3 | 32 |
631
+
632
+ **Search free with no `--style` first.** An agent that runs `--style sticker --free`, gets one
633
+ result, and concludes "IconScout has nothing" will fall straight back to generation — which is
634
+ exactly the waste this section exists to prevent. Narrow with `--style flat` or
635
+ `--style colored-outline` if the raw list is too varied; both keep hundreds of results and both
636
+ read well die-cut. Save `--style sticker` for `--premium` — the last column above is why.
637
+
638
+ ### Illustrated figures are the premium case
639
+
640
+ The doc's own hero example — the grandmother in the Families Echo open — is the weakest free
641
+ query there is: **0** free sticker-style icons, **4** free illustrations, but **3013** premium
642
+ illustrations. Free `illustration` is thin across the board (envelope 16, grandmother 4) while
643
+ premium is enormous (envelope 9087, person 158k). So:
644
+
645
+ - **Objects** (envelope, receipt, calendar, football) → `--free`, no style filter. Plenty.
646
+ - **People and illustrated scenes** → `--asset illustration --premium`, ~$0.02. This is the
647
+ right place to spend the cap, and it is cheaper than a generated sheet.
648
+ - `--asset 3d` for a prop with real depth (free is moderate: envelope 33, person 357),
649
+ `--asset lottie` for something already animated.
650
+
651
+ **Which tier:**
652
+
653
+ | | Cost | Credit line | Use it when |
654
+ | --- | --- | --- | --- |
655
+ | `--free` | **$0** | **required** | Generic OBJECTS, no `--style` filter. The `$0.00` mode. |
656
+ | `--premium` | ~$0.02 each on the wallet | none | People, illustrated scenes, `--style sticker`, or any customer deliverable that can't carry a credit line. ~4 fit inside the $0.10 cap. |
657
+ | `sticker-pack --generate` | $0.01-0.05 per sheet of 3-4 | none | Nothing in the catalog fits, or the sticker is bespoke to the brand |
658
+
659
+ > **⚠️ Confirm the free-tier credit obligation before shipping a customer deliverable.**
660
+ > IconScout's free tier carries an attribution requirement, and the download response returns
661
+ > the credit string in `attribution` with `attribution_required: true`. I could not read
662
+ > IconScout's own licence page to pin down the exact wording or where the credit has to appear
663
+ > (it is behind Cloudflare), so **check it yourself before a paid customer video ships with a
664
+ > free asset in it.** If the credit is awkward for the brand, spend the ~$0.02 on `--premium`
665
+ > — that is four premium stickers inside the same cap that buys two generated sheets, and it
666
+ > is still the cheapest thing in this pipeline.
667
+
668
+ **Then fall through to generation.** A generated die-cut pack is the answer when IconScout has
669
+ nothing that fits — a brand-specific object, a made-up device, a character that has to match a
670
+ particular illustration style. One billed image job produces a whole sheet:
671
+
672
+ ```bash
673
+ vidfarm sticker-pack --generate "<style line>" \
674
+ --items "a;b;c;d" --key-color "#FF00FF" --out-dir stickers
675
+ ```
676
+
677
+ All of the following is generation-only housekeeping — an IconScout asset arrives already
678
+ die-cut and needs none of it:
679
+
680
+ - **3–4 items per sheet.** The sheet is 1024x1024 whatever you ask for, so 8 items is ~200 px each
681
+ and mush on a 1080-wide canvas. 3–4 gives ~400 px. **Buy two sheets** — that is what the cap is for.
682
+ - **Pick the key colour off the BRAND hue.** Magenta eats a purple brand; use orange. Never green.
683
+ - Put **"NO texture, NO grain, NO paper effect, NO shadows"** in the style line.
684
+ - **Always eyeball a LABELLED contact sheet and rename before use.** Item names go off-by-one
685
+ whenever a count mismatch is reported, and several items regularly segment as one blob — split
686
+ it with `scipy.ndimage.label` on a dilated alpha.
687
+ - **Expect soft alpha** and harden it: `a = clip((a-25)/(110-25))*255`, plus a despill
688
+ `spill = max(min(r,b) - g, 0)*0.85; r -= spill; b -= spill`, then re-crop to the new bbox.
689
+ Un-hardened stickers look like a faint watermark over footage, not like a keying bug.
690
+ - **Recolour to the exact brand hex** — mask the saturated pixels and remap scaled by each pixel's
691
+ own brightness, so the shading survives and only the hue changes.
692
+
693
+ **CSS/SVG stickers remain a legitimate fallback** (now the third option, after IconScout and a
694
+ generated sheet) — a shape on a flat brand blob with a 6–10 px
695
+ white die-cut rim, a drop shadow and a −8°..+8° rotation. Use them when the pack came back wrong,
696
+ when the budget is spent, or when the sticker is really a word-mark (`32 TOOLS`, `FREE €0`,
697
+ `ZERO ACCOUNT REQUIRED` all worked better as authored type than they would have as generated art).
698
+ A mixed pack is fine and usually best: **IconScout for anything generic** (envelope, receipt,
699
+ football, calendar, person), generated art for objects only this brand has, authored type for
700
+ word-marks.
701
+
702
+ ### How they move
703
+
704
+ Pop in with overshoot — `back.out(2.2)`, 0.28–0.34 s — and give every sticker a slow idle drift so
705
+ nothing on screen is ever dead still. Never place one over the caption band, over a face, or over a
706
+ number the viewer has to read.
707
+
708
+ ### Stickers as masks
709
+
710
+ A sticker silhouette can be a window onto a real screenshot crop — but only at **≥ 440 px**, and
711
+ only over a **dense, colourful** crop (a whole product card, not page whitespace), or the shape
712
+ reads as a broken asset. Two stacked layers: the sticker as generated behind (keeps the white rim),
713
+ the screenshot in front masked by an **eroded** copy of the alpha (fill holes → erode ~13 px →
714
+ gaussian 1.2) so the rim survives.
715
+
716
+ ---
717
+
718
+ ## 9. Phase 7 — delivery
719
+
720
+ - Final = **watermarked** `work/<slug>/<slug>.mp4`; keep `-clean.mp4` masters for customers.
721
+ - Copy out with numbered, brand-named files so they sort in delivery order.
722
+ - Verify each copy for dimensions, duration and **presence of an audio stream**.
723
+ - **Present the cumulative list every time** — every job shipped so far, not just this batch.
724
+ - Tag anything language-specific in the filename (e.g. `24-cap-horizon-FR`).
725
+
726
+ ---
727
+
728
+ ## 10. Renderer gotchas (all found the hard way)
729
+
730
+ **Silent failures — the render "succeeds" and is wrong:**
731
+
732
+ - **Assets outside the composition root → timeline never runs.** Only `<style>`/`<script>`
733
+ _inside_ the `data-composition-id` root execute, and sibling relative files don't resolve —
734
+ fonts, images and **GSAP itself** may need inlining. Symptom: every frame identical, but
735
+ frame 0 looks fine. _Compare two frames from different scenes, never one._
736
+ - **~450 tweens broke timing globally** — every `.to()` rendered 1.0s early in stills _and_
737
+ the real render; `.set()`s stayed exact. Reducing the tween count fixed it (one clipped
738
+ fully-lit duplicate riding a mask, not one tween per element).
739
+ - **GSAP suppresses `onUpdate` while seeking a paused timeline** — counters lag or never fire.
740
+ Pre-bake values and reveal with `tl.set`.
741
+ - **Timed `<video>` gets no z-index from `data-track-index`** — a sibling with CSS `z-index`
742
+ paints over it and it renders as a blank card.
743
+
744
+ **SVG transforms are unsafe — treat CSS transform props on SVG as broken:**
745
+
746
+ - `transform-box: fill-box` fights GSAP's user-space origin: nodes scale in 20–40px off
747
+ position. Pass explicit `svgOrigin`.
748
+ - `scaleY` on an `<rect>` with `transform-box` **translates it out of position** rather than
749
+ scaling about centre. Tween `attr:{y,height}` instead.
750
+
751
+ **Layout / compositing:**
752
+
753
+ - Smooth alpha ramps in CSS gradients mis-composite into banded colour with black wedges —
754
+ use flat opaque shapes and hard stops.
755
+ - `#root`'s own CSS `background` is never painted — add a full-bleed `<div>`.
756
+ - `clip-path: inset()` is relative to the **element's** border box, not the page.
757
+ - GSAP scale/x tweens break `transform:translateX(-50%)` centring — centre structurally.
758
+ - `<video>` projects need `HF_VIDEO_COVERAGE_THRESHOLD=0` (the gate false-positives at ~94%).
759
+
760
+ **Media:**
761
+
762
+ - Relative `img/*` paths render as broken placeholders in the multi-worker renderer even
763
+ though `stills` looked fine → **base64-inline every image** before the real render.
764
+ - Transparent RGBA PNGs go **black** when encoded as JPEG — composite onto the stage colour first.
765
+ - `<video>` cannot use a data URI at all.
766
+ - A root `<audio>` tag fails the local render — strip it and mux audio on afterwards.
767
+
768
+ **Timing / audio:**
769
+
770
+ - **Never `adelay` the VO** — whisper timings (and your captions) are relative to raw `vo.wav`.
771
+ Use `apad` + `atrim`.
772
+ - Clamp caption fade-out against the next tile: `outAt = max(inAt+.22, min(end+.12, nextIn-.17))`.
773
+ The same rule fixes superimposed scene headlines — exit at `nextIn − 0.18`, duration 0.24,
774
+ ease `power2.out` (a slow-leaving `power2.in` is what causes the overlap).
775
+ - Scene handoffs can leave ~0.3s empty frames — extend each clip ~0.35s into the next.
776
+ - `vidfarm tts` reads stdin — always `</dev/null` inside a loop.
777
+ - Whisper's default model is **English-only** and will hallucinate fluent English on other
778
+ languages — non-English needs `--model large-v3 --language <code>`.
779
+ - Beat-locking: measure the bed's BPM and first downbeat, then trim the bed at a **measured
780
+ downbeat** — `audio.sh`'s default `-ss 2` will knock it out of phase.
781
+
782
+ ---
783
+
784
+ ## 11. Accuracy & judgement rules
785
+
786
+ These protect the customer, and several came from real near-misses:
787
+
788
+ - **Claim only what the site claims.** Not a shipped iOS app if the site says "iOS in the
789
+ works". Not an endorsement if they say "built on X guidelines". No install counts or
790
+ ratings unless quoted exactly.
791
+ - **Never name a real third party in a negative or bias-implying context** in a customer's
792
+ marketing asset — even when the product's own output does it. Find the product's own
793
+ neutral state instead.
794
+ - **Keep humour aimed at the problem, not at a real identifiable business or person.** Use a
795
+ fictional example on an IANA-reserved `.example` domain.
796
+ - **No fear-selling** on safety/family/health products; no medical authority beyond the claim.
797
+ - **Licence obligations stay on screen** (e.g. OpenStreetMap attribution) for the full runtime.
798
+ - **Flag personal data** — real names, faces, emails in a customer's own screenshots are
799
+ already public but should be a conscious decision, not an accident.
800
+ - **Read the whole asset before calling a customer's data inconsistent.** "519 hexes" vs
801
+ "312 hexes" was a national total vs a row inside the same card; the bug was ours.
802
+ - **Collect what's broken on their site as a deliverable.** Across 32 brands we found
803
+ contradictory stats, a demo dataset that undercut the product's own pitch, plaintext API
804
+ keys in a JS bundle, an invisible store badge, a `vercel.app` preview URL in production
805
+ marketing, typos, and an app icon that reads as the wrong country's flag. Customers value
806
+ this more than the video sometimes.
807
+
808
+ ---
809
+
810
+ ## 12. Cost
811
+
812
+ **Budget: <= $0.10 per video** (raised from $0.00 on 2026-08-12, customer directive).
813
+
814
+ Everything except stickers is still free and stays that way: Kokoro TTS (`--engine local`),
815
+ whisper timings, local render, CC0 beds already on disk, the customer's own graphics.
816
+
817
+ **Stickers can now be free too.** `vidfarm iconscout "<thing>" --free` is a $0
818
+ designer catalog that needs no key — so a video whose stickers are all generic objects
819
+ (envelope, receipt, calendar, person) can go back to costing **$0.00** with no loss of quality.
820
+ The price of the free tier is a credit line; see the warning in §8a before putting one in a
821
+ customer deliverable.
822
+
823
+ Spend, in the order you should reach for it:
824
+
825
+ | | Cost | Buys |
826
+ | --- | --- | --- |
827
+ | `iconscout --free` | **$0.00** | generic objects, unlimited, credit line required |
828
+ | `iconscout --premium` | ~$0.02 each | illustrated people/scenes, ~4 inside the cap, no credit line |
829
+ | `sticker-pack --generate` | $0.01-0.05 per sheet | 3-4 bespoke stickers per sheet, no credit line |
830
+
831
+ The cap still buys two generated sheets if you need them — but on most videos you will not,
832
+ because the catalog already has the object. **Search before you generate.**
833
+
834
+ Still never `vidfarm generate image` — a full AI image is a different, unbudgeted job. If one
835
+ would genuinely help, write the prompt to `AI_PROMPTS-<slug>.md` and let the human run it.
836
+
837
+ ---
838
+
839
+ ## 13. What "unique" actually looked like
840
+
841
+ For calibration — 32 concepts, no reuse, all derived from the brand's own material:
842
+
843
+ evidence chain · knowledge graph forking into two front doors · ink bleeds · hex fill ·
844
+ one spectrum bar transformed seven ways · jargon dissolving into plain English · organic
845
+ bloom · map pins dropping · a 3-gate signal funnel · buckets filling 50/30/20 · one pinned
846
+ query bar with the world swapped beneath · a light-sweep painting a blank room · a skills
847
+ orbit collapsing inward · a wardrobe reading you back · two calendars colliding into a rest
848
+ shadow · overlapping arcs interlocking into one morning · a screen recording receding into
849
+ an annotated spec sheet · three silos snapping into one story · brand palettes flipped like
850
+ a record crate · a mountain mirrored in a lake · a table for two widening to seat four · a
851
+ group chat settling up · a thread you send to yourself · a navigator's horizon with three
852
+ bearings · two plates snapping into perfect register · latent prints developing in place ·
853
+ a word prompter on the downbeat · an empty seat at the boardroom table · call and response ·
854
+ a waveform as the only constant while the pictures are disposable · a lock screen waking
855
+ five times across one day · the frame itself split down the middle.