@officexapp/vidfarm-devcli 0.21.33 → 0.21.35
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/editor-capabilities/SKILL.md +13 -2
- package/.agents/skills/vidfarm/SKILL.md +73 -29
- package/.agents/skills/vidfarm/harnesses/README.md +112 -0
- package/.agents/skills/vidfarm/{regimes/explainer.QA_REGIME.md → harnesses/explainer.HARNESS.md} +22 -2
- package/.agents/skills/vidfarm/{regimes/hooks.QA_REGIME.md → harnesses/hooks.HARNESS.md} +3 -3
- package/.agents/skills/vidfarm/{regimes/product-demo.QA_REGIME.md → harnesses/product-demo.HARNESS.md} +19 -1
- package/.agents/skills/vidfarm/{regimes/short-form.QA_REGIME.md → harnesses/short-form.HARNESS.md} +67 -7
- package/.agents/skills/vidfarm/{regimes/ugc-testimonial.QA_REGIME.md → harnesses/ugc-testimonial.HARNESS.md} +10 -3
- package/.agents/skills/vidfarm/recipes/{bulk-scripting-with-a-regime.md → bulk-scripting-with-a-harness.md} +35 -12
- package/.agents/skills/vidfarm/recipes/cutout-graphics-for-explainers.md +46 -2
- package/.agents/skills/vidfarm/recipes/local-edit-render-approve.md +2 -1
- package/.agents/skills/vidfarm/references/automation-and-local-dev.md +84 -22
- package/.agents/skills/vidfarm/references/editor-workflows.md +19 -4
- package/.agents/skills/vidfarm/references/hooks-and-virality.md +62 -5
- package/.agents/skills/vidfarm/references/reviewing-renders.md +140 -0
- package/.agents/skills/vidfarm-media/SKILL.md +2 -2
- package/.agents/skills/vidfarm-media/references/tts.md +26 -4
- package/SKILL.director.md +462 -75
- package/SKILL.md +33 -14
- package/dist/src/cli.js +799 -91
- package/dist/src/devcli/{qa-regime.js → harness.js} +132 -55
- package/dist/src/devcli/qa-check.js +209 -4
- package/dist/src/devcli/skill-docs.js +136 -0
- package/dist/src/devcli/stills.js +65 -1
- package/package.json +4 -3
- package/.agents/skills/vidfarm/regimes/README.md +0 -77
package/SKILL.director.md
CHANGED
|
@@ -35,6 +35,7 @@ vidfarm serve template_<32hex> # local server + browser, opens that
|
|
|
35
35
|
- The API key comes from https://vidfarm.cc/settings and starts with `vf_key_`. Instead of `login`, setting the `VIDFARM_API_KEY` environment variable also works for every command — the CLI reads it from the environment or from a `.env` file in the current directory.
|
|
36
36
|
- No account or key? `vidfarm serve --no-cloud` still gives a fully local editor with free local renders.
|
|
37
37
|
- "Open/run template X locally" is exactly one command: `vidfarm serve <template_id>` (alias: `vidfarm <template_id>`). Do not hand-roll REST or hunt for local `.harness/` files first — `serve` and `pull` create those.
|
|
38
|
+
- **The CLI carries this entire skill offline** — installing the devcli puts a copy of the pack on disk, pinned to that version. `vidfarm skill ls` lists it, `vidfarm skill show <path>` prints one file, and **`vidfarm skill search "<term>"` greps all of it at once**, which is the cheapest way to find the paragraph you need without loading a 650-line reference. No account, no network. It is documentation, not entitlement: the free-local half (clips, hyperframes, `serve` render, `qa`, harnesses, `dedupe`, local TTS/STT) runs offline; AI generation, hosted render, `recycle`, `download-video` and marketplace still need `vidfarm login` and a cloud call.
|
|
38
39
|
|
|
39
40
|
### Entity ID formats
|
|
40
41
|
|
|
@@ -129,7 +130,7 @@ If the user hasn't picked yet and you're about to spend, name the cheaper path a
|
|
|
129
130
|
|
|
130
131
|
Cost mode answers *how much money may I spend*. It does not answer *how much of the user's own hands may I use* — and that second axis moves quality more than the first. **Ask both.** They are independent: every cost mode (`minimize`, `hybrid`, `rich-ai`, `pure-videogen`) runs in either interaction mode.
|
|
131
132
|
|
|
132
|
-
- **interactive** — the user is willing to do a little manual work at fixed checkpoints, and the video gets better for it.
|
|
133
|
+
- **interactive** — the user is willing to do a little manual work at fixed checkpoints, and the video gets better for it. Three checkpoints cover nearly everything: **(1) images** — you write a prompt, they run it in a *free* frontier web generator (meta.ai / ChatGPT / Gemini / a Hugging Face Space) and hand the file back; **(2) raw clips** — you hand over search keywords, they search TikTok/YouTube, download a few with a free online downloader, and point you at the folder; **(3) the voice** — you sample a few narrators and they pick the one the video sounds like (costs them 30 seconds and $0, see below).
|
|
133
134
|
- **autonomous** — you finish end-to-end with zero steps from them: source clips yourself (browser control → `raws scan` → public raws), generate within the budget, or do without.
|
|
134
135
|
|
|
135
136
|
**Why interactive usually wins on quality:** the free tiers of the frontier web image models are typically *better* than what an API-key budget buys per image, and a human eye picks better footage than any keyword scan. In `minimize` the gap is not incremental — it's the difference between **no custom art at all** and **a full sticker pack for $0**.
|
|
@@ -147,6 +148,10 @@ Cost mode answers *how much money may I spend*. It does not answer *how much of
|
|
|
147
148
|
|
|
148
149
|
**In interactive mode, MANUAL IMAGE WORK DEFAULTS TO STICKER PACKS.** Never ask for one graphic per round trip — each hand-off costs the user a context switch and costs you tokens re-reading a file. Ask for **one sheet holding every graphic**, then split it locally for $0. `vidfarm handoff image --theme "<what>" --items "a,b,c"` mints the whole brief (prompt + steps + the free tools + the follow-up command); `--single` when you really do want one subject. When the file comes back: `vidfarm sticker-pack <sheet> --items "a,b,c"`.
|
|
149
150
|
|
|
151
|
+
**In interactive mode, OFFER THE VOICE CHECKPOINT — in every cost mode.** Who the video sounds like is a taste decision, and the default voice is the one choice agents make silently that a director almost always wants a say in. Before narrating, ask *"want to hear a few voices and pick one?"* and sample: `vidfarm voices --sample` (premium) or `vidfarm voices --free --sample` ($0 local). **Sampling costs nothing on either tier** — premium samples are ElevenLabs' own preview clips (a CDN download, not a synthesis call) and free samples render locally — so this checkpoint is just as available in `minimize` as in `hybrid`. Play the files, take their pick, narrate with `--voice <id>`. `vidfarm tts` prints the same nudge on stderr whenever narration would run with no voice named and the mode is interactive (or was never set).
|
|
152
|
+
|
|
153
|
+
**And say where the premium voices come from, because users assume wrong.** The full ElevenLabs catalog is reachable **through vidfarm's own ElevenLabs connection** — no ElevenLabs account, API key, or subscription on the user's side; narration just spends **vidfarm wallet credits** (pennies each). In `hybrid` that is a real option to put on the table next to the free voices, not a locked door. `--own-key` is only for users who already have an ElevenLabs key and would rather bill their own account.
|
|
154
|
+
|
|
150
155
|
**Raw-clip sourcing has a ladder — hand it to the human only at the bottom rung.** In order: (1) **your own browser control**, if you have it — drive the search and download yourself; (2) **Vidfarm cloud** — `vidfarm clipper <url>` / `vidfarm raws scan <url> --cloud` resolves and mines the video for you; (3) the **free public raws catalog** — `vidfarm public-raws --category <shelf>`; (4) **the user**, when you're fully local/keyless or when human taste matters. That last rung is `vidfarm handoff raws --keywords "villa construction,pouring concrete" --platforms tiktok,youtube --purpose "<what the clips are for>"` — it prints the keywords, tells them to google a downloader (a *search*, not a link that rots), and names the import command for when the folder is ready (`vidfarm clipper ./downloads/<file>.mp4`, or `vidfarm raws scan` to mine a long one).
|
|
151
156
|
|
|
152
157
|
## Default stance
|
|
@@ -213,7 +218,7 @@ Present both harnesses to the director, recommend (A) unless they've asked for p
|
|
|
213
218
|
- **The ART must be CLOSED and SOLIDLY FILLED — this is the other half of surviving the key, and the #1 way stickers come back broken.** Ask an image model for "icons on a green plate" and it will happily draw **outline art**: a colored stroke with the shape's interior left as bare plate. It looks perfect on the sheet, and after the key each sticker is a **rim floating around a see-through hole** (an apple-shaped outline with nothing inside it). Same outcome from a *near-plate* fill (the keyer works on tolerance, not exact match), a translucent/glassy/glowing material, or a soft glow fading into the plate. **You cannot key those pixels back — it has to be in the prompt:** *"every object is a closed, solidly filled shape; outlines must enclose an opaque fill of a different color; no outline-only or hollow art; nothing on the art in the plate color or any near-shade of it; fully opaque, no translucency, glow or drop shadow."* `cutout --generate`, `sticker-pack --generate`, `handoff image` and the `create-overlay` primitive **append that clause for you** with the chosen plate hex — write it yourself only when you prompt a generator directly. After the key, both commands report per-item `hole_pct`/`hollow` (console `⚠ N% hollow`, `--json`, `stickers.json`) — a ring or picture frame reads the same way, so it **warns, never blocks**. Flagged and it shouldn't be? Re-generate with the fill clause; a *near*-plate fill can sometimes be rescued with a lower `--tolerance`; one stubborn item can be lifted with `vidfarm mask --crop …` (ONNX matting ignores fill color).
|
|
214
219
|
- **Transparent GIF is supported, for GIF-only surfaces.** `vidfarm sticker-pack … --output-format gif` (stills) and `vidfarm remove-greenscreen <video> --gif` (animated) emit transparent GIFs. GIF alpha is **1-bit**, so edges go hard — fine for chat/forum/Notion sticker surfaces, worse than PNG/WebP/WebM for compositing on a timeline. Prefer PNG/WebP/WebM unless the destination only eats GIF.
|
|
215
220
|
|
|
216
|
-
**Explainer house style — the defaults to build with unless told otherwise.** **White background / light mode** (plain white stage, no gradients, no dark mode, no photo backdrop), **kinetic word-by-word captions** in dark ink on the light stage (`vidfarm captions generate --style word-pop --color "#111111" --active-color "#7C3AED" --background-style plain` — skip outlines/shadows, they're only needed over busy footage), and **female TTS narration** (`vidfarm tts --voice coral` on OpenAI — `nova` for energy, `sage` for calm; `Kore`/`Leda` on Gemini; any ElevenLabs voice via `vidfarm voices`). **Keep it clean and simple** — one idea on screen at a time, two or three cutouts per beat, one accent color, one font, lots of white space; remove before you add. **Illustrations default to simplicity**: flat vector, simple shapes, minimal detail, 2–3 flat colors, no baked-in text — simple art keys cleanly, trims tight, and stays on-style across the whole cast. State the defaults once so the director can override any of them. Full detail: recipe `recipes/cutout-graphics-for-explainers.md` (“House style — the explainer defaults”).
|
|
221
|
+
**Explainer house style — the defaults to build with unless told otherwise.** **White background / light mode** (plain white stage, no gradients, no dark mode, no photo backdrop), **kinetic word-by-word captions** in dark ink on the light stage (`vidfarm captions generate --style word-pop --color "#111111" --active-color "#7C3AED" --background-style plain` — skip outlines/shadows, they're only needed over busy footage), and **female TTS narration** (`vidfarm tts --voice coral` on OpenAI — `nova` for energy, `sage` for calm; `Kore`/`Leda` on Gemini; any ElevenLabs voice via `vidfarm voices`). **Keep it clean and simple** — one idea on screen at a time, two or three cutouts per beat, one accent color, one font, lots of white space; remove before you add. **Illustrations default to simplicity**: flat vector, simple shapes, minimal detail, 2–3 flat colors, no baked-in text — simple art keys cleanly, trims tight, and stays on-style across the whole cast. State the defaults once so the director can override any of them. Full detail: recipe `recipes/cutout-graphics-for-explainers.md` (“House style — the explainer defaults”). **If the director takes the stage off white**, two things stop being optional: every sticker's **white die-cut rim** has to be stripped (on a dark stage it's a glaring halo and the most obvious bot-made artefact in the frame — recipe → “Stickers on a DARK or photographic stage”), and the caption hexes above stop applying — **caption colour, active-word colour and plate are chosen by measuring the composited background behind the caption band**, one treatment per video (`harnesses/short-form.HARNESS.md` → “Caption styling is MEASURED off the background”). Related: **on-screen text and captions must not say the same thing at once** — display text carries the argument, captions carry only what the screen doesn't show.
|
|
217
222
|
|
|
218
223
|
**Landscape footage in a fullscreen vertical explainer — use the blurred plate, never bars.** When an explainer is built on **real filmed footage** and the source is 16:9 (or 4:3) on a 9:16 canvas, do not `contain` it (hard black letterbox bars read as an unfinished export) and do not blindly `cover` it (a wide shot loses its left and right thirds). Duplicate the clip: a full-canvas `cover` copy behind, heavily **gaussian-blurred and faded dark**, plus the sharp copy centered as a hero band — optionally zoomed ~1.3× — with its **top and bottom edges feathered** into the blur. Same clip, same timecode, so it reads as one continuous image with a shallow-depth-of-field plane, fullscreen edge to edge, nothing cropped, and clean dark space for the header and captions. Bake it once with ffmpeg into a single 1080×1920 file (free, local) and place it as one ordinary full-canvas layer — layer blur is not an editor property, so the pre-bake is the path that works in the editor, `serve`, and cloud render alike. Copy-paste ffmpeg + HTML recipes, tuning table, and the failure modes: `references/editor-workflows.md` (“The blurred plate — landscape footage, fullscreen, on a vertical canvas”).
|
|
219
224
|
|
|
@@ -251,6 +256,18 @@ The mechanism is deterministic, not luck: rendering is seek-safe, so frame 0 sho
|
|
|
251
256
|
|
|
252
257
|
Full mechanics and editor verbs: `references/editor-workflows.md` (“The opening frame is the post's thumbnail”); poster-state authoring craft: `hyperframes-creative/references/beat-direction.md`.
|
|
253
258
|
|
|
259
|
+
## Judge the WHOLE video, not the parts you built — and never by one frame
|
|
260
|
+
|
|
261
|
+
**Assume your own finished video has a defect you can't see.** That's the observed base rate, not modesty: across a 32-video batch, *every* first-pass video had a real defect that the agent who built it had already reported as "verified, looks good" — dead space under the content, a placeholder that reads as a failed render, contradictory numbers in one frame, a CTA still animating at the last frame.
|
|
262
|
+
|
|
263
|
+
**The cause is how agents build: part by part, each part correct in isolation.** Scene 3 gets authored while scene 3 is the whole world, so every scene passes on its own and the video fails *as a video* — type size jumps between beats, one scene breathes and the next is crammed, the accent colour drifts, a transition lands like a slap because nothing before it moved that fast, one asset is flat vector and the next is photographic. Nobody watches a scene; they watch the sequence. **So before you ship, look at the whole thing at once as a stranger would**, and ask: is it visually balanced (or top-anchored with a dead band below), is the spacing consistent scene to scene, is there ONE type scale / accent colour / illustration style, does the pacing have a deliberate rhythm instead of N identical beats, is anything jarring at the joins, does any frame make two elements compete for the eye, and — the summary question — **does it look like one person made it in one sitting?** When it doesn't, fix the *system* (define the type scale, the margin, the default beat, the palette, then apply it to every scene) rather than patching the one scene that stood out.
|
|
264
|
+
|
|
265
|
+
The mechanical form is a contact sheet, because one image read shows drift that twelve separate checks can't — and it's one command: **`vidfarm stills ./work --sheet`** renders the frames and tiles them into `stills/contact-sheet.png`. **Read the sheet as an image.**
|
|
266
|
+
|
|
267
|
+
**And never verify a video by a single frame.** Frame 0 is the thumbnail so you judge it alone — but a whole class of render bug (an overlay pass missing `-loop 1`, assets outside the composition root so the timeline never runs) freezes *every* frame while duration, frame count and audio hash all still pass, and frame 0 looks perfect. **Always compare two frames from different scenes; they must differ a lot.** Likewise, verify audio by measurement, never "it sounds good" — you can't hear it.
|
|
268
|
+
|
|
269
|
+
Full method — the holistic checklist, the six defects in observed frequency order, the "does it rest there" test, frozen-render verification, audio measurement, and how to brief a revision pass: **`references/reviewing-renders.md`**.
|
|
270
|
+
|
|
254
271
|
## Web AI chat vs. local devcli — know your surface's limits
|
|
255
272
|
|
|
256
273
|
**Desktop users: this is a skill, not an SDK or required runtime.** It gives an AI agent Vidfarm's operating knowledge; it does not proxy requests, add hidden capabilities, or require the agent to use a particular client. Every cloud workflow described here can be performed with ordinary HTTP against the Vidfarm REST API. For convenience and safer composition/file workflows, recommend the `vidfarm` command from `@officexapp/vidfarm-devcli`: it wraps the same REST routes, handles auth, uploads, polling, pagination, local files, and pull/edit/render loops. Use raw REST whenever the user prefers it, the CLI lacks a new route, or the work belongs in an existing service or automation stack.
|
|
@@ -270,30 +287,63 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
|
|
|
270
287
|
- **Determine the surface before claiming capabilities.** The web chat has only its declared tools and REST routes. It cannot execute arbitrary JavaScript/Python, open a shell, create a local repository script, or use the user's filesystem. Never tell a web-chat user that you ran code or wrote a script unless a dedicated declared tool actually did so. A desktop coding agent has a real shell and filesystem and MAY write/run scripts, perform arbitrary local computations over paginated API results, create reports/CSVs/JSON, edit composition files, and orchestrate long devcli workflows within the user's authorization.
|
|
271
288
|
- **The web AI chat can do all three paintbrushes** — clip raws, author HTML/hyperframe motion, and generate AI media — and it drives edits directly on the live timeline. Keep small-to-medium jobs here: text/caption swaps, a scene or two replaced, single generations, captions, approve/schedule. Just do them.
|
|
272
289
|
- **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. **The test is the native-editor test: could you have made this element with the tools inside TikTok's own editor?** That toolset is font / color / stroke / shadow / tight text box / alignment / rotation / animation presets, plus stickers, emoji, drawn marks and clips — it has no padded capsule, no border, no gradient fill, no blur panel, no card. If you reached past it, cut it. That includes **a single lonely pill around a stat or label** — `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`: being alone doesn't make a badge native, and the only legitimate capsule in a video is the active-word `spotlight`/`karaoke` highlight, which moves with the spoken word. Emphasize a stat the way the editor would instead: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle, or its own beat. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r` — or a `border-radius` over ~8px on anything filled that holds words — stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
|
|
273
|
-
- **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm
|
|
290
|
+
- **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm harness show hooks`.
|
|
291
|
+
- **Then CUT it — every second must earn its place, and most don't.** Assume your first assembly is **30–50% too long**. Run the **deletion test** on every beat: delete it; if the video still makes sense and the payoff still lands, it stays deleted. Whatever survives must serve one of the four charges — "it gives context" is not a charge. Cut on sight: intros/logo stings, the wind-up sentence before the claim ("so I wanted to talk about…"), restatement, inter-sentence silence over ~0.35s, real-time process, establishing shots, reading what's already on screen, and any tail after the last word. **Always ripple the hole closed** (`vidfarm ripple <dir> --at <sec> --delta -<sec>`) — a cut that leaves a gap turns fluff into dead air, which is worse. Density is **not** speed: the held comedic beat, the payoff playing out, and a cue's readability keep their seconds (cut *words*, not the time text is on screen). Length is an **output**, not a plan — a brief that dictates a duration ordered fluff. `vidfarm qa` flags the mechanical half (`dead-air`, `dead-tail`, `slow-scene`); the craft is `references/hooks-and-virality.md` → "Density".
|
|
274
292
|
- **The first frame IS the thumbnail — compose it on purpose.** Frame 0 is a single frame of ~30 in the first second, but it's the poster every feed, share link, and paused player freezes on, so **more people see that one frame than watch the video**. It must never be black, empty, mid-fade, or caught mid-animation: put a real visual at `start:0`, have the hook words already on screen at t=0, and never hang a `fade-black`/`fade-white`/`flash` *entrance* on the **first** clip (junction transitions between later clips are fine — this rule is only about the opening). Verify it, don't assume: devcli `vidfarm stills ./work --at 0` renders that exact frame, and `vidfarm qa` flags a blank or fading open.
|
|
275
|
-
- **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`
|
|
276
|
-
-
|
|
293
|
+
- **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`HARNESS.md`** — because a loop of fifty videos has no human looking at every frame, and the harness is what replaces those eyes. Don't silently ship a one-off when they asked for volume, or drag someone into a harness when they wanted one clip.
|
|
294
|
+
- **"Harness" is a known noun with a known process — recognise it and follow it.** A **harness** is the reusable AI apparatus for ONE format or template: what makes it special, written down as `HARNESS.md` so an agent can reproduce it without the director in the room. It is a first-class artifact — the director owns it, edits it, versions it, and hands it to the next agent. Three phrasings, one artifact:
|
|
295
|
+
- **"create me a harness"** / "set up a harness for this format" → `vidfarm harness init <base> --out ./work/HARNESS.md` (bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo`), then **edit it with them**. The bundled file is a starting point, never a house style; the parts that matter are the ones they add — who the viewer is, the banned vocabulary, the compliance line, the pacing this account actually uses. A harness nobody edited isn't about their videos.
|
|
296
|
+
- **"update the harness for this format/template"** → open the existing `HARNESS.md` and write the new rule in, **with its reason on the same line** (a rule whose "why" is missing gets argued away by the next agent). This is what you do every time a batch teaches you something ("the label-framed hooks all died"): the compositions are disposable, the harness is the artifact that compounds.
|
|
297
|
+
- **"give me the harness for this template_id"** → they mean **the decomposition**: `vidfarm harness derive <templateId|forkId>`. It distils the decompose pass's DNA into an editable `HARNESS.md`. If the template hasn't been decomposed, run `vidfarm decompose` first.
|
|
298
|
+
**A harness mirrors the template JSON's own vocabulary** — `## Viral DNA` (hook / retention / payoff / emotion), `## Visual DNA` (cut rhythm, typography, b-roll, transitions), `## Structural DNA` (the beats, and which are load-bearing), `## Audio DNA` (voice, bed, comedic timing), `## Build DNA` (which paintbrush per beat) — the same strands the decompose pass writes as `viral_dna`, `visual_dna`, and friends. `vidfarm harness show <ref> --dna visual` prints one strand instead of the whole doc.
|
|
299
|
+
**Two halves, and only one is machine-checkable.** The `checks:` front matter is settled deterministically by `vidfarm qa` (duration, aspect, `hook_words_max`, `forbid_text`, …); every `- [ ]` line comes back as a **review item you answer honestly in your report** — never claim a video passed the half the CLI can't judge. Harnesses stack and auto-discover: `vidfarm qa ./work` picks up `./work/HARNESS.md`, `--harness hooks --harness ./brand/HOUSE.md` adds more, and any file of theirs anywhere is valid. Format and strand table: `harnesses/README.md`; scripting-mode detail: `references/automation-and-local-dev.md`. *(Formerly `QA_REGIME.md` — same file, and `vidfarm regime …` still works as an alias.)*
|
|
300
|
+
- **A video is judged as a SEQUENCE, so review it as one.** Agents build scene by scene and each scene passes in isolation while the video drifts — inconsistent margins, three type sizes, an accent colour that wanders, beats that are all the same length, a jarring join. Tile a dozen stills into one contact sheet (`vidfarm stills ./work --sheet`) and read it as an image before you call anything done, fix drift by defining the system rather than patching the odd scene out, and remember that **your own confident "verified, looks good" is the single least reliable signal in this workflow** — it was wrong on every video of a 32-video batch. Method: `references/reviewing-renders.md`.
|
|
277
301
|
- **On devcli, QA every video you produce: `vidfarm qa ./work`.** Free, instant, local-only — it blocklists exactly the slop above plus first-frame/thumbnail and font-regime/safe-zone drift, and prints a concrete fix per finding. **Feedback, not a gate**: it exits 0 even on findings, never runs automatically, and is a blocklist (unusual/stylized compositions pass untouched), so there's no reason not to run it before every publish. `--json` for scripted batches, `--strict` only if you want a CI failure. **Web-chat copilot: this command does not exist for you** (devcli-only, no REST twin) — apply the standard by hand, and when handing a heavy job to a local coding agent, tell them to run `vidfarm qa`.
|
|
278
|
-
- **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
|
|
302
|
+
- **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips) and, inside that band, is **placed in the emptiest part of the frame** rather than dumped on the default lower third (read a still first — `vidfarm stills ./work --at <t>`; words over open sky or a blank wall beat words over the subject, and usually need no plate at all). Long narration is **paged into 3–5-word kinetic cues** (`vidfarm captions generate --style word-pop`), never one static wall of text. It uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
|
|
279
303
|
- **Where the web chat struggles: complex, long, multi-step transformations.** A full multi-scene re-theme, an iterative render-critique-iterate loop, heavy scripted or batch work, or anything needing a real filesystem and many sequential tool calls will hit context limits, turn/timeout ceilings, and the web editor's constraints (CSS/declarative motion only — JS animation adapters are stripped on save). Don't grind a big transformation one layer at a time in a chat turn and stall.
|
|
280
304
|
- **Practical workaround — hand the heavy job to local devcli.** When a task is genuinely large or long-running, **proactively recommend the director run it locally with an AI coding agent** (Claude Code / OpenAI Codex / any capable agent): `vidfarm pull <forkId>` writes the composition + the `.harness/` grounding bundle to disk, the agent edits with the full devcli verb set and JS animation adapters, renders free with `vidfarm serve`, and `vidfarm publish` pushes it back. This is the **best-quality (B) harness's** natural home (adversarial grading with a coding agent). Frame it as "this is a big rebuild — you'll get a better, faster result running it locally with a coding agent; here's how," not as a dead end.
|
|
281
305
|
- **Offer a handoff, do not impersonate the desktop agent.** When web chat reaches that boundary, offer to save a Markdown handoff in My Files containing the objective, selected template/fork IDs, asset paths, grounding, constraints, completed work, and suggested devcli commands. Create it only after the user agrees. The desktop agent should read that document, pull the referenced fork, and then use its actual code/shell capabilities.
|
|
282
306
|
- **Never send the user away just to read knowledge.** Deeper skill knowledge is always a **tool call** away in-place: call `load_skill` (e.g. `load_skill('vidfarm', file='references/editor-workflows.md')`, or a craft pack like `editor-capabilities` / `hyperframes-animation`) to pull the exact reference you need mid-conversation. Only recommend switching surfaces for the WORK (a heavy transformation), never for the information.
|
|
283
307
|
|
|
284
|
-
##
|
|
308
|
+
## File Index — everything in this pack, and when to read it
|
|
285
309
|
|
|
286
|
-
|
|
310
|
+
**This is the complete inventory. Nothing else exists in the pack, and every file here is reachable by name.** Read the narrowest file that answers the question; never preload several. `size` is a context-cost estimate — the four big references are real reads, so pick one deliberately rather than sweeping them.
|
|
287
311
|
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
312
|
+
**References — the broad knowledge domains**
|
|
313
|
+
|
|
314
|
+
| File | Size | Read it when |
|
|
315
|
+
|---|---|---|
|
|
316
|
+
| `references/core-workflows.md` | ~360 ln | Template discovery, auth, fork → render → approve → share, versioning, cost/wallet, marketplace orders, dedupe-before-publish |
|
|
317
|
+
| `references/editor-workflows.md` | ~650 ln | **The biggest read.** Timeline editing, decompose, captions, transitions, motion, AI placement, the caption standard, the editor action verbs |
|
|
318
|
+
| `references/assets-and-sourcing.md` | ~185 ln | Raws hunts, clip scanning, My Files, recurring characters, downloading media off a URL, social recycle |
|
|
319
|
+
| `references/automation-and-local-dev.md` | ~520 ln | **Big.** The whole `vidfarm` command table, REST automation, scripting/bulk mode, `HARNESS.md`, local serve loop, skill packs |
|
|
320
|
+
| `references/primitives.md` | ~475 ln | **Big.** One-shot primitive routes: TTS, STT, music, avatars, overlays, greenscreen, inpaint, background removal, product placement |
|
|
321
|
+
| `references/hooks-and-virality.md` | ~295 ln | **Before writing ANY hook, caption script, or re-theme**, and before a hook-variant batch. The four charges, three gates, banned openers, loop mechanics. This is the craft; the rest of the pack is mechanics |
|
|
322
|
+
| `references/reviewing-renders.md` | ~140 ln | **Before you report a video as done**, or grade someone else's. The holistic pass, the common defects, frozen-render and audio verification |
|
|
323
|
+
| `references/onboarding.md` | ~30 ln | Cold-start interviews, **consultations** (the `brainstorm/*` chain), strategy docs, durable director context |
|
|
324
|
+
| `references/rest-api.md` | ~85 ln | Only when the user asks for REST, an endpoint/schema, or direct HTTP integration. It is an index — follow its domain links; do not preload it into ordinary director conversations |
|
|
325
|
+
|
|
326
|
+
**Recipes — step-by-step procedures. When a recipe matches the task, prefer it over the broad reference.**
|
|
327
|
+
|
|
328
|
+
| File | Size | Read it when |
|
|
329
|
+
|---|---|---|
|
|
330
|
+
| `recipes/find-and-fork-template.md` | ~15 ln | Template selection and the first fork |
|
|
331
|
+
| `recipes/retheme-template.md` | ~15 ln | Full re-theme that preserves the source format's feel |
|
|
332
|
+
| `recipes/local-edit-render-approve.md` | ~20 ln | The local pull → edit → render → approve loop |
|
|
333
|
+
| `recipes/onboard-a-new-director.md` | ~15 ln | New-director onboarding and durable context capture |
|
|
334
|
+
| `recipes/bulk-scripting-with-a-harness.md` | ~100 ln | **Volume**: daily posting, N variants, hook tests — scripting mode with a `HARNESS.md` |
|
|
335
|
+
| `recipes/cutout-graphics-for-explainers.md` | ~265 ln | Building an explainer from sticker/cutout art: the house style, sticker sheets, keying, dark-stage rules |
|
|
336
|
+
|
|
337
|
+
**Harnesses — the `HARNESS.md` format and its bundled bases.** All are readable as-is and copyable with `vidfarm harness init <name>`.
|
|
338
|
+
|
|
339
|
+
| File | Size | Read it when |
|
|
340
|
+
|---|---|---|
|
|
341
|
+
| `harnesses/README.md` | ~110 ln | **Start here for anything harness-shaped**: the three director phrasings, the format, the `checks:` key list, the DNA strand → decompose-JSON map |
|
|
342
|
+
| `harnesses/short-form.HARNESS.md` | ~225 ln | The default base. Also holds the **"Caption styling is MEASURED off the background"** procedure that other files point at |
|
|
343
|
+
| `harnesses/hooks.HARNESS.md` | ~120 ln | Hook-variant batches — chunk-1 legibility, the unguessable test, volume-only anti-patterns. The checkable form of `hooks-and-virality.md` |
|
|
344
|
+
| `harnesses/explainer.HARNESS.md` | ~100 ln | Faceless educational video: one claim, invented visuals |
|
|
345
|
+
| `harnesses/ugc-testimonial.HARNESS.md` | ~90 ln | A person vouching for a product — mostly rules about what NOT to add |
|
|
346
|
+
| `harnesses/product-demo.HARNESS.md` | ~110 ln | Real product doing a real thing; the highest slop-risk format in the catalog |
|
|
297
347
|
|
|
298
348
|
## HyperFrames Skills — Load on Demand
|
|
299
349
|
|
|
@@ -309,9 +359,9 @@ On the web copilot, call `load_skill('<name>')` and load referenced files only w
|
|
|
309
359
|
|
|
310
360
|
HyperFrames authoring and rendering in this package are Vidfarm-native: local work uses the bundled composition toolchain and `vidfarm serve`; cloud work uses Vidfarm render routes. Do not require an external vendor account, repository, publish service, or telemetry endpoint. Keep `HYPERFRAMES_SKIP_SKILLS=1` and `HYPERFRAMES_NO_TELEMETRY=1` in Vidfarm-managed environments so the bundled skills stay pinned and local work does not phone home.
|
|
311
361
|
|
|
312
|
-
## Quick Router
|
|
362
|
+
## Quick Router — from what the user said to what to open
|
|
313
363
|
|
|
314
|
-
Choose the narrowest path that satisfies the request.
|
|
364
|
+
The File Index above says what each file *is*; this says which one a given ask means. Choose the narrowest path that satisfies the request.
|
|
315
365
|
|
|
316
366
|
1. If the user needs help figuring out what to make, **or asks for a "consultation"** (the `brainstorm/*` chain: cold-start interview → awareness stages → angles → hooks), read `references/onboarding.md` first.
|
|
317
367
|
2. If the user already knows the goal and needs a suitable template, read `references/core-workflows.md` and use the template discovery flow.
|
|
@@ -320,7 +370,9 @@ Choose the narrowest path that satisfies the request.
|
|
|
320
370
|
4b. If the task is **“download this video/audio off a website”** (a pasted YouTube / TikTok / Instagram / X post URL the user wants the actual file from), Vidfarm does that for you on a **paid plan** — `POST /api/v1/primitives/videos/download` (or `/audio/download`), devcli `vidfarm download-video <url>` / `vidfarm download-audio <url>`. **Free plan → do not call it; walk the user through opening the URL in Chrome and downloading it from the page, then `vidfarm put-file` the local file in for $0.** Details in `references/assets-and-sourcing.md` → `references/primitives.md`.
|
|
321
371
|
4c. If the task is **“turn this Reddit/X thread, subreddit, or account into a video”** — “tweet to TikTok”, “Reddit to TikTok”, “make a video from this thread”, “what are the top comments saying” — run `vidfarm recycle <source>` (or `POST /api/v1/primitives/social/recycle`) with the URL. It **decomposes** the source into raw JSON (text, comment tree, media URLs, author pics, stats) and hands it back unranked so YOU pick what to remix. **Paid plan; `max_records` is the spend ceiling.** Brokers the reddit-lead-gen / x-lead-gen OfficeX apps, so it waits out their async job for you. Details in `references/assets-and-sourcing.md` → `references/primitives.md`.
|
|
322
372
|
4d. If the task is **“post this again / to several accounts / on another platform”**, or you are about to publish or bulk-produce at all — that is **deduplication**. Run `vidfarm dedupe <mp4> [--variants N]` on the **exported file** (free, local ffmpeg, no re-render), then approve/schedule each variant. **Ask the operator whether they want deduplicated copies, and how many, BEFORE the render/bulk run** — deciding after means paying for a second render. Details in `references/core-workflows.md` → *Deduplicate before you publish* and `references/primitives.md` → *Primitive: media_dedupe*.
|
|
373
|
+
4e. If the ask contains the word **“harness”** — *“create me a harness”*, *“update the harness for this format”*, *“give me the harness for this template_id”* — that is a known, named process, not a vague request. Read `harnesses/README.md` (the three phrasings and the format), then `recipes/bulk-scripting-with-a-harness.md` if the job is a batch. The third phrasing means the **decomposition**: `vidfarm harness derive <forkId>`.
|
|
323
374
|
5. If the task is scripted, local, CI-driven, or `vidfarm serve`-based, read `references/automation-and-local-dev.md`.
|
|
375
|
+
5b. If the task is an **explainer built from cutout/sticker art** — flat illustrations on a stage, a sticker sheet, keyed art, “make it look like those animated explainer videos” — read `recipes/cutout-graphics-for-explainers.md`. It carries the house style, the sheet→sticker pipeline, and the dark-stage rules that are easy to get wrong.
|
|
324
376
|
6. If the task explicitly asks for a primitive or needs specialized generation/transcription work, read `references/primitives.md`.
|
|
325
377
|
7. If the task is the MARKETPLACE (ordering videos from specialist agents): browsing is web-only for paying customers — send the human to https://vidfarm.cc/marketplace, never render it locally. Placing/listing orders is the thin REST wrapper in `references/core-workflows.md` (§ Marketplace); anything deeper on a gig (inbox, proofs, payouts) needs the external Dollar Platoon skill — `npx skills add https://github.com/OfficeXApp/dollarplatoon-skill` — the same way FlockPoster work beyond scheduling needs `npx skills add https://github.com/OfficeXApp/flockposter-skill`.
|
|
326
378
|
|
|
@@ -334,22 +386,14 @@ Choose the narrowest path that satisfies the request.
|
|
|
334
386
|
- Submission routes are generally not idempotent. Especially for renders and expensive primitives, check status before retrying.
|
|
335
387
|
- In the web editor, use CSS/declarative motion only. Script-bearing HTML is stripped or rejected there.
|
|
336
388
|
- **Never render or approve without judging frame 0 as a standalone still.** It is the thumbnail everywhere the post appears; an empty/black opening frame ships a dead post. See “The FIRST FRAME is the thumbnail”.
|
|
337
|
-
|
|
338
|
-
## Recommended Recipes
|
|
339
|
-
|
|
340
|
-
Use these when the user’s task matches the pattern closely.
|
|
341
|
-
|
|
342
|
-
- Template selection and first fork: `recipes/find-and-fork-template.md`
|
|
343
|
-
- Full re-theme while preserving the format’s feel: `recipes/retheme-template.md`
|
|
344
|
-
- Local pull/edit/render/approve loop: `recipes/local-edit-render-approve.md`
|
|
345
|
-
- New-director onboarding and durable context capture: `recipes/onboard-a-new-director.md`
|
|
389
|
+
- **Never judge the VIDEO by one frame, and never report a render as reviewed without the holistic pass.** Compare frames from at least two different scenes (a frozen render passes every other check), read a contact sheet for balance/spacing/style/pacing drift, and state separately what you measured vs. what you judged. See “Judge the WHOLE video”.
|
|
346
390
|
|
|
347
391
|
## Output Posture
|
|
348
392
|
|
|
349
393
|
- Prefer concrete actions over abstract discussion.
|
|
350
394
|
- Name the chosen path explicitly: template reuse, raws hunt, local serve, cloud render, etc.
|
|
351
395
|
- Surface cost tradeoffs before expensive generation.
|
|
352
|
-
- When in doubt between a broad reference and a recipe, start with the recipe.
|
|
396
|
+
- When in doubt between a broad reference and a recipe, start with the recipe — the File Index marks which is which.
|
|
353
397
|
|
|
354
398
|
## Mental model
|
|
355
399
|
|
|
@@ -1230,16 +1274,20 @@ Compositions are authored in HTML, so the single most common way an AI-edited vi
|
|
|
1230
1274
|
- **Emoji inline in text** (sparingly), **sticker/cut-out overlays** on transparent PNG (`create-overlay`), mock social UI when the format calls for it (iMessage bubbles, a TikTok comment card, a fake DM, a countdown/progress bar) — these are native artifacts of the platform, not web furniture.
|
|
1231
1275
|
- **Full-bleed footage** with text sitting directly on it.
|
|
1232
1276
|
|
|
1233
|
-
**On devcli there's a checker: `vidfarm qa <dir|composition.html>`.** Free, instant, local-only — a blocklist pass for everything above plus the font regime and safe zone, with a concrete fix per finding. **Run it on every video you produce.** It is feedback, not a gate (exit 0 even on findings, never runs automatically, `--strict` only if you want a CI failure) and a blocklist, not an allowlist (stylized/hand-made compositions pass untouched — it will not homogenize your videos). No cloud/REST twin: the web copilot enforces this standard by hand. Details in `references/automation-and-local-dev.md` ("`vidfarm qa`").
|
|
1277
|
+
**On devcli there's a checker: `vidfarm qa <dir|composition.html>`.** Free, instant, local-only — a blocklist pass for everything above plus the font regime and safe zone, with a concrete fix per finding. **Run it on every video you produce.** It is feedback, not a gate (exit 0 even on findings, never runs automatically, `--strict` only if you want a CI failure) and a blocklist, not an allowlist (stylized/hand-made compositions pass untouched — it will not homogenize your videos). Every run — including a clean one — ends with a **`▶ NOW WATCH THE VIDEO`** block, because the check never rendered or saw the video and a green tick is not a review; do those steps before you tell anyone the video is done. No cloud/REST twin: the web copilot enforces this standard by hand. Details in `references/automation-and-local-dev.md` ("`vidfarm qa`").
|
|
1234
1278
|
|
|
1235
1279
|
### TikTok-native caption standard (position + font + background) — always adhere
|
|
1236
1280
|
|
|
1237
1281
|
> Captions are also the *delivery system* for three of the four charges: the hook is read before any audio, the loop has to stay on screen, and the payoff number needs its own card. What the words should SAY is in `references/hooks-and-virality.md`; this section is how they must LOOK.
|
|
1238
1282
|
|
|
1239
|
-
Short-form is watched on a phone, and the phone's UI eats the frame's edges. **Never pin on-screen text to the extreme top or bottom** — the top ~8% sits under the status bar / "Following · For You" tabs and the bottom ~15% under the username, caption text, music marquee, and action rail. Text there is literally clipped and reads as amateur.
|
|
1283
|
+
Short-form is watched on a phone, and the phone's UI eats the frame's edges. **Never pin on-screen text to the extreme top or bottom** — the top ~8% sits under the status bar / "Following · For You" tabs and the bottom ~15% under the username, caption text, music marquee, and action rail. Text there is literally clipped and reads as amateur. Four rules, applied to **every** caption/title/overlay you place or inherit:
|
|
1240
1284
|
|
|
1241
|
-
- **Position →
|
|
1242
|
-
-
|
|
1285
|
+
- **Position → the safe zone first, then the EMPTIEST part of the frame.** Two constraints, in that order.
|
|
1286
|
+
- *Hard constraint:* the text box's vertical extent stays inside **~8%–85%** of canvas height (9:16), and wide captions stay clear of the **right ~12%** action rail (a centered box at `x:10 width:80` is safe).
|
|
1287
|
+
- *Judgement call, inside that band:* **put the words where the picture isn't.** `y≈70%` is the `captions generate` default because most footage puts its subject mid-frame — it is a default, not a law. Before you place text, **look at an actual frame** (`vidfarm stills ./work --at <t>`, free) and find the region with the least going on: open sky above a dashboard, a blank wall behind a talking head, an out-of-focus background, an empty tabletop. If nothing else in the video is competing for attention there — no subject, no motion, no product, no second text layer — that is where the caption belongs, even if it means **high-centre at y≈10–25%** instead of a lower third. A caption dropped over the busiest third of the frame (hands on a steering wheel, a face, the product) fights the shot and forces you to armour it with a plate; the same words parked in the sky are legible with no plate at all.
|
|
1288
|
+
- *When you're only rescuing an inherited caption* off a dead-zone edge, preserve its top-vs-bottom anchoring and just pull it inside the band — don't recentre a template you haven't re-read. When **you** are the one placing the text, place it deliberately.
|
|
1289
|
+
- **Size → scaled to the line, not maxed out.** Sizes are PIXELS of a 1080-wide frame: **~36–64px** reads well; below ~28px is unreadable on a phone and **0 is invisible**. Above ~64px is a *hook-word* size — one to three words, on purpose. The failure this catches: a full sentence set at display size runs edge-to-edge, wraps to three lines, and eats a third of the frame, so it has to be armoured with a full-width plate and there is nowhere left to put it. **If a line reaches the frame edges, the fix is a smaller size (or fewer words per cue), not a wider box.** Keep captions to ~2 lines / ~5 words per line; `line_height` 0.95–1.15 for stacked display lines.
|
|
1290
|
+
- **Font → the composition regime.** Use the bundled display fonts only — **Montserrat** (bold default, weight **700–900**), **TikTok Sans**, Abel, Source Code Pro, Yesteryear. Don't request a font the composition doesn't import (it silently falls back to a web-default sans, which is exactly the slop look).
|
|
1243
1291
|
- **Background → one of exactly four valid treatments.** Any text you place uses one of these and nothing else:
|
|
1244
1292
|
|
|
1245
1293
|
| # | Treatment | How to set it | When |
|
|
@@ -1249,8 +1297,19 @@ Short-form is watched on a phone, and the phone's UI eats the frame's edges. **N
|
|
|
1249
1297
|
| 3 | **Highlight pill behind the ACTIVE word only** | `set_captions caption_style:"spotlight"` / `"karaoke"` (+ `caption_highlight_color`) | Hormozi/CapCut word-by-word. **The only legitimate "pill" in a video** — it tracks the spoken word, so it isn't a badge |
|
|
1250
1298
|
| 4 | **Solid band that tightly hugs the text lines** (CapCut "text box") | `background_style:"highlight-solid"` (or `"highlight-translucent"`) + a `background` color | Guaranteed legibility over noisy footage |
|
|
1251
1299
|
|
|
1300
|
+
**Move the text before you armour it.** The plate is the *last* resort, not the default: if the band you picked is busy, first try moving the caption into the calm/empty region the position rule points at — a caption over open sky needs no background at all, and "no plate" is the cleaner, more native look every time you can afford it. Only when the whole frame is busy (or the text has to sit on the subject for meaning) do you reach for treatment 4.
|
|
1301
|
+
|
|
1302
|
+
**Then pick between them by MEASURING the background behind the caption band, not by habit.** Dark-and-calm behind the band (luma < ~70, variation < ~42) → light type, **no plate** (treatment 1/2 — a plate there is a bright slab the design never asked for); bright-and-calm (luma > ~160) → dark type, no plate; busy / mid-tone / moving colour → treatment 4, because nothing else stays readable. The active-word colour has to follow the same call — a deep red that reads on a white plate is unreadable on near-black. **One treatment for the whole video**; styling that flips every few seconds reads as a bug. Procedure, thresholds and how to measure the *composited* value (not the source file): `harnesses/short-form.HARNESS.md` → "Caption styling is MEASURED off the background".
|
|
1303
|
+
|
|
1252
1304
|
Treatment 4 is a **band, not a card**: it hugs the glyphs with minimal padding, corner radius ≤ ~8px, **no border, no drop shadow, no gradient, no blur**, and it wraps *one* text run — never a heading + subheading + URL stacked inside one rounded box. The moment it grows padding, a stroke, or a second element inside it, it has become a web card. Fix it. And the moment its radius goes fully round, it has become a **badge** — treatment 3 is the *only* capsule allowed, and only because it tracks the spoken word. A static "10 hrs / week" in a rounded pill is web furniture; the same words in treatment 1 or 2, bigger and heavier, are a beat.
|
|
1253
1305
|
|
|
1306
|
+
**Long narration → kinetic cues, never a wall of text.** A caption layer is a *page*, not a transcript. The moment a single static text run carries more than ~10–12 words — or sits on screen longer than ~4 seconds while the voice keeps going — it stops being a caption and becomes a paragraph the viewer has to read while also watching the video. Nobody does both; they scroll. Page it instead:
|
|
1307
|
+
|
|
1308
|
+
- **Transcribe and let the tool page it:** `vidfarm captions generate ./work --style word-pop` (or `spotlight` / `karaoke`) splits narration into ~3–5-word cues with real word-level timings, so one short phrase is on screen at a time and the active word tracks the voice. Web copilot twin: the `/primitives/audio/captions` job → `set_captions` (see "Animated captions" below). `--max-words-per-cue` tightens it further.
|
|
1309
|
+
- **The cue count is the readability dial.** Short cues that change with the speech read as *momentum*; one long block reads as homework. Kinetic word-by-word also lets the type be **smaller** (the eye is led to the moving word instead of having to scan a wall), which frees up frame space and usually removes the need for a plate.
|
|
1310
|
+
- **Static text is for the beats that deserve their own moment** — a hook line, a payoff number, a title card. Those are short by nature. Anything spoken should be a caption run, not a static block.
|
|
1311
|
+
- **Exception: verbatim UGC/testimonial captions** stay one plain line at a time (see `harnesses/ugc-testimonial.HARNESS.md`) — the kinetic VFX look is the "made by a marketing team" tell there. Paging still applies; the animation preset doesn't.
|
|
1312
|
+
|
|
1254
1313
|
**A common trap: decomposed templates mirror the source's caption placement**, so a forked meme can arrive with its caption pinned at `top:0` in a non-regime font — and a re-theme prompt ("make this for my tutoring service") is exactly where an agent starts inventing landing-page CTAs and benefit chips because the *subject* is a SaaS product. **Fix to the standard, don't inherit it, and don't import the website's design language into the video.** When placing text yourself (`set_captions`, `set_layer_text`, `set_layer_style`, `add_layer`, devcli `place`/`captions`), set `y` / `font_family` / `font_weight` / `background_style` to the standard from the start.
|
|
1255
1314
|
|
|
1256
1315
|
> Local devcli renders enforce part of this automatically: `renderCompositionLocally` runs `normalizeTikTokCaptionLayout` (src/devcli/composition-edit.ts) on every production, clamping caption/text layers into the 8%–85% safe zone and coercing off-regime primary fonts to Montserrat. It only fixes position and font family — it will happily render your Bootstrap card. Get it right in the composition so the editor preview, the local render, and any cloud render match.
|
|
@@ -1355,7 +1414,7 @@ Most agent-made videos don't fail on polish. They fail on **structure**: no hook
|
|
|
1355
1414
|
|
|
1356
1415
|
This is the harness that fixes it. It is not a style — it's the load-bearing anatomy of anything that travels on TikTok/Reels/Shorts, distilled from grading hundreds of hooks against real funnels. **Run it on one-off videos and on batches alike.** It costs no credits, adds no render time, and it is the single largest quality delta available in this product.
|
|
1357
1416
|
|
|
1358
|
-
The checkable form of this document is the bundled `hooks`
|
|
1417
|
+
The checkable form of this document is the bundled `hooks` harness (`vidfarm harness show hooks`); this file is the craft behind it.
|
|
1359
1418
|
|
|
1360
1419
|
---
|
|
1361
1420
|
|
|
@@ -1368,10 +1427,67 @@ The reason agent videos come out structureless is that the timeline is the fun p
|
|
|
1368
1427
|
3. **Name the payoff.** What is on screen at that moment, and why does it satisfy the promise?
|
|
1369
1428
|
4. **Write the bait.** The final-beat ask, in the video and in the post caption.
|
|
1370
1429
|
5. **Only now build the timeline** — and place the hook text at `start:0` so it's on screen at frame 0 (which is also the thumbnail).
|
|
1371
|
-
6. **Verify the frame and the structure:** `vidfarm stills ./work --at 0` (look at the actual poster) and `vidfarm qa ./work --
|
|
1430
|
+
6. **Verify the frame and the structure:** `vidfarm stills ./work --at 0` (look at the actual poster) and `vidfarm qa ./work --harness hooks` (machine checks + the judgment checklist).
|
|
1372
1431
|
|
|
1373
1432
|
Steps 1–4 are cheap, reversible, and where the entire outcome is decided. Steps 5–6 are where agents want to start.
|
|
1374
1433
|
|
|
1434
|
+
7. **Cut it.** Nothing ships at its first length — see the next section. Assume your first assembly is 30–50% too long and go find the seconds.
|
|
1435
|
+
|
|
1436
|
+
---
|
|
1437
|
+
|
|
1438
|
+
## Density — every second must earn its place, and most don't
|
|
1439
|
+
|
|
1440
|
+
**A viewer's thumb is a hard time limit that resets every second.** They are not "watching your video"; they are re-deciding to stay, ~24 times a second, against an infinite feed of alternatives. A second that carries nothing is not neutral — it is a free exit. This is why the same script cut to 22s outperforms its own 41s version with better footage: fewer exit ramps.
|
|
1441
|
+
|
|
1442
|
+
Agents are structurally bad at this. A model writes a video the way it writes prose — with connective tissue, restatement, a wind-up before the point, a tidy conclusion — and every one of those habits is a hole in the retention curve. **You must cut against your own instinct, and you must cut more than feels right.**
|
|
1443
|
+
|
|
1444
|
+
### The deletion test — the only test that matters
|
|
1445
|
+
|
|
1446
|
+
For every beat, ask: **delete it. Does the video still make sense, and does the payoff still land?** If yes, it stays deleted. Not "trimmed" — deleted. Run this on every scene, every sentence, and every caption before you render, and be honest: the beat you're defending because it took work to make is exactly the one this test exists to kill.
|
|
1447
|
+
|
|
1448
|
+
Second filter for whatever survives: **which of the four charges does this beat serve — hook, loop, payoff, or bait?** A beat that serves none is fluff wearing a costume. "It gives context" is not a charge. "It looks nice" is not a charge.
|
|
1449
|
+
|
|
1450
|
+
### Cut on sight — the standard fluff, in the order it usually appears
|
|
1451
|
+
|
|
1452
|
+
- **Any intro.** Logo sting, title card, brand animation, "welcome back", a beat of black. The video starts at the claim. Frame 0 is the hook (and the thumbnail).
|
|
1453
|
+
- **The wind-up before the point.** "So I wanted to talk about…", "Here's the thing…", "Let me explain…", "In this video I'm going to show you…". Delete the sentence; the next one was the real opening.
|
|
1454
|
+
- **Context before the claim.** Context is beat 2 at the earliest, and usually one clause, not a scene.
|
|
1455
|
+
- **Restatement.** Saying the same thing a second way "so it's clear." It was clear. If it wasn't, fix the first version.
|
|
1456
|
+
- **Dead air in the narration.** Breaths, "um", and any inter-sentence gap over ~0.35s. This alone routinely takes 15–20% off a TTS or talking-head cut.
|
|
1457
|
+
- **Real-time process.** Nobody watches the upload bar. Speed-ramp it, jump-cut it, or show the before and the after and skip the middle.
|
|
1458
|
+
- **Establishing shots.** They know what an office/kitchen/laptop looks like. Open inside the action.
|
|
1459
|
+
- **Reading what's already on screen.** Voice and text should split the work, not duplicate it (the same rule as captions-vs-display-text).
|
|
1460
|
+
- **The tail.** "Thanks for watching", a logo card, an end screen, or footage that keeps rolling after the last word. The bait is the last beat; then it **ends**, hard, on the frame that loops best.
|
|
1461
|
+
- **Filler motion.** A slow pan or Ken Burns that exists because the clip was too short for its slot. Shorten the slot instead.
|
|
1462
|
+
|
|
1463
|
+
### Density is not speed, and this is where over-correcting ruins videos
|
|
1464
|
+
|
|
1465
|
+
Cutting fluff means **removing beats that carry nothing**, never rushing the beats that carry everything. Three things are load-bearing and must keep their seconds:
|
|
1466
|
+
|
|
1467
|
+
- **The comedic beat.** The held pause before a punchline IS the joke. Cutting it saves 0.6s and costs the video.
|
|
1468
|
+
- **The payoff.** It plays, full frame, uninterrupted — ≥5s if that's what it takes. Summarising the payoff to save time is the most expensive cut available.
|
|
1469
|
+
- **A caption's readability.** A cue nobody can finish reading is worse than no cue. If tightening the edit makes text unreadable, cut *words*, not the time they're on screen.
|
|
1470
|
+
|
|
1471
|
+
The target is **information per second**, not seconds. A dense 45s video beats a hollow 20s one; both lose to the same 45s cut to 30s with nothing lost.
|
|
1472
|
+
|
|
1473
|
+
### Length is an output, not a plan
|
|
1474
|
+
|
|
1475
|
+
Don't decide "make it 60 seconds" and then fill 60 seconds — filling is where every one of the fluff patterns above comes from. Build the four charges, cut to the deletion test, and **the length is whatever's left.** If the payoff lands at 0:25, the video ends around 0:27. A brief that dictates a duration is a brief that ordered fluff.
|
|
1476
|
+
|
|
1477
|
+
### How to actually cut it, in Vidfarm
|
|
1478
|
+
|
|
1479
|
+
| Move | devcli | Web copilot |
|
|
1480
|
+
|---|---|---|
|
|
1481
|
+
| Find the dead air | `vidfarm qa ./work` (flags gaps ≥2.5s with nothing on screen, and a tail that keeps rolling after the last word) + read the word timings from `vidfarm captions generate` / `stt` | read `video_context`'s timestamped segments and look for the gaps between them |
|
|
1482
|
+
| Trim one clip's edge | `vidfarm trim ./work --layer <k> --edge start --to-time <sec>` | `editor_action trim_layer` |
|
|
1483
|
+
| Close the hole you just made | `vidfarm ripple ./work --at <sec> --delta -<sec>` (negative = close time, shifts everything downstream) | `editor_action ripple_edit` |
|
|
1484
|
+
| Drop a whole beat | `vidfarm retime`/`remove` the layers, then `ripple` the gap closed | `remove_layer` + `ripple_edit` |
|
|
1485
|
+
| Re-time captions after cutting | re-run `vidfarm captions generate` against the new audio — never hand-shift cues | the `/primitives/audio/captions` job → `set_captions` |
|
|
1486
|
+
|
|
1487
|
+
**Always ripple the gap closed.** A cut that leaves a hole is not a cut; it converts fluff into dead air, which is worse — the viewer now stares at a frozen frame instead of a boring one.
|
|
1488
|
+
|
|
1489
|
+
**Cheap habit that pays every time:** shave the first ~0.5–1s off every sourced clip and the last ~0.5s. People start recording before the action and stop after it, so a montage of raws is carrying a second of nothing per clip by default.
|
|
1490
|
+
|
|
1375
1491
|
---
|
|
1376
1492
|
|
|
1377
1493
|
## Charge 1 — THE HOOK (first 3 seconds)
|
|
@@ -1580,12 +1696,153 @@ A video with replies gets shown again; a video with none dies at its first audie
|
|
|
1580
1696
|
| Check what the source template's hook actually was | `editor_context` → `viral_dna.hook` / `retention` / `payoff` / `emotional_punch` | `.harness/context.json`, `video-context.json` |
|
|
1581
1697
|
| Place the hook at frame 0 | `add_layer` / `set_captions` with `start:0` | `vidfarm set-text ./work --layer hook --text "…"` |
|
|
1582
1698
|
| Look at the poster frame | ask the user to scrub to 0 | `vidfarm stills ./work --at 0` |
|
|
1583
|
-
| Grade the structure | by hand, against this file | `vidfarm qa ./work --
|
|
1584
|
-
| Bulk hook test | hand off to a local agent | `recipes/bulk-scripting-with-a-
|
|
1699
|
+
| Grade the structure | by hand, against this file | `vidfarm qa ./work --harness hooks` |
|
|
1700
|
+
| Bulk hook test | hand off to a local agent | `recipes/bulk-scripting-with-a-harness.md` |
|
|
1585
1701
|
|
|
1586
1702
|
**Re-theming a decomposed template?** `viral_dna` already names the source's hook, retention device, and payoff — that structure is *why the template worked*. Rebuild each charge for the new subject; don't drop the loop because the new topic feels self-explanatory. Flattening a template's loop into a product statement is the single most common way a re-theme kills a format.
|
|
1587
1703
|
|
|
1588
|
-
**The checkable version of everything above:** `vidfarm
|
|
1704
|
+
**The checkable version of everything above:** `vidfarm harness show hooks` — the twelve-item pre-flight checklist is the part you answer honestly on every video, and two items carry most of the weight: *situation, not label* (predicts cold-start survival before you write a word) and *unguessable* (the only item a hook can fail while passing every other one, which is why it ships).
|
|
1705
|
+
|
|
1706
|
+
## Reviewing a render — look at the whole video, and never trust one frame
|
|
1707
|
+
|
|
1708
|
+
**Assume your own finished video has a defect you can't see, because you built it.** This is not humility, it's the observed base rate: across a 32-video bespoke batch, **every single first-pass video had a real defect that the agent who built it had already reported as "verified, looks good."** Dead space under the content, a placeholder that reads as a failed render, two contradictory numbers 200px apart, a CTA still animating when the video ends. None of these are subtle. All of them survived a confident self-review.
|
|
1709
|
+
|
|
1710
|
+
The reason is structural, not sloppiness: **an agent builds a video the way it builds code — part by part, each part correct in isolation.** Scene 3 is written while scene 3 is the whole world. So each scene passes on its own and the video fails as a video: the type jumps two sizes between beats, one scene breathes and the next is crammed to the margins, the accent colour drifts, a transition lands like a slap because nothing before it moved that fast. Nobody watches a scene. They watch the sequence.
|
|
1711
|
+
|
|
1712
|
+
So the review has two jobs, and they need two different passes:
|
|
1713
|
+
|
|
1714
|
+
1. **The holistic pass** — does this read as ONE video, made by one person, on purpose?
|
|
1715
|
+
2. **The defect pass** — is any individual frame broken in one of the six ways frames are usually broken?
|
|
1716
|
+
|
|
1717
|
+
Do them in that order. The holistic pass is the one agents skip, and it is the one that separates "technically correct" from "good."
|
|
1718
|
+
|
|
1719
|
+
---
|
|
1720
|
+
|
|
1721
|
+
## Pass 1 — the holistic pass: watch it as one object
|
|
1722
|
+
|
|
1723
|
+
**Before you look for defects, look at the video the way a stranger will: all at once, start to finish, with no memory of how it was built.** You cannot do this from the code, from the storyboard, or from the scene you just edited — you have to look at the actual frames, in order, side by side.
|
|
1724
|
+
|
|
1725
|
+
```bash
|
|
1726
|
+
# ONE command — renders the stills AND tiles them into ./work/stills/contact-sheet.png.
|
|
1727
|
+
# No MP4 render needed; free, local, in-process.
|
|
1728
|
+
vidfarm stills ./work --sheet # default timestamps = midpoint of each scene clip (cap 8)
|
|
1729
|
+
vidfarm stills ./work --sheet --at 0,2,4,6,8,10,12,14,16,18,20,22 # or pick them yourself
|
|
1730
|
+
# then READ ./work/stills/contact-sheet.png as an image
|
|
1731
|
+
# --sheet-out <file> relocates it; --sheet-width <px> for bigger tiles (default 320)
|
|
1732
|
+
|
|
1733
|
+
# already have the MP4? same idea, straight off the file
|
|
1734
|
+
for t in 0 2 4 6 8 10 12 14 16 18 20 22; do
|
|
1735
|
+
ffmpeg -y -ss $t -i final.mp4 -frames:v 1 "qa/f$(printf %03d $t).png"; done
|
|
1736
|
+
ffmpeg -y -pattern_type glob -i "qa/f*.png" \
|
|
1737
|
+
-vf "scale=300:-1,tile=4x3:margin=6:padding=6:color=0x999999" -frames:v 1 qa/sheet.png
|
|
1738
|
+
```
|
|
1739
|
+
|
|
1740
|
+
**The tile sheet is the point.** One image read, twelve frames, and the eye picks up drift instantly that no per-scene check can see. Read it as an image — not the filenames, not the HTML that produced it. (The still filenames are zero-padded seconds, so a glob stays in chronological order.)
|
|
1741
|
+
|
|
1742
|
+
Then answer these, out loud, in your report:
|
|
1743
|
+
|
|
1744
|
+
- **Balance.** Is weight distributed across the frame, or is every scene top-anchored with an empty band underneath? Does the composition use the canvas, or does it use the top third of the canvas and leave the rest as dead area? A sheet of twelve frames makes a recurring dead zone obvious; one frame at a time never will.
|
|
1745
|
+
- **Fluff, named out loud.** Which beats would you cut? Answer with specific timestamps, not "it's tight". Every tile has to justify its seconds: a frame that repeats the previous one, a scene the video would survive losing, an intro, a tail after the last word, a hold that's just waiting. **Assume 30–50% of the first assembly can go** and name what you'd remove — "nothing to cut" on a first pass is almost always a review that didn't look. Then cut it and `ripple` the hole closed (craft: `references/hooks-and-virality.md` → "Density"; the mechanical half is `vidfarm qa`'s `dead-air` / `dead-tail` / `slow-scene`).
|
|
1746
|
+
- **Spacing and breathing room.** Are margins consistent scene to scene? Does one beat have generous air and the next one crowd the safe zone? Uneven padding across scenes is the single loudest "assembled by a machine" tell, and it's invisible while you're inside any one scene.
|
|
1747
|
+
- **Typographic continuity.** One type system, or three? Headline sizes should belong to a small set (two, maybe three), not be individually chosen per scene. Same for weight, case, and colour. If scene 2's headline is 64px and scene 5's is 41px for no dramatic reason, that's drift, not design.
|
|
1748
|
+
- **Colour and style coherence.** One accent colour, one background treatment, one illustration style. Assets generated or sourced at different moments drift — a flat-vector sticker next to a photographic cutout next to a gradient panel reads as three videos spliced together.
|
|
1749
|
+
- **Rhythm and pacing.** Do scene durations form a deliberate pattern (a fast open, a longer explanation, a fast close), or is every scene the same length because a loop wrote them? Same-length beats are hypnotic in the bad way. Conversely, one 9-second hold in a video of 2-second cuts stalls it dead.
|
|
1750
|
+
- **Nothing jarring at the joins.** Watch each transition specifically. A cut from a dark scene to a white one is a flash in the face; a scale-up entrance immediately after a scale-up exit reads as a stutter; two consecutive scenes whose subjects sit in the same screen position with different content look like a glitch, not a cut. Where a join is harsh, either match the two frames either side of it (colour, position, energy) or make the harshness deliberate and rhythmic.
|
|
1751
|
+
- **One idea per moment.** Across the whole sheet, is there any frame where two things compete for the eye — display text over captions saying the same words, a busy background under type, two headlines superimposed at a handoff? At the video level this shows up as a *density* problem: some beats carry three elements and some carry one.
|
|
1752
|
+
- **Does it look like one person made it in one sitting?** The summary question. If the honest answer is "it looks assembled," name specifically which scenes don't belong and fix them toward the majority, don't average everything.
|
|
1753
|
+
|
|
1754
|
+
**When something is off, fix it globally, not locally.** The instinct after spotting drift is to patch the one scene that stands out. Usually the right fix is to define the rule (two headline sizes, one accent, 8% margins, 2.5s default beat) and apply it across every scene — including the ones that already looked fine. A video is a system; patching one node keeps the system inconsistent.
|
|
1755
|
+
|
|
1756
|
+
**Build order helps too, if you're still building.** Author the shared system first — type scale, palette, margins, motion vocabulary, default beat length — as one thing that every scene reads from, then fill the scenes. Scenes written first and harmonised later almost never fully converge.
|
|
1757
|
+
|
|
1758
|
+
---
|
|
1759
|
+
|
|
1760
|
+
## Pass 2 — the defect pass: what actually goes wrong, in frequency order
|
|
1761
|
+
|
|
1762
|
+
From the same 32-video batch, ranked by how often it happened. Look for these specifically; they are what your own review misses.
|
|
1763
|
+
|
|
1764
|
+
1. **Large flat dead regions.** Content top-anchored with an empty band below it. Agents do this constantly and never notice, because during authoring the element is the subject and the emptiness is just "background."
|
|
1765
|
+
2. **A placeholder empty state that reads as a missing asset.** A big empty dashed rectangle held for two seconds looks exactly like the render failed. If a scene's job is "an empty inbox," it still has to look designed, not broken.
|
|
1766
|
+
3. **Two contradictory numbers in one frame.** Especially on anything data-shaped — a stat in the headline and a different one in the visual beneath it.
|
|
1767
|
+
4. **A CTA or end card still building when the video ends.** The final state must be **settled at least 2 seconds before the last frame**, or the loop-around cuts it off and the ask never lands.
|
|
1768
|
+
5. **Two headlines superimposed at a scene handoff.** The outgoing scene's text hasn't left when the incoming one arrives. Fix at the timing level: exit at `nextIn − 0.18`, duration `0.24`, ease `power2.out` — a slow-leaving `power2.in` is what causes the overlap in the first place.
|
|
1769
|
+
6. **Type colliding with a busy background layer** exactly at the moment it's spoken. Particles, pins, footage detail — legible in the still you checked, unreadable at the second the word lands.
|
|
1770
|
+
|
|
1771
|
+
**If a frame looks empty, check whether it RESTS there.** Sample at 0.2–0.25s intervals through that transition. A transient near-empty wipe frame is fine and normal; anything holding empty for **>0.5s** is a hole in the video.
|
|
1772
|
+
|
|
1773
|
+
```bash
|
|
1774
|
+
vidfarm stills ./work --at 6.0,6.2,6.4,6.6,6.8,7.0 # is it a wipe, or a hole?
|
|
1775
|
+
```
|
|
1776
|
+
|
|
1777
|
+
---
|
|
1778
|
+
|
|
1779
|
+
## Never verify a video by one frame
|
|
1780
|
+
|
|
1781
|
+
**This is the failure mode that survives every check you'd think to run.** Whole classes of render bug produce a video where *every frame is identical* — the timeline never ran — while duration, frame count, file size and audio hash all come out exactly right. Frame 0 looks perfect, so a single-frame check passes and you ship a frozen video.
|
|
1782
|
+
|
|
1783
|
+
Two real causes, both silent:
|
|
1784
|
+
|
|
1785
|
+
- **A watermark/overlay pass without `-loop 1` on a single-frame PNG input.** The frame-sync collapses the whole video onto one frame. Five videos shipped this way before it was caught.
|
|
1786
|
+
- **Assets outside the composition root.** Only `<style>`/`<script>` *inside* the `data-composition-id` root execute, and sibling relative files may not resolve — fonts, images, even the animation library itself. The timeline never starts; frame 0 still renders fine because frame 0 is the static DOM.
|
|
1787
|
+
|
|
1788
|
+
**The rule that catches both: always compare two frames from different scenes.** They must differ a lot. And when you've applied any pass over an existing video (watermark, overlay, dedupe, re-encode), also compare each output frame against **its own** input frame at the same timestamp — that difference should be tiny. Two checks, opposite directions:
|
|
1789
|
+
|
|
1790
|
+
```bash
|
|
1791
|
+
# consecutive/distant frames must DIFFER (motion preserved)
|
|
1792
|
+
vidfarm stills ./work --at 0,4,9,14
|
|
1793
|
+
# after an overlay pass: same timestamp, before vs after, must be NEARLY IDENTICAL
|
|
1794
|
+
ffmpeg -y -ss 7 -i clean.mp4 -frames:v 1 a.png
|
|
1795
|
+
ffmpeg -y -ss 7 -i final.mp4 -frames:v 1 b.png
|
|
1796
|
+
ffmpeg -i a.png -i b.png -filter_complex "psnr" -f null - # very high PSNR = only the mark changed
|
|
1797
|
+
```
|
|
1798
|
+
|
|
1799
|
+
This sits directly beside the frame-0 rule and is its necessary counterweight: **frame 0 is the thumbnail, so judge it alone — but never judge the VIDEO by it.** The two rules are checking different things, and an agent that only internalises the first one has a perfect blind spot for a frozen render.
|
|
1800
|
+
|
|
1801
|
+
---
|
|
1802
|
+
|
|
1803
|
+
## Verify audio by measurement, not by ear
|
|
1804
|
+
|
|
1805
|
+
**You cannot hear the render.** Do not report "the mix sounds good" — you have no way to know it, and it is the claim that most often turns out false. Measure instead.
|
|
1806
|
+
|
|
1807
|
+
- **Speech-over-bed separation: target 12–15 dB.** Measure the RMS of the mix across the spans where words actually occur, minus the RMS of a bed-only stretch. Word spans come free from the transcription you're already running for captions (`vidfarm stt <file> --engine whisper` → word timings).
|
|
1808
|
+
- **Peak below 0 dBFS.** A mix that clips reads as amateur instantly on a phone speaker.
|
|
1809
|
+
- **Beware "separation" numbers computed over the music-only tail** — they measure the wrong thing (bed alone vs. bed alone) and over-report by a wide margin. Don't retune a mix based on one.
|
|
1810
|
+
|
|
1811
|
+
```bash
|
|
1812
|
+
ffmpeg -i final.mp4 -af "volumedetect" -f null - # peak + mean over the whole file
|
|
1813
|
+
ffmpeg -i final.mp4 -ss 3.1 -t 1.4 -af "volumedetect" -f null - # a span where a word is spoken
|
|
1814
|
+
```
|
|
1815
|
+
|
|
1816
|
+
**Narration timing gotchas that produce a correct-looking, wrong-sounding video:**
|
|
1817
|
+
|
|
1818
|
+
- **Never `adelay` the voiceover.** Whisper's word timings — and therefore every caption you generated from them — are relative to the raw `vo.wav`. Delaying the VO desyncs every caption in the video while the file still plays fine. Use `apad` + `atrim` to place it instead.
|
|
1819
|
+
- **Scene handoffs can leave ~0.3s of silence.** Extend each clip's audio ~0.35s into the next.
|
|
1820
|
+
- **Whisper's default model is English-only and will hallucinate fluent English over another language.** Non-English narration needs `--model large-v3 --language <code>`. The output looks like a clean transcript, so this one ships silently.
|
|
1821
|
+
- **`vidfarm tts` reads stdin** — always redirect `</dev/null` when calling it inside a shell loop, or the loop eats its own input.
|
|
1822
|
+
|
|
1823
|
+
---
|
|
1824
|
+
|
|
1825
|
+
## The revision pass — how to fix what review found
|
|
1826
|
+
|
|
1827
|
+
**Spawn a fresh pass rather than re-litigating with the context that produced the defect.** If you're handing fixes to a subagent (or picking the work back up yourself later), the brief that works:
|
|
1828
|
+
|
|
1829
|
+
- **State it as N targeted fixes and nothing else.** "The video is good — you are making three specific fixes." Open-ended "improve it" turns a working video into a different, differently-broken video.
|
|
1830
|
+
- **Edit the generator, not the generated output.** If a script produced `composition.html`, fix the script. Check first that a generator exists — some compositions are hand-authored.
|
|
1831
|
+
- **Back up before overwriting** — keep `<slug>-v1.mp4`. Re-renders are cheap locally; a lost good version isn't.
|
|
1832
|
+
- **Keep audio bit-identical unless audio is the defect.** Reuse the existing `vo.wav` / word timings rather than re-recording; a re-record retimes every caption for no reason.
|
|
1833
|
+
- **Give the PROBLEM, not just your proposed solution.** Repeatedly, the agent handed a described defect found a better fix than the one specified — using an app's own collapsed UI state instead of a redaction box, a type safe-zone solver instead of a scrim, making a document's *arrival* the spectacle instead of cutting the document. Say what's wrong and at what timestamp; let the fix be found.
|
|
1834
|
+
- **Then sweep for the same class of problem** across the rest of the video, and report what else turned up. Defects of a given kind are rarely solitary — they come from a habit.
|
|
1835
|
+
|
|
1836
|
+
---
|
|
1837
|
+
|
|
1838
|
+
## Report both halves honestly
|
|
1839
|
+
|
|
1840
|
+
When you hand back a render, say what you **measured** and what you **judged**, separately:
|
|
1841
|
+
|
|
1842
|
+
- Machine-settled: `vidfarm qa ./work` findings, `vidfarm lint`, durations, peak dBFS, frame-difference checks.
|
|
1843
|
+
- Human-judgment: the holistic pass above — balance, spacing, type continuity, colour coherence, pacing, joins — plus the harness's `- [ ]` review items.
|
|
1844
|
+
|
|
1845
|
+
**Never report a clean pass on the half you didn't actually look at.** A confident "verified, looks good" over an unreviewed video is worse than no review, because it spends the director's trust on nothing — and per the base rate at the top of this file, it is usually wrong.
|
|
1589
1846
|
|
|
1590
1847
|
## Download a video from a website (Vidfarm fetches it for you — paid plans)
|
|
1591
1848
|
|
|
@@ -1814,38 +2071,62 @@ Send a stable `tracer` on export so retries are traceable and filterable in job
|
|
|
1814
2071
|
| | **One-time video** | **Bulk / scripting mode** |
|
|
1815
2072
|
|---|---|---|
|
|
1816
2073
|
| The deliverable | One MP4 you both look at | A loop that produces N videos nobody watches frame-by-frame |
|
|
1817
|
-
| Quality control | Your eyes on the render | **A `
|
|
2074
|
+
| Quality control | Your eyes on the render | **A `HARNESS.md`** — the batch's written standard |
|
|
1818
2075
|
| What you optimize | This video | The *variant axis* (one thing changes; everything else is held) |
|
|
1819
2076
|
| Cost posture | Per-video decisions are fine | Per-video AI spend × N — reuse assets, prefer clip pools |
|
|
1820
2077
|
|
|
1821
|
-
A director who says "make me a video about X" usually wants the first. A director who says "I need to post daily" / "make 20 variants" / "test hooks" wants the second and often doesn't know it has a name. **Offer the upgrade explicitly:** *"Want this as one video, or should we set it up as a repeatable batch? Batches get a
|
|
2078
|
+
A director who says "make me a video about X" usually wants the first. A director who says "I need to post daily" / "make 20 variants" / "test hooks" wants the second and often doesn't know it has a name. **Offer the upgrade explicitly:** *"Want this as one video, or should we set it up as a repeatable batch? Batches get a HARNESS.md so variant #37 is as good as #1."* Don't silently build a one-off when they asked for volume, and don't drag someone into a scripting harness when they wanted one clip.
|
|
1822
2079
|
|
|
1823
|
-
### `
|
|
2080
|
+
### `HARNESS.md` — the reusable AI harness for a format
|
|
1824
2081
|
|
|
1825
|
-
`vidfarm qa`'s built-in rules are **universal** (no HTML slop, the font regime, the thumbnail frame) — the same for everyone, so they live in code. A
|
|
2082
|
+
`vidfarm qa`'s built-in rules are **universal** (no HTML slop, the font regime, the thumbnail frame) — the same for everyone, so they live in code. A harness is the opposite: it's what makes **this** director's **this** format good — their audience, hook shape, banned vocabulary, pacing, compliance line, and the DNA of the template it came from. It can't be hard-coded, so it lives next to the work as Markdown they own and version.
|
|
1826
2083
|
|
|
1827
|
-
**It exists because bulk output loses its human reviewer.** One video gets eyes on every frame; fifty generated in a loop do not. The
|
|
2084
|
+
**It exists because bulk output loses its human reviewer.** One video gets eyes on every frame; fifty generated in a loop do not. The harness is what the loop grades against.
|
|
2085
|
+
|
|
2086
|
+
**Three director phrasings, one artifact:**
|
|
2087
|
+
|
|
2088
|
+
| They say | You run |
|
|
2089
|
+
|---|---|
|
|
2090
|
+
| "create me a harness" | `vidfarm harness init <base> --out ./work/HARNESS.md`, then edit it with them |
|
|
2091
|
+
| "update the harness for this format" | open the file, add the rule **with its reason**, re-run `vidfarm qa` |
|
|
2092
|
+
| "give me the harness for this template_id" | `vidfarm harness derive <templateId\|forkId>` — the **decomposition**, as a harness |
|
|
1828
2093
|
|
|
1829
2094
|
```bash
|
|
1830
|
-
vidfarm
|
|
1831
|
-
vidfarm
|
|
1832
|
-
vidfarm
|
|
1833
|
-
vidfarm
|
|
2095
|
+
vidfarm harness list # the bundled starting points
|
|
2096
|
+
vidfarm harness init short-form --out ./work/HARNESS.md # copy, then EDIT it
|
|
2097
|
+
vidfarm harness derive <forkId> --out ./work/HARNESS.md # a decomposed template → a harness
|
|
2098
|
+
vidfarm harness show ./work/HARNESS.md --dna visual # ONE strand, not the whole doc
|
|
2099
|
+
vidfarm qa ./work # auto-picks up ./work/HARNESS.md
|
|
2100
|
+
vidfarm qa ./work --harness hooks --harness ./brand/HOUSE.md # built-in + your own file — they STACK
|
|
1834
2101
|
```
|
|
1835
2102
|
|
|
1836
|
-
|
|
2103
|
+
**A harness mirrors the template JSON's DNA vocabulary.** Every `## … DNA` heading is indexed under the same key the decompose pass uses, so a derived harness and a hand-written one read the same:
|
|
2104
|
+
|
|
2105
|
+
| Strand | What lives there | Decompose source |
|
|
2106
|
+
|---|---|---|
|
|
2107
|
+
| **Viral DNA** | hook, retention mechanic, payoff, core emotion, contrast | `video-context.json` → `viral_dna` |
|
|
2108
|
+
| **Visual DNA** | cut rhythm, energy curve, caption style/placement, b-roll, transitions | `editor-harness.json` → `pacing` / `typography` / `broll` |
|
|
2109
|
+
| **Structural DNA** | the beats, their roles, which are load-bearing | `editor-harness.json` → `scenes`, `scene-annotations.json` |
|
|
2110
|
+
| **Audio DNA** | voiceover, bed, SFX, comedic timing, intonation | `editor-harness.json` → `audio` / `emotional` |
|
|
2111
|
+
| **Build DNA** | which paintbrush per beat, the free-tier path | `replication-harness.json` |
|
|
1837
2112
|
|
|
1838
|
-
|
|
2113
|
+
`harness derive` writes what the decompose pass actually recorded and marks the rest `unknown` — it never invents a strand to look complete. Treat its output as a **first draft**: the model watched the video, it didn't talk to the customer.
|
|
1839
2114
|
|
|
1840
|
-
|
|
2115
|
+
> Don't confuse `HARNESS.md` with the `.harness/` directory `vidfarm pull` writes. That directory is machine-generated context (`context.json`, `agent-guide.md`), regenerated on every pull — never hand-edit it. `HARNESS.md` is the one the director owns.
|
|
1841
2116
|
|
|
1842
|
-
|
|
2117
|
+
Bundled bases (`vidfarm harness list`, files under `.agents/skills/vidfarm/harnesses/`): **`short-form`** (the default — the four charges hook/loop/payoff/bait + the standalone rule), **`hooks`** (hook-variant batches: chunk-1 legibility, the unguessable test, the anti-patterns that only show up at volume), **`ugc-testimonial`**, **`explainer`**, **`product-demo`**. Each is a *starting point to edit*, never a house style to conform to — the parts that matter most are the parts the director adds. A harness can also be any file anywhere: `--harness ./campaigns/q3/RULES.md` is fully supported, and `VIDFARM_HARNESS=./work/HARNESS.md` sets a default for a whole run.
|
|
2118
|
+
|
|
2119
|
+
**The format is two halves, and the split is deliberate:** a front-matter `checks:` block the CLI settles deterministically (duration, aspect, `hook_words_max`, `forbid_text`, `first_frame_text`, … — full key list in `harnesses/README.md`), and every `- [ ]` checkbox in the body, which comes back as a **review item for you to answer**. "Is the withheld answer one the viewer can't supply themselves?" is a judgment call; a linter claiming to settle it would be lying. **Answer the review items honestly in your report** — the CLI prints them precisely because it can't.
|
|
2120
|
+
|
|
2121
|
+
**Build on it.** When you learn something from a batch ("the label-framed hooks all died"), write it into the harness as a new rule or checklist line. That is the artifact that compounds across runs; the composition files don't.
|
|
2122
|
+
|
|
2123
|
+
### The bulk loop, with the harness in it
|
|
1843
2124
|
|
|
1844
2125
|
```bash
|
|
1845
|
-
vidfarm
|
|
2126
|
+
vidfarm harness init hooks --out ./work/HARNESS.md # once, then edit for this account
|
|
1846
2127
|
for VARIANT in "${VARIANTS[@]}"; do
|
|
1847
2128
|
vidfarm set-text ./work --layer hook --text "$VARIANT"
|
|
1848
|
-
vidfarm qa ./work --json > "qa/$SLUG.json" #
|
|
2129
|
+
vidfarm qa ./work --json > "qa/$SLUG.json" # harness auto-discovered from ./work
|
|
1849
2130
|
jq -e '.ok' "qa/$SLUG.json" >/dev/null || continue # YOUR gate, in YOUR script
|
|
1850
2131
|
vidfarm render "$FORK_ID" --dir ./work --out "renders/$SLUG.mp4"
|
|
1851
2132
|
done
|
|
@@ -2004,7 +2285,7 @@ The licensed harness also carries the **generative build workflow** guidance (ch
|
|
|
2004
2285
|
| `vidfarm sticker-pack [sheet\|url] [--generate "<theme>"] [--items "a,b,c"] [--count <n>] [--dry-run] [--gap <pct>] [--min-area <pct>] [--output-format png\|webp\|gif] [--out-dir <d>]` | **local, free, ffmpeg-only** (no job; only `--generate` bills, ONCE for the whole set) — key + alpha-channel segmentation + per-item trim | **The STICKER-PACK maker — the answer whenever a director asks for "a sticker pack" / prop set / icon set.** A pack is ONE greenscreen sheet holding every item, keyed once and then masked apart: 1/N the cost of N `cutout` calls, and the only way a cast stays on-style. Finds each item **automatically** by segmenting the keyed sheet's alpha into connected islands — no hand-measured `--crop` rects — and writes one snug transparent file per item (named from `--items`, reading order) plus a `stickers.json` manifest. `--dry-run` prints the detected boxes first; `--gap` merges (lower) or splits (raise) items that came out joined/broken; items have **no maximum size** — a full-frame landscape/backdrop is as valid a sticker as a 3% icon. **Plate color is chosen for you:** when generating it reads the subject and moves the plate off any hue the art uses (green → magenta → blue → black → white — a pack of leaves/frogs/money on GREEN would key holes through the art), and when splitting an existing sheet it DETECTS the plate from the sheet's four corners, so a red/purple sheet handed back from a web tool just works. Pin it with `--key-color`/`--preset`, or `--no-auto-key` for plain green. **The ART is made key-safe too:** the generation prompt is auto-appended with "closed, solidly filled shapes, no outline-only/hollow art, nothing in the plate hue or a near-shade, fully opaque, no glow/translucency" — the fix for stickers that come back as a rim around a transparent hole — and after keying each item reports `holes`/`hole_pct`/`hollow` (console `⚠ N% hollow` at ≥20%, plus `--json` and `stickers.json`). It **warns, never blocks** (a ring/frame/donut reads identically); re-generate with the fill clause, or lift that one item with `vidfarm mask --crop …`. `--output-format gif` emits 1-bit-alpha GIFs for GIF-only surfaces. IMAGE-only. Aliases: `stickers`, `sticker-sheet`. See recipe `cutout-graphics-for-explainers.md` → "A sticker pack". |
|
|
2005
2286
|
| `vidfarm tts "…" [--style "…"] [--voice <v>] [--out <file>]` | (LOCAL-FIRST: your own OPENAI/GEMINI/OPENROUTER_API_KEY → audio file on disk; `--cloud` = `POST /api/v1/primitives/audio/speech` + poll, ElevenLabs on the platform key by default, `--own-key` for yours) | text → narration audio; `--cloud --voice <voice_id>` picks an ElevenLabs voice |
|
|
2006
2287
|
| `vidfarm music "<prompt>" [--length <sec>] [--out <f>] [--own-key]` | `POST /api/v1/primitives/music/generate` (polls job) | prompt → music track (ElevenLabs; platform key + wallet by default, `--own-key` for yours) |
|
|
2007
|
-
| `vidfarm voices [--own-key] [--limit N]` | `GET /api/v1/primitives/audio/voices` |
|
|
2288
|
+
| `vidfarm voices [--sample] [--search "…"] [--free\|--all] [--own-key] [--limit N]` | `GET /api/v1/primitives/audio/voices` | **Browse AND sample narration voices.** Default roster = the premium ElevenLabs catalog reached through **vidfarm's own ElevenLabs connection** — the user needs no ElevenLabs account, API key, or subscription; narration is billed as vidfarm wallet credits (pennies each). `--free` = the $0 local Kokoro roster (`--all` = both). `--sample` writes listenable clips to `./voice-samples` (`--sample-count`, `--sample-out`, `--sample-text`) and is **free on both tiers** — premium samples are ElevenLabs' own preview clips, free samples render locally — so it's safe in `minimize`. `--search` filters by name/labels/description. **In interactive mode play the samples and let the USER pick**; autonomous = default a voice and still say they can choose. `--own-key` lists the customer's own ElevenLabs account instead. |
|
|
2008
2289
|
| `vidfarm stt <file\|url> [--out <base>] [--no-diarize]` (alias: `transcribe`) | (LOCAL-FIRST: local ffmpeg demux + your own key; `--cloud` = `POST /api/v1/primitives/audio/transcribe` + poll, ElevenLabs Scribe on the platform key by default, `--own-key` for yours) | video/audio → transcript in BOTH formats: simple subtitles (txt + SRT) and multi-speaker segments (json) |
|
|
2009
2290
|
| `vidfarm place <dir> --src <url\|file> [--at\|--replace]` | (edits local composition.html; local files → serve disk store or temp upload) | drop media (URL **or local file**) into a gap / over a scene |
|
|
2010
2291
|
| `vidfarm captions generate <dir> [--style <preset>] [--audio <f>\|--srt <f>\|--text "…"]` | (LOCAL-FIRST: STT on your own key — OpenAI = real word timestamps — then edits local composition.html) | transcribe narration → animated word-by-word caption cues |
|
|
@@ -2053,11 +2334,12 @@ The licensed harness also carries the **generative build workflow** guidance (ch
|
|
|
2053
2334
|
| `vidfarm raws search "…"` / `raws match "…"` / `raws list` / `raws sources` | (local library; NL→criteria via local agent or provider key) | search/reuse the raws library |
|
|
2054
2335
|
| `vidfarm raws preset list\|run\|save` / `raws export <ids…> --to <dir>` | (local library) | saved queries; copy raw MP4s out |
|
|
2055
2336
|
| `vidfarm lint <dir\|composition.html>` | (local static validation) | pre-publish composition check: timing, overlaps, preset names, media src |
|
|
2056
|
-
| `vidfarm stills <dir> [--at 0,2.5,…]` | (local in-process render of PNG frames) | visually verify an edit without a full render |
|
|
2057
|
-
| `vidfarm qa <dir\|composition.html> [--
|
|
2058
|
-
| `vidfarm
|
|
2337
|
+
| `vidfarm stills <dir> [--at 0,2.5,…] [--sheet]` | (local in-process render of PNG frames) | visually verify an edit without a full render. **`--sheet` also tiles them into one contact sheet** (`<out>/contact-sheet.png`, `--sheet-out`/`--sheet-width` to tune) — the whole-video review pass: read it as ONE image and sequence-level drift (uneven margins, three type sizes, a wandering accent colour, N identical beats, a jarring join) becomes obvious where per-scene checks never see it |
|
|
2338
|
+
| `vidfarm qa <dir\|composition.html> [--harness <name\|path>…] [--json] [--strict]` | (local static QA — **devcli-only**, no cloud/REST twin) | **social-native QA: HTML slop + first frame + font regime. Run it on EVERY video you produce.** `--harness` grades against a HARNESS.md too (stackable). Free, instant, feedback-only |
|
|
2339
|
+
| `vidfarm harness list\|show <ref> [--dna <strand>]\|init <name> [--out <path>]\|derive <forkId\|dir>\|check <dir>` | (local — **devcli-only**) | **HARNESS.md: the reusable AI harness for one format or template.** `init` copies a bundled base to edit; `derive` turns a decomposed template's DNA into one ("give me the harness for this template_id"); `check` is `vidfarm qa` under the harness noun |
|
|
2059
2340
|
| `vidfarm doctor` | (local environment triage) | check ffmpeg/node/keys/agent CLI/poisoned env + list local serve/preview processes before debugging anything else; `--kill-orphans` reaps dead servers squatting ports (fixes the "Waiting for preview server…" hang) |
|
|
2060
2341
|
| `vidfarm skills list\|add <name>\|update` | `GET /skill-pack/index.json` · `/skill-pack/:name/*` | install/refresh skill packs (see "Skill packs — import on demand") |
|
|
2342
|
+
| `vidfarm skill ls\|show <path>\|search "<term>"\|path` | (local — **offline, no account**) | **Read this pack straight off disk.** A full copy ships inside the devcli tarball and is pinned to the installed version. `search` greps all 22 files at once — the cheapest way to find one paragraph without loading a whole reference |
|
|
2061
2343
|
| `vidfarm tts "…" --engine local` / `vidfarm stt <file> --engine whisper` | (keyless LOCAL engines: Kokoro-82M TTS, whisper.cpp STT) | narration + word-timestamp transcripts with zero keys and zero accounts |
|
|
2062
2344
|
| `vidfarm remove-background <video\|image>` | (local ONNX matting — free) | transparent-subject media for occlusion captions/cutouts (arbitrary/messy background; for a FLAT solid background use `remove-background-greenscreen`) |
|
|
2063
2345
|
| `vidfarm capture <url>` | (local headless-Chrome capture) | website screenshots/assets for website-to-video flows |
|
|
@@ -2075,8 +2357,8 @@ The licensed harness also carries the **generative build workflow** guidance (ch
|
|
|
2075
2357
|
vidfarm qa ./work # human-readable findings + verdict
|
|
2076
2358
|
vidfarm qa ./work --json # machine-readable: rule / severity / where / fix
|
|
2077
2359
|
vidfarm qa ./work --strict # ALSO exit 1 on slop (only if you want a CI gate)
|
|
2078
|
-
vidfarm qa ./work --
|
|
2079
|
-
# auto-discovers ./work/
|
|
2360
|
+
vidfarm qa ./work --harness hooks # + grade against a HARNESS.md (repeatable; also
|
|
2361
|
+
# auto-discovers ./work/HARNESS.md)
|
|
2080
2362
|
```
|
|
2081
2363
|
|
|
2082
2364
|
**Run this on every video you produce.** It is free, instant (pure DOM, no ffmpeg/Chrome/network), and it is the only automated check for the thing that most often ruins an agent-made video: **HTML slop**. Compositions are authored in HTML, so an agent's web-page instincts leak straight onto the frame as landing-page furniture that appears on every website and in **zero** real TikToks.
|
|
@@ -2102,13 +2384,22 @@ What it flags:
|
|
|
2102
2384
|
| `font-regime` | error/warn | A text layer in a website body font (Inter/Roboto/Arial/system-ui → **error**) or any family outside the imported regime (Montserrat, TikTok Sans, Abel, Source Code Pro, Yesteryear → warn, it silently falls back at render) |
|
|
2103
2385
|
| `font-size` / `font-weight` | error/warn | `font-size:0` (invisible) is an error; sub-2.6%-of-canvas-width text and weight <600 warn |
|
|
2104
2386
|
| `caption-safe-zone` | warn | Text outside the 8%–85% band on a **portrait** canvas (landscape/square are exempt) |
|
|
2387
|
+
| `caption-oversize` | warn | Display-size type (>7.5% of canvas width) on a line of **5+ words** — it runs edge-to-edge, wraps, covers the frame, and forces a full-width plate. Both signals required, so a giant 2-word hook card passes |
|
|
2388
|
+
| `wall-of-text` | warn | One **static** text layer carrying 14+ words — a paragraph, not a caption. Page it into 3–5-word kinetic cues (`captions generate --style word-pop`). Layers already part of an animated caption run are exempt |
|
|
2389
|
+
| `dead-air` | warn | A gap of **≥2.5s between cues** with nothing on screen to read (needs 3+ cues, so a two-card title sequence is exempt). Dead screen time is a free exit — cut it and `ripple` the hole closed |
|
|
2390
|
+
| `dead-tail` | warn | The video keeps running **>1.5s after the last word** — an outro, an end card, or an untrimmed clip. End on the bait |
|
|
2391
|
+
| `slow-scene` | warn | One clip >6s **and** >2.5× the median clip length — judged against the video's OWN rhythm, so a deliberately slow piece or a single-take talking head passes |
|
|
2105
2392
|
| `thumbnail-blank-open` | error | Nothing on screen at **t=0** — the opening clip starts late, so the poster frame is black |
|
|
2106
2393
|
| `thumbnail-fade-in` | error/warn | An **entrance** transition on the FIRST clip: `fade-black`/`fade-white`/`flash`/`smoke` → **error** (frame 0 is a flat solid); any other preset → warn (frame 0 caught mid-move). Junction transitions on later clips are never flagged |
|
|
2107
2394
|
| `thumbnail-no-hook-text` | warn | The composition has text, but none of it is up at t=0 — the poster carries no hook words. Ignorable when you're deliberately opening on a clean face/product shot |
|
|
2108
2395
|
|
|
2109
2396
|
Every finding carries a concrete `fix` line — the answer is always "say it as timed text on the footage", never just "delete it". Fold `--json` into scripted batch runs to QA N variants at once.
|
|
2110
2397
|
|
|
2111
|
-
**
|
|
2398
|
+
**Every run ends by telling you to go watch the video — that instruction is part of the output, not a footnote.** `vidfarm qa` closes with a `▶ NOW WATCH THE VIDEO — this check never did` block (and a `watch_the_video: { required: true, why, steps[] }` object in `--json`), printed on **clean** runs too, because a green tick on DOM attributes is the single easiest thing to mistake for a reviewed video. The steps are dir-aware and paste-ready: render, `stills --sheet` → *open the contact sheet*, read it as one sequence, judge each caption against its picture, compare two different scenes (a frozen render passes every mechanical check), measure the audio with `volumedetect`, and report what you measured separately from what you judged. **Do them.** Reporting "QA passed" to a director without opening a frame is not a review, and the tool now says so to your face.
|
|
2399
|
+
|
|
2400
|
+
**`vidfarm qa` is a static DOM check — it cannot see the rendered video.** In particular it can tell you a caption is *too big* or *outside the safe zone*, but never whether it sits in the **empty** part of the frame — that needs pixels, so it stays your job: `vidfarm stills <dir> --at <t>`, look, then place (see `references/editor-workflows.md` → "TikTok-native caption standard"). It never looks at pixels, motion, spacing, colour drift, pacing, or the joins between scenes, so a clean `qa` run says nothing about whether the video reads as one coherent piece. That judgment is a separate, mandatory pass: tile stills into a contact sheet, read it as an image, and check balance/spacing/type/colour/rhythm across the whole sequence. It also can't catch a **frozen render** (every frame identical while duration, frame count and audio hash all pass), which is why you compare frames from two different scenes. Full method: `references/reviewing-renders.md`.
|
|
2401
|
+
|
|
2402
|
+
**The two halves, and why the tool only claims one.** Everything above is universal and mechanical. The half that decides whether a *particular* video is any good — is the hook legible cold, does the loop close, is this variant genuinely different from its siblings — is the director's, and it lives in a **`HARNESS.md`** (see "Scripting mode" above). Pass one with `--harness <name|path>` (repeatable, and a `HARNESS.md` sitting next to the composition is picked up automatically): its `checks:` front matter is settled deterministically alongside the built-ins, and its `- [ ]` checklist comes back as **review items you must answer yourself**. `vidfarm qa` deliberately never fakes a verdict on those — a "PASS" it couldn't have earned is worse than no check at all.
|
|
2112
2403
|
|
|
2113
2404
|
## Cost mode — the devcli's money-saving guardrail
|
|
2114
2405
|
|
|
@@ -2133,6 +2424,15 @@ The four modes, quoted as **cost per finished video**. The first two are spend p
|
|
|
2133
2424
|
|
|
2134
2425
|
**Narration defaults to the FREE local voice in minimize AND hybrid.** A bare `vidfarm tts "…"` runs the keyless local Kokoro-82M engine in both of those modes — you no longer have to remember `--engine local`. A run **opts out** of that default by asking for a premium voice (`--style`, `--provider`, `--model`, `--own-key`, or a non-Kokoro `--voice` like `alloy`/`Kore`/an ElevenLabs id), by passing `--cloud`/`--engine byok`, or by being in `rich-ai`/`pure-videogen`. If the local engine isn't installed on the machine (it needs `pip install kokoro-onnx soundfile` + a ~340MB model on first use), the run **falls back** to the user's provider key / cloud instead of failing — it prints the reason on stderr so you can tell the user why the voice changed.
|
|
2135
2426
|
|
|
2427
|
+
**Narration gotchas that ship a correct-looking, wrong-sounding video.** Each of these produces output that passes every structural check:
|
|
2428
|
+
|
|
2429
|
+
- **Never `adelay` the voiceover to position it.** Whisper's word timings — and therefore every caption generated from them — are relative to the raw `vo.wav`. An `adelay` desyncs every caption in the video while the file still plays perfectly. Use `apad` + `atrim`.
|
|
2430
|
+
- **Whisper's default model is English-only and hallucinates fluent English over other languages.** Non-English narration needs `--model large-v3 --language <code>`; without it you get a clean, confident, entirely invented transcript.
|
|
2431
|
+
- **`vidfarm tts` reads stdin** — redirect `</dev/null` when calling it inside a shell loop, or the loop consumes its own input.
|
|
2432
|
+
- **Check the brand/product name's pronunciation** before you render 20 variants with it. TTS engines mangle proper nouns (Kokoro reads *Genki* as "Jenki"); respell it phonetically in the TTS input and confirm with a whisper round-trip — you're already running whisper for the caption timings.
|
|
2433
|
+
- **For a calm, unhurried read, render line by line** and concatenate the takes with measured silences, rather than one continuous pass. A single pass reads rushed however slow the copy is, because the pauses are TTS filler rather than real beats.
|
|
2434
|
+
- **Verify the mix by measurement, not by ear** — you can't hear the render. Target **12–15 dB** of speech-over-bed separation measured across the actual word spans, peak below **0 dBFS**. A "separation" figure computed over the music-only tail measures bed-vs-bed and over-reports badly; don't retune against it. See `references/reviewing-renders.md`.
|
|
2435
|
+
|
|
2136
2436
|
Precedence: `--cost-mode <m>` flag → `VIDFARM_COST_MODE` env → the saved `cost-mode` → default (hybrid, flagged as "not set"). When nothing is saved and a billed op runs, the CLI prints a "no preference set — ask the user" nudge instead of silently spending, so the default posture really is *ask before you spend*.
|
|
2137
2437
|
|
|
2138
2438
|
**Agent-memory handoff.** After the user picks, offer to remember it across sessions — but the destination depends on the agent, so ask: Claude Code → `CLAUDE.md` (or its memory dir); Codex / OpenCode / most others → `AGENTS.md`; or a note file the user names. `vidfarm cost-mode <choice>` already persists the devcli-side preference; agent memory is the extra step that survives a fresh checkout. In the **web app UI** there is no memory file — ask each time unless the user states a standing preference for the session.
|
|
@@ -2208,6 +2508,25 @@ The customer-facing walkthrough (the "VidFarm Walkthrough Tutorial" course) is p
|
|
|
2208
2508
|
|
|
2209
2509
|
Both are public and read-only (no auth). Prefer these to guessing steps — quote the real chapter and link the reader to its `url`. Chapters cover onboarding/setup, the operating funnel (angles/hooks/awareness), each guided edit demo (recaption, product tease, remix-with-raws, actor replacement, animate-static-book, drama series, product promo, motion explainers), sourcing/clipping raws, the wallet, cancellation/refunds, and the developer devcli/scripting/free-mode chapters.
|
|
2210
2510
|
|
|
2511
|
+
## The director pack ships inside the devcli — read it offline
|
|
2512
|
+
|
|
2513
|
+
Installing `@officexapp/vidfarm-devcli` puts a **complete copy of this pack on disk**, pinned to that CLI version. You never have to be online, logged in, or in a project with `.agents/skills/` to read it:
|
|
2514
|
+
|
|
2515
|
+
```bash
|
|
2516
|
+
vidfarm skill ls # every file, with sizes
|
|
2517
|
+
vidfarm skill show primitives # shorthand resolves to references/primitives.md
|
|
2518
|
+
vidfarm skill show harnesses/README.md # or an exact path
|
|
2519
|
+
vidfarm skill search "greenscreen" # grep all of it — find the paragraph, then open that file
|
|
2520
|
+
vidfarm skill path # where the bundled copy lives
|
|
2521
|
+
```
|
|
2522
|
+
|
|
2523
|
+
**Prefer `skill search` over opening a big reference.** `editor-workflows.md` is ~650 lines and `automation-and-local-dev.md` ~520; a grep that returns `references/primitives.md:214` costs almost nothing and tells you exactly which file to load.
|
|
2524
|
+
|
|
2525
|
+
Two things this does NOT mean:
|
|
2526
|
+
|
|
2527
|
+
- **Pinned, not live.** The bundled copy matches the installed CLI — which is the pairing that actually works, since a newer skill against an older binary is the usual cause of *"the skill says to do X but the command 404s"*. For the host's latest, `vidfarm skills add vidfarm` (installs into a project) or `vidfarm skill --print --remote`. When they disagree, update **both halves together**: <https://vidfarm.cc/update.md>.
|
|
2528
|
+
- **Documentation, not entitlement.** Reading about a paid primitive offline does not make it run offline. The **free-local** half genuinely needs nothing — clip hunting, hyperframes, `vidfarm serve` render, `vidfarm qa`, harnesses, `vidfarm dedupe`, Kokoro TTS, whisper STT. The **paid-cloud** half still needs `vidfarm login` and a network call: AI image/video/voice generation, hosted render, `recycle`, `download-video`, marketplace, and the hosted file directory. Tell the director which half a plan lands in *before* you build it.
|
|
2529
|
+
|
|
2211
2530
|
## Skill packs — import on demand (HyperFrames-grade authoring power)
|
|
2212
2531
|
|
|
2213
2532
|
This skill stays lean on purpose. Deep authoring craft lives in **skill packs** — Vidfarm's whitelabel of the open-source `hyperframes` skill suite (same engine as `vidfarm hf` / `vidfarm render`, Vidfarm-branded) plus Vidfarm's own media pack — vendored on the Vidfarm host and installed only when a task needs them. Never install skills from upstream vendor orgs or third-party registries; the vidfarm mirror is the source (`vidfarm skills add <name>` fetches `GET /skill-pack/:name/*` with hash verification into `.agents/skills/` + a `.claude/skills/` link, pinned in `skills-lock.json`; `vidfarm skills list` shows what is available/installed; `vidfarm skills update` refreshes pins).
|
|
@@ -2801,8 +3120,9 @@ Use this when a coding agent is doing the work locally or the user wants a repro
|
|
|
2801
3120
|
3. Read `./work/.harness/agent-guide.md` and `./work/.harness/context.json` before editing.
|
|
2802
3121
|
4. Make deterministic edits to `composition.html` and optionally `composition.json`.
|
|
2803
3122
|
5. Validate with `vidfarm lint` or `vidfarm stills` when useful. **Always look at `vidfarm stills ./work --at 0`** — that frame becomes the thumbnail, so it must not be black, empty, or mid-fade.
|
|
2804
|
-
6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, a lone pill around a static stat/label, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, and flags a blank/fading first frame (the thumbnail). Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
|
|
3123
|
+
6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, a lone pill around a static stat/label, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, flags oversized captions and static walls of text, and flags a blank/fading first frame (the thumbnail). It cannot see pixels, so *where in the frame* the caption sits is still on you — which is why every run ends with a **`▶ NOW WATCH THE VIDEO`** block: render, `vidfarm stills ./work --sheet`, open the contact sheet, and judge each caption against its actual picture. Do that before you report the video as done. Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
|
|
2805
3124
|
7. Render with `vidfarm render <forkId> --dir ./work --wait`.
|
|
3125
|
+
7b. **Review the render as a whole before you approve — this is the step that most changes quality.** `vidfarm qa` and `lint` are static checks on the DOM; neither can see the video. Tile ~12 stills into one contact sheet and read it as an image — `vidfarm stills ./work --sheet` does both in one command (add `--at 0,2,4,…` to pick the timestamps): consistent margins, one type scale, one accent colour, deliberate pacing, no jarring join, no dead band under top-anchored content, end card settled ≥2s before the last frame. Compare frames from two different scenes — a frozen render (overlay pass without `-loop 1`, assets outside the composition root) passes duration, frame-count and audio-hash checks while every frame is identical. Check the mix by measurement, not by ear. Full method + the six most common defects: `references/reviewing-renders.md`.
|
|
2806
3126
|
8. **Ask about deduplication before you approve** — "is this going out more than once (several accounts, another platform, a re-post later)?" If yes, run `vidfarm dedupe ./final.mp4 [--variants N]` on the **exported** MP4 (free, local ffmpeg, no re-render) and approve each variant separately. Asking here rather than after publication is what avoids paying for a second render. See `references/core-workflows.md` → *Deduplicate before you publish*.
|
|
2807
3127
|
9. Approve the finished MP4 with `vidfarm approve --video <url|./final.mp4> --caption "..."`. This prints the shareable `share_url`.
|
|
2808
3128
|
|
|
@@ -2810,15 +3130,15 @@ Use this when a coding agent is doing the work locally or the user wants a repro
|
|
|
2810
3130
|
|
|
2811
3131
|
Prefer this path for batch work, CI-like edits, or when the user wants free local rendering through `vidfarm serve`.
|
|
2812
3132
|
|
|
2813
|
-
## Recipe: Bulk Video Generation (Scripting Mode) with a
|
|
3133
|
+
## Recipe: Bulk Video Generation (Scripting Mode) with a HARNESS.md
|
|
2814
3134
|
|
|
2815
3135
|
Use this when the director wants **volume** — daily posting, hook tests, one video per clip in a pool, N variants of a template. Ask first if you're not sure: *"One video, or should we set this up as a repeatable batch?"* If they want volume, this is the shape.
|
|
2816
3136
|
|
|
2817
|
-
The thing that makes bulk work is not the loop — loops are easy. It's that **nobody is going to watch variant #37 as carefully as variant #1**, so the standard has to be written down before the loop runs. That's the `
|
|
3137
|
+
The thing that makes bulk work is not the loop — loops are easy. It's that **nobody is going to watch variant #37 as carefully as variant #1**, so the standard has to be written down before the loop runs. That's the `HARNESS.md`.
|
|
2818
3138
|
|
|
2819
3139
|
### 0. Read the craft harness first
|
|
2820
3140
|
|
|
2821
|
-
`references/hooks-and-virality.md` — the four charges (hook / loop / payoff / bait), the three gates, and the anti-patterns that only bite at volume. Two of them decide whether this batch is worth running at all: **a different noun is not a different hook** (twenty variants of one sentence with the nouns swapped is one video), and **never point a generator at your grader** (a model writing hooks scored by the same model converges on the rubric, not on what works — scores climb, nothing improves). The
|
|
3141
|
+
`references/hooks-and-virality.md` — the four charges (hook / loop / payoff / bait), the three gates, and the anti-patterns that only bite at volume. Two of them decide whether this batch is worth running at all: **a different noun is not a different hook** (twenty variants of one sentence with the nouns swapped is one video), and **never point a generator at your grader** (a model writing hooks scored by the same model converges on the rubric, not on what works — scores climb, nothing improves). The harness catches defects; it does not rank winners.
|
|
2822
3142
|
|
|
2823
3143
|
### 1. Agree the variant axis — before any code
|
|
2824
3144
|
|
|
@@ -2832,14 +3152,22 @@ vidfarm pull <forkId> --dir ./work # one canonical base fork per batch
|
|
|
2832
3152
|
|
|
2833
3153
|
Read `./work/.harness/agent-guide.md` first, as always.
|
|
2834
3154
|
|
|
2835
|
-
### 3. Install and EDIT the
|
|
3155
|
+
### 3. Install and EDIT the harness
|
|
3156
|
+
|
|
3157
|
+
Two ways in, depending on where the format came from:
|
|
2836
3158
|
|
|
2837
3159
|
```bash
|
|
2838
|
-
|
|
2839
|
-
vidfarm
|
|
3160
|
+
# (a) From a bundled base — when the format is one you're defining
|
|
3161
|
+
vidfarm harness list # short-form | hooks | ugc-testimonial | explainer | product-demo
|
|
3162
|
+
vidfarm harness init hooks --out ./work/HARNESS.md
|
|
3163
|
+
|
|
3164
|
+
# (b) From the template you're batching — when the format is one you're REPLICATING
|
|
3165
|
+
vidfarm harness derive <forkId> --out ./work/HARNESS.md # the decomposition, as a harness
|
|
2840
3166
|
```
|
|
2841
3167
|
|
|
2842
|
-
|
|
3168
|
+
(b) is what a director means by *"give me the harness for this template_id"*: the decompose pass already extracted the template's viral / visual / structural / audio / build DNA, and `derive` folds those strands into the same editable Markdown a bundled base produces. If the fork was never decomposed, run `vidfarm decompose` first.
|
|
3169
|
+
|
|
3170
|
+
Either way, **edit it with the director**. The generated file is a starting point; the parts that matter are the ones they add — who the viewer is, their banned vocabulary, the compliance line, the pacing this account actually uses. A harness nobody edited isn't about their videos. A *derived* harness has the extra failure mode of sounding authoritative: it was written by a model that watched one video, so its "unknown" lines and its confident-but-wrong lines both need a human pass. Existing harness somewhere else on disk? Just point at it: `--harness ./brand/HOUSE_RULES.md`. They stack.
|
|
2843
3171
|
|
|
2844
3172
|
### 4. Source the N cheaply
|
|
2845
3173
|
|
|
@@ -2850,13 +3178,13 @@ vidfarm public-raws --category scroll-stoppers --limit 20 --json > pool.json
|
|
|
2850
3178
|
|
|
2851
3179
|
A curated shelf is a pre-tagged, free, already-hosted clip pool — the cheapest way to get N distinct variants without N downloads or N generation calls.
|
|
2852
3180
|
|
|
2853
|
-
### 5. Loop: edit → QA against the
|
|
3181
|
+
### 5. Loop: edit → QA against the harness → render
|
|
2854
3182
|
|
|
2855
3183
|
```bash
|
|
2856
3184
|
for VARIANT in "${VARIANTS[@]}"; do
|
|
2857
3185
|
SLUG="$(echo "$VARIANT" | tr ' ' '-' | cut -c1-40)"
|
|
2858
3186
|
vidfarm set-text ./work --layer hook --text "$VARIANT"
|
|
2859
|
-
vidfarm qa ./work --json > "qa/$SLUG.json" # ./work/
|
|
3187
|
+
vidfarm qa ./work --json > "qa/$SLUG.json" # ./work/HARNESS.md auto-discovered
|
|
2860
3188
|
jq -e '.ok' "qa/$SLUG.json" >/dev/null || { echo "skipped $SLUG"; continue; }
|
|
2861
3189
|
vidfarm render "$FORK_ID" --dir ./work --out "renders/$SLUG.mp4" --tracer "batch-$SLUG"
|
|
2862
3190
|
done
|
|
@@ -2876,13 +3204,28 @@ done
|
|
|
2876
3204
|
|
|
2877
3205
|
Free, offline, no wallet. Variant 1 is the `standard` preset as authored (skew 2%, zoom 3%, rotate 2°, speed +2%, saturation +4%); later variants get jittered magnitudes and flipped signs, so they differ from the original **and from each other**. One variant per account — two accounts posting the same variant defeats the point. Reuse one `--seed` per source so the batch is reproducible.
|
|
2878
3206
|
|
|
3207
|
+
### 5c. Eyeball the renders — the loop cannot do this for you
|
|
3208
|
+
|
|
3209
|
+
**A batch is exactly where "the agent passed its own broken work" compounds**: nobody is watching variant #37, and `vidfarm qa` is a static DOM check that never sees a rendered pixel. So add one cheap visual pass over the output — a contact sheet per video, read as an image:
|
|
3210
|
+
|
|
3211
|
+
```bash
|
|
3212
|
+
for MP4 in renders/*.mp4; do
|
|
3213
|
+
S="$(basename "$MP4" .mp4)"; mkdir -p "qa/$S"
|
|
3214
|
+
for t in 0 3 6 9 12 15; do ffmpeg -y -ss $t -i "$MP4" -frames:v 1 "qa/$S/f$(printf %03d $t).png"; done
|
|
3215
|
+
ffmpeg -y -pattern_type glob -i "qa/$S/f*.png" \
|
|
3216
|
+
-vf "scale=320:-1,tile=3x2:margin=6:padding=6:color=0x999999" -frames:v 1 "qa/$S-sheet.png"
|
|
3217
|
+
done
|
|
3218
|
+
```
|
|
3219
|
+
|
|
3220
|
+
Read the sheets. In a batch you're looking for two different things: **per-video** defects (dead regions, a placeholder that reads as a failed render, contradictory numbers, an unlanded CTA, superimposed headlines at a handoff) and **cross-video** drift (variants that no longer look like siblings, or look *too* identical to be N distinct posts). Also compare two frames from different scenes in at least a sample of the renders — a systematic frozen-render bug in the loop will produce N broken files that all pass duration/frame-count checks. Full method: `references/reviewing-renders.md`.
|
|
3221
|
+
|
|
2879
3222
|
### 6. Answer the review items — don't skip this
|
|
2880
3223
|
|
|
2881
|
-
The
|
|
3224
|
+
The harness's `- [ ]` checklist comes back on every run because the CLI *can't* settle it. Machine checks catch a 13-word hook or a black first frame; only you can answer "is this variant genuinely different from its siblings?" or "can the viewer guess the withheld answer?" **Report both halves honestly**: what the machine checked, and what you judged. A batch report claiming a clean pass on the judgment half is worse than no report.
|
|
2882
3225
|
|
|
2883
|
-
### 7. Feed what you learn back into the
|
|
3226
|
+
### 7. Feed what you learn back into the harness
|
|
2884
3227
|
|
|
2885
|
-
When the director says "the label-framed hooks all died" or "anything over 30s tanked", write it into `
|
|
3228
|
+
When the director says "the label-framed hooks all died" or "anything over 30s tanked", write it into `HARNESS.md` as a rule or a checklist line — with the reason attached, so the next agent doesn't argue it away. The compositions are disposable; **the harness is the artifact that compounds across batches.**
|
|
2886
3229
|
|
|
2887
3230
|
### Cost note
|
|
2888
3231
|
|
|
@@ -2898,13 +3241,15 @@ The mechanical trio — **generate on a chroma plate → key it out → trim to
|
|
|
2898
3241
|
|
|
2899
3242
|
**Unless the director asks for something else, build every explainer this way. Don't ask, just do it, and mention the defaults once so they can override.** The whole point of the house style is that explainers read as *clean, bright, and easy* — a busy explainer is a failed explainer.
|
|
2900
3243
|
|
|
2901
|
-
- **White background, light mode.** A plain white (or near-white `#FFFFFF`–`#FAFAFA`) stage. No dark mode, no gradients, no photographic backdrop, no texture. Light mode reads cleaner on every feed, keeps cutout stickers legible, and makes flat-vector art look intentional. Set the composition/scene background to white first, before placing anything.
|
|
3244
|
+
- **White background, light mode.** A plain white (or near-white `#FFFFFF`–`#FAFAFA`) stage. No dark mode, no gradients, no photographic backdrop, no texture. Light mode reads cleaner on every feed, keeps cutout stickers legible, and makes flat-vector art look intentional. Set the composition/scene background to white first, before placing anything. **This is the default, not a law** — a director who asks for a dark or photographic stage gets one, but it changes two things mechanically: the stickers need their white die-cut rim stripped, and the caption treatment has to be re-measured. Both are documented in **"Stickers on a DARK or photographic stage"** below.
|
|
2902
3245
|
- **Kinetic captions.** Narration is always captioned word-by-word (`vidfarm captions generate ./work --style word-pop`). Because the stage is white, **override the preset's dark-canvas colors to dark ink on light**:
|
|
2903
3246
|
```
|
|
2904
3247
|
vidfarm captions generate ./work --style word-pop \
|
|
2905
3248
|
--color "#111111" --active-color "#7C3AED" --background-style plain --max-words 4
|
|
2906
3249
|
```
|
|
2907
3250
|
One accent color for the active word, everything else near-black. No outline/stroke, no drop shadow, no pill — those exist to survive busy footage and just add noise on white.
|
|
3251
|
+
|
|
3252
|
+
**Those hexes are the answer for a white stage, not the answer.** They are one instance of a general rule: **caption color, active-word color and plate are chosen by MEASURING the background behind the caption band, never by taste or habit.** On a near-black stage the same flags ship a bright plate the design never asked for and an active word nobody can read. The measurement procedure and the three treatments live in `harnesses/short-form.HARNESS.md` → "Caption styling is measured off the background" — read it before you copy the line above onto anything that isn't white.
|
|
2908
3253
|
- **Female TTS narration.** Default to a warm, friendly **female** voice and say which one you picked: local-first `vidfarm tts "<script>" --voice coral` (OpenAI — `nova` if the script wants more energy, `sage` for calmer), `--voice Kore` or `Leda` on Gemini, or `vidfarm voices` → `vidfarm tts --cloud --voice <voice_id>` on ElevenLabs. Tell the director they can swap it in one flag.
|
|
2909
3254
|
- **Clean and simple wins.** One idea on screen at a time. Two or three cutouts per beat, not eight. Generous white space, one accent color, one font. When in doubt, remove an element rather than add one.
|
|
2910
3255
|
|
|
@@ -2986,6 +3331,48 @@ vidfarm remove-greenscreen ./mascot.mp4 --gif --gif-fps 12 --gif-width 480 # AN
|
|
|
2986
3331
|
|
|
2987
3332
|
GIF alpha is **1-bit** — a pixel is fully opaque or fully gone, so antialiased edges go hard and semi-transparent shadows/glows disappear (`--gif-alpha <0..255>` moves where that line falls). That's the format, not the key. **For anything going onto a composition, prefer PNG/WebP (still) or transparent WebM (clip);** reach for GIF only when the destination demands it.
|
|
2988
3333
|
|
|
3334
|
+
### Stickers on a DARK or photographic stage — strip the white die-cut rim
|
|
3335
|
+
|
|
3336
|
+
The house style above puts stickers on a **white** stage, and on white the thing this section is about is invisible. The moment the stage goes dark, photographic, or coloured, every sticker arrives wearing a **white die-cut rim** — a 4–12px light halo tracing its silhouette — and that halo is the single most obvious "a bot made this" artefact in the frame. The art stops reading as an element in the scene and starts reading as a cutout pasted on top of it.
|
|
3337
|
+
|
|
3338
|
+
**Why the rim is there:** it's a PRINT convention. Real die-cut vinyl needs a white border so the blade has something to cut along, so sticker art is drawn with one, so generators reproduce it. It has no purpose whatsoever in a video composition. **This is a different failure from the `⚠ N% hollow` flag** `sticker-pack` prints — hollow means outline-only art whose interior got keyed away (fix it in the prompt, see above); the rim is extra art that was drawn on purpose and has to be removed after the key.
|
|
3339
|
+
|
|
3340
|
+
**The fix:** delete exactly the band of light pixels **connected to the transparent edge**, by morphological reconstruction inward from the boundary. Interior whites — an eyeball's sclera, a screen highlight, a paper label — are enclosed by linework, so they are not connected to the edge and survive untouched. Local, free, `numpy` + `scipy` + `PIL`:
|
|
3341
|
+
|
|
3342
|
+
```python
|
|
3343
|
+
def strip_rim(img, light=188):
|
|
3344
|
+
a = np.array(img.convert("RGBA")).astype(np.int16)
|
|
3345
|
+
solid = a[..., 3] > 128
|
|
3346
|
+
is_light = (a[..., :3].mean(axis=2) > light) & solid
|
|
3347
|
+
seed = is_light & ndimage.binary_dilation(~solid, iterations=3) # light AND touching transparency
|
|
3348
|
+
if seed.any():
|
|
3349
|
+
rim = ndimage.binary_propagation(seed, mask=is_light) # flood through light only
|
|
3350
|
+
a[..., 3][rim] = 0
|
|
3351
|
+
return Image.fromarray(a.astype(np.uint8), "RGBA")
|
|
3352
|
+
```
|
|
3353
|
+
|
|
3354
|
+
**Test LIGHTNESS, not per-channel whiteness — this is the gotcha that makes the whole thing non-obvious.** The rim is **not white**. It is white **contaminated with the chroma plate**, because the plate fringes into it during keying. Measured on a magenta-plate pack, the outermost solid ring averaged **R≈250, G≈205, B≈250** — the green channel is nowhere near white. A per-channel test like `(rgb > 224).all(axis=2)` therefore misses the rim on every sticker whose edge is even slightly anti-aliased: in testing it stripped **1 of 4** stickers, and because the one it did strip looked *different* from its three siblings, the result read as broken art rather than as a bad threshold. `rgb.mean(axis=2) > ~188` catches all four.
|
|
3355
|
+
|
|
3356
|
+
**Then clean up what stripping leaves behind.** Removing the rim produces two artefacts, and both read to a viewer as "the sticker didn't mask properly":
|
|
3357
|
+
|
|
3358
|
+
1. **A dotted halo** — surviving specks along the old rim edge. Measured **213** and **164** stray connected components of 3–20px each on two different stickers of the same pack.
|
|
3359
|
+
2. **A bright blob** — a large uniform light region the art *enclosed* is no longer visually held in by the rim. Measured at 7,983px (a ring's centre) and 9,582px (a stamp's paper plate).
|
|
3360
|
+
|
|
3361
|
+
```python
|
|
3362
|
+
def clean(img, speck=0.008, blob=0.015, light=200):
|
|
3363
|
+
# 1. drop connected components smaller than `speck` of the largest
|
|
3364
|
+
# 2. for each enclosed light region larger than `blob` of the sticker area:
|
|
3365
|
+
# holes = binary_fill_holes(m) & ~m
|
|
3366
|
+
# if holes.sum() < m.sum() * 0.02: # solid fill, no detail -> it is background
|
|
3367
|
+
# set alpha 0 there
|
|
3368
|
+
```
|
|
3369
|
+
|
|
3370
|
+
**The discriminator is worth remembering on its own: an enclosed light region is only background if it has NO internal detail, and its holes are the giveaway.** An eyeball's sclera is riddled with drawn veins and a pupil → many holes → keep it. A ring's centre is a flat fill → no holes → punch it transparent. Area alone gets this wrong in both directions.
|
|
3371
|
+
|
|
3372
|
+
**When you recolour a pack to a palette** (a duotone or luminance ramp onto a brand accent), **cap the top of the ramp** — e.g. `accent + 0.66·(white − accent)` — so interior whites land as a light tint instead of glaring pure white against a flat two-colour design. Uncapped, every kept interior white becomes the brightest pixel in the frame, which undoes the point of the ramp.
|
|
3373
|
+
|
|
3374
|
+
> **Known gap:** none of this is in the CLI. `vidfarm sticker-pack` / `vidfarm cutout` have no `--strip-rim` (or equivalent) flag today, so on a dark stage you run the two passes above yourself as a local post-step. A `--strip-rim` flag on both commands, defaulting off, is the right long-term home for it.
|
|
3375
|
+
|
|
2989
3376
|
### The guided sequence (prompt harness)
|
|
2990
3377
|
|
|
2991
3378
|
**Step 0 — Decide the cast of stickers.** With the director, list every element the explainer needs as its own cutout: the hero subject, each labelled prop, each icon/arrow/emoji, any mascot, any full-frame backdrop. Each becomes one transparent PNG. Stickers are reusable — generate once, reuse across scenes. **If the cast is more than two or three items, make it a PACK** (one sheet, split locally — see "A sticker pack" above) rather than N separate `cutout` calls.
|
|
@@ -3020,7 +3407,7 @@ GIF alpha is **1-bit** — a pixel is fully opaque or fully gone, so antialiased
|
|
|
3020
3407
|
|
|
3021
3408
|
**Both are image-only.** A moving subject has no single bounding box — matte a video clip with `vidfarm remove-background <video>` or key a flat backdrop with `vidfarm remove-greenscreen <video>` (→ transparent WebM/mov).
|
|
3022
3409
|
|
|
3023
|
-
**Step 2 — Show the director each cutout, get corrections.** Cutouts are cheap to regenerate. Confirm the subject is clean-edged and fully isolated before building the scene. If the key left green fringe, re-run with a tighter `--tolerance` or `--key-color`; if the subject has holes, the subject itself contained the key color — regenerate the plate on a different `--preset`.
|
|
3410
|
+
**Step 2 — Show the director each cutout, get corrections.** Cutouts are cheap to regenerate. Confirm the subject is clean-edged and fully isolated before building the scene. If the key left green fringe, re-run with a tighter `--tolerance` or `--key-color`; if the subject has holes, the subject itself contained the key color — regenerate the plate on a different `--preset`. **If the stage isn't white, strip the white die-cut rim here**, before anything is staged — see "Stickers on a DARK or photographic stage" above.
|
|
3024
3411
|
|
|
3025
3412
|
**Step 3 — Stage them on the composition.** Fork/seed a working composition (`vidfarm pull` or `vidfarm serve`), **set the stage to a white light-mode background first** (house style), then drop each cutout as an **image layer**, sized and positioned deliberately:
|
|
3026
3413
|
```
|