@officexapp/vidfarm-devcli 0.21.33 → 0.21.35
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/editor-capabilities/SKILL.md +13 -2
- package/.agents/skills/vidfarm/SKILL.md +73 -29
- package/.agents/skills/vidfarm/harnesses/README.md +112 -0
- package/.agents/skills/vidfarm/{regimes/explainer.QA_REGIME.md → harnesses/explainer.HARNESS.md} +22 -2
- package/.agents/skills/vidfarm/{regimes/hooks.QA_REGIME.md → harnesses/hooks.HARNESS.md} +3 -3
- package/.agents/skills/vidfarm/{regimes/product-demo.QA_REGIME.md → harnesses/product-demo.HARNESS.md} +19 -1
- package/.agents/skills/vidfarm/{regimes/short-form.QA_REGIME.md → harnesses/short-form.HARNESS.md} +67 -7
- package/.agents/skills/vidfarm/{regimes/ugc-testimonial.QA_REGIME.md → harnesses/ugc-testimonial.HARNESS.md} +10 -3
- package/.agents/skills/vidfarm/recipes/{bulk-scripting-with-a-regime.md → bulk-scripting-with-a-harness.md} +35 -12
- package/.agents/skills/vidfarm/recipes/cutout-graphics-for-explainers.md +46 -2
- package/.agents/skills/vidfarm/recipes/local-edit-render-approve.md +2 -1
- package/.agents/skills/vidfarm/references/automation-and-local-dev.md +84 -22
- package/.agents/skills/vidfarm/references/editor-workflows.md +19 -4
- package/.agents/skills/vidfarm/references/hooks-and-virality.md +62 -5
- package/.agents/skills/vidfarm/references/reviewing-renders.md +140 -0
- package/.agents/skills/vidfarm-media/SKILL.md +2 -2
- package/.agents/skills/vidfarm-media/references/tts.md +26 -4
- package/SKILL.director.md +462 -75
- package/SKILL.md +33 -14
- package/dist/src/cli.js +799 -91
- package/dist/src/devcli/{qa-regime.js → harness.js} +132 -55
- package/dist/src/devcli/qa-check.js +209 -4
- package/dist/src/devcli/skill-docs.js +136 -0
- package/dist/src/devcli/stills.js +65 -1
- package/package.json +4 -3
- package/.agents/skills/vidfarm/regimes/README.md +0 -77
|
@@ -518,16 +518,20 @@ Compositions are authored in HTML, so the single most common way an AI-edited vi
|
|
|
518
518
|
- **Emoji inline in text** (sparingly), **sticker/cut-out overlays** on transparent PNG (`create-overlay`), mock social UI when the format calls for it (iMessage bubbles, a TikTok comment card, a fake DM, a countdown/progress bar) — these are native artifacts of the platform, not web furniture.
|
|
519
519
|
- **Full-bleed footage** with text sitting directly on it.
|
|
520
520
|
|
|
521
|
-
**On devcli there's a checker: `vidfarm qa <dir|composition.html>`.** Free, instant, local-only — a blocklist pass for everything above plus the font regime and safe zone, with a concrete fix per finding. **Run it on every video you produce.** It is feedback, not a gate (exit 0 even on findings, never runs automatically, `--strict` only if you want a CI failure) and a blocklist, not an allowlist (stylized/hand-made compositions pass untouched — it will not homogenize your videos). No cloud/REST twin: the web copilot enforces this standard by hand. Details in `references/automation-and-local-dev.md` ("`vidfarm qa`").
|
|
521
|
+
**On devcli there's a checker: `vidfarm qa <dir|composition.html>`.** Free, instant, local-only — a blocklist pass for everything above plus the font regime and safe zone, with a concrete fix per finding. **Run it on every video you produce.** It is feedback, not a gate (exit 0 even on findings, never runs automatically, `--strict` only if you want a CI failure) and a blocklist, not an allowlist (stylized/hand-made compositions pass untouched — it will not homogenize your videos). Every run — including a clean one — ends with a **`▶ NOW WATCH THE VIDEO`** block, because the check never rendered or saw the video and a green tick is not a review; do those steps before you tell anyone the video is done. No cloud/REST twin: the web copilot enforces this standard by hand. Details in `references/automation-and-local-dev.md` ("`vidfarm qa`").
|
|
522
522
|
|
|
523
523
|
### TikTok-native caption standard (position + font + background) — always adhere
|
|
524
524
|
|
|
525
525
|
> Captions are also the *delivery system* for three of the four charges: the hook is read before any audio, the loop has to stay on screen, and the payoff number needs its own card. What the words should SAY is in `references/hooks-and-virality.md`; this section is how they must LOOK.
|
|
526
526
|
|
|
527
|
-
Short-form is watched on a phone, and the phone's UI eats the frame's edges. **Never pin on-screen text to the extreme top or bottom** — the top ~8% sits under the status bar / "Following · For You" tabs and the bottom ~15% under the username, caption text, music marquee, and action rail. Text there is literally clipped and reads as amateur.
|
|
527
|
+
Short-form is watched on a phone, and the phone's UI eats the frame's edges. **Never pin on-screen text to the extreme top or bottom** — the top ~8% sits under the status bar / "Following · For You" tabs and the bottom ~15% under the username, caption text, music marquee, and action rail. Text there is literally clipped and reads as amateur. Four rules, applied to **every** caption/title/overlay you place or inherit:
|
|
528
528
|
|
|
529
|
-
- **Position →
|
|
530
|
-
-
|
|
529
|
+
- **Position → the safe zone first, then the EMPTIEST part of the frame.** Two constraints, in that order.
|
|
530
|
+
- *Hard constraint:* the text box's vertical extent stays inside **~8%–85%** of canvas height (9:16), and wide captions stay clear of the **right ~12%** action rail (a centered box at `x:10 width:80` is safe).
|
|
531
|
+
- *Judgement call, inside that band:* **put the words where the picture isn't.** `y≈70%` is the `captions generate` default because most footage puts its subject mid-frame — it is a default, not a law. Before you place text, **look at an actual frame** (`vidfarm stills ./work --at <t>`, free) and find the region with the least going on: open sky above a dashboard, a blank wall behind a talking head, an out-of-focus background, an empty tabletop. If nothing else in the video is competing for attention there — no subject, no motion, no product, no second text layer — that is where the caption belongs, even if it means **high-centre at y≈10–25%** instead of a lower third. A caption dropped over the busiest third of the frame (hands on a steering wheel, a face, the product) fights the shot and forces you to armour it with a plate; the same words parked in the sky are legible with no plate at all.
|
|
532
|
+
- *When you're only rescuing an inherited caption* off a dead-zone edge, preserve its top-vs-bottom anchoring and just pull it inside the band — don't recentre a template you haven't re-read. When **you** are the one placing the text, place it deliberately.
|
|
533
|
+
- **Size → scaled to the line, not maxed out.** Sizes are PIXELS of a 1080-wide frame: **~36–64px** reads well; below ~28px is unreadable on a phone and **0 is invisible**. Above ~64px is a *hook-word* size — one to three words, on purpose. The failure this catches: a full sentence set at display size runs edge-to-edge, wraps to three lines, and eats a third of the frame, so it has to be armoured with a full-width plate and there is nowhere left to put it. **If a line reaches the frame edges, the fix is a smaller size (or fewer words per cue), not a wider box.** Keep captions to ~2 lines / ~5 words per line; `line_height` 0.95–1.15 for stacked display lines.
|
|
534
|
+
- **Font → the composition regime.** Use the bundled display fonts only — **Montserrat** (bold default, weight **700–900**), **TikTok Sans**, Abel, Source Code Pro, Yesteryear. Don't request a font the composition doesn't import (it silently falls back to a web-default sans, which is exactly the slop look).
|
|
531
535
|
- **Background → one of exactly four valid treatments.** Any text you place uses one of these and nothing else:
|
|
532
536
|
|
|
533
537
|
| # | Treatment | How to set it | When |
|
|
@@ -537,8 +541,19 @@ Short-form is watched on a phone, and the phone's UI eats the frame's edges. **N
|
|
|
537
541
|
| 3 | **Highlight pill behind the ACTIVE word only** | `set_captions caption_style:"spotlight"` / `"karaoke"` (+ `caption_highlight_color`) | Hormozi/CapCut word-by-word. **The only legitimate "pill" in a video** — it tracks the spoken word, so it isn't a badge |
|
|
538
542
|
| 4 | **Solid band that tightly hugs the text lines** (CapCut "text box") | `background_style:"highlight-solid"` (or `"highlight-translucent"`) + a `background` color | Guaranteed legibility over noisy footage |
|
|
539
543
|
|
|
544
|
+
**Move the text before you armour it.** The plate is the *last* resort, not the default: if the band you picked is busy, first try moving the caption into the calm/empty region the position rule points at — a caption over open sky needs no background at all, and "no plate" is the cleaner, more native look every time you can afford it. Only when the whole frame is busy (or the text has to sit on the subject for meaning) do you reach for treatment 4.
|
|
545
|
+
|
|
546
|
+
**Then pick between them by MEASURING the background behind the caption band, not by habit.** Dark-and-calm behind the band (luma < ~70, variation < ~42) → light type, **no plate** (treatment 1/2 — a plate there is a bright slab the design never asked for); bright-and-calm (luma > ~160) → dark type, no plate; busy / mid-tone / moving colour → treatment 4, because nothing else stays readable. The active-word colour has to follow the same call — a deep red that reads on a white plate is unreadable on near-black. **One treatment for the whole video**; styling that flips every few seconds reads as a bug. Procedure, thresholds and how to measure the *composited* value (not the source file): `harnesses/short-form.HARNESS.md` → "Caption styling is MEASURED off the background".
|
|
547
|
+
|
|
540
548
|
Treatment 4 is a **band, not a card**: it hugs the glyphs with minimal padding, corner radius ≤ ~8px, **no border, no drop shadow, no gradient, no blur**, and it wraps *one* text run — never a heading + subheading + URL stacked inside one rounded box. The moment it grows padding, a stroke, or a second element inside it, it has become a web card. Fix it. And the moment its radius goes fully round, it has become a **badge** — treatment 3 is the *only* capsule allowed, and only because it tracks the spoken word. A static "10 hrs / week" in a rounded pill is web furniture; the same words in treatment 1 or 2, bigger and heavier, are a beat.
|
|
541
549
|
|
|
550
|
+
**Long narration → kinetic cues, never a wall of text.** A caption layer is a *page*, not a transcript. The moment a single static text run carries more than ~10–12 words — or sits on screen longer than ~4 seconds while the voice keeps going — it stops being a caption and becomes a paragraph the viewer has to read while also watching the video. Nobody does both; they scroll. Page it instead:
|
|
551
|
+
|
|
552
|
+
- **Transcribe and let the tool page it:** `vidfarm captions generate ./work --style word-pop` (or `spotlight` / `karaoke`) splits narration into ~3–5-word cues with real word-level timings, so one short phrase is on screen at a time and the active word tracks the voice. Web copilot twin: the `/primitives/audio/captions` job → `set_captions` (see "Animated captions" below). `--max-words-per-cue` tightens it further.
|
|
553
|
+
- **The cue count is the readability dial.** Short cues that change with the speech read as *momentum*; one long block reads as homework. Kinetic word-by-word also lets the type be **smaller** (the eye is led to the moving word instead of having to scan a wall), which frees up frame space and usually removes the need for a plate.
|
|
554
|
+
- **Static text is for the beats that deserve their own moment** — a hook line, a payoff number, a title card. Those are short by nature. Anything spoken should be a caption run, not a static block.
|
|
555
|
+
- **Exception: verbatim UGC/testimonial captions** stay one plain line at a time (see `harnesses/ugc-testimonial.HARNESS.md`) — the kinetic VFX look is the "made by a marketing team" tell there. Paging still applies; the animation preset doesn't.
|
|
556
|
+
|
|
542
557
|
**A common trap: decomposed templates mirror the source's caption placement**, so a forked meme can arrive with its caption pinned at `top:0` in a non-regime font — and a re-theme prompt ("make this for my tutoring service") is exactly where an agent starts inventing landing-page CTAs and benefit chips because the *subject* is a SaaS product. **Fix to the standard, don't inherit it, and don't import the website's design language into the video.** When placing text yourself (`set_captions`, `set_layer_text`, `set_layer_style`, `add_layer`, devcli `place`/`captions`), set `y` / `font_family` / `font_weight` / `background_style` to the standard from the start.
|
|
543
558
|
|
|
544
559
|
> Local devcli renders enforce part of this automatically: `renderCompositionLocally` runs `normalizeTikTokCaptionLayout` (src/devcli/composition-edit.ts) on every production, clamping caption/text layers into the 8%–85% safe zone and coercing off-regime primary fonts to Montserrat. It only fixes position and font family — it will happily render your Bootstrap card. Get it right in the composition so the editor preview, the local render, and any cloud render match.
|
|
@@ -4,7 +4,7 @@ Most agent-made videos don't fail on polish. They fail on **structure**: no hook
|
|
|
4
4
|
|
|
5
5
|
This is the harness that fixes it. It is not a style — it's the load-bearing anatomy of anything that travels on TikTok/Reels/Shorts, distilled from grading hundreds of hooks against real funnels. **Run it on one-off videos and on batches alike.** It costs no credits, adds no render time, and it is the single largest quality delta available in this product.
|
|
6
6
|
|
|
7
|
-
The checkable form of this document is the bundled `hooks`
|
|
7
|
+
The checkable form of this document is the bundled `hooks` harness (`vidfarm harness show hooks`); this file is the craft behind it.
|
|
8
8
|
|
|
9
9
|
---
|
|
10
10
|
|
|
@@ -17,10 +17,67 @@ The reason agent videos come out structureless is that the timeline is the fun p
|
|
|
17
17
|
3. **Name the payoff.** What is on screen at that moment, and why does it satisfy the promise?
|
|
18
18
|
4. **Write the bait.** The final-beat ask, in the video and in the post caption.
|
|
19
19
|
5. **Only now build the timeline** — and place the hook text at `start:0` so it's on screen at frame 0 (which is also the thumbnail).
|
|
20
|
-
6. **Verify the frame and the structure:** `vidfarm stills ./work --at 0` (look at the actual poster) and `vidfarm qa ./work --
|
|
20
|
+
6. **Verify the frame and the structure:** `vidfarm stills ./work --at 0` (look at the actual poster) and `vidfarm qa ./work --harness hooks` (machine checks + the judgment checklist).
|
|
21
21
|
|
|
22
22
|
Steps 1–4 are cheap, reversible, and where the entire outcome is decided. Steps 5–6 are where agents want to start.
|
|
23
23
|
|
|
24
|
+
7. **Cut it.** Nothing ships at its first length — see the next section. Assume your first assembly is 30–50% too long and go find the seconds.
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## Density — every second must earn its place, and most don't
|
|
29
|
+
|
|
30
|
+
**A viewer's thumb is a hard time limit that resets every second.** They are not "watching your video"; they are re-deciding to stay, ~24 times a second, against an infinite feed of alternatives. A second that carries nothing is not neutral — it is a free exit. This is why the same script cut to 22s outperforms its own 41s version with better footage: fewer exit ramps.
|
|
31
|
+
|
|
32
|
+
Agents are structurally bad at this. A model writes a video the way it writes prose — with connective tissue, restatement, a wind-up before the point, a tidy conclusion — and every one of those habits is a hole in the retention curve. **You must cut against your own instinct, and you must cut more than feels right.**
|
|
33
|
+
|
|
34
|
+
### The deletion test — the only test that matters
|
|
35
|
+
|
|
36
|
+
For every beat, ask: **delete it. Does the video still make sense, and does the payoff still land?** If yes, it stays deleted. Not "trimmed" — deleted. Run this on every scene, every sentence, and every caption before you render, and be honest: the beat you're defending because it took work to make is exactly the one this test exists to kill.
|
|
37
|
+
|
|
38
|
+
Second filter for whatever survives: **which of the four charges does this beat serve — hook, loop, payoff, or bait?** A beat that serves none is fluff wearing a costume. "It gives context" is not a charge. "It looks nice" is not a charge.
|
|
39
|
+
|
|
40
|
+
### Cut on sight — the standard fluff, in the order it usually appears
|
|
41
|
+
|
|
42
|
+
- **Any intro.** Logo sting, title card, brand animation, "welcome back", a beat of black. The video starts at the claim. Frame 0 is the hook (and the thumbnail).
|
|
43
|
+
- **The wind-up before the point.** "So I wanted to talk about…", "Here's the thing…", "Let me explain…", "In this video I'm going to show you…". Delete the sentence; the next one was the real opening.
|
|
44
|
+
- **Context before the claim.** Context is beat 2 at the earliest, and usually one clause, not a scene.
|
|
45
|
+
- **Restatement.** Saying the same thing a second way "so it's clear." It was clear. If it wasn't, fix the first version.
|
|
46
|
+
- **Dead air in the narration.** Breaths, "um", and any inter-sentence gap over ~0.35s. This alone routinely takes 15–20% off a TTS or talking-head cut.
|
|
47
|
+
- **Real-time process.** Nobody watches the upload bar. Speed-ramp it, jump-cut it, or show the before and the after and skip the middle.
|
|
48
|
+
- **Establishing shots.** They know what an office/kitchen/laptop looks like. Open inside the action.
|
|
49
|
+
- **Reading what's already on screen.** Voice and text should split the work, not duplicate it (the same rule as captions-vs-display-text).
|
|
50
|
+
- **The tail.** "Thanks for watching", a logo card, an end screen, or footage that keeps rolling after the last word. The bait is the last beat; then it **ends**, hard, on the frame that loops best.
|
|
51
|
+
- **Filler motion.** A slow pan or Ken Burns that exists because the clip was too short for its slot. Shorten the slot instead.
|
|
52
|
+
|
|
53
|
+
### Density is not speed, and this is where over-correcting ruins videos
|
|
54
|
+
|
|
55
|
+
Cutting fluff means **removing beats that carry nothing**, never rushing the beats that carry everything. Three things are load-bearing and must keep their seconds:
|
|
56
|
+
|
|
57
|
+
- **The comedic beat.** The held pause before a punchline IS the joke. Cutting it saves 0.6s and costs the video.
|
|
58
|
+
- **The payoff.** It plays, full frame, uninterrupted — ≥5s if that's what it takes. Summarising the payoff to save time is the most expensive cut available.
|
|
59
|
+
- **A caption's readability.** A cue nobody can finish reading is worse than no cue. If tightening the edit makes text unreadable, cut *words*, not the time they're on screen.
|
|
60
|
+
|
|
61
|
+
The target is **information per second**, not seconds. A dense 45s video beats a hollow 20s one; both lose to the same 45s cut to 30s with nothing lost.
|
|
62
|
+
|
|
63
|
+
### Length is an output, not a plan
|
|
64
|
+
|
|
65
|
+
Don't decide "make it 60 seconds" and then fill 60 seconds — filling is where every one of the fluff patterns above comes from. Build the four charges, cut to the deletion test, and **the length is whatever's left.** If the payoff lands at 0:25, the video ends around 0:27. A brief that dictates a duration is a brief that ordered fluff.
|
|
66
|
+
|
|
67
|
+
### How to actually cut it, in Vidfarm
|
|
68
|
+
|
|
69
|
+
| Move | devcli | Web copilot |
|
|
70
|
+
|---|---|---|
|
|
71
|
+
| Find the dead air | `vidfarm qa ./work` (flags gaps ≥2.5s with nothing on screen, and a tail that keeps rolling after the last word) + read the word timings from `vidfarm captions generate` / `stt` | read `video_context`'s timestamped segments and look for the gaps between them |
|
|
72
|
+
| Trim one clip's edge | `vidfarm trim ./work --layer <k> --edge start --to-time <sec>` | `editor_action trim_layer` |
|
|
73
|
+
| Close the hole you just made | `vidfarm ripple ./work --at <sec> --delta -<sec>` (negative = close time, shifts everything downstream) | `editor_action ripple_edit` |
|
|
74
|
+
| Drop a whole beat | `vidfarm retime`/`remove` the layers, then `ripple` the gap closed | `remove_layer` + `ripple_edit` |
|
|
75
|
+
| Re-time captions after cutting | re-run `vidfarm captions generate` against the new audio — never hand-shift cues | the `/primitives/audio/captions` job → `set_captions` |
|
|
76
|
+
|
|
77
|
+
**Always ripple the gap closed.** A cut that leaves a hole is not a cut; it converts fluff into dead air, which is worse — the viewer now stares at a frozen frame instead of a boring one.
|
|
78
|
+
|
|
79
|
+
**Cheap habit that pays every time:** shave the first ~0.5–1s off every sourced clip and the last ~0.5s. People start recording before the action and stop after it, so a montage of raws is carrying a second of nothing per clip by default.
|
|
80
|
+
|
|
24
81
|
---
|
|
25
82
|
|
|
26
83
|
## Charge 1 — THE HOOK (first 3 seconds)
|
|
@@ -229,9 +286,9 @@ A video with replies gets shown again; a video with none dies at its first audie
|
|
|
229
286
|
| Check what the source template's hook actually was | `editor_context` → `viral_dna.hook` / `retention` / `payoff` / `emotional_punch` | `.harness/context.json`, `video-context.json` |
|
|
230
287
|
| Place the hook at frame 0 | `add_layer` / `set_captions` with `start:0` | `vidfarm set-text ./work --layer hook --text "…"` |
|
|
231
288
|
| Look at the poster frame | ask the user to scrub to 0 | `vidfarm stills ./work --at 0` |
|
|
232
|
-
| Grade the structure | by hand, against this file | `vidfarm qa ./work --
|
|
233
|
-
| Bulk hook test | hand off to a local agent | `recipes/bulk-scripting-with-a-
|
|
289
|
+
| Grade the structure | by hand, against this file | `vidfarm qa ./work --harness hooks` |
|
|
290
|
+
| Bulk hook test | hand off to a local agent | `recipes/bulk-scripting-with-a-harness.md` |
|
|
234
291
|
|
|
235
292
|
**Re-theming a decomposed template?** `viral_dna` already names the source's hook, retention device, and payoff — that structure is *why the template worked*. Rebuild each charge for the new subject; don't drop the loop because the new topic feels self-explanatory. Flattening a template's loop into a product statement is the single most common way a re-theme kills a format.
|
|
236
293
|
|
|
237
|
-
**The checkable version of everything above:** `vidfarm
|
|
294
|
+
**The checkable version of everything above:** `vidfarm harness show hooks` — the twelve-item pre-flight checklist is the part you answer honestly on every video, and two items carry most of the weight: *situation, not label* (predicts cold-start survival before you write a word) and *unguessable* (the only item a hook can fail while passing every other one, which is why it ships).
|
|
@@ -0,0 +1,140 @@
|
|
|
1
|
+
## Reviewing a render — look at the whole video, and never trust one frame
|
|
2
|
+
|
|
3
|
+
**Assume your own finished video has a defect you can't see, because you built it.** This is not humility, it's the observed base rate: across a 32-video bespoke batch, **every single first-pass video had a real defect that the agent who built it had already reported as "verified, looks good."** Dead space under the content, a placeholder that reads as a failed render, two contradictory numbers 200px apart, a CTA still animating when the video ends. None of these are subtle. All of them survived a confident self-review.
|
|
4
|
+
|
|
5
|
+
The reason is structural, not sloppiness: **an agent builds a video the way it builds code — part by part, each part correct in isolation.** Scene 3 is written while scene 3 is the whole world. So each scene passes on its own and the video fails as a video: the type jumps two sizes between beats, one scene breathes and the next is crammed to the margins, the accent colour drifts, a transition lands like a slap because nothing before it moved that fast. Nobody watches a scene. They watch the sequence.
|
|
6
|
+
|
|
7
|
+
So the review has two jobs, and they need two different passes:
|
|
8
|
+
|
|
9
|
+
1. **The holistic pass** — does this read as ONE video, made by one person, on purpose?
|
|
10
|
+
2. **The defect pass** — is any individual frame broken in one of the six ways frames are usually broken?
|
|
11
|
+
|
|
12
|
+
Do them in that order. The holistic pass is the one agents skip, and it is the one that separates "technically correct" from "good."
|
|
13
|
+
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
## Pass 1 — the holistic pass: watch it as one object
|
|
17
|
+
|
|
18
|
+
**Before you look for defects, look at the video the way a stranger will: all at once, start to finish, with no memory of how it was built.** You cannot do this from the code, from the storyboard, or from the scene you just edited — you have to look at the actual frames, in order, side by side.
|
|
19
|
+
|
|
20
|
+
```bash
|
|
21
|
+
# ONE command — renders the stills AND tiles them into ./work/stills/contact-sheet.png.
|
|
22
|
+
# No MP4 render needed; free, local, in-process.
|
|
23
|
+
vidfarm stills ./work --sheet # default timestamps = midpoint of each scene clip (cap 8)
|
|
24
|
+
vidfarm stills ./work --sheet --at 0,2,4,6,8,10,12,14,16,18,20,22 # or pick them yourself
|
|
25
|
+
# then READ ./work/stills/contact-sheet.png as an image
|
|
26
|
+
# --sheet-out <file> relocates it; --sheet-width <px> for bigger tiles (default 320)
|
|
27
|
+
|
|
28
|
+
# already have the MP4? same idea, straight off the file
|
|
29
|
+
for t in 0 2 4 6 8 10 12 14 16 18 20 22; do
|
|
30
|
+
ffmpeg -y -ss $t -i final.mp4 -frames:v 1 "qa/f$(printf %03d $t).png"; done
|
|
31
|
+
ffmpeg -y -pattern_type glob -i "qa/f*.png" \
|
|
32
|
+
-vf "scale=300:-1,tile=4x3:margin=6:padding=6:color=0x999999" -frames:v 1 qa/sheet.png
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
**The tile sheet is the point.** One image read, twelve frames, and the eye picks up drift instantly that no per-scene check can see. Read it as an image — not the filenames, not the HTML that produced it. (The still filenames are zero-padded seconds, so a glob stays in chronological order.)
|
|
36
|
+
|
|
37
|
+
Then answer these, out loud, in your report:
|
|
38
|
+
|
|
39
|
+
- **Balance.** Is weight distributed across the frame, or is every scene top-anchored with an empty band underneath? Does the composition use the canvas, or does it use the top third of the canvas and leave the rest as dead area? A sheet of twelve frames makes a recurring dead zone obvious; one frame at a time never will.
|
|
40
|
+
- **Fluff, named out loud.** Which beats would you cut? Answer with specific timestamps, not "it's tight". Every tile has to justify its seconds: a frame that repeats the previous one, a scene the video would survive losing, an intro, a tail after the last word, a hold that's just waiting. **Assume 30–50% of the first assembly can go** and name what you'd remove — "nothing to cut" on a first pass is almost always a review that didn't look. Then cut it and `ripple` the hole closed (craft: `references/hooks-and-virality.md` → "Density"; the mechanical half is `vidfarm qa`'s `dead-air` / `dead-tail` / `slow-scene`).
|
|
41
|
+
- **Spacing and breathing room.** Are margins consistent scene to scene? Does one beat have generous air and the next one crowd the safe zone? Uneven padding across scenes is the single loudest "assembled by a machine" tell, and it's invisible while you're inside any one scene.
|
|
42
|
+
- **Typographic continuity.** One type system, or three? Headline sizes should belong to a small set (two, maybe three), not be individually chosen per scene. Same for weight, case, and colour. If scene 2's headline is 64px and scene 5's is 41px for no dramatic reason, that's drift, not design.
|
|
43
|
+
- **Colour and style coherence.** One accent colour, one background treatment, one illustration style. Assets generated or sourced at different moments drift — a flat-vector sticker next to a photographic cutout next to a gradient panel reads as three videos spliced together.
|
|
44
|
+
- **Rhythm and pacing.** Do scene durations form a deliberate pattern (a fast open, a longer explanation, a fast close), or is every scene the same length because a loop wrote them? Same-length beats are hypnotic in the bad way. Conversely, one 9-second hold in a video of 2-second cuts stalls it dead.
|
|
45
|
+
- **Nothing jarring at the joins.** Watch each transition specifically. A cut from a dark scene to a white one is a flash in the face; a scale-up entrance immediately after a scale-up exit reads as a stutter; two consecutive scenes whose subjects sit in the same screen position with different content look like a glitch, not a cut. Where a join is harsh, either match the two frames either side of it (colour, position, energy) or make the harshness deliberate and rhythmic.
|
|
46
|
+
- **One idea per moment.** Across the whole sheet, is there any frame where two things compete for the eye — display text over captions saying the same words, a busy background under type, two headlines superimposed at a handoff? At the video level this shows up as a *density* problem: some beats carry three elements and some carry one.
|
|
47
|
+
- **Does it look like one person made it in one sitting?** The summary question. If the honest answer is "it looks assembled," name specifically which scenes don't belong and fix them toward the majority, don't average everything.
|
|
48
|
+
|
|
49
|
+
**When something is off, fix it globally, not locally.** The instinct after spotting drift is to patch the one scene that stands out. Usually the right fix is to define the rule (two headline sizes, one accent, 8% margins, 2.5s default beat) and apply it across every scene — including the ones that already looked fine. A video is a system; patching one node keeps the system inconsistent.
|
|
50
|
+
|
|
51
|
+
**Build order helps too, if you're still building.** Author the shared system first — type scale, palette, margins, motion vocabulary, default beat length — as one thing that every scene reads from, then fill the scenes. Scenes written first and harmonised later almost never fully converge.
|
|
52
|
+
|
|
53
|
+
---
|
|
54
|
+
|
|
55
|
+
## Pass 2 — the defect pass: what actually goes wrong, in frequency order
|
|
56
|
+
|
|
57
|
+
From the same 32-video batch, ranked by how often it happened. Look for these specifically; they are what your own review misses.
|
|
58
|
+
|
|
59
|
+
1. **Large flat dead regions.** Content top-anchored with an empty band below it. Agents do this constantly and never notice, because during authoring the element is the subject and the emptiness is just "background."
|
|
60
|
+
2. **A placeholder empty state that reads as a missing asset.** A big empty dashed rectangle held for two seconds looks exactly like the render failed. If a scene's job is "an empty inbox," it still has to look designed, not broken.
|
|
61
|
+
3. **Two contradictory numbers in one frame.** Especially on anything data-shaped — a stat in the headline and a different one in the visual beneath it.
|
|
62
|
+
4. **A CTA or end card still building when the video ends.** The final state must be **settled at least 2 seconds before the last frame**, or the loop-around cuts it off and the ask never lands.
|
|
63
|
+
5. **Two headlines superimposed at a scene handoff.** The outgoing scene's text hasn't left when the incoming one arrives. Fix at the timing level: exit at `nextIn − 0.18`, duration `0.24`, ease `power2.out` — a slow-leaving `power2.in` is what causes the overlap in the first place.
|
|
64
|
+
6. **Type colliding with a busy background layer** exactly at the moment it's spoken. Particles, pins, footage detail — legible in the still you checked, unreadable at the second the word lands.
|
|
65
|
+
|
|
66
|
+
**If a frame looks empty, check whether it RESTS there.** Sample at 0.2–0.25s intervals through that transition. A transient near-empty wipe frame is fine and normal; anything holding empty for **>0.5s** is a hole in the video.
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
vidfarm stills ./work --at 6.0,6.2,6.4,6.6,6.8,7.0 # is it a wipe, or a hole?
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
---
|
|
73
|
+
|
|
74
|
+
## Never verify a video by one frame
|
|
75
|
+
|
|
76
|
+
**This is the failure mode that survives every check you'd think to run.** Whole classes of render bug produce a video where *every frame is identical* — the timeline never ran — while duration, frame count, file size and audio hash all come out exactly right. Frame 0 looks perfect, so a single-frame check passes and you ship a frozen video.
|
|
77
|
+
|
|
78
|
+
Two real causes, both silent:
|
|
79
|
+
|
|
80
|
+
- **A watermark/overlay pass without `-loop 1` on a single-frame PNG input.** The frame-sync collapses the whole video onto one frame. Five videos shipped this way before it was caught.
|
|
81
|
+
- **Assets outside the composition root.** Only `<style>`/`<script>` *inside* the `data-composition-id` root execute, and sibling relative files may not resolve — fonts, images, even the animation library itself. The timeline never starts; frame 0 still renders fine because frame 0 is the static DOM.
|
|
82
|
+
|
|
83
|
+
**The rule that catches both: always compare two frames from different scenes.** They must differ a lot. And when you've applied any pass over an existing video (watermark, overlay, dedupe, re-encode), also compare each output frame against **its own** input frame at the same timestamp — that difference should be tiny. Two checks, opposite directions:
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
# consecutive/distant frames must DIFFER (motion preserved)
|
|
87
|
+
vidfarm stills ./work --at 0,4,9,14
|
|
88
|
+
# after an overlay pass: same timestamp, before vs after, must be NEARLY IDENTICAL
|
|
89
|
+
ffmpeg -y -ss 7 -i clean.mp4 -frames:v 1 a.png
|
|
90
|
+
ffmpeg -y -ss 7 -i final.mp4 -frames:v 1 b.png
|
|
91
|
+
ffmpeg -i a.png -i b.png -filter_complex "psnr" -f null - # very high PSNR = only the mark changed
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
This sits directly beside the frame-0 rule and is its necessary counterweight: **frame 0 is the thumbnail, so judge it alone — but never judge the VIDEO by it.** The two rules are checking different things, and an agent that only internalises the first one has a perfect blind spot for a frozen render.
|
|
95
|
+
|
|
96
|
+
---
|
|
97
|
+
|
|
98
|
+
## Verify audio by measurement, not by ear
|
|
99
|
+
|
|
100
|
+
**You cannot hear the render.** Do not report "the mix sounds good" — you have no way to know it, and it is the claim that most often turns out false. Measure instead.
|
|
101
|
+
|
|
102
|
+
- **Speech-over-bed separation: target 12–15 dB.** Measure the RMS of the mix across the spans where words actually occur, minus the RMS of a bed-only stretch. Word spans come free from the transcription you're already running for captions (`vidfarm stt <file> --engine whisper` → word timings).
|
|
103
|
+
- **Peak below 0 dBFS.** A mix that clips reads as amateur instantly on a phone speaker.
|
|
104
|
+
- **Beware "separation" numbers computed over the music-only tail** — they measure the wrong thing (bed alone vs. bed alone) and over-report by a wide margin. Don't retune a mix based on one.
|
|
105
|
+
|
|
106
|
+
```bash
|
|
107
|
+
ffmpeg -i final.mp4 -af "volumedetect" -f null - # peak + mean over the whole file
|
|
108
|
+
ffmpeg -i final.mp4 -ss 3.1 -t 1.4 -af "volumedetect" -f null - # a span where a word is spoken
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
**Narration timing gotchas that produce a correct-looking, wrong-sounding video:**
|
|
112
|
+
|
|
113
|
+
- **Never `adelay` the voiceover.** Whisper's word timings — and therefore every caption you generated from them — are relative to the raw `vo.wav`. Delaying the VO desyncs every caption in the video while the file still plays fine. Use `apad` + `atrim` to place it instead.
|
|
114
|
+
- **Scene handoffs can leave ~0.3s of silence.** Extend each clip's audio ~0.35s into the next.
|
|
115
|
+
- **Whisper's default model is English-only and will hallucinate fluent English over another language.** Non-English narration needs `--model large-v3 --language <code>`. The output looks like a clean transcript, so this one ships silently.
|
|
116
|
+
- **`vidfarm tts` reads stdin** — always redirect `</dev/null` when calling it inside a shell loop, or the loop eats its own input.
|
|
117
|
+
|
|
118
|
+
---
|
|
119
|
+
|
|
120
|
+
## The revision pass — how to fix what review found
|
|
121
|
+
|
|
122
|
+
**Spawn a fresh pass rather than re-litigating with the context that produced the defect.** If you're handing fixes to a subagent (or picking the work back up yourself later), the brief that works:
|
|
123
|
+
|
|
124
|
+
- **State it as N targeted fixes and nothing else.** "The video is good — you are making three specific fixes." Open-ended "improve it" turns a working video into a different, differently-broken video.
|
|
125
|
+
- **Edit the generator, not the generated output.** If a script produced `composition.html`, fix the script. Check first that a generator exists — some compositions are hand-authored.
|
|
126
|
+
- **Back up before overwriting** — keep `<slug>-v1.mp4`. Re-renders are cheap locally; a lost good version isn't.
|
|
127
|
+
- **Keep audio bit-identical unless audio is the defect.** Reuse the existing `vo.wav` / word timings rather than re-recording; a re-record retimes every caption for no reason.
|
|
128
|
+
- **Give the PROBLEM, not just your proposed solution.** Repeatedly, the agent handed a described defect found a better fix than the one specified — using an app's own collapsed UI state instead of a redaction box, a type safe-zone solver instead of a scrim, making a document's *arrival* the spectacle instead of cutting the document. Say what's wrong and at what timestamp; let the fix be found.
|
|
129
|
+
- **Then sweep for the same class of problem** across the rest of the video, and report what else turned up. Defects of a given kind are rarely solitary — they come from a habit.
|
|
130
|
+
|
|
131
|
+
---
|
|
132
|
+
|
|
133
|
+
## Report both halves honestly
|
|
134
|
+
|
|
135
|
+
When you hand back a render, say what you **measured** and what you **judged**, separately:
|
|
136
|
+
|
|
137
|
+
- Machine-settled: `vidfarm qa ./work` findings, `vidfarm lint`, durations, peak dBFS, frame-difference checks.
|
|
138
|
+
- Human-judgment: the holistic pass above — balance, spacing, type continuity, colour coherence, pacing, joins — plus the harness's `- [ ]` review items.
|
|
139
|
+
|
|
140
|
+
**Never report a clean pass on the half you didn't actually look at.** A confident "verified, looks good" over an unreviewed video is worse than no review, because it spends the director's trust on nothing — and per the base rate at the top of this file, it is usually wrong.
|
|
@@ -73,8 +73,8 @@ openrouter key; music always needs ElevenLabs (own key or platform).
|
|
|
73
73
|
| Need | Command / route | Notes |
|
|
74
74
|
| --- | --- | --- |
|
|
75
75
|
| **Music** (bed, beat, jingle, song, score) | `vidfarm music "upbeat lo-fi beat" --length 30` · `POST /api/v1/primitives/music/generate` | ElevenLabs. `use_wallet_credits` default true (platform key + wallet); `--own-key` = your ElevenLabs key. `music_length_ms` ≤ 300000 (5 min). Place as its own `<audio>` layer ~0.1–0.2 under narration. |
|
|
76
|
-
| **Narration** (default) | `vidfarm tts "…" --cloud` · `POST /api/v1/primitives/audio/speech` |
|
|
77
|
-
| **
|
|
76
|
+
| **Narration** (default) | `vidfarm tts "…" --cloud` · `POST /api/v1/primitives/audio/speech` | Premium ElevenLabs **through vidfarm's own ElevenLabs connection** — the user needs NO ElevenLabs account or API key; it's billed as vidfarm wallet credits. Pick a voice with `--voice <voice_id>` (browse below). `--own-key` for your own ElevenLabs/BYOK key. Local-first `vidfarm tts` (no `--cloud`) still runs on your env openai/gemini key. |
|
|
77
|
+
| **Browse + SAMPLE voices** | `vidfarm voices [--sample] [--search "…"] [--free\|--all]` · `GET /api/v1/primitives/audio/voices` | The premium catalog reached over vidfarm's connection — **no ElevenLabs signup needed, wallet credits only** — plus `--free` for the $0 local Kokoro roster. `--sample` writes listenable clips to `./voice-samples`, **free on both tiers** (preview CDN clips + local renders), so it's safe in `minimize`. **Interactive mode: play samples and let the USER choose. Autonomous: default a sensible voice and still say they can pick from many.** |
|
|
78
78
|
| Narration, zero keys / cost-saving | `vidfarm tts "…" --out narration.wav` (free local by default in `minimize`/`hybrid`) · or `npx hyperframes tts "…" -v af_heart --json` | Kokoro-82M, local, WAV + duration in JSON. Fixed voice presets, no `--style`. |
|
|
79
79
|
| **Transcript + SRT** | `vidfarm stt <file\|url> --cloud` · `POST /api/v1/primitives/audio/transcribe` | Default = ElevenLabs Scribe (native diarization + real word timestamps), wallet-billed. `--own-key`/BYOK: gemini labels speakers, openai/whisper-1 gives real word timings. |
|
|
80
80
|
| Reword existing narration in the (approximate) original voice | `POST /api/v1/primitives/audio/regenerate-speech` | Listens, profiles the speaker (needs a Gemini key), rewords, regenerates with the closest preset voice + matched style. Approximation, never a clone. Details: `references/tts.md` |
|
|
@@ -32,13 +32,35 @@ vidfarm music "chill lo-fi hip hop beat with jazzy piano" --length 30 --json
|
|
|
32
32
|
## Voices — `vidfarm voices`
|
|
33
33
|
|
|
34
34
|
```bash
|
|
35
|
-
vidfarm voices
|
|
36
|
-
vidfarm voices --
|
|
35
|
+
vidfarm voices # premium ElevenLabs catalog, via VIDFARM'S OWN connection
|
|
36
|
+
vidfarm voices --sample # download 6 preview clips to ./voice-samples — FREE
|
|
37
|
+
vidfarm voices --search "british narrator" --limit 10
|
|
38
|
+
vidfarm voices --free --sample # the $0 local Kokoro voices, rendered locally — also FREE
|
|
39
|
+
vidfarm voices --all # both rosters
|
|
40
|
+
vidfarm voices --own-key # the customer's own ElevenLabs account voices
|
|
37
41
|
```
|
|
38
42
|
|
|
43
|
+
**The premium voices do NOT require an ElevenLabs account.** This is the single most under-told
|
|
44
|
+
thing in the whole audio surface: vidfarm holds its own ElevenLabs connection, so any user —
|
|
45
|
+
free-tier, no API key, no ElevenLabs subscription — can narrate with the full premium catalog and
|
|
46
|
+
simply pay **vidfarm wallet credits** (pennies per narration). Say that out loud when you offer
|
|
47
|
+
voices; don't let the user think "premium voice" means "go sign up for ElevenLabs first".
|
|
48
|
+
`--own-key` (`?use_wallet_credits=false`) is the opt-out for users who already have a key and would
|
|
49
|
+
rather bill their own account.
|
|
50
|
+
|
|
51
|
+
**Sampling is free on both tiers**, so it is safe even in cost mode `minimize`: premium samples are
|
|
52
|
+
ElevenLabs' own static preview clips (a CDN download, not a synthesis call) and free samples render
|
|
53
|
+
on the local engine. `--sample-count N`, `--sample-out <dir>`, `--sample-text "<line>"` tune it.
|
|
54
|
+
Free samples need the local Kokoro deps installed; if they're missing the command says so and the
|
|
55
|
+
premium samples still work.
|
|
56
|
+
|
|
39
57
|
- `GET /api/v1/primitives/audio/voices` (`?use_wallet_credits=false` for the user's own key).
|
|
40
|
-
- **
|
|
41
|
-
the
|
|
58
|
+
- **In interactive mode, the user picks the voice — by ear, not from a list of names.** Sample a
|
|
59
|
+
handful, hand over the files, let them choose, then pass their `voice_id` to `tts --voice`. In
|
|
60
|
+
autonomous mode default a sensible voice and still tell them they can pick from many (surface a
|
|
61
|
+
few names + the returned `voice_library_url`).
|
|
62
|
+
- `vidfarm tts` prints the same reminder on stderr whenever narration is about to run with no
|
|
63
|
+
`--voice` and the mode is interactive (or was never set).
|
|
42
64
|
|
|
43
65
|
## Preflight
|
|
44
66
|
|