@officexapp/vidfarm-devcli 0.21.33 → 0.21.35
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/editor-capabilities/SKILL.md +13 -2
- package/.agents/skills/vidfarm/SKILL.md +73 -29
- package/.agents/skills/vidfarm/harnesses/README.md +112 -0
- package/.agents/skills/vidfarm/{regimes/explainer.QA_REGIME.md → harnesses/explainer.HARNESS.md} +22 -2
- package/.agents/skills/vidfarm/{regimes/hooks.QA_REGIME.md → harnesses/hooks.HARNESS.md} +3 -3
- package/.agents/skills/vidfarm/{regimes/product-demo.QA_REGIME.md → harnesses/product-demo.HARNESS.md} +19 -1
- package/.agents/skills/vidfarm/{regimes/short-form.QA_REGIME.md → harnesses/short-form.HARNESS.md} +67 -7
- package/.agents/skills/vidfarm/{regimes/ugc-testimonial.QA_REGIME.md → harnesses/ugc-testimonial.HARNESS.md} +10 -3
- package/.agents/skills/vidfarm/recipes/{bulk-scripting-with-a-regime.md → bulk-scripting-with-a-harness.md} +35 -12
- package/.agents/skills/vidfarm/recipes/cutout-graphics-for-explainers.md +46 -2
- package/.agents/skills/vidfarm/recipes/local-edit-render-approve.md +2 -1
- package/.agents/skills/vidfarm/references/automation-and-local-dev.md +84 -22
- package/.agents/skills/vidfarm/references/editor-workflows.md +19 -4
- package/.agents/skills/vidfarm/references/hooks-and-virality.md +62 -5
- package/.agents/skills/vidfarm/references/reviewing-renders.md +140 -0
- package/.agents/skills/vidfarm-media/SKILL.md +2 -2
- package/.agents/skills/vidfarm-media/references/tts.md +26 -4
- package/SKILL.director.md +462 -75
- package/SKILL.md +33 -14
- package/dist/src/cli.js +799 -91
- package/dist/src/devcli/{qa-regime.js → harness.js} +132 -55
- package/dist/src/devcli/qa-check.js +209 -4
- package/dist/src/devcli/skill-docs.js +136 -0
- package/dist/src/devcli/stills.js +65 -1
- package/package.json +4 -3
- package/.agents/skills/vidfarm/regimes/README.md +0 -77
package/.agents/skills/vidfarm/{regimes/short-form.QA_REGIME.md → harnesses/short-form.HARNESS.md}
RENAMED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: short-form
|
|
3
|
-
video_type: general short-form social video (TikTok / Reels / Shorts) — the default base
|
|
3
|
+
video_type: general short-form social video (TikTok / Reels / Shorts) — the default base harness
|
|
4
4
|
checks:
|
|
5
5
|
duration_sec: 8-90
|
|
6
6
|
aspect: 9:16
|
|
@@ -11,16 +11,19 @@ checks:
|
|
|
11
11
|
font_regime: required
|
|
12
12
|
max_text_cards: 3
|
|
13
13
|
max_simultaneous_text: 2
|
|
14
|
+
max_words_per_cue: 12
|
|
15
|
+
max_dead_air_sec: 2.5
|
|
16
|
+
max_tail_sec: 1.5
|
|
14
17
|
max_scene_sec: 8
|
|
15
18
|
---
|
|
16
19
|
|
|
17
|
-
# Short-Form
|
|
20
|
+
# Short-Form Harness
|
|
18
21
|
|
|
19
|
-
The default base. Start here, copy it next to your work, then **delete what doesn't apply and add what makes your format yours** — a
|
|
22
|
+
The default base. Start here, copy it next to your work, then **delete what doesn't apply and add what makes your format yours** — a harness you didn't edit is a harness that isn't about your videos.
|
|
20
23
|
|
|
21
24
|
> **Part I** is the anatomy: the four charges that decide whether a video travels.
|
|
22
25
|
> **Part II** is the rules that keep it credible.
|
|
23
|
-
> Run the pre-flight checklist before you build, and `vidfarm qa <dir> --
|
|
26
|
+
> Run the pre-flight checklist before you build, and `vidfarm qa <dir> --harness ./HARNESS.md` before you publish. An unchecked box is a rewrite, not a fix in the edit.
|
|
24
27
|
|
|
25
28
|
## Part 0 — who this is for (fill this in yourself)
|
|
26
29
|
|
|
@@ -57,7 +60,7 @@ Three seconds, not five. The decision is made before you finish the first senten
|
|
|
57
60
|
|
|
58
61
|
**Banned openings:** throat-clearing ("Hey guys", "So I wanted to talk about…") · a logo, title card, fade from black, or a beat of silence · any sentence whose subject arrives in the second half · context before the claim (context is beat 2).
|
|
59
62
|
|
|
60
|
-
⚠️ **Frame 0 is the hook AND the thumbnail.** No black open, no fade, subject in frame, caption already legible. See the `hooks`
|
|
63
|
+
⚠️ **Frame 0 is the hook AND the thumbnail.** No black open, no fade, subject in frame, caption already legible. See the `hooks` harness for the full chunk-1 craft.
|
|
61
64
|
|
|
62
65
|
### 🔄 Curiosity loop — retention
|
|
63
66
|
|
|
@@ -94,12 +97,55 @@ Not an accessibility afterthought: captions are how the hook, the loop, and the
|
|
|
94
97
|
|
|
95
98
|
- **Verbatim, every word.** Paraphrased captions desync from the voice and read as fake.
|
|
96
99
|
- **Weight 700–900, inside the 8–85% safe zone.** Below 700 disappears against footage.
|
|
97
|
-
- **One to three words per line, one line at a time.** A block of full sentences doesn't get read.
|
|
100
|
+
- **One to three words per line, one line at a time.** A block of full sentences doesn't get read. Long narration is paged into 3–5-word kinetic cues (`captions generate --style word-pop|spotlight`), never held as one static block.
|
|
101
|
+
- **Placed in the quietest region of the frame**, measured off a still — not dropped on the default lower third.
|
|
98
102
|
- **Cards are timed text over footage** — never a card UI, table, chip row, or frosted panel (`vidfarm qa` flags those as slop).
|
|
99
103
|
- **Max ~3 standalone cards per video:** one for the loop, one for the payoff, one for the bait.
|
|
100
104
|
|
|
105
|
+
#### Caption PLACEMENT is measured off the frame, before styling is decided
|
|
106
|
+
|
|
107
|
+
The safe zone (8–85%) says where text is *allowed*; it does not say where text *belongs*. Inside that band, the caption goes in the region of the frame with the least competing for attention — and that region is found by looking at a frame, not by defaulting to `y:70`.
|
|
108
|
+
|
|
109
|
+
- **Grab the frame and read it.** `vidfarm stills <dir> --at <t>` is free. Split the safe band into thirds and ask which one is quietest: open sky, a blank wall, a defocused background, an empty tabletop, dead space above or below the subject. That third gets the words.
|
|
110
|
+
- **A caption over the busy third is a self-inflicted wound.** It collides with the subject, so it needs a plate to survive, so the plate becomes a bright slab, so the frame now has two things fighting instead of one. Moving the text 40% up the frame solves all three at once. Nothing else on screen competing → the placement is purely a legibility choice, so make it the *readable* one.
|
|
111
|
+
- **Placement decides the plate, not the other way round.** Choose the position first, then measure the band you actually chose (below). Empty sky measures calm → no plate. Reaching for treatment 4 before you've tried moving the text is the mistake.
|
|
112
|
+
- **Size is part of placement.** A line that runs frame-edge to frame-edge has no placement left to choose. If the words don't fit in the quiet region at ~36–64px, cut words or shrink the type — don't widen the box.
|
|
113
|
+
|
|
114
|
+
#### Caption styling is MEASURED off the background, never hardcoded
|
|
115
|
+
|
|
116
|
+
A white rounded caption plate copied from another video onto a near-black stage is a bright slab the design never asked for — it dominates the frame and reads as a UI element pasted over the video. So measure what is actually behind the caption band, then pick one of three treatments:
|
|
117
|
+
|
|
118
|
+
| Background behind the caption band | Treatment |
|
|
119
|
+
|---|---|
|
|
120
|
+
| Dark and calm (luma < ~70, variation < ~42) | **Light type, NO plate** — the plate is dead weight |
|
|
121
|
+
| Bright and calm (luma > ~160, variation < ~42) | **Dark type, NO plate** |
|
|
122
|
+
| Busy / mid-tone / moving colour | **Plate** — nothing else stays readable |
|
|
123
|
+
|
|
124
|
+
- **The active-word highlight colour follows the treatment.** A deep red active word is fine on a white plate and unreadable against near-black; on a dark treatment it has to become a bright accent. Pick `--color` / `--active-color` / `--background-style` on `vidfarm captions generate` by measuring, not by taste.
|
|
125
|
+
- **ONE treatment per video, not per caption.** Styling that flips every few seconds reads as a bug, not as intelligence. Re-measure per scene only if the footage genuinely changes character.
|
|
126
|
+
- **Measure the COMPOSITED value, not the source file.** When the background is a moving image, walk the caption band's rect back through the background's own transform (Ken Burns scale/pan) into the source image, crop, then apply the composition's own darkening stack — gradient overlay, vignette, layer opacity — so the number is what the viewer sees. Sampling the raw source is the mistake: it reported a mid-grey for a stage that renders near-black.
|
|
127
|
+
- **Verify it for free in a rendered frame.** On the reference video the model predicted `luma 11, variation 5` and the rendered file measured `luma 16, variation 7` — same bucket, and it correctly shipped plate-free. Grab a frame where no caption is showing and measure the band there.
|
|
128
|
+
|
|
129
|
+
#### On-screen text and captions must not say the same thing at once
|
|
130
|
+
|
|
131
|
+
Animated text and diagrams on the page competing with voiceover captions is a real, nameable defect: the two are fighting for the same visual space and the same attention, as if they aren't aware of each other. **Display text carries the argument; captions carry only the parts of the narration the screen does NOT show.** When both render the same words, the viewer reads one line twice in two sizes while hearing it once, and the frame feels like two videos overlaid.
|
|
132
|
+
|
|
133
|
+
The mechanical fix: compare each caption phrase against the words on screen in that scene and **drop the caption when overlap is ≥60% of its content words** (words longer than 2 chars, case- and punctuation-normalised). On the reference video that suppressed **10 of 27 tiles**, and the hook and end card came out **entirely caption-free** — that is correct, not a bug, and worth saying out loud because an author who doesn't expect it will think captions broke. Second-order benefit: the surviving captions inherit wider time windows (the narrowest tile went from **0.41s to 0.71s**), so it buys readability as well as calm.
|
|
134
|
+
|
|
101
135
|
## Part II — the rules
|
|
102
136
|
|
|
137
|
+
### Rule 0 — every second earns its place, or it gets cut
|
|
138
|
+
|
|
139
|
+
The thumb re-decides continuously; a second carrying nothing is a free exit. Assume the first assembly is **30–50% too long** and go find the seconds — agents write videos like prose (wind-up, restatement, tidy conclusion) and every one of those habits is a hole in the retention curve.
|
|
140
|
+
|
|
141
|
+
- **The deletion test, on every beat:** delete it — does the video still make sense and does the payoff still land? Then it stays deleted. Anything that survives must serve one of the four charges; "it gives context" is not a charge.
|
|
142
|
+
- **Cut on sight:** intros/logo stings, the wind-up sentence before the claim, restatement, inter-sentence silence over ~0.35s, real-time process, establishing shots, reading what's already on screen, the outro tail, and filler pans that exist because a clip was short.
|
|
143
|
+
- **Always ripple the hole closed** (`vidfarm ripple <dir> --at <sec> --delta -<sec>`) — a cut that leaves a gap converts fluff into dead air, which is worse.
|
|
144
|
+
- **Density is not speed.** The held comedic beat, the payoff playing out, and a cue's readability keep their seconds. Cut *words*, never the time text needs to be read.
|
|
145
|
+
- **Length is an output.** Build the charges, cut, and ship whatever's left. A brief that dictates a duration ordered fluff.
|
|
146
|
+
|
|
147
|
+
Machine half: `max_dead_air_sec` / `max_tail_sec` / `max_scene_sec` above, plus `vidfarm qa`'s `dead-air`, `dead-tail`, `slow-scene`. Craft: `references/hooks-and-virality.md` → "Density".
|
|
148
|
+
|
|
103
149
|
### Rule 1 — every video is standalone. There is no part two.
|
|
104
150
|
|
|
105
151
|
A viewer arriving mid-scroll with zero context must get a complete, useful video. **You do not control the order** — your video 12 is most people's video 1, and if one breaks out it breaks out *alone*. A cross-video cliffhanger converts your one winner into a dead end. Sharing an angle, a look, or a set of beliefs across the catalog is the strategy; *dependency* is what's banned. Test: hand it to someone who knows nothing — "wait, what is this?" and "where's the rest?" are both failures.
|
|
@@ -147,9 +193,23 @@ Verbatim captions in the font regime and the safe zone · frame 0 works as hook
|
|
|
147
193
|
|
|
148
194
|
**Production**
|
|
149
195
|
- [ ] Verbatim captions, weight 700–900, safe zone, one line at a time
|
|
196
|
+
- [ ] Caption position was chosen off an actual still — it sits in the quietest region of the frame, not on top of the subject, and no line runs edge-to-edge
|
|
197
|
+
- [ ] Anything longer than a phrase is paged into kinetic cues rather than held as a static block
|
|
198
|
+
- [ ] Caption colour / active colour / plate were chosen by measuring the composited background behind the band, not copied from another video
|
|
199
|
+
- [ ] One caption treatment for the whole video, and the active-word colour is readable on it
|
|
200
|
+
- [ ] No caption repeats the words already on screen in that scene (≥60% overlap → drop the caption)
|
|
201
|
+
- [ ] The deletion test was actually run: name the beat you cut, or why nothing could go
|
|
202
|
+
- [ ] No stretch of the video is there to fill time — no intro, no wind-up, no restatement, no tail after the last word
|
|
150
203
|
- [ ] ≤3 standalone cards, no slop furniture, `vidfarm qa` otherwise clean
|
|
151
204
|
- [ ] In a batch: this variant differs from its siblings by more than one noun
|
|
152
205
|
|
|
206
|
+
**Whole-video review** — on the render, not the plan; scenes built one at a time pass individually and drift as a sequence
|
|
207
|
+
- [ ] A contact sheet of ~12 stills was read as an image: consistent margins, one type scale, one accent colour, one illustration style
|
|
208
|
+
- [ ] Pacing is deliberate — not N identically-long beats, and no single hold that stalls it — and no join is jarring
|
|
209
|
+
- [ ] No dead band under top-anchored content; no frame rests empty >0.5s; the end card is settled ≥2s before the last frame
|
|
210
|
+
- [ ] Frames from two different scenes were compared (a frozen render passes duration, frame-count and audio-hash checks)
|
|
211
|
+
- [ ] Audio verified by measurement — ~12–15 dB speech-over-bed, peak <0 dBFS — not by "it sounds fine"
|
|
212
|
+
|
|
153
213
|
## Diagnosing a flop — read the charges, not the video
|
|
154
214
|
|
|
155
215
|
| What the numbers say | Weak charge | Fix |
|
|
@@ -160,4 +220,4 @@ Verbatim captions in the font regime and the safe zone · frame 0 works as hook
|
|
|
160
220
|
| Good retention, no comments | 🎣 Bait | No ask, or the ask was "thoughts?" |
|
|
161
221
|
| Comments, but hostile | 🎣 Bait | Ragebait or an over-claim |
|
|
162
222
|
|
|
163
|
-
**Vary one charge at a time.** A batch where everything changed at once teaches you nothing — which is the entire point of running scripting mode against a
|
|
223
|
+
**Vary one charge at a time.** A batch where everything changed at once teaches you nothing — which is the entire point of running scripting mode against a harness instead of just generating volume.
|
|
@@ -13,9 +13,9 @@ checks:
|
|
|
13
13
|
max_simultaneous_text: 1
|
|
14
14
|
---
|
|
15
15
|
|
|
16
|
-
# UGC / Testimonial
|
|
16
|
+
# UGC / Testimonial Harness
|
|
17
17
|
|
|
18
|
-
For the format where a person talks to camera about a product they use. The entire value of the format is that it **doesn't look produced** — so most of this
|
|
18
|
+
For the format where a person talks to camera about a product they use. The entire value of the format is that it **doesn't look produced** — so most of this harness is about what NOT to add.
|
|
19
19
|
|
|
20
20
|
## The one test
|
|
21
21
|
|
|
@@ -33,7 +33,7 @@ If the answer needs the brand's permission, budget, or logo, it isn't UGC — it
|
|
|
33
33
|
| **Honest limit** | before the close | What it doesn't do / who it isn't for. **The credibility beat** |
|
|
34
34
|
| **Close** | final beat | What they'd tell a friend. Not a CTA read |
|
|
35
35
|
|
|
36
|
-
**The product enters second, never first.** A testimonial that opens on the product is a commercial. It opens on the *situation* — and situations are also what survives a cold algorithm (see the `hooks`
|
|
36
|
+
**The product enters second, never first.** A testimonial that opens on the product is a commercial. It opens on the *situation* — and situations are also what survives a cold algorithm (see the `hooks` harness).
|
|
37
37
|
|
|
38
38
|
## Rules
|
|
39
39
|
|
|
@@ -80,3 +80,10 @@ If you're generating variants with different speakers/avatars, hold the honest-l
|
|
|
80
80
|
- [ ] Any paid/gifted/affiliate relationship is disclosed on screen and in the caption
|
|
81
81
|
- [ ] Frame 0 shows the speaker mid-sentence with a legible caption
|
|
82
82
|
- [ ] In a batch: this variant's *situation* differs from its siblings, not just its wording
|
|
83
|
+
|
|
84
|
+
**Whole-video review** — on the render, not the plan. The bar here is *un*-produced, so drift shows up as the opposite defect: one scene that suddenly looks designed breaks the whole illusion.
|
|
85
|
+
- [ ] A contact sheet of ~12 stills was read as an image — nothing in it looks more produced than the rest
|
|
86
|
+
- [ ] Caption treatment, framing and lighting are consistent end to end; no join is jarring
|
|
87
|
+
- [ ] No frame rests empty >0.5s, and the close is settled ≥2s before the last frame
|
|
88
|
+
- [ ] Frames from two different scenes were compared (a frozen render passes duration, frame-count and audio-hash checks)
|
|
89
|
+
- [ ] Audio verified by measurement — peak <0 dBFS, speech clearly above any bed — not by "it sounds fine"
|
|
@@ -1,12 +1,12 @@
|
|
|
1
|
-
## Recipe: Bulk Video Generation (Scripting Mode) with a
|
|
1
|
+
## Recipe: Bulk Video Generation (Scripting Mode) with a HARNESS.md
|
|
2
2
|
|
|
3
3
|
Use this when the director wants **volume** — daily posting, hook tests, one video per clip in a pool, N variants of a template. Ask first if you're not sure: *"One video, or should we set this up as a repeatable batch?"* If they want volume, this is the shape.
|
|
4
4
|
|
|
5
|
-
The thing that makes bulk work is not the loop — loops are easy. It's that **nobody is going to watch variant #37 as carefully as variant #1**, so the standard has to be written down before the loop runs. That's the `
|
|
5
|
+
The thing that makes bulk work is not the loop — loops are easy. It's that **nobody is going to watch variant #37 as carefully as variant #1**, so the standard has to be written down before the loop runs. That's the `HARNESS.md`.
|
|
6
6
|
|
|
7
7
|
### 0. Read the craft harness first
|
|
8
8
|
|
|
9
|
-
`references/hooks-and-virality.md` — the four charges (hook / loop / payoff / bait), the three gates, and the anti-patterns that only bite at volume. Two of them decide whether this batch is worth running at all: **a different noun is not a different hook** (twenty variants of one sentence with the nouns swapped is one video), and **never point a generator at your grader** (a model writing hooks scored by the same model converges on the rubric, not on what works — scores climb, nothing improves). The
|
|
9
|
+
`references/hooks-and-virality.md` — the four charges (hook / loop / payoff / bait), the three gates, and the anti-patterns that only bite at volume. Two of them decide whether this batch is worth running at all: **a different noun is not a different hook** (twenty variants of one sentence with the nouns swapped is one video), and **never point a generator at your grader** (a model writing hooks scored by the same model converges on the rubric, not on what works — scores climb, nothing improves). The harness catches defects; it does not rank winners.
|
|
10
10
|
|
|
11
11
|
### 1. Agree the variant axis — before any code
|
|
12
12
|
|
|
@@ -20,14 +20,22 @@ vidfarm pull <forkId> --dir ./work # one canonical base fork per batch
|
|
|
20
20
|
|
|
21
21
|
Read `./work/.harness/agent-guide.md` first, as always.
|
|
22
22
|
|
|
23
|
-
### 3. Install and EDIT the
|
|
23
|
+
### 3. Install and EDIT the harness
|
|
24
|
+
|
|
25
|
+
Two ways in, depending on where the format came from:
|
|
24
26
|
|
|
25
27
|
```bash
|
|
26
|
-
|
|
27
|
-
vidfarm
|
|
28
|
+
# (a) From a bundled base — when the format is one you're defining
|
|
29
|
+
vidfarm harness list # short-form | hooks | ugc-testimonial | explainer | product-demo
|
|
30
|
+
vidfarm harness init hooks --out ./work/HARNESS.md
|
|
31
|
+
|
|
32
|
+
# (b) From the template you're batching — when the format is one you're REPLICATING
|
|
33
|
+
vidfarm harness derive <forkId> --out ./work/HARNESS.md # the decomposition, as a harness
|
|
28
34
|
```
|
|
29
35
|
|
|
30
|
-
|
|
36
|
+
(b) is what a director means by *"give me the harness for this template_id"*: the decompose pass already extracted the template's viral / visual / structural / audio / build DNA, and `derive` folds those strands into the same editable Markdown a bundled base produces. If the fork was never decomposed, run `vidfarm decompose` first.
|
|
37
|
+
|
|
38
|
+
Either way, **edit it with the director**. The generated file is a starting point; the parts that matter are the ones they add — who the viewer is, their banned vocabulary, the compliance line, the pacing this account actually uses. A harness nobody edited isn't about their videos. A *derived* harness has the extra failure mode of sounding authoritative: it was written by a model that watched one video, so its "unknown" lines and its confident-but-wrong lines both need a human pass. Existing harness somewhere else on disk? Just point at it: `--harness ./brand/HOUSE_RULES.md`. They stack.
|
|
31
39
|
|
|
32
40
|
### 4. Source the N cheaply
|
|
33
41
|
|
|
@@ -38,13 +46,13 @@ vidfarm public-raws --category scroll-stoppers --limit 20 --json > pool.json
|
|
|
38
46
|
|
|
39
47
|
A curated shelf is a pre-tagged, free, already-hosted clip pool — the cheapest way to get N distinct variants without N downloads or N generation calls.
|
|
40
48
|
|
|
41
|
-
### 5. Loop: edit → QA against the
|
|
49
|
+
### 5. Loop: edit → QA against the harness → render
|
|
42
50
|
|
|
43
51
|
```bash
|
|
44
52
|
for VARIANT in "${VARIANTS[@]}"; do
|
|
45
53
|
SLUG="$(echo "$VARIANT" | tr ' ' '-' | cut -c1-40)"
|
|
46
54
|
vidfarm set-text ./work --layer hook --text "$VARIANT"
|
|
47
|
-
vidfarm qa ./work --json > "qa/$SLUG.json" # ./work/
|
|
55
|
+
vidfarm qa ./work --json > "qa/$SLUG.json" # ./work/HARNESS.md auto-discovered
|
|
48
56
|
jq -e '.ok' "qa/$SLUG.json" >/dev/null || { echo "skipped $SLUG"; continue; }
|
|
49
57
|
vidfarm render "$FORK_ID" --dir ./work --out "renders/$SLUG.mp4" --tracer "batch-$SLUG"
|
|
50
58
|
done
|
|
@@ -64,13 +72,28 @@ done
|
|
|
64
72
|
|
|
65
73
|
Free, offline, no wallet. Variant 1 is the `standard` preset as authored (skew 2%, zoom 3%, rotate 2°, speed +2%, saturation +4%); later variants get jittered magnitudes and flipped signs, so they differ from the original **and from each other**. One variant per account — two accounts posting the same variant defeats the point. Reuse one `--seed` per source so the batch is reproducible.
|
|
66
74
|
|
|
75
|
+
### 5c. Eyeball the renders — the loop cannot do this for you
|
|
76
|
+
|
|
77
|
+
**A batch is exactly where "the agent passed its own broken work" compounds**: nobody is watching variant #37, and `vidfarm qa` is a static DOM check that never sees a rendered pixel. So add one cheap visual pass over the output — a contact sheet per video, read as an image:
|
|
78
|
+
|
|
79
|
+
```bash
|
|
80
|
+
for MP4 in renders/*.mp4; do
|
|
81
|
+
S="$(basename "$MP4" .mp4)"; mkdir -p "qa/$S"
|
|
82
|
+
for t in 0 3 6 9 12 15; do ffmpeg -y -ss $t -i "$MP4" -frames:v 1 "qa/$S/f$(printf %03d $t).png"; done
|
|
83
|
+
ffmpeg -y -pattern_type glob -i "qa/$S/f*.png" \
|
|
84
|
+
-vf "scale=320:-1,tile=3x2:margin=6:padding=6:color=0x999999" -frames:v 1 "qa/$S-sheet.png"
|
|
85
|
+
done
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
Read the sheets. In a batch you're looking for two different things: **per-video** defects (dead regions, a placeholder that reads as a failed render, contradictory numbers, an unlanded CTA, superimposed headlines at a handoff) and **cross-video** drift (variants that no longer look like siblings, or look *too* identical to be N distinct posts). Also compare two frames from different scenes in at least a sample of the renders — a systematic frozen-render bug in the loop will produce N broken files that all pass duration/frame-count checks. Full method: `references/reviewing-renders.md`.
|
|
89
|
+
|
|
67
90
|
### 6. Answer the review items — don't skip this
|
|
68
91
|
|
|
69
|
-
The
|
|
92
|
+
The harness's `- [ ]` checklist comes back on every run because the CLI *can't* settle it. Machine checks catch a 13-word hook or a black first frame; only you can answer "is this variant genuinely different from its siblings?" or "can the viewer guess the withheld answer?" **Report both halves honestly**: what the machine checked, and what you judged. A batch report claiming a clean pass on the judgment half is worse than no report.
|
|
70
93
|
|
|
71
|
-
### 7. Feed what you learn back into the
|
|
94
|
+
### 7. Feed what you learn back into the harness
|
|
72
95
|
|
|
73
|
-
When the director says "the label-framed hooks all died" or "anything over 30s tanked", write it into `
|
|
96
|
+
When the director says "the label-framed hooks all died" or "anything over 30s tanked", write it into `HARNESS.md` as a rule or a checklist line — with the reason attached, so the next agent doesn't argue it away. The compositions are disposable; **the harness is the artifact that compounds across batches.**
|
|
74
97
|
|
|
75
98
|
### Cost note
|
|
76
99
|
|
|
@@ -8,13 +8,15 @@ The mechanical trio — **generate on a chroma plate → key it out → trim to
|
|
|
8
8
|
|
|
9
9
|
**Unless the director asks for something else, build every explainer this way. Don't ask, just do it, and mention the defaults once so they can override.** The whole point of the house style is that explainers read as *clean, bright, and easy* — a busy explainer is a failed explainer.
|
|
10
10
|
|
|
11
|
-
- **White background, light mode.** A plain white (or near-white `#FFFFFF`–`#FAFAFA`) stage. No dark mode, no gradients, no photographic backdrop, no texture. Light mode reads cleaner on every feed, keeps cutout stickers legible, and makes flat-vector art look intentional. Set the composition/scene background to white first, before placing anything.
|
|
11
|
+
- **White background, light mode.** A plain white (or near-white `#FFFFFF`–`#FAFAFA`) stage. No dark mode, no gradients, no photographic backdrop, no texture. Light mode reads cleaner on every feed, keeps cutout stickers legible, and makes flat-vector art look intentional. Set the composition/scene background to white first, before placing anything. **This is the default, not a law** — a director who asks for a dark or photographic stage gets one, but it changes two things mechanically: the stickers need their white die-cut rim stripped, and the caption treatment has to be re-measured. Both are documented in **"Stickers on a DARK or photographic stage"** below.
|
|
12
12
|
- **Kinetic captions.** Narration is always captioned word-by-word (`vidfarm captions generate ./work --style word-pop`). Because the stage is white, **override the preset's dark-canvas colors to dark ink on light**:
|
|
13
13
|
```
|
|
14
14
|
vidfarm captions generate ./work --style word-pop \
|
|
15
15
|
--color "#111111" --active-color "#7C3AED" --background-style plain --max-words 4
|
|
16
16
|
```
|
|
17
17
|
One accent color for the active word, everything else near-black. No outline/stroke, no drop shadow, no pill — those exist to survive busy footage and just add noise on white.
|
|
18
|
+
|
|
19
|
+
**Those hexes are the answer for a white stage, not the answer.** They are one instance of a general rule: **caption color, active-word color and plate are chosen by MEASURING the background behind the caption band, never by taste or habit.** On a near-black stage the same flags ship a bright plate the design never asked for and an active word nobody can read. The measurement procedure and the three treatments live in `harnesses/short-form.HARNESS.md` → "Caption styling is measured off the background" — read it before you copy the line above onto anything that isn't white.
|
|
18
20
|
- **Female TTS narration.** Default to a warm, friendly **female** voice and say which one you picked: local-first `vidfarm tts "<script>" --voice coral` (OpenAI — `nova` if the script wants more energy, `sage` for calmer), `--voice Kore` or `Leda` on Gemini, or `vidfarm voices` → `vidfarm tts --cloud --voice <voice_id>` on ElevenLabs. Tell the director they can swap it in one flag.
|
|
19
21
|
- **Clean and simple wins.** One idea on screen at a time. Two or three cutouts per beat, not eight. Generous white space, one accent color, one font. When in doubt, remove an element rather than add one.
|
|
20
22
|
|
|
@@ -96,6 +98,48 @@ vidfarm remove-greenscreen ./mascot.mp4 --gif --gif-fps 12 --gif-width 480 # AN
|
|
|
96
98
|
|
|
97
99
|
GIF alpha is **1-bit** — a pixel is fully opaque or fully gone, so antialiased edges go hard and semi-transparent shadows/glows disappear (`--gif-alpha <0..255>` moves where that line falls). That's the format, not the key. **For anything going onto a composition, prefer PNG/WebP (still) or transparent WebM (clip);** reach for GIF only when the destination demands it.
|
|
98
100
|
|
|
101
|
+
### Stickers on a DARK or photographic stage — strip the white die-cut rim
|
|
102
|
+
|
|
103
|
+
The house style above puts stickers on a **white** stage, and on white the thing this section is about is invisible. The moment the stage goes dark, photographic, or coloured, every sticker arrives wearing a **white die-cut rim** — a 4–12px light halo tracing its silhouette — and that halo is the single most obvious "a bot made this" artefact in the frame. The art stops reading as an element in the scene and starts reading as a cutout pasted on top of it.
|
|
104
|
+
|
|
105
|
+
**Why the rim is there:** it's a PRINT convention. Real die-cut vinyl needs a white border so the blade has something to cut along, so sticker art is drawn with one, so generators reproduce it. It has no purpose whatsoever in a video composition. **This is a different failure from the `⚠ N% hollow` flag** `sticker-pack` prints — hollow means outline-only art whose interior got keyed away (fix it in the prompt, see above); the rim is extra art that was drawn on purpose and has to be removed after the key.
|
|
106
|
+
|
|
107
|
+
**The fix:** delete exactly the band of light pixels **connected to the transparent edge**, by morphological reconstruction inward from the boundary. Interior whites — an eyeball's sclera, a screen highlight, a paper label — are enclosed by linework, so they are not connected to the edge and survive untouched. Local, free, `numpy` + `scipy` + `PIL`:
|
|
108
|
+
|
|
109
|
+
```python
|
|
110
|
+
def strip_rim(img, light=188):
|
|
111
|
+
a = np.array(img.convert("RGBA")).astype(np.int16)
|
|
112
|
+
solid = a[..., 3] > 128
|
|
113
|
+
is_light = (a[..., :3].mean(axis=2) > light) & solid
|
|
114
|
+
seed = is_light & ndimage.binary_dilation(~solid, iterations=3) # light AND touching transparency
|
|
115
|
+
if seed.any():
|
|
116
|
+
rim = ndimage.binary_propagation(seed, mask=is_light) # flood through light only
|
|
117
|
+
a[..., 3][rim] = 0
|
|
118
|
+
return Image.fromarray(a.astype(np.uint8), "RGBA")
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
**Test LIGHTNESS, not per-channel whiteness — this is the gotcha that makes the whole thing non-obvious.** The rim is **not white**. It is white **contaminated with the chroma plate**, because the plate fringes into it during keying. Measured on a magenta-plate pack, the outermost solid ring averaged **R≈250, G≈205, B≈250** — the green channel is nowhere near white. A per-channel test like `(rgb > 224).all(axis=2)` therefore misses the rim on every sticker whose edge is even slightly anti-aliased: in testing it stripped **1 of 4** stickers, and because the one it did strip looked *different* from its three siblings, the result read as broken art rather than as a bad threshold. `rgb.mean(axis=2) > ~188` catches all four.
|
|
122
|
+
|
|
123
|
+
**Then clean up what stripping leaves behind.** Removing the rim produces two artefacts, and both read to a viewer as "the sticker didn't mask properly":
|
|
124
|
+
|
|
125
|
+
1. **A dotted halo** — surviving specks along the old rim edge. Measured **213** and **164** stray connected components of 3–20px each on two different stickers of the same pack.
|
|
126
|
+
2. **A bright blob** — a large uniform light region the art *enclosed* is no longer visually held in by the rim. Measured at 7,983px (a ring's centre) and 9,582px (a stamp's paper plate).
|
|
127
|
+
|
|
128
|
+
```python
|
|
129
|
+
def clean(img, speck=0.008, blob=0.015, light=200):
|
|
130
|
+
# 1. drop connected components smaller than `speck` of the largest
|
|
131
|
+
# 2. for each enclosed light region larger than `blob` of the sticker area:
|
|
132
|
+
# holes = binary_fill_holes(m) & ~m
|
|
133
|
+
# if holes.sum() < m.sum() * 0.02: # solid fill, no detail -> it is background
|
|
134
|
+
# set alpha 0 there
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
**The discriminator is worth remembering on its own: an enclosed light region is only background if it has NO internal detail, and its holes are the giveaway.** An eyeball's sclera is riddled with drawn veins and a pupil → many holes → keep it. A ring's centre is a flat fill → no holes → punch it transparent. Area alone gets this wrong in both directions.
|
|
138
|
+
|
|
139
|
+
**When you recolour a pack to a palette** (a duotone or luminance ramp onto a brand accent), **cap the top of the ramp** — e.g. `accent + 0.66·(white − accent)` — so interior whites land as a light tint instead of glaring pure white against a flat two-colour design. Uncapped, every kept interior white becomes the brightest pixel in the frame, which undoes the point of the ramp.
|
|
140
|
+
|
|
141
|
+
> **Known gap:** none of this is in the CLI. `vidfarm sticker-pack` / `vidfarm cutout` have no `--strip-rim` (or equivalent) flag today, so on a dark stage you run the two passes above yourself as a local post-step. A `--strip-rim` flag on both commands, defaulting off, is the right long-term home for it.
|
|
142
|
+
|
|
99
143
|
### The guided sequence (prompt harness)
|
|
100
144
|
|
|
101
145
|
**Step 0 — Decide the cast of stickers.** With the director, list every element the explainer needs as its own cutout: the hero subject, each labelled prop, each icon/arrow/emoji, any mascot, any full-frame backdrop. Each becomes one transparent PNG. Stickers are reusable — generate once, reuse across scenes. **If the cast is more than two or three items, make it a PACK** (one sheet, split locally — see "A sticker pack" above) rather than N separate `cutout` calls.
|
|
@@ -130,7 +174,7 @@ GIF alpha is **1-bit** — a pixel is fully opaque or fully gone, so antialiased
|
|
|
130
174
|
|
|
131
175
|
**Both are image-only.** A moving subject has no single bounding box — matte a video clip with `vidfarm remove-background <video>` or key a flat backdrop with `vidfarm remove-greenscreen <video>` (→ transparent WebM/mov).
|
|
132
176
|
|
|
133
|
-
**Step 2 — Show the director each cutout, get corrections.** Cutouts are cheap to regenerate. Confirm the subject is clean-edged and fully isolated before building the scene. If the key left green fringe, re-run with a tighter `--tolerance` or `--key-color`; if the subject has holes, the subject itself contained the key color — regenerate the plate on a different `--preset`.
|
|
177
|
+
**Step 2 — Show the director each cutout, get corrections.** Cutouts are cheap to regenerate. Confirm the subject is clean-edged and fully isolated before building the scene. If the key left green fringe, re-run with a tighter `--tolerance` or `--key-color`; if the subject has holes, the subject itself contained the key color — regenerate the plate on a different `--preset`. **If the stage isn't white, strip the white die-cut rim here**, before anything is staged — see "Stickers on a DARK or photographic stage" above.
|
|
134
178
|
|
|
135
179
|
**Step 3 — Stage them on the composition.** Fork/seed a working composition (`vidfarm pull` or `vidfarm serve`), **set the stage to a white light-mode background first** (house style), then drop each cutout as an **image layer**, sized and positioned deliberately:
|
|
136
180
|
```
|
|
@@ -7,8 +7,9 @@ Use this when a coding agent is doing the work locally or the user wants a repro
|
|
|
7
7
|
3. Read `./work/.harness/agent-guide.md` and `./work/.harness/context.json` before editing.
|
|
8
8
|
4. Make deterministic edits to `composition.html` and optionally `composition.json`.
|
|
9
9
|
5. Validate with `vidfarm lint` or `vidfarm stills` when useful. **Always look at `vidfarm stills ./work --at 0`** — that frame becomes the thumbnail, so it must not be black, empty, or mid-fade.
|
|
10
|
-
6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, a lone pill around a static stat/label, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, and flags a blank/fading first frame (the thumbnail). Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
|
|
10
|
+
6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, a lone pill around a static stat/label, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, flags oversized captions and static walls of text, and flags a blank/fading first frame (the thumbnail). It cannot see pixels, so *where in the frame* the caption sits is still on you — which is why every run ends with a **`▶ NOW WATCH THE VIDEO`** block: render, `vidfarm stills ./work --sheet`, open the contact sheet, and judge each caption against its actual picture. Do that before you report the video as done. Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
|
|
11
11
|
7. Render with `vidfarm render <forkId> --dir ./work --wait`.
|
|
12
|
+
7b. **Review the render as a whole before you approve — this is the step that most changes quality.** `vidfarm qa` and `lint` are static checks on the DOM; neither can see the video. Tile ~12 stills into one contact sheet and read it as an image — `vidfarm stills ./work --sheet` does both in one command (add `--at 0,2,4,…` to pick the timestamps): consistent margins, one type scale, one accent colour, deliberate pacing, no jarring join, no dead band under top-anchored content, end card settled ≥2s before the last frame. Compare frames from two different scenes — a frozen render (overlay pass without `-loop 1`, assets outside the composition root) passes duration, frame-count and audio-hash checks while every frame is identical. Check the mix by measurement, not by ear. Full method + the six most common defects: `references/reviewing-renders.md`.
|
|
12
13
|
8. **Ask about deduplication before you approve** — "is this going out more than once (several accounts, another platform, a re-post later)?" If yes, run `vidfarm dedupe ./final.mp4 [--variants N]` on the **exported** MP4 (free, local ffmpeg, no re-render) and approve each variant separately. Asking here rather than after publication is what avoids paying for a second render. See `references/core-workflows.md` → *Deduplicate before you publish*.
|
|
13
14
|
9. Approve the finished MP4 with `vidfarm approve --video <url|./final.mp4> --caption "..."`. This prints the shareable `share_url`.
|
|
14
15
|
|
|
@@ -41,38 +41,62 @@ Send a stable `tracer` on export so retries are traceable and filterable in job
|
|
|
41
41
|
| | **One-time video** | **Bulk / scripting mode** |
|
|
42
42
|
|---|---|---|
|
|
43
43
|
| The deliverable | One MP4 you both look at | A loop that produces N videos nobody watches frame-by-frame |
|
|
44
|
-
| Quality control | Your eyes on the render | **A `
|
|
44
|
+
| Quality control | Your eyes on the render | **A `HARNESS.md`** — the batch's written standard |
|
|
45
45
|
| What you optimize | This video | The *variant axis* (one thing changes; everything else is held) |
|
|
46
46
|
| Cost posture | Per-video decisions are fine | Per-video AI spend × N — reuse assets, prefer clip pools |
|
|
47
47
|
|
|
48
|
-
A director who says "make me a video about X" usually wants the first. A director who says "I need to post daily" / "make 20 variants" / "test hooks" wants the second and often doesn't know it has a name. **Offer the upgrade explicitly:** *"Want this as one video, or should we set it up as a repeatable batch? Batches get a
|
|
48
|
+
A director who says "make me a video about X" usually wants the first. A director who says "I need to post daily" / "make 20 variants" / "test hooks" wants the second and often doesn't know it has a name. **Offer the upgrade explicitly:** *"Want this as one video, or should we set it up as a repeatable batch? Batches get a HARNESS.md so variant #37 is as good as #1."* Don't silently build a one-off when they asked for volume, and don't drag someone into a scripting harness when they wanted one clip.
|
|
49
49
|
|
|
50
|
-
### `
|
|
50
|
+
### `HARNESS.md` — the reusable AI harness for a format
|
|
51
51
|
|
|
52
|
-
`vidfarm qa`'s built-in rules are **universal** (no HTML slop, the font regime, the thumbnail frame) — the same for everyone, so they live in code. A
|
|
52
|
+
`vidfarm qa`'s built-in rules are **universal** (no HTML slop, the font regime, the thumbnail frame) — the same for everyone, so they live in code. A harness is the opposite: it's what makes **this** director's **this** format good — their audience, hook shape, banned vocabulary, pacing, compliance line, and the DNA of the template it came from. It can't be hard-coded, so it lives next to the work as Markdown they own and version.
|
|
53
53
|
|
|
54
|
-
**It exists because bulk output loses its human reviewer.** One video gets eyes on every frame; fifty generated in a loop do not. The
|
|
54
|
+
**It exists because bulk output loses its human reviewer.** One video gets eyes on every frame; fifty generated in a loop do not. The harness is what the loop grades against.
|
|
55
|
+
|
|
56
|
+
**Three director phrasings, one artifact:**
|
|
57
|
+
|
|
58
|
+
| They say | You run |
|
|
59
|
+
|---|---|
|
|
60
|
+
| "create me a harness" | `vidfarm harness init <base> --out ./work/HARNESS.md`, then edit it with them |
|
|
61
|
+
| "update the harness for this format" | open the file, add the rule **with its reason**, re-run `vidfarm qa` |
|
|
62
|
+
| "give me the harness for this template_id" | `vidfarm harness derive <templateId\|forkId>` — the **decomposition**, as a harness |
|
|
55
63
|
|
|
56
64
|
```bash
|
|
57
|
-
vidfarm
|
|
58
|
-
vidfarm
|
|
59
|
-
vidfarm
|
|
60
|
-
vidfarm
|
|
65
|
+
vidfarm harness list # the bundled starting points
|
|
66
|
+
vidfarm harness init short-form --out ./work/HARNESS.md # copy, then EDIT it
|
|
67
|
+
vidfarm harness derive <forkId> --out ./work/HARNESS.md # a decomposed template → a harness
|
|
68
|
+
vidfarm harness show ./work/HARNESS.md --dna visual # ONE strand, not the whole doc
|
|
69
|
+
vidfarm qa ./work # auto-picks up ./work/HARNESS.md
|
|
70
|
+
vidfarm qa ./work --harness hooks --harness ./brand/HOUSE.md # built-in + your own file — they STACK
|
|
61
71
|
```
|
|
62
72
|
|
|
63
|
-
|
|
73
|
+
**A harness mirrors the template JSON's DNA vocabulary.** Every `## … DNA` heading is indexed under the same key the decompose pass uses, so a derived harness and a hand-written one read the same:
|
|
74
|
+
|
|
75
|
+
| Strand | What lives there | Decompose source |
|
|
76
|
+
|---|---|---|
|
|
77
|
+
| **Viral DNA** | hook, retention mechanic, payoff, core emotion, contrast | `video-context.json` → `viral_dna` |
|
|
78
|
+
| **Visual DNA** | cut rhythm, energy curve, caption style/placement, b-roll, transitions | `editor-harness.json` → `pacing` / `typography` / `broll` |
|
|
79
|
+
| **Structural DNA** | the beats, their roles, which are load-bearing | `editor-harness.json` → `scenes`, `scene-annotations.json` |
|
|
80
|
+
| **Audio DNA** | voiceover, bed, SFX, comedic timing, intonation | `editor-harness.json` → `audio` / `emotional` |
|
|
81
|
+
| **Build DNA** | which paintbrush per beat, the free-tier path | `replication-harness.json` |
|
|
82
|
+
|
|
83
|
+
`harness derive` writes what the decompose pass actually recorded and marks the rest `unknown` — it never invents a strand to look complete. Treat its output as a **first draft**: the model watched the video, it didn't talk to the customer.
|
|
84
|
+
|
|
85
|
+
> Don't confuse `HARNESS.md` with the `.harness/` directory `vidfarm pull` writes. That directory is machine-generated context (`context.json`, `agent-guide.md`), regenerated on every pull — never hand-edit it. `HARNESS.md` is the one the director owns.
|
|
86
|
+
|
|
87
|
+
Bundled bases (`vidfarm harness list`, files under `.agents/skills/vidfarm/harnesses/`): **`short-form`** (the default — the four charges hook/loop/payoff/bait + the standalone rule), **`hooks`** (hook-variant batches: chunk-1 legibility, the unguessable test, the anti-patterns that only show up at volume), **`ugc-testimonial`**, **`explainer`**, **`product-demo`**. Each is a *starting point to edit*, never a house style to conform to — the parts that matter most are the parts the director adds. A harness can also be any file anywhere: `--harness ./campaigns/q3/RULES.md` is fully supported, and `VIDFARM_HARNESS=./work/HARNESS.md` sets a default for a whole run.
|
|
64
88
|
|
|
65
|
-
**The format is two halves, and the split is deliberate:** a front-matter `checks:` block the CLI settles deterministically (duration, aspect, `hook_words_max`, `forbid_text`, `first_frame_text`, … — full key list in `
|
|
89
|
+
**The format is two halves, and the split is deliberate:** a front-matter `checks:` block the CLI settles deterministically (duration, aspect, `hook_words_max`, `forbid_text`, `first_frame_text`, … — full key list in `harnesses/README.md`), and every `- [ ]` checkbox in the body, which comes back as a **review item for you to answer**. "Is the withheld answer one the viewer can't supply themselves?" is a judgment call; a linter claiming to settle it would be lying. **Answer the review items honestly in your report** — the CLI prints them precisely because it can't.
|
|
66
90
|
|
|
67
|
-
**Build on it.** When you learn something from a batch ("the label-framed hooks all died"), write it into the
|
|
91
|
+
**Build on it.** When you learn something from a batch ("the label-framed hooks all died"), write it into the harness as a new rule or checklist line. That is the artifact that compounds across runs; the composition files don't.
|
|
68
92
|
|
|
69
|
-
### The bulk loop, with the
|
|
93
|
+
### The bulk loop, with the harness in it
|
|
70
94
|
|
|
71
95
|
```bash
|
|
72
|
-
vidfarm
|
|
96
|
+
vidfarm harness init hooks --out ./work/HARNESS.md # once, then edit for this account
|
|
73
97
|
for VARIANT in "${VARIANTS[@]}"; do
|
|
74
98
|
vidfarm set-text ./work --layer hook --text "$VARIANT"
|
|
75
|
-
vidfarm qa ./work --json > "qa/$SLUG.json" #
|
|
99
|
+
vidfarm qa ./work --json > "qa/$SLUG.json" # harness auto-discovered from ./work
|
|
76
100
|
jq -e '.ok' "qa/$SLUG.json" >/dev/null || continue # YOUR gate, in YOUR script
|
|
77
101
|
vidfarm render "$FORK_ID" --dir ./work --out "renders/$SLUG.mp4"
|
|
78
102
|
done
|
|
@@ -231,7 +255,7 @@ The licensed harness also carries the **generative build workflow** guidance (ch
|
|
|
231
255
|
| `vidfarm sticker-pack [sheet\|url] [--generate "<theme>"] [--items "a,b,c"] [--count <n>] [--dry-run] [--gap <pct>] [--min-area <pct>] [--output-format png\|webp\|gif] [--out-dir <d>]` | **local, free, ffmpeg-only** (no job; only `--generate` bills, ONCE for the whole set) — key + alpha-channel segmentation + per-item trim | **The STICKER-PACK maker — the answer whenever a director asks for "a sticker pack" / prop set / icon set.** A pack is ONE greenscreen sheet holding every item, keyed once and then masked apart: 1/N the cost of N `cutout` calls, and the only way a cast stays on-style. Finds each item **automatically** by segmenting the keyed sheet's alpha into connected islands — no hand-measured `--crop` rects — and writes one snug transparent file per item (named from `--items`, reading order) plus a `stickers.json` manifest. `--dry-run` prints the detected boxes first; `--gap` merges (lower) or splits (raise) items that came out joined/broken; items have **no maximum size** — a full-frame landscape/backdrop is as valid a sticker as a 3% icon. **Plate color is chosen for you:** when generating it reads the subject and moves the plate off any hue the art uses (green → magenta → blue → black → white — a pack of leaves/frogs/money on GREEN would key holes through the art), and when splitting an existing sheet it DETECTS the plate from the sheet's four corners, so a red/purple sheet handed back from a web tool just works. Pin it with `--key-color`/`--preset`, or `--no-auto-key` for plain green. **The ART is made key-safe too:** the generation prompt is auto-appended with "closed, solidly filled shapes, no outline-only/hollow art, nothing in the plate hue or a near-shade, fully opaque, no glow/translucency" — the fix for stickers that come back as a rim around a transparent hole — and after keying each item reports `holes`/`hole_pct`/`hollow` (console `⚠ N% hollow` at ≥20%, plus `--json` and `stickers.json`). It **warns, never blocks** (a ring/frame/donut reads identically); re-generate with the fill clause, or lift that one item with `vidfarm mask --crop …`. `--output-format gif` emits 1-bit-alpha GIFs for GIF-only surfaces. IMAGE-only. Aliases: `stickers`, `sticker-sheet`. See recipe `cutout-graphics-for-explainers.md` → "A sticker pack". |
|
|
232
256
|
| `vidfarm tts "…" [--style "…"] [--voice <v>] [--out <file>]` | (LOCAL-FIRST: your own OPENAI/GEMINI/OPENROUTER_API_KEY → audio file on disk; `--cloud` = `POST /api/v1/primitives/audio/speech` + poll, ElevenLabs on the platform key by default, `--own-key` for yours) | text → narration audio; `--cloud --voice <voice_id>` picks an ElevenLabs voice |
|
|
233
257
|
| `vidfarm music "<prompt>" [--length <sec>] [--out <f>] [--own-key]` | `POST /api/v1/primitives/music/generate` (polls job) | prompt → music track (ElevenLabs; platform key + wallet by default, `--own-key` for yours) |
|
|
234
|
-
| `vidfarm voices [--own-key] [--limit N]` | `GET /api/v1/primitives/audio/voices` |
|
|
258
|
+
| `vidfarm voices [--sample] [--search "…"] [--free\|--all] [--own-key] [--limit N]` | `GET /api/v1/primitives/audio/voices` | **Browse AND sample narration voices.** Default roster = the premium ElevenLabs catalog reached through **vidfarm's own ElevenLabs connection** — the user needs no ElevenLabs account, API key, or subscription; narration is billed as vidfarm wallet credits (pennies each). `--free` = the $0 local Kokoro roster (`--all` = both). `--sample` writes listenable clips to `./voice-samples` (`--sample-count`, `--sample-out`, `--sample-text`) and is **free on both tiers** — premium samples are ElevenLabs' own preview clips, free samples render locally — so it's safe in `minimize`. `--search` filters by name/labels/description. **In interactive mode play the samples and let the USER pick**; autonomous = default a voice and still say they can choose. `--own-key` lists the customer's own ElevenLabs account instead. |
|
|
235
259
|
| `vidfarm stt <file\|url> [--out <base>] [--no-diarize]` (alias: `transcribe`) | (LOCAL-FIRST: local ffmpeg demux + your own key; `--cloud` = `POST /api/v1/primitives/audio/transcribe` + poll, ElevenLabs Scribe on the platform key by default, `--own-key` for yours) | video/audio → transcript in BOTH formats: simple subtitles (txt + SRT) and multi-speaker segments (json) |
|
|
236
260
|
| `vidfarm place <dir> --src <url\|file> [--at\|--replace]` | (edits local composition.html; local files → serve disk store or temp upload) | drop media (URL **or local file**) into a gap / over a scene |
|
|
237
261
|
| `vidfarm captions generate <dir> [--style <preset>] [--audio <f>\|--srt <f>\|--text "…"]` | (LOCAL-FIRST: STT on your own key — OpenAI = real word timestamps — then edits local composition.html) | transcribe narration → animated word-by-word caption cues |
|
|
@@ -280,11 +304,12 @@ The licensed harness also carries the **generative build workflow** guidance (ch
|
|
|
280
304
|
| `vidfarm raws search "…"` / `raws match "…"` / `raws list` / `raws sources` | (local library; NL→criteria via local agent or provider key) | search/reuse the raws library |
|
|
281
305
|
| `vidfarm raws preset list\|run\|save` / `raws export <ids…> --to <dir>` | (local library) | saved queries; copy raw MP4s out |
|
|
282
306
|
| `vidfarm lint <dir\|composition.html>` | (local static validation) | pre-publish composition check: timing, overlaps, preset names, media src |
|
|
283
|
-
| `vidfarm stills <dir> [--at 0,2.5,…]` | (local in-process render of PNG frames) | visually verify an edit without a full render |
|
|
284
|
-
| `vidfarm qa <dir\|composition.html> [--
|
|
285
|
-
| `vidfarm
|
|
307
|
+
| `vidfarm stills <dir> [--at 0,2.5,…] [--sheet]` | (local in-process render of PNG frames) | visually verify an edit without a full render. **`--sheet` also tiles them into one contact sheet** (`<out>/contact-sheet.png`, `--sheet-out`/`--sheet-width` to tune) — the whole-video review pass: read it as ONE image and sequence-level drift (uneven margins, three type sizes, a wandering accent colour, N identical beats, a jarring join) becomes obvious where per-scene checks never see it |
|
|
308
|
+
| `vidfarm qa <dir\|composition.html> [--harness <name\|path>…] [--json] [--strict]` | (local static QA — **devcli-only**, no cloud/REST twin) | **social-native QA: HTML slop + first frame + font regime. Run it on EVERY video you produce.** `--harness` grades against a HARNESS.md too (stackable). Free, instant, feedback-only |
|
|
309
|
+
| `vidfarm harness list\|show <ref> [--dna <strand>]\|init <name> [--out <path>]\|derive <forkId\|dir>\|check <dir>` | (local — **devcli-only**) | **HARNESS.md: the reusable AI harness for one format or template.** `init` copies a bundled base to edit; `derive` turns a decomposed template's DNA into one ("give me the harness for this template_id"); `check` is `vidfarm qa` under the harness noun |
|
|
286
310
|
| `vidfarm doctor` | (local environment triage) | check ffmpeg/node/keys/agent CLI/poisoned env + list local serve/preview processes before debugging anything else; `--kill-orphans` reaps dead servers squatting ports (fixes the "Waiting for preview server…" hang) |
|
|
287
311
|
| `vidfarm skills list\|add <name>\|update` | `GET /skill-pack/index.json` · `/skill-pack/:name/*` | install/refresh skill packs (see "Skill packs — import on demand") |
|
|
312
|
+
| `vidfarm skill ls\|show <path>\|search "<term>"\|path` | (local — **offline, no account**) | **Read this pack straight off disk.** A full copy ships inside the devcli tarball and is pinned to the installed version. `search` greps all 22 files at once — the cheapest way to find one paragraph without loading a whole reference |
|
|
288
313
|
| `vidfarm tts "…" --engine local` / `vidfarm stt <file> --engine whisper` | (keyless LOCAL engines: Kokoro-82M TTS, whisper.cpp STT) | narration + word-timestamp transcripts with zero keys and zero accounts |
|
|
289
314
|
| `vidfarm remove-background <video\|image>` | (local ONNX matting — free) | transparent-subject media for occlusion captions/cutouts (arbitrary/messy background; for a FLAT solid background use `remove-background-greenscreen`) |
|
|
290
315
|
| `vidfarm capture <url>` | (local headless-Chrome capture) | website screenshots/assets for website-to-video flows |
|
|
@@ -302,8 +327,8 @@ The licensed harness also carries the **generative build workflow** guidance (ch
|
|
|
302
327
|
vidfarm qa ./work # human-readable findings + verdict
|
|
303
328
|
vidfarm qa ./work --json # machine-readable: rule / severity / where / fix
|
|
304
329
|
vidfarm qa ./work --strict # ALSO exit 1 on slop (only if you want a CI gate)
|
|
305
|
-
vidfarm qa ./work --
|
|
306
|
-
# auto-discovers ./work/
|
|
330
|
+
vidfarm qa ./work --harness hooks # + grade against a HARNESS.md (repeatable; also
|
|
331
|
+
# auto-discovers ./work/HARNESS.md)
|
|
307
332
|
```
|
|
308
333
|
|
|
309
334
|
**Run this on every video you produce.** It is free, instant (pure DOM, no ffmpeg/Chrome/network), and it is the only automated check for the thing that most often ruins an agent-made video: **HTML slop**. Compositions are authored in HTML, so an agent's web-page instincts leak straight onto the frame as landing-page furniture that appears on every website and in **zero** real TikToks.
|
|
@@ -329,13 +354,22 @@ What it flags:
|
|
|
329
354
|
| `font-regime` | error/warn | A text layer in a website body font (Inter/Roboto/Arial/system-ui → **error**) or any family outside the imported regime (Montserrat, TikTok Sans, Abel, Source Code Pro, Yesteryear → warn, it silently falls back at render) |
|
|
330
355
|
| `font-size` / `font-weight` | error/warn | `font-size:0` (invisible) is an error; sub-2.6%-of-canvas-width text and weight <600 warn |
|
|
331
356
|
| `caption-safe-zone` | warn | Text outside the 8%–85% band on a **portrait** canvas (landscape/square are exempt) |
|
|
357
|
+
| `caption-oversize` | warn | Display-size type (>7.5% of canvas width) on a line of **5+ words** — it runs edge-to-edge, wraps, covers the frame, and forces a full-width plate. Both signals required, so a giant 2-word hook card passes |
|
|
358
|
+
| `wall-of-text` | warn | One **static** text layer carrying 14+ words — a paragraph, not a caption. Page it into 3–5-word kinetic cues (`captions generate --style word-pop`). Layers already part of an animated caption run are exempt |
|
|
359
|
+
| `dead-air` | warn | A gap of **≥2.5s between cues** with nothing on screen to read (needs 3+ cues, so a two-card title sequence is exempt). Dead screen time is a free exit — cut it and `ripple` the hole closed |
|
|
360
|
+
| `dead-tail` | warn | The video keeps running **>1.5s after the last word** — an outro, an end card, or an untrimmed clip. End on the bait |
|
|
361
|
+
| `slow-scene` | warn | One clip >6s **and** >2.5× the median clip length — judged against the video's OWN rhythm, so a deliberately slow piece or a single-take talking head passes |
|
|
332
362
|
| `thumbnail-blank-open` | error | Nothing on screen at **t=0** — the opening clip starts late, so the poster frame is black |
|
|
333
363
|
| `thumbnail-fade-in` | error/warn | An **entrance** transition on the FIRST clip: `fade-black`/`fade-white`/`flash`/`smoke` → **error** (frame 0 is a flat solid); any other preset → warn (frame 0 caught mid-move). Junction transitions on later clips are never flagged |
|
|
334
364
|
| `thumbnail-no-hook-text` | warn | The composition has text, but none of it is up at t=0 — the poster carries no hook words. Ignorable when you're deliberately opening on a clean face/product shot |
|
|
335
365
|
|
|
336
366
|
Every finding carries a concrete `fix` line — the answer is always "say it as timed text on the footage", never just "delete it". Fold `--json` into scripted batch runs to QA N variants at once.
|
|
337
367
|
|
|
338
|
-
**
|
|
368
|
+
**Every run ends by telling you to go watch the video — that instruction is part of the output, not a footnote.** `vidfarm qa` closes with a `▶ NOW WATCH THE VIDEO — this check never did` block (and a `watch_the_video: { required: true, why, steps[] }` object in `--json`), printed on **clean** runs too, because a green tick on DOM attributes is the single easiest thing to mistake for a reviewed video. The steps are dir-aware and paste-ready: render, `stills --sheet` → *open the contact sheet*, read it as one sequence, judge each caption against its picture, compare two different scenes (a frozen render passes every mechanical check), measure the audio with `volumedetect`, and report what you measured separately from what you judged. **Do them.** Reporting "QA passed" to a director without opening a frame is not a review, and the tool now says so to your face.
|
|
369
|
+
|
|
370
|
+
**`vidfarm qa` is a static DOM check — it cannot see the rendered video.** In particular it can tell you a caption is *too big* or *outside the safe zone*, but never whether it sits in the **empty** part of the frame — that needs pixels, so it stays your job: `vidfarm stills <dir> --at <t>`, look, then place (see `references/editor-workflows.md` → "TikTok-native caption standard"). It never looks at pixels, motion, spacing, colour drift, pacing, or the joins between scenes, so a clean `qa` run says nothing about whether the video reads as one coherent piece. That judgment is a separate, mandatory pass: tile stills into a contact sheet, read it as an image, and check balance/spacing/type/colour/rhythm across the whole sequence. It also can't catch a **frozen render** (every frame identical while duration, frame count and audio hash all pass), which is why you compare frames from two different scenes. Full method: `references/reviewing-renders.md`.
|
|
371
|
+
|
|
372
|
+
**The two halves, and why the tool only claims one.** Everything above is universal and mechanical. The half that decides whether a *particular* video is any good — is the hook legible cold, does the loop close, is this variant genuinely different from its siblings — is the director's, and it lives in a **`HARNESS.md`** (see "Scripting mode" above). Pass one with `--harness <name|path>` (repeatable, and a `HARNESS.md` sitting next to the composition is picked up automatically): its `checks:` front matter is settled deterministically alongside the built-ins, and its `- [ ]` checklist comes back as **review items you must answer yourself**. `vidfarm qa` deliberately never fakes a verdict on those — a "PASS" it couldn't have earned is worse than no check at all.
|
|
339
373
|
|
|
340
374
|
## Cost mode — the devcli's money-saving guardrail
|
|
341
375
|
|
|
@@ -360,6 +394,15 @@ The four modes, quoted as **cost per finished video**. The first two are spend p
|
|
|
360
394
|
|
|
361
395
|
**Narration defaults to the FREE local voice in minimize AND hybrid.** A bare `vidfarm tts "…"` runs the keyless local Kokoro-82M engine in both of those modes — you no longer have to remember `--engine local`. A run **opts out** of that default by asking for a premium voice (`--style`, `--provider`, `--model`, `--own-key`, or a non-Kokoro `--voice` like `alloy`/`Kore`/an ElevenLabs id), by passing `--cloud`/`--engine byok`, or by being in `rich-ai`/`pure-videogen`. If the local engine isn't installed on the machine (it needs `pip install kokoro-onnx soundfile` + a ~340MB model on first use), the run **falls back** to the user's provider key / cloud instead of failing — it prints the reason on stderr so you can tell the user why the voice changed.
|
|
362
396
|
|
|
397
|
+
**Narration gotchas that ship a correct-looking, wrong-sounding video.** Each of these produces output that passes every structural check:
|
|
398
|
+
|
|
399
|
+
- **Never `adelay` the voiceover to position it.** Whisper's word timings — and therefore every caption generated from them — are relative to the raw `vo.wav`. An `adelay` desyncs every caption in the video while the file still plays perfectly. Use `apad` + `atrim`.
|
|
400
|
+
- **Whisper's default model is English-only and hallucinates fluent English over other languages.** Non-English narration needs `--model large-v3 --language <code>`; without it you get a clean, confident, entirely invented transcript.
|
|
401
|
+
- **`vidfarm tts` reads stdin** — redirect `</dev/null` when calling it inside a shell loop, or the loop consumes its own input.
|
|
402
|
+
- **Check the brand/product name's pronunciation** before you render 20 variants with it. TTS engines mangle proper nouns (Kokoro reads *Genki* as "Jenki"); respell it phonetically in the TTS input and confirm with a whisper round-trip — you're already running whisper for the caption timings.
|
|
403
|
+
- **For a calm, unhurried read, render line by line** and concatenate the takes with measured silences, rather than one continuous pass. A single pass reads rushed however slow the copy is, because the pauses are TTS filler rather than real beats.
|
|
404
|
+
- **Verify the mix by measurement, not by ear** — you can't hear the render. Target **12–15 dB** of speech-over-bed separation measured across the actual word spans, peak below **0 dBFS**. A "separation" figure computed over the music-only tail measures bed-vs-bed and over-reports badly; don't retune against it. See `references/reviewing-renders.md`.
|
|
405
|
+
|
|
363
406
|
Precedence: `--cost-mode <m>` flag → `VIDFARM_COST_MODE` env → the saved `cost-mode` → default (hybrid, flagged as "not set"). When nothing is saved and a billed op runs, the CLI prints a "no preference set — ask the user" nudge instead of silently spending, so the default posture really is *ask before you spend*.
|
|
364
407
|
|
|
365
408
|
**Agent-memory handoff.** After the user picks, offer to remember it across sessions — but the destination depends on the agent, so ask: Claude Code → `CLAUDE.md` (or its memory dir); Codex / OpenCode / most others → `AGENTS.md`; or a note file the user names. `vidfarm cost-mode <choice>` already persists the devcli-side preference; agent memory is the extra step that survives a fresh checkout. In the **web app UI** there is no memory file — ask each time unless the user states a standing preference for the session.
|
|
@@ -435,6 +478,25 @@ The customer-facing walkthrough (the "VidFarm Walkthrough Tutorial" course) is p
|
|
|
435
478
|
|
|
436
479
|
Both are public and read-only (no auth). Prefer these to guessing steps — quote the real chapter and link the reader to its `url`. Chapters cover onboarding/setup, the operating funnel (angles/hooks/awareness), each guided edit demo (recaption, product tease, remix-with-raws, actor replacement, animate-static-book, drama series, product promo, motion explainers), sourcing/clipping raws, the wallet, cancellation/refunds, and the developer devcli/scripting/free-mode chapters.
|
|
437
480
|
|
|
481
|
+
## The director pack ships inside the devcli — read it offline
|
|
482
|
+
|
|
483
|
+
Installing `@officexapp/vidfarm-devcli` puts a **complete copy of this pack on disk**, pinned to that CLI version. You never have to be online, logged in, or in a project with `.agents/skills/` to read it:
|
|
484
|
+
|
|
485
|
+
```bash
|
|
486
|
+
vidfarm skill ls # every file, with sizes
|
|
487
|
+
vidfarm skill show primitives # shorthand resolves to references/primitives.md
|
|
488
|
+
vidfarm skill show harnesses/README.md # or an exact path
|
|
489
|
+
vidfarm skill search "greenscreen" # grep all of it — find the paragraph, then open that file
|
|
490
|
+
vidfarm skill path # where the bundled copy lives
|
|
491
|
+
```
|
|
492
|
+
|
|
493
|
+
**Prefer `skill search` over opening a big reference.** `editor-workflows.md` is ~650 lines and `automation-and-local-dev.md` ~520; a grep that returns `references/primitives.md:214` costs almost nothing and tells you exactly which file to load.
|
|
494
|
+
|
|
495
|
+
Two things this does NOT mean:
|
|
496
|
+
|
|
497
|
+
- **Pinned, not live.** The bundled copy matches the installed CLI — which is the pairing that actually works, since a newer skill against an older binary is the usual cause of *"the skill says to do X but the command 404s"*. For the host's latest, `vidfarm skills add vidfarm` (installs into a project) or `vidfarm skill --print --remote`. When they disagree, update **both halves together**: <https://vidfarm.cc/update.md>.
|
|
498
|
+
- **Documentation, not entitlement.** Reading about a paid primitive offline does not make it run offline. The **free-local** half genuinely needs nothing — clip hunting, hyperframes, `vidfarm serve` render, `vidfarm qa`, harnesses, `vidfarm dedupe`, Kokoro TTS, whisper STT. The **paid-cloud** half still needs `vidfarm login` and a network call: AI image/video/voice generation, hosted render, `recycle`, `download-video`, marketplace, and the hosted file directory. Tell the director which half a plan lands in *before* you build it.
|
|
499
|
+
|
|
438
500
|
## Skill packs — import on demand (HyperFrames-grade authoring power)
|
|
439
501
|
|
|
440
502
|
This skill stays lean on purpose. Deep authoring craft lives in **skill packs** — Vidfarm's whitelabel of the open-source `hyperframes` skill suite (same engine as `vidfarm hf` / `vidfarm render`, Vidfarm-branded) plus Vidfarm's own media pack — vendored on the Vidfarm host and installed only when a task needs them. Never install skills from upstream vendor orgs or third-party registries; the vidfarm mirror is the source (`vidfarm skills add <name>` fetches `GET /skill-pack/:name/*` with hash verification into `.agents/skills/` + a `.claude/skills/` link, pinned in `skills-lock.json`; `vidfarm skills list` shows what is available/installed; `vidfarm skills update` refreshes pins).
|