@officexapp/vidfarm-devcli 0.21.34 → 0.21.36
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/editor-capabilities/SKILL.md +14 -3
- package/.agents/skills/vidfarm/SKILL.md +66 -33
- package/.agents/skills/vidfarm/harnesses/README.md +112 -0
- package/.agents/skills/vidfarm/{regimes/explainer.QA_REGIME.md → harnesses/explainer.HARNESS.md} +3 -2
- package/.agents/skills/vidfarm/{regimes/hooks.QA_REGIME.md → harnesses/hooks.HARNESS.md} +3 -3
- package/.agents/skills/vidfarm/{regimes/product-demo.QA_REGIME.md → harnesses/product-demo.HARNESS.md} +1 -1
- package/.agents/skills/vidfarm/{regimes/short-form.QA_REGIME.md → harnesses/short-form.HARNESS.md} +39 -10
- package/.agents/skills/vidfarm/{regimes/ugc-testimonial.QA_REGIME.md → harnesses/ugc-testimonial.HARNESS.md} +3 -3
- package/.agents/skills/vidfarm/recipes/{bulk-scripting-with-a-regime.md → bulk-scripting-with-a-harness.md} +20 -12
- package/.agents/skills/vidfarm/recipes/cutout-graphics-for-explainers.md +43 -13
- package/.agents/skills/vidfarm/recipes/local-edit-render-approve.md +1 -1
- package/.agents/skills/vidfarm/references/automation-and-local-dev.md +77 -26
- package/.agents/skills/vidfarm/references/editor-workflows.md +18 -5
- package/.agents/skills/vidfarm/references/hooks-and-virality.md +65 -7
- package/.agents/skills/vidfarm/references/reviewing-renders.md +2 -1
- package/.agents/skills/vidfarm-media/SKILL.md +2 -2
- package/.agents/skills/vidfarm-media/references/tts.md +26 -4
- package/SKILL.director.md +292 -98
- package/SKILL.md +33 -15
- package/dist/src/cli.js +1200 -141
- package/dist/src/devcli/handoff.js +54 -33
- package/dist/src/devcli/{qa-regime.js → harness.js} +132 -55
- package/dist/src/devcli/plate-key.js +698 -0
- package/dist/src/devcli/qa-check.js +209 -4
- package/dist/src/devcli/skill-docs.js +136 -0
- package/dist/src/devcli/sticker-pack.js +48 -0
- package/package.json +6 -4
- package/.agents/skills/vidfarm/regimes/README.md +0 -79
|
@@ -84,12 +84,22 @@ Every video you touch has four charges in series, and **you write them before yo
|
|
|
84
84
|
1. 🪝 **Hook** — the opening line, as text, on screen at `start:0`. A **complete clause** (subject + verb), no jargon, naming a **situation** ("I've quit six businesses") not a label ("anonymity"). Caption chunk 1 is read before any audio — muted autoplay is the default viewing condition, so the text hook outworks the spoken one. Banned openings: throat-clearing, a logo, a title card, a fade from black, context before the claim.
|
|
85
85
|
2. 🔄 **Loop** — one open question by 0:10, stated **on screen**, closing **inside this video** (name the timestamp; if you can't, there's no loop). The withheld answer must be one the viewer **can't supply themselves** — a loop whose answer they can guess passes every mechanical check and dies in the field.
|
|
86
86
|
3. 😍 **Payoff** — shown, not summarized, landing before the final beat. The payoff is not the CTA.
|
|
87
|
-
4. 🎣 **Bait** — one ask, final beat, and tell the user to put it in the post caption too.
|
|
87
|
+
4. 🎣 **Bait** — one ask, final beat, and tell the user to put it in the post caption too. A keyword comment ask ("comment CLIPPER and I'll send the breakdown") is standard and allowed. Never "follow for part two", ragebait, or an earnings/health claim traded for the reply.
|
|
88
88
|
|
|
89
89
|
**On a re-theme this is the thing you protect.** `editor_context` → `viral_dna.hook` / `retention` / `payoff` / `emotional_punch` tells you what the source's charges were — that structure is *why the template worked*. Rebuild each charge for the new subject; flattening the loop into a product statement is the most common way a re-theme kills a format.
|
|
90
90
|
|
|
91
91
|
Full craft (the three gates, situations-vs-labels with worked fixes, loop mechanics, compliance, diagnosis-by-charge): `load_skill('vidfarm', file='references/hooks-and-virality.md')`. Read it before writing hook copy or re-theming.
|
|
92
92
|
|
|
93
|
+
## Every second earns its place — cut hard (hard constraint)
|
|
94
|
+
|
|
95
|
+
The thumb re-decides continuously, so a second that carries nothing is a free exit. **Every edit you make should be shorter than what you started with unless the director asked for more.** When you finish a change, look at the timeline you produced and ask what comes out.
|
|
96
|
+
|
|
97
|
+
- **Deletion test, per beat:** delete it — does the video still make sense and does the payoff still land? Then leave it deleted. Whatever survives must serve one of the four charges; "it gives context" is not a charge.
|
|
98
|
+
- **Cut on sight:** intro/logo/title cards, the wind-up line before the claim, restatement, real-time process, establishing shots, captions reading what's already on screen, and any footage after the last word.
|
|
99
|
+
- **Close the hole.** `editor_action ripple_edit` with a negative `delta_start` at the cut point — never leave a gap. Dead air is worse than the boring beat you removed: `editor_context` layer timings are how you spot both (a stretch with no text layer covering it, or a last clip that outruns the last cue).
|
|
100
|
+
- **Not speed.** Keep the held comedic beat, the payoff playing out, and enough time to read each cue. To tighten a talky stretch, cut *words* (`set_layer_text`), not the seconds text is on screen.
|
|
101
|
+
- **Say what you cut.** When you report back, name the beats you removed and the new duration. If the director asked for something longer, say plainly what is now filling the extra time.
|
|
102
|
+
|
|
93
103
|
## The FIRST FRAME is the thumbnail (hard constraint)
|
|
94
104
|
|
|
95
105
|
Frame 0 is a single frame of ~30 in the first second, and it outweighs all of them: every feed card, share link, embed, and paused player freezes on it, so **more people see that one frame than watch the video**. Whatever you change, check what `t=0` looks like before you call the job done.
|
|
@@ -121,12 +131,13 @@ You author into HTML, which makes it dangerously easy to build a **web page inst
|
|
|
121
131
|
|
|
122
132
|
**Font + background regime (every caption/title, via `set_captions` / `set_layer_style` / `add_layer`):**
|
|
123
133
|
- **Font:** Montserrat (default) or TikTok Sans / Abel / Source Code Pro / Yesteryear — a family the composition actually imports, or it silently falls back to the slop sans. Weight **700–900**. `font_size` in px of the render canvas: **~36–64px** on a 1080-wide frame; never <28, never 0 (invisible). ~2 lines, ~5 words per line; `line_height` 0.95–1.15.
|
|
124
|
-
- **Position:** inside the **8%–85%** vertical safe zone (phone UI clips the edges) and clear of the right ~12% action rail — a centered box at `x:10 width:80` is safe. Lower-third ≈ `y:70`; a "POV:" top line ≈ `y:8`, never `y:0`.
|
|
134
|
+
- **Position:** inside the **8%–85%** vertical safe zone (phone UI clips the edges) and clear of the right ~12% action rail — a centered box at `x:10 width:80` is safe. Lower-third ≈ `y:70`; a "POV:" top line ≈ `y:8`, never `y:0`. **Inside that band, put the words where the picture ISN'T** — read the actual frame (the scene's media, or ask the user for a still) and park the caption in the emptiest region with nothing competing for attention (open sky, a blank wall, a defocused background), even if that means `y:12` instead of a lower third. `y:70` is a default, not a law. Text over the busy third of the frame fights the shot and then needs a plate to survive; the same words in the empty third need none.
|
|
135
|
+
- **Length:** a caption layer is a *page*, not a transcript. Past ~10–12 words (or ~4s on screen while the voice keeps talking) it's a wall of text nobody reads — page it into 3–5-word kinetic cues via `set_captions` (`spotlight`/`karaoke`/`word-pop`), which also lets the type be smaller and usually removes the need for a plate. Static text is for hook lines, payoff numbers, and title cards, which are short by nature.
|
|
125
136
|
- **Background — exactly one of four:** `background_style:"outline"` (stroke, the default look) · `"plain"` (bare + soft shadow) · an **active-word highlight pill** via `set_captions caption_style:"spotlight"|"karaoke"` (the *only* legitimate pill anywhere in the frame — it tracks the spoken word; a static label never gets one) · `"highlight-solid"`/`"highlight-translucent"` as a band that **hugs** the text (radius ≤~8px, no border, no shadow, no gradient, no blur, one text run — never a heading+subheading+URL stacked inside it). Anything else is a web card.
|
|
126
137
|
|
|
127
138
|
**There is no QA tool for you.** The devcli ships `vidfarm qa <dir>` — a free local blocklist pass over exactly the rules above — but it is **devcli-only with no REST twin**, so in the web editor you enforce this by reading your own output. When you hand a heavy job off to a local coding agent, tell them to run `vidfarm qa ./work` before rendering.
|
|
128
139
|
|
|
129
|
-
**If the user wants VOLUME, say so and hand it off.** "I need to post daily", "make 20 versions", "test these hooks" is **bulk/scripting mode**, not twenty turns of editor chat: a pinned base fork, a loop varying one thing per variant, and a **`
|
|
140
|
+
**If the user wants VOLUME, say so and hand it off.** "I need to post daily", "make 20 versions", "test these hooks" is **bulk/scripting mode**, not twenty turns of editor chat: a pinned base fork, a loop varying one thing per variant, and a **`HARNESS.md`** — the user's own written quality standard, the reusable AI harness for that format, which exists because nobody reviews variant #37 as carefully as #1. You can't run that loop (no shell, no filesystem), so name the pattern, offer the My Files handoff, and tell them the local agent should run `vidfarm harness init <base> --out ./work/HARNESS.md` (or `vidfarm harness derive <forkId>` to distil one from a decomposed template), edit it with them, and gate the batch on `vidfarm qa ./work`. If they already have a harness file, its rules are still worth reading into your own edits here.
|
|
130
141
|
|
|
131
142
|
Deeper rationale and the devcli-side twins live in `vidfarm` → `references/editor-workflows.md` ("Social-native visual standard" / "TikTok-native caption standard").
|
|
132
143
|
|
|
@@ -35,6 +35,7 @@ vidfarm serve template_<32hex> # local server + browser, opens that
|
|
|
35
35
|
- The API key comes from https://vidfarm.cc/settings and starts with `vf_key_`. Instead of `login`, setting the `VIDFARM_API_KEY` environment variable also works for every command — the CLI reads it from the environment or from a `.env` file in the current directory.
|
|
36
36
|
- No account or key? `vidfarm serve --no-cloud` still gives a fully local editor with free local renders.
|
|
37
37
|
- "Open/run template X locally" is exactly one command: `vidfarm serve <template_id>` (alias: `vidfarm <template_id>`). Do not hand-roll REST or hunt for local `.harness/` files first — `serve` and `pull` create those.
|
|
38
|
+
- **The CLI carries this entire skill offline** — installing the devcli puts a copy of the pack on disk, pinned to that version. `vidfarm skill ls` lists it, `vidfarm skill show <path>` prints one file, and **`vidfarm skill search "<term>"` greps all of it at once**, which is the cheapest way to find the paragraph you need without loading a 650-line reference. No account, no network. It is documentation, not entitlement: the free-local half (clips, hyperframes, `serve` render, `qa`, harnesses, `dedupe`, local TTS/STT) runs offline; AI generation, hosted render, `recycle`, `download-video` and marketplace still need `vidfarm login` and a cloud call.
|
|
38
39
|
|
|
39
40
|
### Entity ID formats
|
|
40
41
|
|
|
@@ -129,7 +130,7 @@ If the user hasn't picked yet and you're about to spend, name the cheaper path a
|
|
|
129
130
|
|
|
130
131
|
Cost mode answers *how much money may I spend*. It does not answer *how much of the user's own hands may I use* — and that second axis moves quality more than the first. **Ask both.** They are independent: every cost mode (`minimize`, `hybrid`, `rich-ai`, `pure-videogen`) runs in either interaction mode.
|
|
131
132
|
|
|
132
|
-
- **interactive** — the user is willing to do a little manual work at fixed checkpoints, and the video gets better for it.
|
|
133
|
+
- **interactive** — the user is willing to do a little manual work at fixed checkpoints, and the video gets better for it. Three checkpoints cover nearly everything: **(1) images** — you write a prompt, they run it in a *free* frontier web generator (meta.ai / ChatGPT / Gemini / a Hugging Face Space) and hand the file back; **(2) raw clips** — you hand over search keywords, they search TikTok/YouTube, download a few with a free online downloader, and point you at the folder; **(3) the voice** — you sample a few narrators and they pick the one the video sounds like (costs them 30 seconds and $0, see below).
|
|
133
134
|
- **autonomous** — you finish end-to-end with zero steps from them: source clips yourself (browser control → `raws scan` → public raws), generate within the budget, or do without.
|
|
134
135
|
|
|
135
136
|
**Why interactive usually wins on quality:** the free tiers of the frontier web image models are typically *better* than what an API-key budget buys per image, and a human eye picks better footage than any keyword scan. In `minimize` the gap is not incremental — it's the difference between **no custom art at all** and **a full sticker pack for $0**.
|
|
@@ -147,6 +148,10 @@ Cost mode answers *how much money may I spend*. It does not answer *how much of
|
|
|
147
148
|
|
|
148
149
|
**In interactive mode, MANUAL IMAGE WORK DEFAULTS TO STICKER PACKS.** Never ask for one graphic per round trip — each hand-off costs the user a context switch and costs you tokens re-reading a file. Ask for **one sheet holding every graphic**, then split it locally for $0. `vidfarm handoff image --theme "<what>" --items "a,b,c"` mints the whole brief (prompt + steps + the free tools + the follow-up command); `--single` when you really do want one subject. When the file comes back: `vidfarm sticker-pack <sheet> --items "a,b,c"`.
|
|
149
150
|
|
|
151
|
+
**In interactive mode, OFFER THE VOICE CHECKPOINT — in every cost mode.** Who the video sounds like is a taste decision, and the default voice is the one choice agents make silently that a director almost always wants a say in. Before narrating, ask *"want to hear a few voices and pick one?"* and sample: `vidfarm voices --sample` (premium) or `vidfarm voices --free --sample` ($0 local). **Sampling costs nothing on either tier** — premium samples are ElevenLabs' own preview clips (a CDN download, not a synthesis call) and free samples render locally — so this checkpoint is just as available in `minimize` as in `hybrid`. Play the files, take their pick, narrate with `--voice <id>`. `vidfarm tts` prints the same nudge on stderr whenever narration would run with no voice named and the mode is interactive (or was never set).
|
|
152
|
+
|
|
153
|
+
**And say where the premium voices come from, because users assume wrong.** The full ElevenLabs catalog is reachable **through vidfarm's own ElevenLabs connection** — no ElevenLabs account, API key, or subscription on the user's side; narration just spends **vidfarm wallet credits** (pennies each). In `hybrid` that is a real option to put on the table next to the free voices, not a locked door. `--own-key` is only for users who already have an ElevenLabs key and would rather bill their own account.
|
|
154
|
+
|
|
150
155
|
**Raw-clip sourcing has a ladder — hand it to the human only at the bottom rung.** In order: (1) **your own browser control**, if you have it — drive the search and download yourself; (2) **Vidfarm cloud** — `vidfarm clipper <url>` / `vidfarm raws scan <url> --cloud` resolves and mines the video for you; (3) the **free public raws catalog** — `vidfarm public-raws --category <shelf>`; (4) **the user**, when you're fully local/keyless or when human taste matters. That last rung is `vidfarm handoff raws --keywords "villa construction,pouring concrete" --platforms tiktok,youtube --purpose "<what the clips are for>"` — it prints the keywords, tells them to google a downloader (a *search*, not a link that rots), and names the import command for when the folder is ready (`vidfarm clipper ./downloads/<file>.mp4`, or `vidfarm raws scan` to mine a long one).
|
|
151
156
|
|
|
152
157
|
## Default stance
|
|
@@ -204,16 +209,20 @@ Present both harnesses to the director, recommend (A) unless they've asked for p
|
|
|
204
209
|
|
|
205
210
|
**Explainer/cutout videos — the transparent-sticker workflow.** For explainers (a subject "on stage" while labels, arrows, and props pop in around it), the cheap workhorse is a **transparent cutout sticker**: `vidfarm cutout --generate "<subject>"` AI-generates the graphic on a chroma plate, keys it out, **and trims the canvas down to the subject's true min width/height** — one free local ffmpeg step, no mostly-empty PNG to fight with — then `vidfarm place` + `vidfarm keyframes` scale/position and animate it (zoom, grow, shake, drift). It's the same "generate a reusable element once, then reuse it" thrift as the cheap harness, tuned for stickers. Full guided harness: recipe `recipes/cutout-graphics-for-explainers.md`; placement + zoom/grow/shake/move motion recipes: `references/editor-workflows.md` (“Cutout graphics for explainers”).
|
|
206
211
|
|
|
207
|
-
**"Make me a sticker pack" = ONE greenscreen sheet of many items, then masked apart — and `vidfarm sticker-pack` is that whole loop.** A sticker pack is never one graphic; it's a *set* (props, icons, reactions, characters, backdrops) that must share one art style. Generating them one at a time is both expensive (N image jobs) and inconsistent (N independent styles), so the move is the opposite: **generate a single image holding every item, laid out on a flat greenscreen plate, then cut each item out locally for $0.** `vidfarm sticker-pack --generate "<theme>" --items "a,b,c"` does all of it — one billed image job for the whole set, then a free local key, an **automatic** alpha-segmentation that finds each item (no hand-measured `--crop` rects), a per-item trim to its true bounding box, and a `stickers.json` manifest. Already have a greenscreen sheet? `vidfarm sticker-pack ./sheet.png` cuts it up for **$0**. Use `--dry-run` to eyeball the detected boxes first; `--gap` merges/splits items that came out joined or broken; `vidfarm mask <sheet> --crop …` is the manual fallback for one stubborn item.
|
|
212
|
+
**"Make me a sticker pack" = ONE greenscreen sheet of many items, then masked apart — and `vidfarm sticker-pack` is that whole loop.** A sticker pack is never one graphic; it's a *set* (props, icons, reactions, characters, backdrops) that must share one art style. Generating them one at a time is both expensive (N image jobs) and inconsistent (N independent styles), so the move is the opposite: **generate a single image holding every item, laid out on a flat greenscreen plate, then cut each item out locally for $0.** `vidfarm sticker-pack --generate "<theme>" --items "a,b,c"` does all of it — one billed image job for the whole set, then a free local key, an **automatic** alpha-segmentation that finds each item (no hand-measured `--crop` rects), a per-item trim to its true bounding box, and a `stickers.json` manifest. The sheet comes in two shapes: **zoned** (a grid of color panels, one plate color per item — the default for 2+ named items, and what frees the art from a single banned hue) and **flat** (the classic one-color plate, for a model that can't follow a color-block grid). Already have a greenscreen sheet? `vidfarm sticker-pack ./sheet.png` cuts it up for **$0**. Use `--dry-run` to eyeball the detected boxes first; `--gap` merges/splits items that came out joined or broken; `vidfarm mask <sheet> --crop …` is the manual fallback for one stubborn item.
|
|
208
213
|
|
|
209
214
|
- **Stickers are not necessarily small.** A sticker is *any* transparent element you place and animate — an icon, a mascot, a prop, a character, and equally **a full-width landscape, skyline, or backdrop** that fills the frame. `sticker-pack` filters speckle only; it has no maximum item size. Ask for the big pieces in the same sheet as the small ones.
|
|
210
215
|
- **Stickers are usually animated, not pasted.** Once placed, animate each one with `vidfarm keyframes` presets (`pop-in`, `float`, `shake`, `grow`, `slide-in-left`, `drift`) — that's HTML/CSS canvas motion, deterministic, free, and identical in preview and render. Layer moves up (pop-in, then idle float) for real life. See `references/editor-workflows.md` → "Cutout graphics for explainers".
|
|
211
216
|
- **A sticker can carry its OWN motion too.** A *moving* subject has no single bounding box, so it isn't a PNG: key the clip with `vidfarm remove-greenscreen <video>` → transparent WebM (browser/editor-playable, the right choice on a composition).
|
|
212
|
-
- **The
|
|
213
|
-
- **
|
|
217
|
+
- **The key is CONNECTIVITY-based, so "the art can't use the plate color" is no longer true — only its OUTER EDGE can't.** `sticker-pack`/`cutout` (and `remove-greenscreen <image> --smart`) don't delete every pixel that looks like the plate. They flood-fill the plate **inward from the edge of the sheet** and delete only background that **reaches** that edge. A green leaf inside a mascot, a plate-colored eye, an outline shape whose interior was left as bare plate — none of it is reachable, so none of it is deleted. Edges are feathered and the plate is **un-mixed out of each edge pixel individually** (real alpha math, not a global `despill`), which is what kills the green fringe a flat key leaves. What still matters: the item's **silhouette** must be a different color from its own plate, and nothing may **fade** into the plate (no soft glow, blur or drop shadow on the background). Say the win out loud when it matters — the console prints *"kept N plate-colored pixels INSIDE the art that a flat key would have punched out."* `--key-mode flat` restores the old plain-chromakey behaviour (the cloud path's exact filter chain — use it to reproduce a cloud render, or as a simple fallback).
|
|
218
|
+
- **Plate color is still chosen for you, and it still matters for the silhouette.** When generating, `sticker-pack`/`cutout` read the subject and move the plate off any hue it mentions — green → magenta (`#FF00FF`) → blue (`#0047BB`) → black → white — printing which plate they picked and why. When splitting a sheet you already have, they **read the plate off the sheet itself**, so a red/purple/blue sheet from a web generator just works. Pin it with `--key-color "#FF00FF"` / `--preset magenta`, or `--no-auto-key` for plain green.
|
|
219
|
+
- **ONE PLATE COLOR PER STICKER — `--sheet-mode zoned`.** The real fix for "our art has to be simple because of the greenscreen" is to stop giving a whole sheet one background. A **zoned** sheet is a grid of solid color **panels**, one item per panel, each panel's plate chosen against **that item**: a green frog on magenta beside a pink flower on green, in one image job. Each panel is keyed independently with its own color (read back off that panel's own corners, because models drift the hue they were asked for), and item art may then use **any palette at all — including the color of a different panel**. Bonus: names stop being guessed from reading order — panel N holds the item you asked for in panel N, so `--items` maps exactly, and `stickers.json` records each sticker's `panel` and `plate`. `--sheet-mode auto` (the default) zones a generation of 2+ named items and stays flat otherwise. Reading a zoned sheet you already have: `--zones auto` (default — recovers the grid from the sheet's own edges) or `--zones 3x2`.
|
|
220
|
+
- **The FLAT single-color sheet is still first-class — use it for a weaker image model.** Not every model can hold a color-block grid; a cheap or small one will paint one background regardless of the prompt. That path is fully supported: `--sheet-mode flat` asks for the classic one-plate sheet, and even if you asked for zones, keying **detects a model that ignored the grid** (a panel with no plate to remove, or one whose item filled it corner to corner) and **automatically re-keys the sheet as one plate**, telling you it did. So zoning can't strand you — worst case you're back on the simple method, still for $0.
|
|
221
|
+
- **Painterly, soft, furry or glassy art → `--refine`.** Chroma keying of any kind needs a crisp silhouette. When the art has genuinely soft edges (watercolor, fur, glow, glass, a cast shadow), add `--refine`: after the keyer locates each item, that item is re-cut from the **un-keyed** sheet with the local ONNX matting model (free, ~1–2s each), which is color-blind and handles soft mattes. It sanity-checks each matte and falls back to the keyed cut per item if the model didn't find a subject — flat vector art is exactly the case where matting fails and the chroma cut is better, so don't reach for `--refine` by default.
|
|
222
|
+
- **What the ART still has to honor: a crisp silhouette, sealed shapes, and gaps between items.** The connectivity key removes the old bans on hollow art and plate-colored fills, but three prompt rules are still load-bearing: **(1)** every object's **outer edge** is a clearly different color from its own plate and is **crisp** — no glow, blur, mist or drop shadow fading into the background; **(2)** enclosed areas are **sealed by the artwork**, because a gap in an outline lets the background flow in and the fill really does get keyed; **(3)** items are separated by a clear margin of plate — two touching items segment as ONE sticker. `cutout --generate`, `sticker-pack --generate`, `handoff image` and the `create-overlay` primitive **append the right clause for you** (a relaxed one for the smart keyer, the strict "closed, solidly filled, nothing in a near-plate shade" one when `--key-mode flat` is in play) — write it yourself only when prompting a generator directly. Both commands still report per-item `hole_pct`/`hollow` (console `⚠ N% hollow`, `--json`, `stickers.json`); under the smart keyer a flagged item is *usually real* (a ring, frame, donut, or pieces with background between them), so it **warns, never blocks**. One stubborn item can always be lifted with `vidfarm mask --crop …` (ONNX matting ignores color entirely).
|
|
214
223
|
- **Transparent GIF is supported, for GIF-only surfaces.** `vidfarm sticker-pack … --output-format gif` (stills) and `vidfarm remove-greenscreen <video> --gif` (animated) emit transparent GIFs. GIF alpha is **1-bit**, so edges go hard — fine for chat/forum/Notion sticker surfaces, worse than PNG/WebP/WebM for compositing on a timeline. Prefer PNG/WebP/WebM unless the destination only eats GIF.
|
|
215
224
|
|
|
216
|
-
**Explainer house style — the defaults to build with unless told otherwise.** **White background / light mode** (plain white stage, no gradients, no dark mode, no photo backdrop), **kinetic word-by-word captions** in dark ink on the light stage (`vidfarm captions generate --style word-pop --color "#111111" --active-color "#7C3AED" --background-style plain` — skip outlines/shadows, they're only needed over busy footage), and **female TTS narration** (`vidfarm tts --voice coral` on OpenAI — `nova` for energy, `sage` for calm; `Kore`/`Leda` on Gemini; any ElevenLabs voice via `vidfarm voices`). **Keep it clean and simple** — one idea on screen at a time, two or three cutouts per beat, one accent color, one font, lots of white space; remove before you add. **Illustrations default to simplicity**: flat vector, simple shapes, minimal detail, 2–3 flat colors, no baked-in text — simple art keys cleanly, trims tight, and stays on-style across the whole cast. State the defaults once so the director can override any of them. Full detail: recipe `recipes/cutout-graphics-for-explainers.md` (“House style — the explainer defaults”). **If the director takes the stage off white**, two things stop being optional: every sticker's **white die-cut rim** has to be stripped (on a dark stage it's a glaring halo and the most obvious bot-made artefact in the frame — recipe → “Stickers on a DARK or photographic stage”), and the caption hexes above stop applying — **caption colour, active-word colour and plate are chosen by measuring the composited background behind the caption band**, one treatment per video (`
|
|
225
|
+
**Explainer house style — the defaults to build with unless told otherwise.** **White background / light mode** (plain white stage, no gradients, no dark mode, no photo backdrop), **kinetic word-by-word captions** in dark ink on the light stage (`vidfarm captions generate --style word-pop --color "#111111" --active-color "#7C3AED" --background-style plain` — skip outlines/shadows, they're only needed over busy footage), and **female TTS narration** (`vidfarm tts --voice coral` on OpenAI — `nova` for energy, `sage` for calm; `Kore`/`Leda` on Gemini; any ElevenLabs voice via `vidfarm voices`). **Keep it clean and simple** — one idea on screen at a time, two or three cutouts per beat, one accent color, one font, lots of white space; remove before you add. **Illustrations default to simplicity**: flat vector, simple shapes, minimal detail, 2–3 flat colors, no baked-in text — simple art keys cleanly, trims tight, and stays on-style across the whole cast. State the defaults once so the director can override any of them. Full detail: recipe `recipes/cutout-graphics-for-explainers.md` (“House style — the explainer defaults”). **If the director takes the stage off white**, two things stop being optional: every sticker's **white die-cut rim** has to be stripped (on a dark stage it's a glaring halo and the most obvious bot-made artefact in the frame — recipe → “Stickers on a DARK or photographic stage”), and the caption hexes above stop applying — **caption colour, active-word colour and plate are chosen by measuring the composited background behind the caption band**, one treatment per video (`harnesses/short-form.HARNESS.md` → “Caption styling is MEASURED off the background”). Related: **on-screen text and captions must not say the same thing at once** — display text carries the argument, captions carry only what the screen doesn't show.
|
|
217
226
|
|
|
218
227
|
**Landscape footage in a fullscreen vertical explainer — use the blurred plate, never bars.** When an explainer is built on **real filmed footage** and the source is 16:9 (or 4:3) on a 9:16 canvas, do not `contain` it (hard black letterbox bars read as an unfinished export) and do not blindly `cover` it (a wide shot loses its left and right thirds). Duplicate the clip: a full-canvas `cover` copy behind, heavily **gaussian-blurred and faded dark**, plus the sharp copy centered as a hero band — optionally zoomed ~1.3× — with its **top and bottom edges feathered** into the blur. Same clip, same timecode, so it reads as one continuous image with a shallow-depth-of-field plane, fullscreen edge to edge, nothing cropped, and clean dark space for the header and captions. Bake it once with ffmpeg into a single 1080×1920 file (free, local) and place it as one ordinary full-canvas layer — layer blur is not an editor property, so the pre-bake is the path that works in the editor, `serve`, and cloud render alike. Copy-paste ffmpeg + HTML recipes, tuning table, and the failure modes: `references/editor-workflows.md` (“The blurred plate — landscape footage, fullscreen, on a vertical canvas”).
|
|
219
228
|
|
|
@@ -282,32 +291,63 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
|
|
|
282
291
|
- **Determine the surface before claiming capabilities.** The web chat has only its declared tools and REST routes. It cannot execute arbitrary JavaScript/Python, open a shell, create a local repository script, or use the user's filesystem. Never tell a web-chat user that you ran code or wrote a script unless a dedicated declared tool actually did so. A desktop coding agent has a real shell and filesystem and MAY write/run scripts, perform arbitrary local computations over paginated API results, create reports/CSVs/JSON, edit composition files, and orchestrate long devcli workflows within the user's authorization.
|
|
283
292
|
- **The web AI chat can do all three paintbrushes** — clip raws, author HTML/hyperframe motion, and generate AI media — and it drives edits directly on the live timeline. Keep small-to-medium jobs here: text/caption swaps, a scene or two replaced, single generations, captions, approve/schedule. Just do them.
|
|
284
293
|
- **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. **The test is the native-editor test: could you have made this element with the tools inside TikTok's own editor?** That toolset is font / color / stroke / shadow / tight text box / alignment / rotation / animation presets, plus stickers, emoji, drawn marks and clips — it has no padded capsule, no border, no gradient fill, no blur panel, no card. If you reached past it, cut it. That includes **a single lonely pill around a stat or label** — `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`: being alone doesn't make a badge native, and the only legitimate capsule in a video is the active-word `spotlight`/`karaoke` highlight, which moves with the spoken word. Emphasize a stat the way the editor would instead: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle, or its own beat. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r` — or a `border-radius` over ~8px on anything filled that holds words — stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
|
|
285
|
-
- **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm
|
|
294
|
+
- **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm harness show hooks`.
|
|
295
|
+
- **Then CUT it — every second must earn its place, and most don't.** Assume your first assembly is **30–50% too long**. Run the **deletion test** on every beat: delete it; if the video still makes sense and the payoff still lands, it stays deleted. Whatever survives must serve one of the four charges — "it gives context" is not a charge. Cut on sight: intros/logo stings, the wind-up sentence before the claim ("so I wanted to talk about…"), restatement, inter-sentence silence over ~0.35s, real-time process, establishing shots, reading what's already on screen, and any tail after the last word. **Always ripple the hole closed** (`vidfarm ripple <dir> --at <sec> --delta -<sec>`) — a cut that leaves a gap turns fluff into dead air, which is worse. Density is **not** speed: the held comedic beat, the payoff playing out, and a cue's readability keep their seconds (cut *words*, not the time text is on screen). Length is an **output**, not a plan — a brief that dictates a duration ordered fluff. `vidfarm qa` flags the mechanical half (`dead-air`, `dead-tail`, `slow-scene`); the craft is `references/hooks-and-virality.md` → "Density".
|
|
286
296
|
- **The first frame IS the thumbnail — compose it on purpose.** Frame 0 is a single frame of ~30 in the first second, but it's the poster every feed, share link, and paused player freezes on, so **more people see that one frame than watch the video**. It must never be black, empty, mid-fade, or caught mid-animation: put a real visual at `start:0`, have the hook words already on screen at t=0, and never hang a `fade-black`/`fade-white`/`flash` *entrance* on the **first** clip (junction transitions between later clips are fine — this rule is only about the opening). Verify it, don't assume: devcli `vidfarm stills ./work --at 0` renders that exact frame, and `vidfarm qa` flags a blank or fading open.
|
|
287
|
-
- **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`
|
|
288
|
-
-
|
|
297
|
+
- **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`HARNESS.md`** — because a loop of fifty videos has no human looking at every frame, and the harness is what replaces those eyes. Don't silently ship a one-off when they asked for volume, or drag someone into a harness when they wanted one clip.
|
|
298
|
+
- **"Harness" is a known noun with a known process — recognise it and follow it.** A **harness** is the reusable AI apparatus for ONE format or template: what makes it special, written down as `HARNESS.md` so an agent can reproduce it without the director in the room. It is a first-class artifact — the director owns it, edits it, versions it, and hands it to the next agent. Three phrasings, one artifact:
|
|
299
|
+
- **"create me a harness"** / "set up a harness for this format" → `vidfarm harness init <base> --out ./work/HARNESS.md` (bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo`), then **edit it with them**. The bundled file is a starting point, never a house style; the parts that matter are the ones they add — who the viewer is, the banned vocabulary, the compliance line, the pacing this account actually uses. A harness nobody edited isn't about their videos.
|
|
300
|
+
- **"update the harness for this format/template"** → open the existing `HARNESS.md` and write the new rule in, **with its reason on the same line** (a rule whose "why" is missing gets argued away by the next agent). This is what you do every time a batch teaches you something ("the label-framed hooks all died"): the compositions are disposable, the harness is the artifact that compounds.
|
|
301
|
+
- **"give me the harness for this template_id"** → they mean **the decomposition**: `vidfarm harness derive <templateId|forkId>`. It distils the decompose pass's DNA into an editable `HARNESS.md`. If the template hasn't been decomposed, run `vidfarm decompose` first.
|
|
302
|
+
**A harness mirrors the template JSON's own vocabulary** — `## Viral DNA` (hook / retention / payoff / emotion), `## Visual DNA` (cut rhythm, typography, b-roll, transitions), `## Structural DNA` (the beats, and which are load-bearing), `## Audio DNA` (voice, bed, comedic timing), `## Build DNA` (which paintbrush per beat) — the same strands the decompose pass writes as `viral_dna`, `visual_dna`, and friends. `vidfarm harness show <ref> --dna visual` prints one strand instead of the whole doc.
|
|
303
|
+
**Two halves, and only one is machine-checkable.** The `checks:` front matter is settled deterministically by `vidfarm qa` (duration, aspect, `hook_words_max`, `forbid_text`, …); every `- [ ]` line comes back as a **review item you answer honestly in your report** — never claim a video passed the half the CLI can't judge. Harnesses stack and auto-discover: `vidfarm qa ./work` picks up `./work/HARNESS.md`, `--harness hooks --harness ./brand/HOUSE.md` adds more, and any file of theirs anywhere is valid. Format and strand table: `harnesses/README.md`; scripting-mode detail: `references/automation-and-local-dev.md`. *(Formerly `QA_REGIME.md` — same file, and `vidfarm regime …` still works as an alias.)*
|
|
289
304
|
- **A video is judged as a SEQUENCE, so review it as one.** Agents build scene by scene and each scene passes in isolation while the video drifts — inconsistent margins, three type sizes, an accent colour that wanders, beats that are all the same length, a jarring join. Tile a dozen stills into one contact sheet (`vidfarm stills ./work --sheet`) and read it as an image before you call anything done, fix drift by defining the system rather than patching the odd scene out, and remember that **your own confident "verified, looks good" is the single least reliable signal in this workflow** — it was wrong on every video of a 32-video batch. Method: `references/reviewing-renders.md`.
|
|
290
305
|
- **On devcli, QA every video you produce: `vidfarm qa ./work`.** Free, instant, local-only — it blocklists exactly the slop above plus first-frame/thumbnail and font-regime/safe-zone drift, and prints a concrete fix per finding. **Feedback, not a gate**: it exits 0 even on findings, never runs automatically, and is a blocklist (unusual/stylized compositions pass untouched), so there's no reason not to run it before every publish. `--json` for scripted batches, `--strict` only if you want a CI failure. **Web-chat copilot: this command does not exist for you** (devcli-only, no REST twin) — apply the standard by hand, and when handing a heavy job to a local coding agent, tell them to run `vidfarm qa`.
|
|
291
|
-
- **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
|
|
306
|
+
- **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips) and, inside that band, is **placed in the emptiest part of the frame** rather than dumped on the default lower third (read a still first — `vidfarm stills ./work --at <t>`; words over open sky or a blank wall beat words over the subject, and usually need no plate at all). Long narration is **paged into 3–5-word kinetic cues** (`vidfarm captions generate --style word-pop`), never one static wall of text. It uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
|
|
292
307
|
- **Where the web chat struggles: complex, long, multi-step transformations.** A full multi-scene re-theme, an iterative render-critique-iterate loop, heavy scripted or batch work, or anything needing a real filesystem and many sequential tool calls will hit context limits, turn/timeout ceilings, and the web editor's constraints (CSS/declarative motion only — JS animation adapters are stripped on save). Don't grind a big transformation one layer at a time in a chat turn and stall.
|
|
293
308
|
- **Practical workaround — hand the heavy job to local devcli.** When a task is genuinely large or long-running, **proactively recommend the director run it locally with an AI coding agent** (Claude Code / OpenAI Codex / any capable agent): `vidfarm pull <forkId>` writes the composition + the `.harness/` grounding bundle to disk, the agent edits with the full devcli verb set and JS animation adapters, renders free with `vidfarm serve`, and `vidfarm publish` pushes it back. This is the **best-quality (B) harness's** natural home (adversarial grading with a coding agent). Frame it as "this is a big rebuild — you'll get a better, faster result running it locally with a coding agent; here's how," not as a dead end.
|
|
294
309
|
- **Offer a handoff, do not impersonate the desktop agent.** When web chat reaches that boundary, offer to save a Markdown handoff in My Files containing the objective, selected template/fork IDs, asset paths, grounding, constraints, completed work, and suggested devcli commands. Create it only after the user agrees. The desktop agent should read that document, pull the referenced fork, and then use its actual code/shell capabilities.
|
|
295
310
|
- **Never send the user away just to read knowledge.** Deeper skill knowledge is always a **tool call** away in-place: call `load_skill` (e.g. `load_skill('vidfarm', file='references/editor-workflows.md')`, or a craft pack like `editor-capabilities` / `hyperframes-animation`) to pull the exact reference you need mid-conversation. Only recommend switching surfaces for the WORK (a heavy transformation), never for the information.
|
|
296
311
|
|
|
297
|
-
##
|
|
312
|
+
## File Index — everything in this pack, and when to read it
|
|
313
|
+
|
|
314
|
+
**This is the complete inventory. Nothing else exists in the pack, and every file here is reachable by name.** Read the narrowest file that answers the question; never preload several. `size` is a context-cost estimate — the four big references are real reads, so pick one deliberately rather than sweeping them.
|
|
298
315
|
|
|
299
|
-
|
|
316
|
+
**References — the broad knowledge domains**
|
|
317
|
+
|
|
318
|
+
| File | Size | Read it when |
|
|
319
|
+
|---|---|---|
|
|
320
|
+
| `references/core-workflows.md` | ~360 ln | Template discovery, auth, fork → render → approve → share, versioning, cost/wallet, marketplace orders, dedupe-before-publish |
|
|
321
|
+
| `references/editor-workflows.md` | ~650 ln | **The biggest read.** Timeline editing, decompose, captions, transitions, motion, AI placement, the caption standard, the editor action verbs |
|
|
322
|
+
| `references/assets-and-sourcing.md` | ~185 ln | Raws hunts, clip scanning, My Files, recurring characters, downloading media off a URL, social recycle |
|
|
323
|
+
| `references/automation-and-local-dev.md` | ~520 ln | **Big.** The whole `vidfarm` command table, REST automation, scripting/bulk mode, `HARNESS.md`, local serve loop, skill packs |
|
|
324
|
+
| `references/primitives.md` | ~475 ln | **Big.** One-shot primitive routes: TTS, STT, music, avatars, overlays, greenscreen, inpaint, background removal, product placement |
|
|
325
|
+
| `references/hooks-and-virality.md` | ~295 ln | **Before writing ANY hook, caption script, or re-theme**, and before a hook-variant batch. The four charges, three gates, banned openers, loop mechanics. This is the craft; the rest of the pack is mechanics |
|
|
326
|
+
| `references/reviewing-renders.md` | ~140 ln | **Before you report a video as done**, or grade someone else's. The holistic pass, the common defects, frozen-render and audio verification |
|
|
327
|
+
| `references/onboarding.md` | ~30 ln | Cold-start interviews, **consultations** (the `brainstorm/*` chain), strategy docs, durable director context |
|
|
328
|
+
| `references/rest-api.md` | ~85 ln | Only when the user asks for REST, an endpoint/schema, or direct HTTP integration. It is an index — follow its domain links; do not preload it into ordinary director conversations |
|
|
329
|
+
|
|
330
|
+
**Recipes — step-by-step procedures. When a recipe matches the task, prefer it over the broad reference.**
|
|
331
|
+
|
|
332
|
+
| File | Size | Read it when |
|
|
333
|
+
|---|---|---|
|
|
334
|
+
| `recipes/find-and-fork-template.md` | ~15 ln | Template selection and the first fork |
|
|
335
|
+
| `recipes/retheme-template.md` | ~15 ln | Full re-theme that preserves the source format's feel |
|
|
336
|
+
| `recipes/local-edit-render-approve.md` | ~20 ln | The local pull → edit → render → approve loop |
|
|
337
|
+
| `recipes/onboard-a-new-director.md` | ~15 ln | New-director onboarding and durable context capture |
|
|
338
|
+
| `recipes/bulk-scripting-with-a-harness.md` | ~100 ln | **Volume**: daily posting, N variants, hook tests — scripting mode with a `HARNESS.md` |
|
|
339
|
+
| `recipes/cutout-graphics-for-explainers.md` | ~265 ln | Building an explainer from sticker/cutout art: the house style, sticker sheets, keying, dark-stage rules |
|
|
300
340
|
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
-
|
|
307
|
-
|
|
308
|
-
|
|
309
|
-
|
|
310
|
-
|
|
341
|
+
**Harnesses — the `HARNESS.md` format and its bundled bases.** All are readable as-is and copyable with `vidfarm harness init <name>`.
|
|
342
|
+
|
|
343
|
+
| File | Size | Read it when |
|
|
344
|
+
|---|---|---|
|
|
345
|
+
| `harnesses/README.md` | ~110 ln | **Start here for anything harness-shaped**: the three director phrasings, the format, the `checks:` key list, the DNA strand → decompose-JSON map |
|
|
346
|
+
| `harnesses/short-form.HARNESS.md` | ~225 ln | The default base. Also holds the **"Caption styling is MEASURED off the background"** procedure that other files point at |
|
|
347
|
+
| `harnesses/hooks.HARNESS.md` | ~120 ln | Hook-variant batches — chunk-1 legibility, the unguessable test, volume-only anti-patterns. The checkable form of `hooks-and-virality.md` |
|
|
348
|
+
| `harnesses/explainer.HARNESS.md` | ~100 ln | Faceless educational video: one claim, invented visuals |
|
|
349
|
+
| `harnesses/ugc-testimonial.HARNESS.md` | ~90 ln | A person vouching for a product — mostly rules about what NOT to add |
|
|
350
|
+
| `harnesses/product-demo.HARNESS.md` | ~110 ln | Real product doing a real thing; the highest slop-risk format in the catalog |
|
|
311
351
|
|
|
312
352
|
## HyperFrames Skills — Load on Demand
|
|
313
353
|
|
|
@@ -323,9 +363,9 @@ On the web copilot, call `load_skill('<name>')` and load referenced files only w
|
|
|
323
363
|
|
|
324
364
|
HyperFrames authoring and rendering in this package are Vidfarm-native: local work uses the bundled composition toolchain and `vidfarm serve`; cloud work uses Vidfarm render routes. Do not require an external vendor account, repository, publish service, or telemetry endpoint. Keep `HYPERFRAMES_SKIP_SKILLS=1` and `HYPERFRAMES_NO_TELEMETRY=1` in Vidfarm-managed environments so the bundled skills stay pinned and local work does not phone home.
|
|
325
365
|
|
|
326
|
-
## Quick Router
|
|
366
|
+
## Quick Router — from what the user said to what to open
|
|
327
367
|
|
|
328
|
-
Choose the narrowest path that satisfies the request.
|
|
368
|
+
The File Index above says what each file *is*; this says which one a given ask means. Choose the narrowest path that satisfies the request.
|
|
329
369
|
|
|
330
370
|
1. If the user needs help figuring out what to make, **or asks for a "consultation"** (the `brainstorm/*` chain: cold-start interview → awareness stages → angles → hooks), read `references/onboarding.md` first.
|
|
331
371
|
2. If the user already knows the goal and needs a suitable template, read `references/core-workflows.md` and use the template discovery flow.
|
|
@@ -334,7 +374,9 @@ Choose the narrowest path that satisfies the request.
|
|
|
334
374
|
4b. If the task is **“download this video/audio off a website”** (a pasted YouTube / TikTok / Instagram / X post URL the user wants the actual file from), Vidfarm does that for you on a **paid plan** — `POST /api/v1/primitives/videos/download` (or `/audio/download`), devcli `vidfarm download-video <url>` / `vidfarm download-audio <url>`. **Free plan → do not call it; walk the user through opening the URL in Chrome and downloading it from the page, then `vidfarm put-file` the local file in for $0.** Details in `references/assets-and-sourcing.md` → `references/primitives.md`.
|
|
335
375
|
4c. If the task is **“turn this Reddit/X thread, subreddit, or account into a video”** — “tweet to TikTok”, “Reddit to TikTok”, “make a video from this thread”, “what are the top comments saying” — run `vidfarm recycle <source>` (or `POST /api/v1/primitives/social/recycle`) with the URL. It **decomposes** the source into raw JSON (text, comment tree, media URLs, author pics, stats) and hands it back unranked so YOU pick what to remix. **Paid plan; `max_records` is the spend ceiling.** Brokers the reddit-lead-gen / x-lead-gen OfficeX apps, so it waits out their async job for you. Details in `references/assets-and-sourcing.md` → `references/primitives.md`.
|
|
336
376
|
4d. If the task is **“post this again / to several accounts / on another platform”**, or you are about to publish or bulk-produce at all — that is **deduplication**. Run `vidfarm dedupe <mp4> [--variants N]` on the **exported file** (free, local ffmpeg, no re-render), then approve/schedule each variant. **Ask the operator whether they want deduplicated copies, and how many, BEFORE the render/bulk run** — deciding after means paying for a second render. Details in `references/core-workflows.md` → *Deduplicate before you publish* and `references/primitives.md` → *Primitive: media_dedupe*.
|
|
377
|
+
4e. If the ask contains the word **“harness”** — *“create me a harness”*, *“update the harness for this format”*, *“give me the harness for this template_id”* — that is a known, named process, not a vague request. Read `harnesses/README.md` (the three phrasings and the format), then `recipes/bulk-scripting-with-a-harness.md` if the job is a batch. The third phrasing means the **decomposition**: `vidfarm harness derive <forkId>`.
|
|
337
378
|
5. If the task is scripted, local, CI-driven, or `vidfarm serve`-based, read `references/automation-and-local-dev.md`.
|
|
379
|
+
5b. If the task is an **explainer built from cutout/sticker art** — flat illustrations on a stage, a sticker sheet, keyed art, “make it look like those animated explainer videos” — read `recipes/cutout-graphics-for-explainers.md`. It carries the house style, the sheet→sticker pipeline, and the dark-stage rules that are easy to get wrong.
|
|
338
380
|
6. If the task explicitly asks for a primitive or needs specialized generation/transcription work, read `references/primitives.md`.
|
|
339
381
|
7. If the task is the MARKETPLACE (ordering videos from specialist agents): browsing is web-only for paying customers — send the human to https://vidfarm.cc/marketplace, never render it locally. Placing/listing orders is the thin REST wrapper in `references/core-workflows.md` (§ Marketplace); anything deeper on a gig (inbox, proofs, payouts) needs the external Dollar Platoon skill — `npx skills add https://github.com/OfficeXApp/dollarplatoon-skill` — the same way FlockPoster work beyond scheduling needs `npx skills add https://github.com/OfficeXApp/flockposter-skill`.
|
|
340
382
|
|
|
@@ -350,18 +392,9 @@ Choose the narrowest path that satisfies the request.
|
|
|
350
392
|
- **Never render or approve without judging frame 0 as a standalone still.** It is the thumbnail everywhere the post appears; an empty/black opening frame ships a dead post. See “The FIRST FRAME is the thumbnail”.
|
|
351
393
|
- **Never judge the VIDEO by one frame, and never report a render as reviewed without the holistic pass.** Compare frames from at least two different scenes (a frozen render passes every other check), read a contact sheet for balance/spacing/style/pacing drift, and state separately what you measured vs. what you judged. See “Judge the WHOLE video”.
|
|
352
394
|
|
|
353
|
-
## Recommended Recipes
|
|
354
|
-
|
|
355
|
-
Use these when the user’s task matches the pattern closely.
|
|
356
|
-
|
|
357
|
-
- Template selection and first fork: `recipes/find-and-fork-template.md`
|
|
358
|
-
- Full re-theme while preserving the format’s feel: `recipes/retheme-template.md`
|
|
359
|
-
- Local pull/edit/render/approve loop: `recipes/local-edit-render-approve.md`
|
|
360
|
-
- New-director onboarding and durable context capture: `recipes/onboard-a-new-director.md`
|
|
361
|
-
|
|
362
395
|
## Output Posture
|
|
363
396
|
|
|
364
397
|
- Prefer concrete actions over abstract discussion.
|
|
365
398
|
- Name the chosen path explicitly: template reuse, raws hunt, local serve, cloud render, etc.
|
|
366
399
|
- Surface cost tradeoffs before expensive generation.
|
|
367
|
-
- When in doubt between a broad reference and a recipe, start with the recipe.
|
|
400
|
+
- When in doubt between a broad reference and a recipe, start with the recipe — the File Index marks which is which.
|
|
@@ -0,0 +1,112 @@
|
|
|
1
|
+
# HARNESS.md — the reusable AI harness for one format or template
|
|
2
|
+
|
|
3
|
+
A **harness** is the written apparatus that lets an agent reproduce what makes a format good, over and over, without the director in the room. It is the unit this product is organised around, and three different director phrasings all mean it:
|
|
4
|
+
|
|
5
|
+
| They say | They mean | You run |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| "create me a harness" | a new one for a format we're about to produce at volume | `vidfarm harness init <base> --out ./work/HARNESS.md`, then **edit it with them** |
|
|
8
|
+
| "update the harness for this format/template" | edit the existing `HARNESS.md` — a rule learned, a beat that changed, vocabulary that died | open the file, add the rule **with its reason**, re-run `vidfarm qa` |
|
|
9
|
+
| "give me the harness for this template_id" | **the decomposition** — what makes *that* template special, distilled | `vidfarm harness derive <templateId\|forkId>` |
|
|
10
|
+
|
|
11
|
+
All three produce the same artifact: one Markdown file the director owns, edits, versions, and hands to the next agent. A derived one is still a first draft — the decompose pass watched the video, it didn't talk to the customer.
|
|
12
|
+
|
|
13
|
+
> Formerly called `QA_REGIME.md`. Same file, better name — a harness is not only a QA pass, it's the whole reproduction kit. `vidfarm regime …` still works as an alias and existing `QA_REGIME.md` files are still auto-discovered, but nothing writes that name any more.
|
|
14
|
+
|
|
15
|
+
## Using one
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
vidfarm harness list # what ships with the CLI
|
|
19
|
+
vidfarm harness show hooks # read one
|
|
20
|
+
vidfarm harness show ./work/HARNESS.md --dna visual # ONE strand, not the whole doc
|
|
21
|
+
vidfarm harness init short-form --out ./work/HARNESS.md # copy it next to your work, then EDIT it
|
|
22
|
+
vidfarm harness derive <forkId> --out ./work/HARNESS.md # a decomposed template → a harness
|
|
23
|
+
|
|
24
|
+
vidfarm qa ./work # auto-uses ./work/HARNESS.md if present
|
|
25
|
+
vidfarm qa ./work --harness hooks # a built-in by name
|
|
26
|
+
vidfarm qa ./work --harness ./brand/HOUSE_RULES.md # any file, anywhere
|
|
27
|
+
vidfarm qa ./work --harness short-form --harness ./work/HARNESS.md # they STACK
|
|
28
|
+
vidfarm qa ./work --json # checks + review items, for a scripted batch
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Harnesses compose: a shared house harness plus a per-campaign one is the intended shape. `--no-harness` skips auto-discovery; `VIDFARM_HARNESS` sets a default for a whole scripting run. `vidfarm harness check <dir>` is the same grader under the noun the director used.
|
|
32
|
+
|
|
33
|
+
## The format
|
|
34
|
+
|
|
35
|
+
Plain Markdown, with three machine-readable affordances:
|
|
36
|
+
|
|
37
|
+
**1. Optional front matter with a `checks:` block** — the assertions the CLI settles deterministically from the composition, instantly, with no AI and no network:
|
|
38
|
+
|
|
39
|
+
```markdown
|
|
40
|
+
---
|
|
41
|
+
name: my-house-style
|
|
42
|
+
video_type: what this harness is for
|
|
43
|
+
derived_from: decompose # set by `harness derive`; absent on hand-written ones
|
|
44
|
+
source_template_id: tpl_... # ditto
|
|
45
|
+
checks:
|
|
46
|
+
duration_sec: 8-34 # also "<=34", ">=8", or "30"
|
|
47
|
+
aspect: 9:16 # "9:16|1:1" to allow several
|
|
48
|
+
first_frame_visual: required
|
|
49
|
+
first_frame_text: required | forbidden
|
|
50
|
+
hook_words_max: 7
|
|
51
|
+
text_by_sec: 1.0
|
|
52
|
+
audio: required | forbidden
|
|
53
|
+
captions: required
|
|
54
|
+
font_regime: required
|
|
55
|
+
safe_zone: required
|
|
56
|
+
scenes: 3-12
|
|
57
|
+
max_scene_sec: 8
|
|
58
|
+
max_text_cards: 3
|
|
59
|
+
max_simultaneous_text: 2
|
|
60
|
+
max_words_per_cue: 12 # longest single text run — the wall-of-text dial
|
|
61
|
+
max_dead_air_sec: 2.5 # widest hole between cues with nothing to read
|
|
62
|
+
max_tail_sec: 1.5 # screen time still running after the last word
|
|
63
|
+
forbid_text: ["link in bio", "comment below"]
|
|
64
|
+
require_text: []
|
|
65
|
+
---
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
Unknown keys are reported and ignored, never silently dropped.
|
|
69
|
+
|
|
70
|
+
**2. Any `- [ ]` checkbox line** in the body becomes a **review item** — a question handed back for the agent or the human to answer. "Is the withheld answer one the viewer can't supply themselves?" is a judgment call; pretending a linter settles it would be a lie.
|
|
71
|
+
|
|
72
|
+
**3. Any heading containing "DNA"** is indexed as a **strand**, keyed the same way the decompose pass keys the template JSON — `## Viral DNA` → `viral_dna`, `## Visual DNA` → `visual_dna`. That's what makes "show me the visual DNA of this format" and "show me the `visual_dna` of this template" the same question. `vidfarm harness show <ref> --dna visual` prints one strand instead of the whole doc.
|
|
73
|
+
|
|
74
|
+
Everything else is prose the agent reads for context. That split is the whole design: the CLI is honest about which half it can enforce, and it never passes a video on the strength of the half it can't.
|
|
75
|
+
|
|
76
|
+
## The strands
|
|
77
|
+
|
|
78
|
+
A harness doesn't have to carry all of these, but this is the layout `harness derive` writes and the one the bundled bases follow — mirroring the decompose JSON so a derived harness and a hand-written one read the same:
|
|
79
|
+
|
|
80
|
+
| Strand | What lives there | Decompose source |
|
|
81
|
+
|---|---|---|
|
|
82
|
+
| **Viral DNA** | hook, retention mechanic, payoff, core emotion, contrast — *why it travelled* | `video-context.json` → `viral_dna` |
|
|
83
|
+
| **Visual DNA** | cut rhythm, energy curve, caption style + placement, font character, b-roll reliance, transitions | `editor-harness.json` → `pacing` / `typography` / `broll` / `transitions` |
|
|
84
|
+
| **Structural DNA** | the beats, their roles, which are load-bearing and must not be reskinned past | `editor-harness.json` → `scenes` + `scene-annotations.json` → `must_preserve` |
|
|
85
|
+
| **Audio DNA** | voiceover, bed, SFX, comedic timing, intonation | `editor-harness.json` → `audio` / `emotional` |
|
|
86
|
+
| **Build DNA** | which paintbrush per beat (raw_clip / hyperframes / reusable_asset / ai_gen), the free-tier path | `replication-harness.json` |
|
|
87
|
+
|
|
88
|
+
## Writing your own
|
|
89
|
+
|
|
90
|
+
Start from the closest built-in (`vidfarm harness init <name>`) or from a real template (`vidfarm harness derive <forkId>`), then **delete what doesn't apply and add what makes your format yours**. A harness you didn't edit isn't about your videos.
|
|
91
|
+
|
|
92
|
+
Good harnesses tend to have: a **Part 0** naming the viewer in one line (the thing that decides everything else), the **DNA strands** for the format's anatomy, **rules with the reason attached** — a rule whose "why" is missing gets argued away by the next agent that reads it — and a **pre-flight checklist** of `- [ ]` items, which is the part the CLI hands back on every run.
|
|
93
|
+
|
|
94
|
+
Keep the checklist short enough that answering it honestly is cheaper than skipping it.
|
|
95
|
+
|
|
96
|
+
**Give every harness a "whole-video review" block, and put it last.** The bundled ones all have one. Front-matter `checks:` grade the composition's structure and `vidfarm qa` grades its DOM — neither can see the finished video, and the defects that actually ship are sequence-level: margins that shift scene to scene, three type sizes, an accent colour that wanders, N identically-long beats, a jarring join, a dead band under top-anchored content. Those come from how the video was built (one scene at a time, each correct in isolation), so they are invisible to every per-scene check *and* to the agent that built it — across a 32-video batch, every first-pass video had a real defect its own author had already called "verified, looks good." The review block is what forces the contact-sheet pass that catches them. Method: `references/reviewing-renders.md`.
|
|
97
|
+
|
|
98
|
+
## Where a harness lives
|
|
99
|
+
|
|
100
|
+
Next to the work: `./work/HARNESS.md`, auto-discovered by `vidfarm qa ./work`. Not in this repo — the bundled ones under `harnesses/` are *starting points*, and editing them instead of copying them means the next director inherits your campaign's rules.
|
|
101
|
+
|
|
102
|
+
The `.harness/` directory a `vidfarm pull` writes is a different thing: machine-generated context (`context.json`, `agent-guide.md`) regenerated on every pull. Never hand-edit it. `HARNESS.md` is the one you own.
|
|
103
|
+
|
|
104
|
+
## Built-ins
|
|
105
|
+
|
|
106
|
+
| Name | For |
|
|
107
|
+
|---|---|
|
|
108
|
+
| `short-form` | The general default: the four charges (hook / loop / payoff / bait) + the standalone rule. Start here |
|
|
109
|
+
| `hooks` | Hook-variant batches — chunk-1 legibility, the unguessable test, the anti-patterns that only appear at volume |
|
|
110
|
+
| `ugc-testimonial` | A person vouching for a product. Mostly rules about what NOT to add |
|
|
111
|
+
| `explainer` | Faceless educational video: one claim, invented visuals |
|
|
112
|
+
| `product-demo` | Real product doing a real thing — the highest slop-risk format in the catalog |
|
package/.agents/skills/vidfarm/{regimes/explainer.QA_REGIME.md → harnesses/explainer.HARNESS.md}
RENAMED
|
@@ -10,9 +10,10 @@ checks:
|
|
|
10
10
|
font_regime: required
|
|
11
11
|
max_scene_sec: 8
|
|
12
12
|
max_simultaneous_text: 1
|
|
13
|
+
max_words_per_cue: 12
|
|
13
14
|
---
|
|
14
15
|
|
|
15
|
-
# Explainer
|
|
16
|
+
# Explainer Harness
|
|
16
17
|
|
|
17
18
|
For faceless educational video: one idea, explained, with visuals that are *invented* (typography, diagrams, data, abstract motion) rather than captured. No presenter, so the structure has to carry everything a face would.
|
|
18
19
|
|
|
@@ -62,7 +63,7 @@ This format is the one most exposed to the defect, because it *invents* its visu
|
|
|
62
63
|
|
|
63
64
|
The mechanical form: compare each caption phrase against the words on screen in that scene and **drop the caption when overlap is ≥60% of its content words** (words longer than 2 chars, case- and punctuation-normalised). On a reference build that suppressed **10 of 27 tiles**, and the typographic hook and end card came out **entirely caption-free** — correct, not a bug. The survivors also get wider time windows (narrowest tile **0.41s → 0.71s**), so it improves readability too.
|
|
64
65
|
|
|
65
|
-
**Caption styling itself is measured, not hardcoded.** A caption plate copied from another video onto this video's stage is a slab the design never asked for. Measure the composited background behind the caption band and choose light-type-no-plate / dark-type-no-plate / plate accordingly, one treatment for the whole video — the procedure and thresholds are in `short-form.
|
|
66
|
+
**Caption styling itself is measured, not hardcoded.** A caption plate copied from another video onto this video's stage is a slab the design never asked for. Measure the composited background behind the caption band and choose light-type-no-plate / dark-type-no-plate / plate accordingly, one treatment for the whole video — the procedure and thresholds are in `short-form.HARNESS.md` → "Caption styling is MEASURED off the background".
|
|
66
67
|
|
|
67
68
|
### Rule 7 — production floor
|
|
68
69
|
|
|
@@ -12,11 +12,11 @@ checks:
|
|
|
12
12
|
max_simultaneous_text: 1
|
|
13
13
|
---
|
|
14
14
|
|
|
15
|
-
# Hooks
|
|
15
|
+
# Hooks Harness
|
|
16
16
|
|
|
17
|
-
Operating rules for short-form hooks that have to survive a cold algorithm and convert into a funnel. Use this
|
|
17
|
+
Operating rules for short-form hooks that have to survive a cold algorithm and convert into a funnel. Use this harness when the thing you are bulk-generating **is the hook** — same body, N openings — which is the highest-leverage variant axis there is.
|
|
18
18
|
|
|
19
|
-
Copy this file next to your work (`vidfarm
|
|
19
|
+
Copy this file next to your work (`vidfarm harness init hooks --out ./work/HARNESS.md`) and edit it. The parts that matter most to you are the parts you add.
|
|
20
20
|
|
|
21
21
|
*(Craft reference: the vidfarm skill's `references/hooks-and-virality.md`. This file is its checkable form — copy and edit it per account.)*
|
|
22
22
|
|
|
@@ -16,7 +16,7 @@ checks:
|
|
|
16
16
|
- learn more
|
|
17
17
|
---
|
|
18
18
|
|
|
19
|
-
# Product Demo
|
|
19
|
+
# Product Demo Harness
|
|
20
20
|
|
|
21
21
|
For showing a real product doing a real thing. This is the format with the **highest slop risk in the entire catalog**, because the subject matter is a website — so the author's web instincts and the product's own design language both push toward putting a landing page on the timeline.
|
|
22
22
|
|