@officexapp/vidfarm-devcli 0.21.33 → 0.21.35
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/editor-capabilities/SKILL.md +13 -2
- package/.agents/skills/vidfarm/SKILL.md +73 -29
- package/.agents/skills/vidfarm/harnesses/README.md +112 -0
- package/.agents/skills/vidfarm/{regimes/explainer.QA_REGIME.md → harnesses/explainer.HARNESS.md} +22 -2
- package/.agents/skills/vidfarm/{regimes/hooks.QA_REGIME.md → harnesses/hooks.HARNESS.md} +3 -3
- package/.agents/skills/vidfarm/{regimes/product-demo.QA_REGIME.md → harnesses/product-demo.HARNESS.md} +19 -1
- package/.agents/skills/vidfarm/{regimes/short-form.QA_REGIME.md → harnesses/short-form.HARNESS.md} +67 -7
- package/.agents/skills/vidfarm/{regimes/ugc-testimonial.QA_REGIME.md → harnesses/ugc-testimonial.HARNESS.md} +10 -3
- package/.agents/skills/vidfarm/recipes/{bulk-scripting-with-a-regime.md → bulk-scripting-with-a-harness.md} +35 -12
- package/.agents/skills/vidfarm/recipes/cutout-graphics-for-explainers.md +46 -2
- package/.agents/skills/vidfarm/recipes/local-edit-render-approve.md +2 -1
- package/.agents/skills/vidfarm/references/automation-and-local-dev.md +84 -22
- package/.agents/skills/vidfarm/references/editor-workflows.md +19 -4
- package/.agents/skills/vidfarm/references/hooks-and-virality.md +62 -5
- package/.agents/skills/vidfarm/references/reviewing-renders.md +140 -0
- package/.agents/skills/vidfarm-media/SKILL.md +2 -2
- package/.agents/skills/vidfarm-media/references/tts.md +26 -4
- package/SKILL.director.md +462 -75
- package/SKILL.md +33 -14
- package/dist/src/cli.js +799 -91
- package/dist/src/devcli/{qa-regime.js → harness.js} +132 -55
- package/dist/src/devcli/qa-check.js +209 -4
- package/dist/src/devcli/skill-docs.js +136 -0
- package/dist/src/devcli/stills.js +65 -1
- package/package.json +4 -3
- package/.agents/skills/vidfarm/regimes/README.md +0 -77
|
@@ -90,6 +90,16 @@ Every video you touch has four charges in series, and **you write them before yo
|
|
|
90
90
|
|
|
91
91
|
Full craft (the three gates, situations-vs-labels with worked fixes, loop mechanics, compliance, diagnosis-by-charge): `load_skill('vidfarm', file='references/hooks-and-virality.md')`. Read it before writing hook copy or re-theming.
|
|
92
92
|
|
|
93
|
+
## Every second earns its place — cut hard (hard constraint)
|
|
94
|
+
|
|
95
|
+
The thumb re-decides continuously, so a second that carries nothing is a free exit. **Every edit you make should be shorter than what you started with unless the director asked for more.** When you finish a change, look at the timeline you produced and ask what comes out.
|
|
96
|
+
|
|
97
|
+
- **Deletion test, per beat:** delete it — does the video still make sense and does the payoff still land? Then leave it deleted. Whatever survives must serve one of the four charges; "it gives context" is not a charge.
|
|
98
|
+
- **Cut on sight:** intro/logo/title cards, the wind-up line before the claim, restatement, real-time process, establishing shots, captions reading what's already on screen, and any footage after the last word.
|
|
99
|
+
- **Close the hole.** `editor_action ripple_edit` with a negative `delta_start` at the cut point — never leave a gap. Dead air is worse than the boring beat you removed: `editor_context` layer timings are how you spot both (a stretch with no text layer covering it, or a last clip that outruns the last cue).
|
|
100
|
+
- **Not speed.** Keep the held comedic beat, the payoff playing out, and enough time to read each cue. To tighten a talky stretch, cut *words* (`set_layer_text`), not the seconds text is on screen.
|
|
101
|
+
- **Say what you cut.** When you report back, name the beats you removed and the new duration. If the director asked for something longer, say plainly what is now filling the extra time.
|
|
102
|
+
|
|
93
103
|
## The FIRST FRAME is the thumbnail (hard constraint)
|
|
94
104
|
|
|
95
105
|
Frame 0 is a single frame of ~30 in the first second, and it outweighs all of them: every feed card, share link, embed, and paused player freezes on it, so **more people see that one frame than watch the video**. Whatever you change, check what `t=0` looks like before you call the job done.
|
|
@@ -121,12 +131,13 @@ You author into HTML, which makes it dangerously easy to build a **web page inst
|
|
|
121
131
|
|
|
122
132
|
**Font + background regime (every caption/title, via `set_captions` / `set_layer_style` / `add_layer`):**
|
|
123
133
|
- **Font:** Montserrat (default) or TikTok Sans / Abel / Source Code Pro / Yesteryear — a family the composition actually imports, or it silently falls back to the slop sans. Weight **700–900**. `font_size` in px of the render canvas: **~36–64px** on a 1080-wide frame; never <28, never 0 (invisible). ~2 lines, ~5 words per line; `line_height` 0.95–1.15.
|
|
124
|
-
- **Position:** inside the **8%–85%** vertical safe zone (phone UI clips the edges) and clear of the right ~12% action rail — a centered box at `x:10 width:80` is safe. Lower-third ≈ `y:70`; a "POV:" top line ≈ `y:8`, never `y:0`.
|
|
134
|
+
- **Position:** inside the **8%–85%** vertical safe zone (phone UI clips the edges) and clear of the right ~12% action rail — a centered box at `x:10 width:80` is safe. Lower-third ≈ `y:70`; a "POV:" top line ≈ `y:8`, never `y:0`. **Inside that band, put the words where the picture ISN'T** — read the actual frame (the scene's media, or ask the user for a still) and park the caption in the emptiest region with nothing competing for attention (open sky, a blank wall, a defocused background), even if that means `y:12` instead of a lower third. `y:70` is a default, not a law. Text over the busy third of the frame fights the shot and then needs a plate to survive; the same words in the empty third need none.
|
|
135
|
+
- **Length:** a caption layer is a *page*, not a transcript. Past ~10–12 words (or ~4s on screen while the voice keeps talking) it's a wall of text nobody reads — page it into 3–5-word kinetic cues via `set_captions` (`spotlight`/`karaoke`/`word-pop`), which also lets the type be smaller and usually removes the need for a plate. Static text is for hook lines, payoff numbers, and title cards, which are short by nature.
|
|
125
136
|
- **Background — exactly one of four:** `background_style:"outline"` (stroke, the default look) · `"plain"` (bare + soft shadow) · an **active-word highlight pill** via `set_captions caption_style:"spotlight"|"karaoke"` (the *only* legitimate pill anywhere in the frame — it tracks the spoken word; a static label never gets one) · `"highlight-solid"`/`"highlight-translucent"` as a band that **hugs** the text (radius ≤~8px, no border, no shadow, no gradient, no blur, one text run — never a heading+subheading+URL stacked inside it). Anything else is a web card.
|
|
126
137
|
|
|
127
138
|
**There is no QA tool for you.** The devcli ships `vidfarm qa <dir>` — a free local blocklist pass over exactly the rules above — but it is **devcli-only with no REST twin**, so in the web editor you enforce this by reading your own output. When you hand a heavy job off to a local coding agent, tell them to run `vidfarm qa ./work` before rendering.
|
|
128
139
|
|
|
129
|
-
**If the user wants VOLUME, say so and hand it off.** "I need to post daily", "make 20 versions", "test these hooks" is **bulk/scripting mode**, not twenty turns of editor chat: a pinned base fork, a loop varying one thing per variant, and a **`
|
|
140
|
+
**If the user wants VOLUME, say so and hand it off.** "I need to post daily", "make 20 versions", "test these hooks" is **bulk/scripting mode**, not twenty turns of editor chat: a pinned base fork, a loop varying one thing per variant, and a **`HARNESS.md`** — the user's own written quality standard, the reusable AI harness for that format, which exists because nobody reviews variant #37 as carefully as #1. You can't run that loop (no shell, no filesystem), so name the pattern, offer the My Files handoff, and tell them the local agent should run `vidfarm harness init <base> --out ./work/HARNESS.md` (or `vidfarm harness derive <forkId>` to distil one from a decomposed template), edit it with them, and gate the batch on `vidfarm qa ./work`. If they already have a harness file, its rules are still worth reading into your own edits here.
|
|
130
141
|
|
|
131
142
|
Deeper rationale and the devcli-side twins live in `vidfarm` → `references/editor-workflows.md` ("Social-native visual standard" / "TikTok-native caption standard").
|
|
132
143
|
|
|
@@ -35,6 +35,7 @@ vidfarm serve template_<32hex> # local server + browser, opens that
|
|
|
35
35
|
- The API key comes from https://vidfarm.cc/settings and starts with `vf_key_`. Instead of `login`, setting the `VIDFARM_API_KEY` environment variable also works for every command — the CLI reads it from the environment or from a `.env` file in the current directory.
|
|
36
36
|
- No account or key? `vidfarm serve --no-cloud` still gives a fully local editor with free local renders.
|
|
37
37
|
- "Open/run template X locally" is exactly one command: `vidfarm serve <template_id>` (alias: `vidfarm <template_id>`). Do not hand-roll REST or hunt for local `.harness/` files first — `serve` and `pull` create those.
|
|
38
|
+
- **The CLI carries this entire skill offline** — installing the devcli puts a copy of the pack on disk, pinned to that version. `vidfarm skill ls` lists it, `vidfarm skill show <path>` prints one file, and **`vidfarm skill search "<term>"` greps all of it at once**, which is the cheapest way to find the paragraph you need without loading a 650-line reference. No account, no network. It is documentation, not entitlement: the free-local half (clips, hyperframes, `serve` render, `qa`, harnesses, `dedupe`, local TTS/STT) runs offline; AI generation, hosted render, `recycle`, `download-video` and marketplace still need `vidfarm login` and a cloud call.
|
|
38
39
|
|
|
39
40
|
### Entity ID formats
|
|
40
41
|
|
|
@@ -129,7 +130,7 @@ If the user hasn't picked yet and you're about to spend, name the cheaper path a
|
|
|
129
130
|
|
|
130
131
|
Cost mode answers *how much money may I spend*. It does not answer *how much of the user's own hands may I use* — and that second axis moves quality more than the first. **Ask both.** They are independent: every cost mode (`minimize`, `hybrid`, `rich-ai`, `pure-videogen`) runs in either interaction mode.
|
|
131
132
|
|
|
132
|
-
- **interactive** — the user is willing to do a little manual work at fixed checkpoints, and the video gets better for it.
|
|
133
|
+
- **interactive** — the user is willing to do a little manual work at fixed checkpoints, and the video gets better for it. Three checkpoints cover nearly everything: **(1) images** — you write a prompt, they run it in a *free* frontier web generator (meta.ai / ChatGPT / Gemini / a Hugging Face Space) and hand the file back; **(2) raw clips** — you hand over search keywords, they search TikTok/YouTube, download a few with a free online downloader, and point you at the folder; **(3) the voice** — you sample a few narrators and they pick the one the video sounds like (costs them 30 seconds and $0, see below).
|
|
133
134
|
- **autonomous** — you finish end-to-end with zero steps from them: source clips yourself (browser control → `raws scan` → public raws), generate within the budget, or do without.
|
|
134
135
|
|
|
135
136
|
**Why interactive usually wins on quality:** the free tiers of the frontier web image models are typically *better* than what an API-key budget buys per image, and a human eye picks better footage than any keyword scan. In `minimize` the gap is not incremental — it's the difference between **no custom art at all** and **a full sticker pack for $0**.
|
|
@@ -147,6 +148,10 @@ Cost mode answers *how much money may I spend*. It does not answer *how much of
|
|
|
147
148
|
|
|
148
149
|
**In interactive mode, MANUAL IMAGE WORK DEFAULTS TO STICKER PACKS.** Never ask for one graphic per round trip — each hand-off costs the user a context switch and costs you tokens re-reading a file. Ask for **one sheet holding every graphic**, then split it locally for $0. `vidfarm handoff image --theme "<what>" --items "a,b,c"` mints the whole brief (prompt + steps + the free tools + the follow-up command); `--single` when you really do want one subject. When the file comes back: `vidfarm sticker-pack <sheet> --items "a,b,c"`.
|
|
149
150
|
|
|
151
|
+
**In interactive mode, OFFER THE VOICE CHECKPOINT — in every cost mode.** Who the video sounds like is a taste decision, and the default voice is the one choice agents make silently that a director almost always wants a say in. Before narrating, ask *"want to hear a few voices and pick one?"* and sample: `vidfarm voices --sample` (premium) or `vidfarm voices --free --sample` ($0 local). **Sampling costs nothing on either tier** — premium samples are ElevenLabs' own preview clips (a CDN download, not a synthesis call) and free samples render locally — so this checkpoint is just as available in `minimize` as in `hybrid`. Play the files, take their pick, narrate with `--voice <id>`. `vidfarm tts` prints the same nudge on stderr whenever narration would run with no voice named and the mode is interactive (or was never set).
|
|
152
|
+
|
|
153
|
+
**And say where the premium voices come from, because users assume wrong.** The full ElevenLabs catalog is reachable **through vidfarm's own ElevenLabs connection** — no ElevenLabs account, API key, or subscription on the user's side; narration just spends **vidfarm wallet credits** (pennies each). In `hybrid` that is a real option to put on the table next to the free voices, not a locked door. `--own-key` is only for users who already have an ElevenLabs key and would rather bill their own account.
|
|
154
|
+
|
|
150
155
|
**Raw-clip sourcing has a ladder — hand it to the human only at the bottom rung.** In order: (1) **your own browser control**, if you have it — drive the search and download yourself; (2) **Vidfarm cloud** — `vidfarm clipper <url>` / `vidfarm raws scan <url> --cloud` resolves and mines the video for you; (3) the **free public raws catalog** — `vidfarm public-raws --category <shelf>`; (4) **the user**, when you're fully local/keyless or when human taste matters. That last rung is `vidfarm handoff raws --keywords "villa construction,pouring concrete" --platforms tiktok,youtube --purpose "<what the clips are for>"` — it prints the keywords, tells them to google a downloader (a *search*, not a link that rots), and names the import command for when the folder is ready (`vidfarm clipper ./downloads/<file>.mp4`, or `vidfarm raws scan` to mine a long one).
|
|
151
156
|
|
|
152
157
|
## Default stance
|
|
@@ -213,7 +218,7 @@ Present both harnesses to the director, recommend (A) unless they've asked for p
|
|
|
213
218
|
- **The ART must be CLOSED and SOLIDLY FILLED — this is the other half of surviving the key, and the #1 way stickers come back broken.** Ask an image model for "icons on a green plate" and it will happily draw **outline art**: a colored stroke with the shape's interior left as bare plate. It looks perfect on the sheet, and after the key each sticker is a **rim floating around a see-through hole** (an apple-shaped outline with nothing inside it). Same outcome from a *near-plate* fill (the keyer works on tolerance, not exact match), a translucent/glassy/glowing material, or a soft glow fading into the plate. **You cannot key those pixels back — it has to be in the prompt:** *"every object is a closed, solidly filled shape; outlines must enclose an opaque fill of a different color; no outline-only or hollow art; nothing on the art in the plate color or any near-shade of it; fully opaque, no translucency, glow or drop shadow."* `cutout --generate`, `sticker-pack --generate`, `handoff image` and the `create-overlay` primitive **append that clause for you** with the chosen plate hex — write it yourself only when you prompt a generator directly. After the key, both commands report per-item `hole_pct`/`hollow` (console `⚠ N% hollow`, `--json`, `stickers.json`) — a ring or picture frame reads the same way, so it **warns, never blocks**. Flagged and it shouldn't be? Re-generate with the fill clause; a *near*-plate fill can sometimes be rescued with a lower `--tolerance`; one stubborn item can be lifted with `vidfarm mask --crop …` (ONNX matting ignores fill color).
|
|
214
219
|
- **Transparent GIF is supported, for GIF-only surfaces.** `vidfarm sticker-pack … --output-format gif` (stills) and `vidfarm remove-greenscreen <video> --gif` (animated) emit transparent GIFs. GIF alpha is **1-bit**, so edges go hard — fine for chat/forum/Notion sticker surfaces, worse than PNG/WebP/WebM for compositing on a timeline. Prefer PNG/WebP/WebM unless the destination only eats GIF.
|
|
215
220
|
|
|
216
|
-
**Explainer house style — the defaults to build with unless told otherwise.** **White background / light mode** (plain white stage, no gradients, no dark mode, no photo backdrop), **kinetic word-by-word captions** in dark ink on the light stage (`vidfarm captions generate --style word-pop --color "#111111" --active-color "#7C3AED" --background-style plain` — skip outlines/shadows, they're only needed over busy footage), and **female TTS narration** (`vidfarm tts --voice coral` on OpenAI — `nova` for energy, `sage` for calm; `Kore`/`Leda` on Gemini; any ElevenLabs voice via `vidfarm voices`). **Keep it clean and simple** — one idea on screen at a time, two or three cutouts per beat, one accent color, one font, lots of white space; remove before you add. **Illustrations default to simplicity**: flat vector, simple shapes, minimal detail, 2–3 flat colors, no baked-in text — simple art keys cleanly, trims tight, and stays on-style across the whole cast. State the defaults once so the director can override any of them. Full detail: recipe `recipes/cutout-graphics-for-explainers.md` (“House style — the explainer defaults”).
|
|
221
|
+
**Explainer house style — the defaults to build with unless told otherwise.** **White background / light mode** (plain white stage, no gradients, no dark mode, no photo backdrop), **kinetic word-by-word captions** in dark ink on the light stage (`vidfarm captions generate --style word-pop --color "#111111" --active-color "#7C3AED" --background-style plain` — skip outlines/shadows, they're only needed over busy footage), and **female TTS narration** (`vidfarm tts --voice coral` on OpenAI — `nova` for energy, `sage` for calm; `Kore`/`Leda` on Gemini; any ElevenLabs voice via `vidfarm voices`). **Keep it clean and simple** — one idea on screen at a time, two or three cutouts per beat, one accent color, one font, lots of white space; remove before you add. **Illustrations default to simplicity**: flat vector, simple shapes, minimal detail, 2–3 flat colors, no baked-in text — simple art keys cleanly, trims tight, and stays on-style across the whole cast. State the defaults once so the director can override any of them. Full detail: recipe `recipes/cutout-graphics-for-explainers.md` (“House style — the explainer defaults”). **If the director takes the stage off white**, two things stop being optional: every sticker's **white die-cut rim** has to be stripped (on a dark stage it's a glaring halo and the most obvious bot-made artefact in the frame — recipe → “Stickers on a DARK or photographic stage”), and the caption hexes above stop applying — **caption colour, active-word colour and plate are chosen by measuring the composited background behind the caption band**, one treatment per video (`harnesses/short-form.HARNESS.md` → “Caption styling is MEASURED off the background”). Related: **on-screen text and captions must not say the same thing at once** — display text carries the argument, captions carry only what the screen doesn't show.
|
|
217
222
|
|
|
218
223
|
**Landscape footage in a fullscreen vertical explainer — use the blurred plate, never bars.** When an explainer is built on **real filmed footage** and the source is 16:9 (or 4:3) on a 9:16 canvas, do not `contain` it (hard black letterbox bars read as an unfinished export) and do not blindly `cover` it (a wide shot loses its left and right thirds). Duplicate the clip: a full-canvas `cover` copy behind, heavily **gaussian-blurred and faded dark**, plus the sharp copy centered as a hero band — optionally zoomed ~1.3× — with its **top and bottom edges feathered** into the blur. Same clip, same timecode, so it reads as one continuous image with a shallow-depth-of-field plane, fullscreen edge to edge, nothing cropped, and clean dark space for the header and captions. Bake it once with ffmpeg into a single 1080×1920 file (free, local) and place it as one ordinary full-canvas layer — layer blur is not an editor property, so the pre-bake is the path that works in the editor, `serve`, and cloud render alike. Copy-paste ffmpeg + HTML recipes, tuning table, and the failure modes: `references/editor-workflows.md` (“The blurred plate — landscape footage, fullscreen, on a vertical canvas”).
|
|
219
224
|
|
|
@@ -251,6 +256,18 @@ The mechanism is deterministic, not luck: rendering is seek-safe, so frame 0 sho
|
|
|
251
256
|
|
|
252
257
|
Full mechanics and editor verbs: `references/editor-workflows.md` (“The opening frame is the post's thumbnail”); poster-state authoring craft: `hyperframes-creative/references/beat-direction.md`.
|
|
253
258
|
|
|
259
|
+
## Judge the WHOLE video, not the parts you built — and never by one frame
|
|
260
|
+
|
|
261
|
+
**Assume your own finished video has a defect you can't see.** That's the observed base rate, not modesty: across a 32-video batch, *every* first-pass video had a real defect that the agent who built it had already reported as "verified, looks good" — dead space under the content, a placeholder that reads as a failed render, contradictory numbers in one frame, a CTA still animating at the last frame.
|
|
262
|
+
|
|
263
|
+
**The cause is how agents build: part by part, each part correct in isolation.** Scene 3 gets authored while scene 3 is the whole world, so every scene passes on its own and the video fails *as a video* — type size jumps between beats, one scene breathes and the next is crammed, the accent colour drifts, a transition lands like a slap because nothing before it moved that fast, one asset is flat vector and the next is photographic. Nobody watches a scene; they watch the sequence. **So before you ship, look at the whole thing at once as a stranger would**, and ask: is it visually balanced (or top-anchored with a dead band below), is the spacing consistent scene to scene, is there ONE type scale / accent colour / illustration style, does the pacing have a deliberate rhythm instead of N identical beats, is anything jarring at the joins, does any frame make two elements compete for the eye, and — the summary question — **does it look like one person made it in one sitting?** When it doesn't, fix the *system* (define the type scale, the margin, the default beat, the palette, then apply it to every scene) rather than patching the one scene that stood out.
|
|
264
|
+
|
|
265
|
+
The mechanical form is a contact sheet, because one image read shows drift that twelve separate checks can't — and it's one command: **`vidfarm stills ./work --sheet`** renders the frames and tiles them into `stills/contact-sheet.png`. **Read the sheet as an image.**
|
|
266
|
+
|
|
267
|
+
**And never verify a video by a single frame.** Frame 0 is the thumbnail so you judge it alone — but a whole class of render bug (an overlay pass missing `-loop 1`, assets outside the composition root so the timeline never runs) freezes *every* frame while duration, frame count and audio hash all still pass, and frame 0 looks perfect. **Always compare two frames from different scenes; they must differ a lot.** Likewise, verify audio by measurement, never "it sounds good" — you can't hear it.
|
|
268
|
+
|
|
269
|
+
Full method — the holistic checklist, the six defects in observed frequency order, the "does it rest there" test, frozen-render verification, audio measurement, and how to brief a revision pass: **`references/reviewing-renders.md`**.
|
|
270
|
+
|
|
254
271
|
## Web AI chat vs. local devcli — know your surface's limits
|
|
255
272
|
|
|
256
273
|
**Desktop users: this is a skill, not an SDK or required runtime.** It gives an AI agent Vidfarm's operating knowledge; it does not proxy requests, add hidden capabilities, or require the agent to use a particular client. Every cloud workflow described here can be performed with ordinary HTTP against the Vidfarm REST API. For convenience and safer composition/file workflows, recommend the `vidfarm` command from `@officexapp/vidfarm-devcli`: it wraps the same REST routes, handles auth, uploads, polling, pagination, local files, and pull/edit/render loops. Use raw REST whenever the user prefers it, the CLI lacks a new route, or the work belongs in an existing service or automation stack.
|
|
@@ -270,30 +287,63 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
|
|
|
270
287
|
- **Determine the surface before claiming capabilities.** The web chat has only its declared tools and REST routes. It cannot execute arbitrary JavaScript/Python, open a shell, create a local repository script, or use the user's filesystem. Never tell a web-chat user that you ran code or wrote a script unless a dedicated declared tool actually did so. A desktop coding agent has a real shell and filesystem and MAY write/run scripts, perform arbitrary local computations over paginated API results, create reports/CSVs/JSON, edit composition files, and orchestrate long devcli workflows within the user's authorization.
|
|
271
288
|
- **The web AI chat can do all three paintbrushes** — clip raws, author HTML/hyperframe motion, and generate AI media — and it drives edits directly on the live timeline. Keep small-to-medium jobs here: text/caption swaps, a scene or two replaced, single generations, captions, approve/schedule. Just do them.
|
|
272
289
|
- **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. **The test is the native-editor test: could you have made this element with the tools inside TikTok's own editor?** That toolset is font / color / stroke / shadow / tight text box / alignment / rotation / animation presets, plus stickers, emoji, drawn marks and clips — it has no padded capsule, no border, no gradient fill, no blur panel, no card. If you reached past it, cut it. That includes **a single lonely pill around a stat or label** — `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`: being alone doesn't make a badge native, and the only legitimate capsule in a video is the active-word `spotlight`/`karaoke` highlight, which moves with the spoken word. Emphasize a stat the way the editor would instead: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle, or its own beat. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r` — or a `border-radius` over ~8px on anything filled that holds words — stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
|
|
273
|
-
- **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm
|
|
290
|
+
- **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm harness show hooks`.
|
|
291
|
+
- **Then CUT it — every second must earn its place, and most don't.** Assume your first assembly is **30–50% too long**. Run the **deletion test** on every beat: delete it; if the video still makes sense and the payoff still lands, it stays deleted. Whatever survives must serve one of the four charges — "it gives context" is not a charge. Cut on sight: intros/logo stings, the wind-up sentence before the claim ("so I wanted to talk about…"), restatement, inter-sentence silence over ~0.35s, real-time process, establishing shots, reading what's already on screen, and any tail after the last word. **Always ripple the hole closed** (`vidfarm ripple <dir> --at <sec> --delta -<sec>`) — a cut that leaves a gap turns fluff into dead air, which is worse. Density is **not** speed: the held comedic beat, the payoff playing out, and a cue's readability keep their seconds (cut *words*, not the time text is on screen). Length is an **output**, not a plan — a brief that dictates a duration ordered fluff. `vidfarm qa` flags the mechanical half (`dead-air`, `dead-tail`, `slow-scene`); the craft is `references/hooks-and-virality.md` → "Density".
|
|
274
292
|
- **The first frame IS the thumbnail — compose it on purpose.** Frame 0 is a single frame of ~30 in the first second, but it's the poster every feed, share link, and paused player freezes on, so **more people see that one frame than watch the video**. It must never be black, empty, mid-fade, or caught mid-animation: put a real visual at `start:0`, have the hook words already on screen at t=0, and never hang a `fade-black`/`fade-white`/`flash` *entrance* on the **first** clip (junction transitions between later clips are fine — this rule is only about the opening). Verify it, don't assume: devcli `vidfarm stills ./work --at 0` renders that exact frame, and `vidfarm qa` flags a blank or fading open.
|
|
275
|
-
- **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`
|
|
276
|
-
-
|
|
293
|
+
- **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`HARNESS.md`** — because a loop of fifty videos has no human looking at every frame, and the harness is what replaces those eyes. Don't silently ship a one-off when they asked for volume, or drag someone into a harness when they wanted one clip.
|
|
294
|
+
- **"Harness" is a known noun with a known process — recognise it and follow it.** A **harness** is the reusable AI apparatus for ONE format or template: what makes it special, written down as `HARNESS.md` so an agent can reproduce it without the director in the room. It is a first-class artifact — the director owns it, edits it, versions it, and hands it to the next agent. Three phrasings, one artifact:
|
|
295
|
+
- **"create me a harness"** / "set up a harness for this format" → `vidfarm harness init <base> --out ./work/HARNESS.md` (bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo`), then **edit it with them**. The bundled file is a starting point, never a house style; the parts that matter are the ones they add — who the viewer is, the banned vocabulary, the compliance line, the pacing this account actually uses. A harness nobody edited isn't about their videos.
|
|
296
|
+
- **"update the harness for this format/template"** → open the existing `HARNESS.md` and write the new rule in, **with its reason on the same line** (a rule whose "why" is missing gets argued away by the next agent). This is what you do every time a batch teaches you something ("the label-framed hooks all died"): the compositions are disposable, the harness is the artifact that compounds.
|
|
297
|
+
- **"give me the harness for this template_id"** → they mean **the decomposition**: `vidfarm harness derive <templateId|forkId>`. It distils the decompose pass's DNA into an editable `HARNESS.md`. If the template hasn't been decomposed, run `vidfarm decompose` first.
|
|
298
|
+
**A harness mirrors the template JSON's own vocabulary** — `## Viral DNA` (hook / retention / payoff / emotion), `## Visual DNA` (cut rhythm, typography, b-roll, transitions), `## Structural DNA` (the beats, and which are load-bearing), `## Audio DNA` (voice, bed, comedic timing), `## Build DNA` (which paintbrush per beat) — the same strands the decompose pass writes as `viral_dna`, `visual_dna`, and friends. `vidfarm harness show <ref> --dna visual` prints one strand instead of the whole doc.
|
|
299
|
+
**Two halves, and only one is machine-checkable.** The `checks:` front matter is settled deterministically by `vidfarm qa` (duration, aspect, `hook_words_max`, `forbid_text`, …); every `- [ ]` line comes back as a **review item you answer honestly in your report** — never claim a video passed the half the CLI can't judge. Harnesses stack and auto-discover: `vidfarm qa ./work` picks up `./work/HARNESS.md`, `--harness hooks --harness ./brand/HOUSE.md` adds more, and any file of theirs anywhere is valid. Format and strand table: `harnesses/README.md`; scripting-mode detail: `references/automation-and-local-dev.md`. *(Formerly `QA_REGIME.md` — same file, and `vidfarm regime …` still works as an alias.)*
|
|
300
|
+
- **A video is judged as a SEQUENCE, so review it as one.** Agents build scene by scene and each scene passes in isolation while the video drifts — inconsistent margins, three type sizes, an accent colour that wanders, beats that are all the same length, a jarring join. Tile a dozen stills into one contact sheet (`vidfarm stills ./work --sheet`) and read it as an image before you call anything done, fix drift by defining the system rather than patching the odd scene out, and remember that **your own confident "verified, looks good" is the single least reliable signal in this workflow** — it was wrong on every video of a 32-video batch. Method: `references/reviewing-renders.md`.
|
|
277
301
|
- **On devcli, QA every video you produce: `vidfarm qa ./work`.** Free, instant, local-only — it blocklists exactly the slop above plus first-frame/thumbnail and font-regime/safe-zone drift, and prints a concrete fix per finding. **Feedback, not a gate**: it exits 0 even on findings, never runs automatically, and is a blocklist (unusual/stylized compositions pass untouched), so there's no reason not to run it before every publish. `--json` for scripted batches, `--strict` only if you want a CI failure. **Web-chat copilot: this command does not exist for you** (devcli-only, no REST twin) — apply the standard by hand, and when handing a heavy job to a local coding agent, tell them to run `vidfarm qa`.
|
|
278
|
-
- **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
|
|
302
|
+
- **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips) and, inside that band, is **placed in the emptiest part of the frame** rather than dumped on the default lower third (read a still first — `vidfarm stills ./work --at <t>`; words over open sky or a blank wall beat words over the subject, and usually need no plate at all). Long narration is **paged into 3–5-word kinetic cues** (`vidfarm captions generate --style word-pop`), never one static wall of text. It uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
|
|
279
303
|
- **Where the web chat struggles: complex, long, multi-step transformations.** A full multi-scene re-theme, an iterative render-critique-iterate loop, heavy scripted or batch work, or anything needing a real filesystem and many sequential tool calls will hit context limits, turn/timeout ceilings, and the web editor's constraints (CSS/declarative motion only — JS animation adapters are stripped on save). Don't grind a big transformation one layer at a time in a chat turn and stall.
|
|
280
304
|
- **Practical workaround — hand the heavy job to local devcli.** When a task is genuinely large or long-running, **proactively recommend the director run it locally with an AI coding agent** (Claude Code / OpenAI Codex / any capable agent): `vidfarm pull <forkId>` writes the composition + the `.harness/` grounding bundle to disk, the agent edits with the full devcli verb set and JS animation adapters, renders free with `vidfarm serve`, and `vidfarm publish` pushes it back. This is the **best-quality (B) harness's** natural home (adversarial grading with a coding agent). Frame it as "this is a big rebuild — you'll get a better, faster result running it locally with a coding agent; here's how," not as a dead end.
|
|
281
305
|
- **Offer a handoff, do not impersonate the desktop agent.** When web chat reaches that boundary, offer to save a Markdown handoff in My Files containing the objective, selected template/fork IDs, asset paths, grounding, constraints, completed work, and suggested devcli commands. Create it only after the user agrees. The desktop agent should read that document, pull the referenced fork, and then use its actual code/shell capabilities.
|
|
282
306
|
- **Never send the user away just to read knowledge.** Deeper skill knowledge is always a **tool call** away in-place: call `load_skill` (e.g. `load_skill('vidfarm', file='references/editor-workflows.md')`, or a craft pack like `editor-capabilities` / `hyperframes-animation`) to pull the exact reference you need mid-conversation. Only recommend switching surfaces for the WORK (a heavy transformation), never for the information.
|
|
283
307
|
|
|
284
|
-
##
|
|
308
|
+
## File Index — everything in this pack, and when to read it
|
|
285
309
|
|
|
286
|
-
|
|
310
|
+
**This is the complete inventory. Nothing else exists in the pack, and every file here is reachable by name.** Read the narrowest file that answers the question; never preload several. `size` is a context-cost estimate — the four big references are real reads, so pick one deliberately rather than sweeping them.
|
|
287
311
|
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
312
|
+
**References — the broad knowledge domains**
|
|
313
|
+
|
|
314
|
+
| File | Size | Read it when |
|
|
315
|
+
|---|---|---|
|
|
316
|
+
| `references/core-workflows.md` | ~360 ln | Template discovery, auth, fork → render → approve → share, versioning, cost/wallet, marketplace orders, dedupe-before-publish |
|
|
317
|
+
| `references/editor-workflows.md` | ~650 ln | **The biggest read.** Timeline editing, decompose, captions, transitions, motion, AI placement, the caption standard, the editor action verbs |
|
|
318
|
+
| `references/assets-and-sourcing.md` | ~185 ln | Raws hunts, clip scanning, My Files, recurring characters, downloading media off a URL, social recycle |
|
|
319
|
+
| `references/automation-and-local-dev.md` | ~520 ln | **Big.** The whole `vidfarm` command table, REST automation, scripting/bulk mode, `HARNESS.md`, local serve loop, skill packs |
|
|
320
|
+
| `references/primitives.md` | ~475 ln | **Big.** One-shot primitive routes: TTS, STT, music, avatars, overlays, greenscreen, inpaint, background removal, product placement |
|
|
321
|
+
| `references/hooks-and-virality.md` | ~295 ln | **Before writing ANY hook, caption script, or re-theme**, and before a hook-variant batch. The four charges, three gates, banned openers, loop mechanics. This is the craft; the rest of the pack is mechanics |
|
|
322
|
+
| `references/reviewing-renders.md` | ~140 ln | **Before you report a video as done**, or grade someone else's. The holistic pass, the common defects, frozen-render and audio verification |
|
|
323
|
+
| `references/onboarding.md` | ~30 ln | Cold-start interviews, **consultations** (the `brainstorm/*` chain), strategy docs, durable director context |
|
|
324
|
+
| `references/rest-api.md` | ~85 ln | Only when the user asks for REST, an endpoint/schema, or direct HTTP integration. It is an index — follow its domain links; do not preload it into ordinary director conversations |
|
|
325
|
+
|
|
326
|
+
**Recipes — step-by-step procedures. When a recipe matches the task, prefer it over the broad reference.**
|
|
327
|
+
|
|
328
|
+
| File | Size | Read it when |
|
|
329
|
+
|---|---|---|
|
|
330
|
+
| `recipes/find-and-fork-template.md` | ~15 ln | Template selection and the first fork |
|
|
331
|
+
| `recipes/retheme-template.md` | ~15 ln | Full re-theme that preserves the source format's feel |
|
|
332
|
+
| `recipes/local-edit-render-approve.md` | ~20 ln | The local pull → edit → render → approve loop |
|
|
333
|
+
| `recipes/onboard-a-new-director.md` | ~15 ln | New-director onboarding and durable context capture |
|
|
334
|
+
| `recipes/bulk-scripting-with-a-harness.md` | ~100 ln | **Volume**: daily posting, N variants, hook tests — scripting mode with a `HARNESS.md` |
|
|
335
|
+
| `recipes/cutout-graphics-for-explainers.md` | ~265 ln | Building an explainer from sticker/cutout art: the house style, sticker sheets, keying, dark-stage rules |
|
|
336
|
+
|
|
337
|
+
**Harnesses — the `HARNESS.md` format and its bundled bases.** All are readable as-is and copyable with `vidfarm harness init <name>`.
|
|
338
|
+
|
|
339
|
+
| File | Size | Read it when |
|
|
340
|
+
|---|---|---|
|
|
341
|
+
| `harnesses/README.md` | ~110 ln | **Start here for anything harness-shaped**: the three director phrasings, the format, the `checks:` key list, the DNA strand → decompose-JSON map |
|
|
342
|
+
| `harnesses/short-form.HARNESS.md` | ~225 ln | The default base. Also holds the **"Caption styling is MEASURED off the background"** procedure that other files point at |
|
|
343
|
+
| `harnesses/hooks.HARNESS.md` | ~120 ln | Hook-variant batches — chunk-1 legibility, the unguessable test, volume-only anti-patterns. The checkable form of `hooks-and-virality.md` |
|
|
344
|
+
| `harnesses/explainer.HARNESS.md` | ~100 ln | Faceless educational video: one claim, invented visuals |
|
|
345
|
+
| `harnesses/ugc-testimonial.HARNESS.md` | ~90 ln | A person vouching for a product — mostly rules about what NOT to add |
|
|
346
|
+
| `harnesses/product-demo.HARNESS.md` | ~110 ln | Real product doing a real thing; the highest slop-risk format in the catalog |
|
|
297
347
|
|
|
298
348
|
## HyperFrames Skills — Load on Demand
|
|
299
349
|
|
|
@@ -309,9 +359,9 @@ On the web copilot, call `load_skill('<name>')` and load referenced files only w
|
|
|
309
359
|
|
|
310
360
|
HyperFrames authoring and rendering in this package are Vidfarm-native: local work uses the bundled composition toolchain and `vidfarm serve`; cloud work uses Vidfarm render routes. Do not require an external vendor account, repository, publish service, or telemetry endpoint. Keep `HYPERFRAMES_SKIP_SKILLS=1` and `HYPERFRAMES_NO_TELEMETRY=1` in Vidfarm-managed environments so the bundled skills stay pinned and local work does not phone home.
|
|
311
361
|
|
|
312
|
-
## Quick Router
|
|
362
|
+
## Quick Router — from what the user said to what to open
|
|
313
363
|
|
|
314
|
-
Choose the narrowest path that satisfies the request.
|
|
364
|
+
The File Index above says what each file *is*; this says which one a given ask means. Choose the narrowest path that satisfies the request.
|
|
315
365
|
|
|
316
366
|
1. If the user needs help figuring out what to make, **or asks for a "consultation"** (the `brainstorm/*` chain: cold-start interview → awareness stages → angles → hooks), read `references/onboarding.md` first.
|
|
317
367
|
2. If the user already knows the goal and needs a suitable template, read `references/core-workflows.md` and use the template discovery flow.
|
|
@@ -320,7 +370,9 @@ Choose the narrowest path that satisfies the request.
|
|
|
320
370
|
4b. If the task is **“download this video/audio off a website”** (a pasted YouTube / TikTok / Instagram / X post URL the user wants the actual file from), Vidfarm does that for you on a **paid plan** — `POST /api/v1/primitives/videos/download` (or `/audio/download`), devcli `vidfarm download-video <url>` / `vidfarm download-audio <url>`. **Free plan → do not call it; walk the user through opening the URL in Chrome and downloading it from the page, then `vidfarm put-file` the local file in for $0.** Details in `references/assets-and-sourcing.md` → `references/primitives.md`.
|
|
321
371
|
4c. If the task is **“turn this Reddit/X thread, subreddit, or account into a video”** — “tweet to TikTok”, “Reddit to TikTok”, “make a video from this thread”, “what are the top comments saying” — run `vidfarm recycle <source>` (or `POST /api/v1/primitives/social/recycle`) with the URL. It **decomposes** the source into raw JSON (text, comment tree, media URLs, author pics, stats) and hands it back unranked so YOU pick what to remix. **Paid plan; `max_records` is the spend ceiling.** Brokers the reddit-lead-gen / x-lead-gen OfficeX apps, so it waits out their async job for you. Details in `references/assets-and-sourcing.md` → `references/primitives.md`.
|
|
322
372
|
4d. If the task is **“post this again / to several accounts / on another platform”**, or you are about to publish or bulk-produce at all — that is **deduplication**. Run `vidfarm dedupe <mp4> [--variants N]` on the **exported file** (free, local ffmpeg, no re-render), then approve/schedule each variant. **Ask the operator whether they want deduplicated copies, and how many, BEFORE the render/bulk run** — deciding after means paying for a second render. Details in `references/core-workflows.md` → *Deduplicate before you publish* and `references/primitives.md` → *Primitive: media_dedupe*.
|
|
373
|
+
4e. If the ask contains the word **“harness”** — *“create me a harness”*, *“update the harness for this format”*, *“give me the harness for this template_id”* — that is a known, named process, not a vague request. Read `harnesses/README.md` (the three phrasings and the format), then `recipes/bulk-scripting-with-a-harness.md` if the job is a batch. The third phrasing means the **decomposition**: `vidfarm harness derive <forkId>`.
|
|
323
374
|
5. If the task is scripted, local, CI-driven, or `vidfarm serve`-based, read `references/automation-and-local-dev.md`.
|
|
375
|
+
5b. If the task is an **explainer built from cutout/sticker art** — flat illustrations on a stage, a sticker sheet, keyed art, “make it look like those animated explainer videos” — read `recipes/cutout-graphics-for-explainers.md`. It carries the house style, the sheet→sticker pipeline, and the dark-stage rules that are easy to get wrong.
|
|
324
376
|
6. If the task explicitly asks for a primitive or needs specialized generation/transcription work, read `references/primitives.md`.
|
|
325
377
|
7. If the task is the MARKETPLACE (ordering videos from specialist agents): browsing is web-only for paying customers — send the human to https://vidfarm.cc/marketplace, never render it locally. Placing/listing orders is the thin REST wrapper in `references/core-workflows.md` (§ Marketplace); anything deeper on a gig (inbox, proofs, payouts) needs the external Dollar Platoon skill — `npx skills add https://github.com/OfficeXApp/dollarplatoon-skill` — the same way FlockPoster work beyond scheduling needs `npx skills add https://github.com/OfficeXApp/flockposter-skill`.
|
|
326
378
|
|
|
@@ -334,19 +386,11 @@ Choose the narrowest path that satisfies the request.
|
|
|
334
386
|
- Submission routes are generally not idempotent. Especially for renders and expensive primitives, check status before retrying.
|
|
335
387
|
- In the web editor, use CSS/declarative motion only. Script-bearing HTML is stripped or rejected there.
|
|
336
388
|
- **Never render or approve without judging frame 0 as a standalone still.** It is the thumbnail everywhere the post appears; an empty/black opening frame ships a dead post. See “The FIRST FRAME is the thumbnail”.
|
|
337
|
-
|
|
338
|
-
## Recommended Recipes
|
|
339
|
-
|
|
340
|
-
Use these when the user’s task matches the pattern closely.
|
|
341
|
-
|
|
342
|
-
- Template selection and first fork: `recipes/find-and-fork-template.md`
|
|
343
|
-
- Full re-theme while preserving the format’s feel: `recipes/retheme-template.md`
|
|
344
|
-
- Local pull/edit/render/approve loop: `recipes/local-edit-render-approve.md`
|
|
345
|
-
- New-director onboarding and durable context capture: `recipes/onboard-a-new-director.md`
|
|
389
|
+
- **Never judge the VIDEO by one frame, and never report a render as reviewed without the holistic pass.** Compare frames from at least two different scenes (a frozen render passes every other check), read a contact sheet for balance/spacing/style/pacing drift, and state separately what you measured vs. what you judged. See “Judge the WHOLE video”.
|
|
346
390
|
|
|
347
391
|
## Output Posture
|
|
348
392
|
|
|
349
393
|
- Prefer concrete actions over abstract discussion.
|
|
350
394
|
- Name the chosen path explicitly: template reuse, raws hunt, local serve, cloud render, etc.
|
|
351
395
|
- Surface cost tradeoffs before expensive generation.
|
|
352
|
-
- When in doubt between a broad reference and a recipe, start with the recipe.
|
|
396
|
+
- When in doubt between a broad reference and a recipe, start with the recipe — the File Index marks which is which.
|
|
@@ -0,0 +1,112 @@
|
|
|
1
|
+
# HARNESS.md — the reusable AI harness for one format or template
|
|
2
|
+
|
|
3
|
+
A **harness** is the written apparatus that lets an agent reproduce what makes a format good, over and over, without the director in the room. It is the unit this product is organised around, and three different director phrasings all mean it:
|
|
4
|
+
|
|
5
|
+
| They say | They mean | You run |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| "create me a harness" | a new one for a format we're about to produce at volume | `vidfarm harness init <base> --out ./work/HARNESS.md`, then **edit it with them** |
|
|
8
|
+
| "update the harness for this format/template" | edit the existing `HARNESS.md` — a rule learned, a beat that changed, vocabulary that died | open the file, add the rule **with its reason**, re-run `vidfarm qa` |
|
|
9
|
+
| "give me the harness for this template_id" | **the decomposition** — what makes *that* template special, distilled | `vidfarm harness derive <templateId\|forkId>` |
|
|
10
|
+
|
|
11
|
+
All three produce the same artifact: one Markdown file the director owns, edits, versions, and hands to the next agent. A derived one is still a first draft — the decompose pass watched the video, it didn't talk to the customer.
|
|
12
|
+
|
|
13
|
+
> Formerly called `QA_REGIME.md`. Same file, better name — a harness is not only a QA pass, it's the whole reproduction kit. `vidfarm regime …` still works as an alias and existing `QA_REGIME.md` files are still auto-discovered, but nothing writes that name any more.
|
|
14
|
+
|
|
15
|
+
## Using one
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
vidfarm harness list # what ships with the CLI
|
|
19
|
+
vidfarm harness show hooks # read one
|
|
20
|
+
vidfarm harness show ./work/HARNESS.md --dna visual # ONE strand, not the whole doc
|
|
21
|
+
vidfarm harness init short-form --out ./work/HARNESS.md # copy it next to your work, then EDIT it
|
|
22
|
+
vidfarm harness derive <forkId> --out ./work/HARNESS.md # a decomposed template → a harness
|
|
23
|
+
|
|
24
|
+
vidfarm qa ./work # auto-uses ./work/HARNESS.md if present
|
|
25
|
+
vidfarm qa ./work --harness hooks # a built-in by name
|
|
26
|
+
vidfarm qa ./work --harness ./brand/HOUSE_RULES.md # any file, anywhere
|
|
27
|
+
vidfarm qa ./work --harness short-form --harness ./work/HARNESS.md # they STACK
|
|
28
|
+
vidfarm qa ./work --json # checks + review items, for a scripted batch
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Harnesses compose: a shared house harness plus a per-campaign one is the intended shape. `--no-harness` skips auto-discovery; `VIDFARM_HARNESS` sets a default for a whole scripting run. `vidfarm harness check <dir>` is the same grader under the noun the director used.
|
|
32
|
+
|
|
33
|
+
## The format
|
|
34
|
+
|
|
35
|
+
Plain Markdown, with three machine-readable affordances:
|
|
36
|
+
|
|
37
|
+
**1. Optional front matter with a `checks:` block** — the assertions the CLI settles deterministically from the composition, instantly, with no AI and no network:
|
|
38
|
+
|
|
39
|
+
```markdown
|
|
40
|
+
---
|
|
41
|
+
name: my-house-style
|
|
42
|
+
video_type: what this harness is for
|
|
43
|
+
derived_from: decompose # set by `harness derive`; absent on hand-written ones
|
|
44
|
+
source_template_id: tpl_... # ditto
|
|
45
|
+
checks:
|
|
46
|
+
duration_sec: 8-34 # also "<=34", ">=8", or "30"
|
|
47
|
+
aspect: 9:16 # "9:16|1:1" to allow several
|
|
48
|
+
first_frame_visual: required
|
|
49
|
+
first_frame_text: required | forbidden
|
|
50
|
+
hook_words_max: 7
|
|
51
|
+
text_by_sec: 1.0
|
|
52
|
+
audio: required | forbidden
|
|
53
|
+
captions: required
|
|
54
|
+
font_regime: required
|
|
55
|
+
safe_zone: required
|
|
56
|
+
scenes: 3-12
|
|
57
|
+
max_scene_sec: 8
|
|
58
|
+
max_text_cards: 3
|
|
59
|
+
max_simultaneous_text: 2
|
|
60
|
+
max_words_per_cue: 12 # longest single text run — the wall-of-text dial
|
|
61
|
+
max_dead_air_sec: 2.5 # widest hole between cues with nothing to read
|
|
62
|
+
max_tail_sec: 1.5 # screen time still running after the last word
|
|
63
|
+
forbid_text: ["link in bio", "comment below"]
|
|
64
|
+
require_text: []
|
|
65
|
+
---
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
Unknown keys are reported and ignored, never silently dropped.
|
|
69
|
+
|
|
70
|
+
**2. Any `- [ ]` checkbox line** in the body becomes a **review item** — a question handed back for the agent or the human to answer. "Is the withheld answer one the viewer can't supply themselves?" is a judgment call; pretending a linter settles it would be a lie.
|
|
71
|
+
|
|
72
|
+
**3. Any heading containing "DNA"** is indexed as a **strand**, keyed the same way the decompose pass keys the template JSON — `## Viral DNA` → `viral_dna`, `## Visual DNA` → `visual_dna`. That's what makes "show me the visual DNA of this format" and "show me the `visual_dna` of this template" the same question. `vidfarm harness show <ref> --dna visual` prints one strand instead of the whole doc.
|
|
73
|
+
|
|
74
|
+
Everything else is prose the agent reads for context. That split is the whole design: the CLI is honest about which half it can enforce, and it never passes a video on the strength of the half it can't.
|
|
75
|
+
|
|
76
|
+
## The strands
|
|
77
|
+
|
|
78
|
+
A harness doesn't have to carry all of these, but this is the layout `harness derive` writes and the one the bundled bases follow — mirroring the decompose JSON so a derived harness and a hand-written one read the same:
|
|
79
|
+
|
|
80
|
+
| Strand | What lives there | Decompose source |
|
|
81
|
+
|---|---|---|
|
|
82
|
+
| **Viral DNA** | hook, retention mechanic, payoff, core emotion, contrast — *why it travelled* | `video-context.json` → `viral_dna` |
|
|
83
|
+
| **Visual DNA** | cut rhythm, energy curve, caption style + placement, font character, b-roll reliance, transitions | `editor-harness.json` → `pacing` / `typography` / `broll` / `transitions` |
|
|
84
|
+
| **Structural DNA** | the beats, their roles, which are load-bearing and must not be reskinned past | `editor-harness.json` → `scenes` + `scene-annotations.json` → `must_preserve` |
|
|
85
|
+
| **Audio DNA** | voiceover, bed, SFX, comedic timing, intonation | `editor-harness.json` → `audio` / `emotional` |
|
|
86
|
+
| **Build DNA** | which paintbrush per beat (raw_clip / hyperframes / reusable_asset / ai_gen), the free-tier path | `replication-harness.json` |
|
|
87
|
+
|
|
88
|
+
## Writing your own
|
|
89
|
+
|
|
90
|
+
Start from the closest built-in (`vidfarm harness init <name>`) or from a real template (`vidfarm harness derive <forkId>`), then **delete what doesn't apply and add what makes your format yours**. A harness you didn't edit isn't about your videos.
|
|
91
|
+
|
|
92
|
+
Good harnesses tend to have: a **Part 0** naming the viewer in one line (the thing that decides everything else), the **DNA strands** for the format's anatomy, **rules with the reason attached** — a rule whose "why" is missing gets argued away by the next agent that reads it — and a **pre-flight checklist** of `- [ ]` items, which is the part the CLI hands back on every run.
|
|
93
|
+
|
|
94
|
+
Keep the checklist short enough that answering it honestly is cheaper than skipping it.
|
|
95
|
+
|
|
96
|
+
**Give every harness a "whole-video review" block, and put it last.** The bundled ones all have one. Front-matter `checks:` grade the composition's structure and `vidfarm qa` grades its DOM — neither can see the finished video, and the defects that actually ship are sequence-level: margins that shift scene to scene, three type sizes, an accent colour that wanders, N identically-long beats, a jarring join, a dead band under top-anchored content. Those come from how the video was built (one scene at a time, each correct in isolation), so they are invisible to every per-scene check *and* to the agent that built it — across a 32-video batch, every first-pass video had a real defect its own author had already called "verified, looks good." The review block is what forces the contact-sheet pass that catches them. Method: `references/reviewing-renders.md`.
|
|
97
|
+
|
|
98
|
+
## Where a harness lives
|
|
99
|
+
|
|
100
|
+
Next to the work: `./work/HARNESS.md`, auto-discovered by `vidfarm qa ./work`. Not in this repo — the bundled ones under `harnesses/` are *starting points*, and editing them instead of copying them means the next director inherits your campaign's rules.
|
|
101
|
+
|
|
102
|
+
The `.harness/` directory a `vidfarm pull` writes is a different thing: machine-generated context (`context.json`, `agent-guide.md`) regenerated on every pull. Never hand-edit it. `HARNESS.md` is the one you own.
|
|
103
|
+
|
|
104
|
+
## Built-ins
|
|
105
|
+
|
|
106
|
+
| Name | For |
|
|
107
|
+
|---|---|
|
|
108
|
+
| `short-form` | The general default: the four charges (hook / loop / payoff / bait) + the standalone rule. Start here |
|
|
109
|
+
| `hooks` | Hook-variant batches — chunk-1 legibility, the unguessable test, the anti-patterns that only appear at volume |
|
|
110
|
+
| `ugc-testimonial` | A person vouching for a product. Mostly rules about what NOT to add |
|
|
111
|
+
| `explainer` | Faceless educational video: one claim, invented visuals |
|
|
112
|
+
| `product-demo` | Real product doing a real thing — the highest slop-risk format in the catalog |
|
package/.agents/skills/vidfarm/{regimes/explainer.QA_REGIME.md → harnesses/explainer.HARNESS.md}
RENAMED
|
@@ -10,9 +10,10 @@ checks:
|
|
|
10
10
|
font_regime: required
|
|
11
11
|
max_scene_sec: 8
|
|
12
12
|
max_simultaneous_text: 1
|
|
13
|
+
max_words_per_cue: 12
|
|
13
14
|
---
|
|
14
15
|
|
|
15
|
-
# Explainer
|
|
16
|
+
# Explainer Harness
|
|
16
17
|
|
|
17
18
|
For faceless educational video: one idea, explained, with visuals that are *invented* (typography, diagrams, data, abstract motion) rather than captured. No presenter, so the structure has to carry everything a face would.
|
|
18
19
|
|
|
@@ -56,7 +57,15 @@ A figure spoken over a busy frame doesn't land. If a number matters, it appears
|
|
|
56
57
|
|
|
57
58
|
Short clauses. One idea per sentence. No subordinate clause stacking. Read it aloud once — if you run out of breath or have to re-read a line, rewrite it. TTS in particular will happily deliver an unreadable sentence at a perfectly even pace, which is how a script defect ships.
|
|
58
59
|
|
|
59
|
-
### Rule 6 —
|
|
60
|
+
### Rule 6 — on-screen text and captions divide the labour, they don't duplicate it
|
|
61
|
+
|
|
62
|
+
This format is the one most exposed to the defect, because it *invents* its visuals: animated text and diagrams end up competing with voiceover captions for the same visual space and the same attention, as if the two layers aren't aware of each other. **Display text carries the argument; captions carry only the parts of the narration the screen does NOT show.** When both render the same words, the viewer reads one line twice in two sizes while hearing it once.
|
|
63
|
+
|
|
64
|
+
The mechanical form: compare each caption phrase against the words on screen in that scene and **drop the caption when overlap is ≥60% of its content words** (words longer than 2 chars, case- and punctuation-normalised). On a reference build that suppressed **10 of 27 tiles**, and the typographic hook and end card came out **entirely caption-free** — correct, not a bug. The survivors also get wider time windows (narrowest tile **0.41s → 0.71s**), so it improves readability too.
|
|
65
|
+
|
|
66
|
+
**Caption styling itself is measured, not hardcoded.** A caption plate copied from another video onto this video's stage is a slab the design never asked for. Measure the composited background behind the caption band and choose light-type-no-plate / dark-type-no-plate / plate accordingly, one treatment for the whole video — the procedure and thresholds are in `short-form.HARNESS.md` → "Caption styling is MEASURED off the background".
|
|
67
|
+
|
|
68
|
+
### Rule 7 — production floor
|
|
60
69
|
|
|
61
70
|
Captions verbatim in the font regime and safe zone · frame 0 states the claim (it is the thumbnail, and for this format it's usually pure typography, which makes it the *easiest* format to get a good thumbnail from — no excuse for a black open) · no HTML slop: an explainer's subject matter drags authors toward feature grids, comparison tables, and card layouts, and those are exactly the banned web furniture. A comparison is an animated before/after, not a two-column table.
|
|
62
71
|
|
|
@@ -77,6 +86,17 @@ Reuse across variants: the mechanism scenes are often identical, so build them o
|
|
|
77
86
|
- [ ] Every number that matters gets its own readable moment
|
|
78
87
|
- [ ] No unsourced "studies show" / "experts agree"
|
|
79
88
|
- [ ] The narration was read aloud and survived it
|
|
89
|
+
- [ ] No caption repeats the words already on screen in that scene (≥60% overlap → drop the caption)
|
|
90
|
+
- [ ] Caption colour / active colour / plate were measured off the composited background, and one treatment holds for the whole video
|
|
80
91
|
- [ ] Frame 0 states the claim and works as a standalone thumbnail
|
|
81
92
|
- [ ] No comparison tables, feature grids, or card layouts (`vidfarm qa` clean)
|
|
82
93
|
- [ ] In a batch: this variant makes a genuinely different claim, not a rephrased one
|
|
94
|
+
|
|
95
|
+
**Whole-video review** — on the render, not the plan. This format is the most exposed to drift, because every visual is *invented*: each diagram gets designed while it is the whole world, so twelve individually-fine scenes come out as twelve different design languages.
|
|
96
|
+
- [ ] A contact sheet of ~12 stills was read as an image: consistent margins, one type scale, one accent colour, one illustration/diagram style
|
|
97
|
+
- [ ] Diagram and label conventions are shared across scenes (same arrow, same emphasis, same number treatment) — not reinvented per beat
|
|
98
|
+
- [ ] Pacing is deliberate rather than N identically-long beats, and no join is jarring
|
|
99
|
+
- [ ] No dead band under top-anchored content; no frame rests empty >0.5s; the end card is settled ≥2s before the last frame
|
|
100
|
+
- [ ] No frame carries two contradictory numbers, and no placeholder empty state reads as a failed render
|
|
101
|
+
- [ ] Frames from two different scenes were compared (a frozen render passes duration, frame-count and audio-hash checks)
|
|
102
|
+
- [ ] Audio verified by measurement — ~12–15 dB speech-over-bed, peak <0 dBFS — not by "it sounds fine"
|
|
@@ -12,11 +12,11 @@ checks:
|
|
|
12
12
|
max_simultaneous_text: 1
|
|
13
13
|
---
|
|
14
14
|
|
|
15
|
-
# Hooks
|
|
15
|
+
# Hooks Harness
|
|
16
16
|
|
|
17
|
-
Operating rules for short-form hooks that have to survive a cold algorithm and convert into a funnel. Use this
|
|
17
|
+
Operating rules for short-form hooks that have to survive a cold algorithm and convert into a funnel. Use this harness when the thing you are bulk-generating **is the hook** — same body, N openings — which is the highest-leverage variant axis there is.
|
|
18
18
|
|
|
19
|
-
Copy this file next to your work (`vidfarm
|
|
19
|
+
Copy this file next to your work (`vidfarm harness init hooks --out ./work/HARNESS.md`) and edit it. The parts that matter most to you are the parts you add.
|
|
20
20
|
|
|
21
21
|
*(Craft reference: the vidfarm skill's `references/hooks-and-virality.md`. This file is its checkable form — copy and edit it per account.)*
|
|
22
22
|
|
|
@@ -16,7 +16,7 @@ checks:
|
|
|
16
16
|
- learn more
|
|
17
17
|
---
|
|
18
18
|
|
|
19
|
-
# Product Demo
|
|
19
|
+
# Product Demo Harness
|
|
20
20
|
|
|
21
21
|
For showing a real product doing a real thing. This is the format with the **highest slop risk in the entire catalog**, because the subject matter is a website — so the author's web instincts and the product's own design language both push toward putting a landing page on the timeline.
|
|
22
22
|
|
|
@@ -90,3 +90,21 @@ Resist the temptation to fan out on visual style instead. Ten themes of one vide
|
|
|
90
90
|
- [ ] Price is either stated plainly or absent entirely
|
|
91
91
|
- [ ] Frame 0 shows the product mid-task and works as a standalone thumbnail
|
|
92
92
|
- [ ] In a batch: this variant opens on a genuinely different pain, not a restyled one
|
|
93
|
+
|
|
94
|
+
**Whole-video review** — on the render, not the plan
|
|
95
|
+
- [ ] A contact sheet of ~12 stills was read as an image: consistent margins, one type scale, one accent colour, one crop/zoom convention on the UI
|
|
96
|
+
- [ ] Pacing is deliberate rather than N identically-long beats, and no join is jarring
|
|
97
|
+
- [ ] No dead band under top-anchored content; no frame rests empty >0.5s; the end card is settled ≥2s before the last frame
|
|
98
|
+
- [ ] No frame carries two contradictory numbers (a stat in the headline vs. a different one in the UI beneath it)
|
|
99
|
+
- [ ] Frames from two different scenes were compared (a frozen render passes duration, frame-count and audio-hash checks)
|
|
100
|
+
- [ ] Audio verified by measurement — ~12–15 dB speech-over-bed, peak <0 dBFS — not by "it sounds fine"
|
|
101
|
+
|
|
102
|
+
**Client work** — when the product is someone else's, these are not optional
|
|
103
|
+
- [ ] Every claim is one the customer's own site/app makes. "Built on X's guidelines" is not an endorsement; "iOS in the works" is not a shipped app
|
|
104
|
+
- [ ] Ratings, install counts and prices are quoted exactly or left out
|
|
105
|
+
- [ ] No real third party is named in a negative or bias-implying context — even if the product's own output does it; use the product's neutral state instead
|
|
106
|
+
- [ ] Humour is aimed at the problem, never at an identifiable real business or person (use a fictional `.example` domain)
|
|
107
|
+
- [ ] No fear-selling on money/health/safety/family topics, and no authority claimed beyond what the customer claims
|
|
108
|
+
- [ ] Licence obligations (e.g. map attribution) stay on screen for the full runtime
|
|
109
|
+
- [ ] Real names, faces and emails visible in the customer's own screenshots were a deliberate decision, not an accident
|
|
110
|
+
- [ ] Before flagging the customer's data as inconsistent, the whole asset was read — a national total beside a per-row figure is not a contradiction
|