@officexapp/vidfarm-devcli 0.21.39 → 0.21.43

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (30) hide show
  1. package/.agents/skills/editor-capabilities/SKILL.md +16 -0
  2. package/.agents/skills/vidfarm/SKILL.md +31 -6
  3. package/.agents/skills/vidfarm/harnesses/README.md +1 -0
  4. package/.agents/skills/vidfarm/harnesses/explainer.HARNESS.md +11 -0
  5. package/.agents/skills/vidfarm/harnesses/product-demo.HARNESS.md +2 -0
  6. package/.agents/skills/vidfarm/harnesses/product-explainer.HARNESS.md +242 -0
  7. package/.agents/skills/vidfarm/recipes/bulk-scripting-with-a-harness.md +1 -1
  8. package/.agents/skills/vidfarm/recipes/cutout-graphics-for-explainers.md +2 -0
  9. package/.agents/skills/vidfarm/recipes/onboard-a-new-director.md +10 -8
  10. package/.agents/skills/vidfarm/references/automation-and-local-dev.md +21 -8
  11. package/.agents/skills/vidfarm/references/content-ideas.md +113 -0
  12. package/.agents/skills/vidfarm/references/editor-workflows.md +36 -3
  13. package/.agents/skills/vidfarm/references/onboarding.md +64 -1
  14. package/.agents/skills/vidfarm/references/primitives.md +67 -0
  15. package/.agents/skills/vidfarm-media/SKILL.md +50 -0
  16. package/SKILL.director.md +347 -27
  17. package/SKILL.md +6 -3
  18. package/crowdsourcing.md +49 -1
  19. package/dist/src/cli.js +579 -19
  20. package/dist/src/devcli/consult.js +389 -0
  21. package/dist/src/devcli/cost-mode.js +8 -0
  22. package/dist/src/devcli/experiments.js +10 -5
  23. package/dist/src/devcli/qa-check.js +74 -0
  24. package/dist/src/devcli/skill-docs.js +103 -0
  25. package/dist/src/services/brainstorm-prompts.js +132 -0
  26. package/experimental/unique-product-explainers.md +855 -0
  27. package/experiments.md +33 -2
  28. package/package.json +10 -1
  29. package/src/assets/SELLING_AWARENESS_STAGES.md +579 -0
  30. package/src/assets/SELLING_WITH_HOOKS.md +377 -0
@@ -77,6 +77,22 @@ Audio is natively **multi-track**. The timeline mixes UNLIMITED simultaneous `<a
77
77
  - **The headline move — split a combined original when recreating.** When the user recreates a template whose ORIGINAL had music + narration baked into ONE audio track, do NOT reproduce a single combined bed. Rebuild it as TWO independent tracks: a fresh narration track (`/api/v1/primitives/audio/speech`, or same-voice reword via `/api/v1/primitives/audio/regenerate-speech`) at ~1.0, and a separate real music track at ~0.1–0.2 — then mute or `remove_layer` the original combined source-audio layer so the old voice doesn't play under the new one. This hands the user independent voice/music volume and is the elegant workaround for AI TTS being unable to emit narration+music in one file.
78
78
  - **Honesty (ties to the create-media rules):** you cannot un-mix / stem-separate the original's baked audio — the two tracks are BUILT from a fresh narration track PLUS a real music file (owned / user-provided / `browse_files` across `/files` and `/raws`), never a faked "music" layer and never the voice track duplicated. There is no music-generation primitive.
79
79
 
80
+ ## Orient the cold viewer in the first 3 seconds (hard constraint)
81
+
82
+ The hook makes a stranger *want* to watch; orientation makes watching *possible*. The viewer has no context, did not choose this video, and has never heard of the subject — so by **~3s** they must be able to say **what kind of thing this is** (the category noun), **who it is for**, and **why it is on their screen** (the situation). The failure is not a bad first frame; it is a good video that **begins at beat two**, and the author can't see it because the author already knows.
83
+
84
+ Signatures — each means rebuilding the first beat, not polishing it:
85
+
86
+ - **A pronoun with no referent** — "it just works", "this changes everything", "here's how they do it".
87
+ - **Starting at step three** — the process already running, the dashboard already full, the metaphor mid-payoff.
88
+ - **A late subject** — an abstract open whose meaning lands at 6s spends the seconds that decide whether anyone reaches 6s.
89
+ - **Insider vocabulary or an acronym** in the first line.
90
+ - **A detail crop** that reads as texture until you know the whole.
91
+
92
+ Replace it with both channels in one beat: an **easy image** (one large subject, already moving, legible at a glance and at thumbnail scale) and an **easy line** (one clause, ≤12 words, everyday words, concrete noun + verb, the **category named**, brand name said once). Give the situation, not the label. **It costs one sentence, not one beat** — it replaces the wind-up line and never licenses a logo, title card, or fade from black.
93
+
94
+ Test on the render: play the first 3s only to someone with no context. "Something about audio" is a fail. Full rule: `load_skill('vidfarm', file='references/editor-workflows.md')` → "Orient the cold viewer".
95
+
80
96
  ## The four charges — structure before polish (hard constraint)
81
97
 
82
98
  Every video you touch has four charges in series, and **you write them before you start moving layers**. Editing is the fun part, so it gets done first and the words get retrofitted — that's how a beautifully-edited video ends up with nothing to stop for.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: vidfarm
3
- description: Use Vidfarm as a director. Run a strategy **consultation** (the `brainstorm/*` chain — cold-start interview, awareness stages, persuasive angles, hooks, product placement). Browse/add inspiration videos, browse the free public raws catalog BY CATEGORY (curated shelves like scroll-stoppers/greenscreen/reaction — the cheapest way to source footage for one video, and a ready-made clip pool for bulk scripting N variants), fork a template into a composition, edit it in the Trackpad Editor (timeline-based like Premiere/DaVinci), auto-decompose source video into scenes, render to MP4, approve into a shareable post, and schedule it. Includes login, provider keys, discovery, versioning, uploads/downloads, and billing. Every step is available as raw REST; `vidfarm-devcli` wraps those routes and composes the file-backed scripting flows.
3
+ description: Use Vidfarm as a director. Run a strategy **consultation** (the `brainstorm/*` chain — cold-start interview, awareness stages, persuasive angles, hooks, product placement). Answer "give me content ideas" / "what should I post" / "I need 30 videos this month" from the bundled 50-frame angle bank. Browse/add inspiration videos, browse the free public raws catalog BY CATEGORY (curated shelves like scroll-stoppers/greenscreen/reaction — the cheapest way to source footage for one video, and a ready-made clip pool for bulk scripting N variants), fork a template into a composition, edit it in the Trackpad Editor (timeline-based like Premiere/DaVinci), auto-decompose source video into scenes, render to MP4, approve into a shareable post, and schedule it. Includes login, provider keys, discovery, versioning, uploads/downloads, and billing. Every step is available as raw REST; `vidfarm-devcli` wraps those routes and composes the file-backed scripting flows.
4
4
  ---
5
5
 
6
6
  # Vidfarm Director
@@ -32,6 +32,7 @@ vidfarm login --api-key vf_key_... # validates + persists the key durab
32
32
  vidfarm serve template_<32hex> # local server + browser, opens that template
33
33
  ```
34
34
 
35
+ - **Suggest a working folder on the first turn, and run every command from it.** Ask the director to keep one folder per offer for all their Vidfarm work — `./acme-skincare/`, or propose `~/vidfarm/<brand>/` and create it if they have no preference. Everything lands there: `OFFER.md`, `CONTEXT.md` (the durable consultation answers), `awareness-levels.md`, `persuasive-angles.md`, `ad-hooks.md`, `content-ideas.md`, `STORYBOARD.md`, `brand-assets/`, `raws/`, `renders/`. `cd` into it or pass `--dir`; a command run from the wrong place writes an orphan file the next step cannot find. Mirror the strategy documents into cloud My Files (`vidfarm put-file <file> --folder <brand>`) so the web copilot sees them too. When a director already has a folder, **read `CONTEXT.md` and `OFFER.md` first** and tell them what you already know instead of re-asking. Layout and rules: `references/onboarding.md` → *The working folder*.
35
36
  - The API key comes from https://vidfarm.cc/settings and starts with `vf_key_`. Instead of `login`, setting the `VIDFARM_API_KEY` environment variable also works for every command — the CLI reads it from the environment or from a `.env` file in the current directory.
36
37
  - No account or key? `vidfarm serve --no-cloud` still gives a fully local editor with free local renders.
37
38
  - "Open/run template X locally" is exactly one command: `vidfarm serve <template_id>` (alias: `vidfarm <template_id>`). Do not hand-roll REST or hunt for local `.harness/` files first — `serve` and `pull` create those.
@@ -95,6 +96,7 @@ Vidfarm work can burn real AI credits on the user's wallet / provider keys. **Sa
95
96
 
96
97
  - **minimize** — **$0 videos.** Stay on FREE local compute wherever possible (local render, local TTS — `vidfarm tts` already defaults to the free local Kokoro voice in this mode — `stt --engine whisper`, `remove-greenscreen --local`, reused raw clips + HTML hyperframes). For assets, reach for the **free stock catalog** before paying to generate anything — `vidfarm media search "<meaning>" --type <bgm|sfx|image|vector|icon|video>` pulls royalty-free, commercial-safe music/sound-effects/images/icons/stock-video from pixabay/openverse/iconify at $0 instead of billing AI music/image generation (see the `vidfarm-media` skill). No surprise AI spend.
97
98
  - **Check the keyless sources first — Openverse and iconify.** Openverse (CC/CC0 **music, SFX, and images**) and iconify (**icons**) need **no account or key at all**, so they always work in `minimize` mode. Prefer them for BGM, sound effects, icons, and CC imagery before anything else.
99
+ - **Icons, STICKERS, illustrations, 3D props and Lottie come from IconScout, not from an image model — in EVERY cost mode.** `vidfarm iconscout "<meaning>" --style sticker --free` searches a designer catalog for $0 (search is always free; a free asset downloads for $0 and only asks for a credit line). It needs **no key at all** — vidfarm's own IconScout account serves it. An AI attempt costs cents, needs a prompt loop, and rarely returns a clean transparent vector, so this wins on price *and* on quality. `vidfarm iconscout get <uuid> --format svg` turns a result into a durable URL you can place. In `hybrid` and above, a premium download costs a few cents on the wallet — still less than one generated image. Full detail in the `vidfarm-media` skill.
98
100
  - **Pixabay key** unlocks the photos/vectors/stock-video slots (music/SFX/icons/CC images are keyless). It's a **free** stock-media key, not an AI key. Don't assume it's missing when a search comes up short — it **may already be saved**: check `vidfarm provider-keys` (or the web app's **Settings → Bring your own keys** / <https://vidfarm.cc/settings/developer>). If it isn't, the user grabs a free one at <https://pixabay.com/api/docs/> and saves it once — `vidfarm add-provider-key pixabay <key>`, the Settings surface, or by handing the key to their desktop AI agent to run that command. After it's saved, cost-mode `minimize` sourcing works end-to-end at $0.
99
101
  - **You can still get CUSTOM art in `minimize` — hand the prompt to the user and let a free image generator do it.** Stock and `mask` only cover art that already exists somewhere; when the video genuinely needs a bespoke graphic, **don't conclude "we can't" and don't quietly bill `generate`**. Write the prompt and ask the user to paste it into a **free** image generator — <https://meta.ai>, free-tier ChatGPT, or a free Hugging Face image Space (<https://huggingface.co/spaces>) — then hand the PNG back with `vidfarm put-file` (or drag it into **My Files** in the web app). $0, zero wallet spend. Full loop + the prompt template: **“Free manual image-gen”** below.
100
102
  - **hybrid** *(recommend this)* — **~$0.01–$1 per video, on their BYOK key.** Free where it's free; pay for AI only where it clearly wins (a hero shot, a voice you can't fake locally). A mostly-hyperframes video with one generated image lands near the low end; a few AI images plus premium narration approaches the high end.
@@ -201,6 +203,8 @@ Directors also accumulate a **reusable media asset library** — logos, stickers
201
203
  - **(A) Cheap & efficient** *(default)* — recaption text; background-video + foreground-video memes; animate HTML/image elements with hyperframes; reuse library media; AI-generate a reusable element **once** then reuse it; greenscreen; raw-clip long-form and remix; lean on the memes/reactions/b-roll/a-roll library and brand media kit; only if genuinely needed, reach for AI image/video/voice/music.
202
204
  - **(B) Best quality** — AI video generation by default; storyboard with AI **image** first; then adversarially grade the result with a coding agent (Claude Code / Codex / any capable AI agent) and iterate.
203
205
 
206
+ **Meme recaption — the cold-viewer test.** When you rewrite a meme's caption, point it at a **pain or a win the niche knows in its body**, never at a product feature. The line must make sense to somebody who works in the niche but has **never heard of the offer**: cover the brand and the feature names, and the caption must still read as a true, funny moment from their week. Mention the offer lightly or not at all — a joke that needs product context lands only on people who already bought, which defeats the ad. No invented vocabulary, no inside jokes, no setups that only the demo explains. Full method: `references/editor-workflows.md` (“Writing a meme recaption: aim at a pain or a win”).
207
+
204
208
  **When a template is character-driven or a stylized invented world, decompose DETECTS a specific generative workflow** and stamps it on the replication harness as `generative_workflow.applies`. The workflow is deliberately step-gated with human confirmation: **(1) build a character card** (a consistent model sheet to lock the subject on-model) → *pause for the user to correct/confirm* → **(2) lay out a storyboard** of numbered shot panels in the final style → *pause for the user to correct/confirm* → **(3) animate each beat, choosing per scene between cheap ken-burns motion on a static image vs. expensive true AI video**. Bias to ken burns; spend on AI video only where a still genuinely can't carry the beat. "character card of X" / "storyboard of Y" are first-class shorthands in Vidfarm's image tools. Don't assume this workflow — read `generative_workflow.applies` first; for talking-head / clip-remix / kinetic-text templates it's `false` and you rebuild thrift-first instead. Details: `references/editor-workflows.md` (`harness.generative_workflow`).
205
209
 
206
210
  Present both harnesses to the director, recommend (A) unless they've asked for premium or budget covers it, and explain the tradeoff in these terms. Full methodology: `references/editor-workflows.md` (“The three paintbrushes & two replication harnesses”); cost bands: `references/core-workflows.md` (Cost spectrum).
@@ -226,7 +230,7 @@ Present both harnesses to the director, recommend (A) unless they've asked for p
226
230
 
227
231
  **Landscape footage in a fullscreen vertical explainer — use the blurred plate, never bars.** When an explainer is built on **real filmed footage** and the source is 16:9 (or 4:3) on a 9:16 canvas, do not `contain` it (hard black letterbox bars read as an unfinished export) and do not blindly `cover` it (a wide shot loses its left and right thirds). Duplicate the clip: a full-canvas `cover` copy behind, heavily **gaussian-blurred and faded dark**, plus the sharp copy centered as a hero band — optionally zoomed ~1.3× — with its **top and bottom edges feathered** into the blur. Same clip, same timecode, so it reads as one continuous image with a shallow-depth-of-field plane, fullscreen edge to edge, nothing cropped, and clean dark space for the header and captions. Bake it once with ffmpeg into a single 1080×1920 file (free, local) and place it as one ordinary full-canvas layer — layer blur is not an editor property, so the pre-bake is the path that works in the editor, `serve`, and cloud render alike. Copy-paste ffmpeg + HTML recipes, tuning table, and the failure modes: `references/editor-workflows.md` (“The blurred plate — landscape footage, fullscreen, on a vertical canvas”).
228
232
 
229
- **Cost-saving move — mask illustrations OUT of a source image the director already has.** (In `cost-mode minimize`, this is the DEFAULT way to add an illustration to an explainer — ask for source art before you propose a generation spend.) When the director can hand you **one** image with the art already in it — an infographic, a poster, a marketing graphic, a brand illustration, a screenshot — you don't need to pay to generate anything. `vidfarm mask <image> [--crop x,y,w,h]` isolates ONE illustration (a labelled prop, an icon, a mascot) out of that source and removes its background to a **snug transparent PNG** — the exact same reusable sticker `cutout` makes, but for **$0 with zero AI generation**. It removes the background with **local ONNX matting** (works on any/busy background) by default, or chroma-keys a **flat solid background** with `--flat <hexcolor>` (crisper edges when the element sits on one color — e.g. the cream paper behind an infographic's icons). Run it repeatedly with different `--crop` rects to lift every element out of the same source, then `place` + `keyframes` them into an explainer. **Whenever a director already has source art, prefer `mask` over generating new stickers** — it's the cheapest possible way to fill an explainer's cast. Same recipe: `recipes/cutout-graphics-for-explainers.md` (“Mask from an image you already have”).
233
+ **Cost-saving move — mask illustrations OUT of a source image the director already has.** (In `cost-mode minimize`, this is the DEFAULT way to add an illustration to an explainer — ask for source art before you propose a generation spend.) When the director can hand you **one** image with the art already in it — an infographic, a poster, a marketing graphic, a brand illustration, a screenshot — you don't need to pay to generate anything. `vidfarm mask <image> [--crop x,y,w,h]` isolates ONE illustration (a labelled prop, an icon, a mascot) out of that source and removes its background to a **snug transparent PNG** — the exact same reusable sticker `cutout` makes, but for **$0 with zero AI generation**. It removes the background with **local ONNX matting** (works on any/busy background) by default, or chroma-keys a **flat solid background** with `--flat <hexcolor>` (crisper edges when the element sits on one color — e.g. the cream paper behind an infographic's icons). Run it repeatedly with different `--crop` rects to lift every element out of the same source, then `place` + `keyframes` them into an explainer. **Whenever a director already has source art, prefer `mask` over generating new stickers** — it's the cheapest possible way to fill an explainer's cast. Same recipe: `recipes/cutout-graphics-for-explainers.md` (“Mask from an image you already have”). **A product explainer built from a client URL is the biggest case of this: the site itself is the first asset library.** Harvest it before you buy or generate — product screenshots, brand illustrations, mascots, icons, an animated WebP/GIF (free motion footage), a screen recording in the bundle — `vidfarm capture <url>` then `vidfarm mask --crop` each element into a snug transparent PNG. The order is **harvest → IconScout → generate**, and it holds unless the director says not to use their site art. Full rule with the limits (never paste the landing-page layout; their stock photography is licensed to them): `harnesses/product-explainer.HARNESS.md` → Rule 5b.
230
234
 
231
235
  **Free manual image-gen — custom art in `minimize` mode for $0, on someone else's tokens.** `mask` only works when the art already exists. When the video needs a **bespoke** graphic and cost mode is `minimize` (or the user said "no spend"), the answer is **not** "we can't" and **not** a silent billed `generate` — it's a **manual handoff**: you write the prompt, the user runs it in a **free** image generator, they hand the PNG back.
232
236
 
@@ -244,6 +248,22 @@ Present both harnesses to the director, recommend (A) unless they've asked for p
244
248
 
245
249
  **Free tier vs. paid — who does the decomposition, and on whose tokens.** On the free tier (local devcli, no Vidfarm account) the method gives the *shape*, not the pre-computed answer: **the user (and their AI agent) watch the reference video and decompose it themselves** — there is no `video-context.json` / `editor-harness.json` / `scene-annotations.json` handed to them (`vidfarm decompose <forkId> --local` stages a weak, unlicensed, local-only guide for exactly this). **Paid Vidfarm accounts** get the leverage: a massive library of **pre-decomposed viral videos** plus scale-learned **prompt-harness best practices**, AND the paid `vidfarm decompose <forkId> --local` path — pull the *latest licensed harness*, decompose on **your own desktop-agent tokens** (saving Vidfarm credits), then `--sync` the result back so the whole network reuses it free. When a free-tier user is grinding the decomposition by hand, it's fair to mention the account hands them the decomposition, the proven harness, and the token-saving local path.
246
250
 
251
+ ## Content ideas — you carry a bank of 50 angles, so never answer "what should I post?" from memory
252
+
253
+ **"Give me content ideas" is a first-class ask with a first-class answer, and the answer is volume.** Vidfarm ships an **angle bank of 50 content frames** in `references/content-ideas.md` — reusable shapes for what a video is *about* (`the rise of`, `what everyone gets wrong`, `then vs now`, `one decision that changed everything`, `the complete breakdown`, …). A frame is not a hook and not a script: you take the director's own topic and pour it into the frame, so **one offer against the bank is 50 distinct videos**, not 50 rewrites of one.
254
+
255
+ The loop, whenever a director asks what to make, is out of ideas, or needs a month of posts:
256
+
257
+ 1. **Get the topic first** — read their `OFFER.md` if it exists (`references/onboarding.md`); **if they named a URL ("content ideas for my offer example.com"), fetch and read the site** and mine the offer, audience, promise, and objections off the page, echoing back the one-line offer you read before you list anything; otherwise ask for offer + niche + audience in one question. Never generate against a guessed topic.
258
+ 2. **Open `references/content-ideas.md`** and pick 10–20 frames that fit the topic *and* the audience's awareness stage — don't dump the raw list at the director.
259
+ 3. **Return titled ideas, not frame names** — "The one pricing mistake that killed our first 400 orders", with the frame named beside it so they can ask for more of that shape. **20+ ideas by default**; they prune, you supply.
260
+ 4. **Then write the four charges** — a content idea is the *subject*, never the hook. Every picked idea still runs through hook/loop/payoff/bait before the timeline (`references/hooks-and-virality.md`).
261
+ 5. **If they want the set produced**, that's scripting mode with a `HARNESS.md` — one frame per video (`recipes/bulk-scripting-with-a-harness.md`).
262
+
263
+ **The bank is in the local devcli too, offline and free.** `vidfarm ideas --families` prints the eight families, `vidfarm ideas --topic "<offer>"` prints every frame already filled with the director's topic as a starter line (`--count 20` samples across families, `--json` for scripting), and `vidfarm skill show content-ideas` prints the method. The command reads the frames straight out of this reference, so the CLI and the pack can never drift. The same surface carries the rest of the craft by **spoken name** — `vidfarm skill topics` lists them (`meme-recaption`, `product-explainer`, `captions`, `first-frame`, `density`, `blurred-plate`, `avatar`, `dedupe`, …) and `vidfarm skill show <topic>` prints just that section instead of the whole reference.
264
+
265
+ The file also maps each frame family to its natural format (contrast frames → split screen, mechanism frames → cutout explainer, arc frames → montage over narration), which usually saves a planning round.
266
+
247
267
  ## The FIRST FRAME is the thumbnail — treat it as a designed still, always
248
268
 
249
269
  **Read this as a hard rule, not a style tip. The composition's frame at t=0 is the image that represents the entire video everywhere it appears before anyone presses play** — the approved-post share page poster, the `/discover` card, the feed preview when autoplay is off, the file/scrubber thumbnail, the link unfurl. **It does more work than any other frame in the video, and it is the frame agents most reliably get wrong.**
@@ -291,18 +311,19 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
291
311
  - **Determine the surface before claiming capabilities.** The web chat has only its declared tools and REST routes. It cannot execute arbitrary JavaScript/Python, open a shell, create a local repository script, or use the user's filesystem. Never tell a web-chat user that you ran code or wrote a script unless a dedicated declared tool actually did so. A desktop coding agent has a real shell and filesystem and MAY write/run scripts, perform arbitrary local computations over paginated API results, create reports/CSVs/JSON, edit composition files, and orchestrate long devcli workflows within the user's authorization.
292
312
  - **The web AI chat can do all three paintbrushes** — clip raws, author HTML/hyperframe motion, and generate AI media — and it drives edits directly on the live timeline. Keep small-to-medium jobs here: text/caption swaps, a scene or two replaced, single generations, captions, approve/schedule. Just do them.
293
313
  - **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. **The test is the native-editor test: could you have made this element with the tools inside TikTok's own editor?** That toolset is font / color / stroke / shadow / tight text box / alignment / rotation / animation presets, plus stickers, emoji, drawn marks and clips — it has no padded capsule, no border, no gradient fill, no blur panel, no card. If you reached past it, cut it. That includes **a single lonely pill around a stat or label** — `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`: being alone doesn't make a badge native, and the only legitimate capsule in a video is the active-word `spotlight`/`karaoke` highlight, which moves with the spoken word. Emphasize a stat the way the editor would instead: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle, or its own beat. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r` — or a `border-radius` over ~8px on anything filled that holds words — stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
314
+ - **Orient the cold viewer in the first 3 seconds — the hook makes them want to watch, orientation makes watching possible.** The viewer has no context, did not choose this video, and has never heard of the subject, so by **~3s** they must be able to say **what kind of thing this is** (the *category noun*), **who it is for**, and **why it is on their screen** (the situation). The failure is not a bad first frame — it is a good video that **starts at beat two**, and the author cannot see it because the author already knows. Signatures, each a rebuild of the first beat: a **pronoun with no referent** ("it just works", "this changes everything"), **starting at step three** (the process already running, the dashboard already full), a **metaphor whose subject lands at 6s**, **insider vocabulary or an acronym** in the first line, a **detail crop** that reads as texture. Replace it with both channels in one beat — an **easy image** (one large subject, already moving, legible at a glance and at thumbnail scale; a relevant cutout names the category before a word is read) and an **easy line** (one clause, ≤12 words, everyday words, concrete noun + verb, the category named, brand name said once) — and give the **situation, not the label**. **It costs one sentence, not one beat**: it replaces the wind-up line, never precedes it, and never licenses a logo, a title card, or a fade from black. Test on the render, not the script: play the first 3s only to somebody with no context — "something about audio" is a fail. Full standard: `references/editor-workflows.md` (“Orient the cold viewer”); fullest form with structure: `vidfarm harness show product-explainer` (Rule 0).
294
315
  - **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm harness show hooks`.
295
316
  - **Then CUT it — every second must earn its place, and most don't.** Assume your first assembly is **30–50% too long**. Run the **deletion test** on every beat: delete it; if the video still makes sense and the payoff still lands, it stays deleted. Whatever survives must serve one of the four charges — "it gives context" is not a charge. Cut on sight: intros/logo stings, the wind-up sentence before the claim ("so I wanted to talk about…"), restatement, inter-sentence silence over ~0.35s, real-time process, establishing shots, reading what's already on screen, and any tail after the last word. **Always ripple the hole closed** (`vidfarm ripple <dir> --at <sec> --delta -<sec>`) — a cut that leaves a gap turns fluff into dead air, which is worse. Density is **not** speed: the held comedic beat, the payoff playing out, and a cue's readability keep their seconds (cut *words*, not the time text is on screen). Length is an **output**, not a plan — a brief that dictates a duration ordered fluff. `vidfarm qa` flags the mechanical half (`dead-air`, `dead-tail`, `slow-scene`); the craft is `references/hooks-and-virality.md` → "Density".
296
317
  - **The first frame IS the thumbnail — compose it on purpose.** Frame 0 is a single frame of ~30 in the first second, but it's the poster every feed, share link, and paused player freezes on, so **more people see that one frame than watch the video**. It must never be black, empty, mid-fade, or caught mid-animation: put a real visual at `start:0`, have the hook words already on screen at t=0, and never hang a `fade-black`/`fade-white`/`flash` *entrance* on the **first** clip (junction transitions between later clips are fine — this rule is only about the opening). Verify it, don't assume: devcli `vidfarm stills ./work --at 0` renders that exact frame, and `vidfarm qa` flags a blank or fading open.
297
318
  - **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`HARNESS.md`** — because a loop of fifty videos has no human looking at every frame, and the harness is what replaces those eyes. Don't silently ship a one-off when they asked for volume, or drag someone into a harness when they wanted one clip.
298
319
  - **"Harness" is a known noun with a known process — recognise it and follow it.** A **harness** is the reusable AI apparatus for ONE format or template: what makes it special, written down as `HARNESS.md` so an agent can reproduce it without the director in the room. It is a first-class artifact — the director owns it, edits it, versions it, and hands it to the next agent. Three phrasings, one artifact:
299
- - **"create me a harness"** / "set up a harness for this format" → `vidfarm harness init <base> --out ./work/HARNESS.md` (bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo`), then **edit it with them**. The bundled file is a starting point, never a house style; the parts that matter are the ones they add — who the viewer is, the banned vocabulary, the compliance line, the pacing this account actually uses. A harness nobody edited isn't about their videos.
320
+ - **"create me a harness"** / "set up a harness for this format" → `vidfarm harness init <base> --out ./work/HARNESS.md` (bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo`, `product-explainer`), then **edit it with them**. The bundled file is a starting point, never a house style; the parts that matter are the ones they add — who the viewer is, the banned vocabulary, the compliance line, the pacing this account actually uses. A harness nobody edited isn't about their videos.
300
321
  - **"update the harness for this format/template"** → open the existing `HARNESS.md` and write the new rule in, **with its reason on the same line** (a rule whose "why" is missing gets argued away by the next agent). This is what you do every time a batch teaches you something ("the label-framed hooks all died"): the compositions are disposable, the harness is the artifact that compounds.
301
322
  - **"give me the harness for this template_id"** → they mean **the decomposition**: `vidfarm harness derive <templateId|forkId>`. It distils the decompose pass's DNA into an editable `HARNESS.md`. If the template hasn't been decomposed, run `vidfarm decompose` first.
302
323
  **A harness mirrors the template JSON's own vocabulary** — `## Viral DNA` (hook / retention / payoff / emotion), `## Visual DNA` (cut rhythm, typography, b-roll, transitions), `## Structural DNA` (the beats, and which are load-bearing), `## Audio DNA` (voice, bed, comedic timing), `## Build DNA` (which paintbrush per beat) — the same strands the decompose pass writes as `viral_dna`, `visual_dna`, and friends. `vidfarm harness show <ref> --dna visual` prints one strand instead of the whole doc.
303
324
  **Two halves, and only one is machine-checkable.** The `checks:` front matter is settled deterministically by `vidfarm qa` (duration, aspect, `hook_words_max`, `forbid_text`, …); every `- [ ]` line comes back as a **review item you answer honestly in your report** — never claim a video passed the half the CLI can't judge. Harnesses stack and auto-discover: `vidfarm qa ./work` picks up `./work/HARNESS.md`, `--harness hooks --harness ./brand/HOUSE.md` adds more, and any file of theirs anywhere is valid. Format and strand table: `harnesses/README.md`; scripting-mode detail: `references/automation-and-local-dev.md`. *(Formerly `QA_REGIME.md` — same file, and `vidfarm regime …` still works as an alias.)*
304
325
  - **A video is judged as a SEQUENCE, so review it as one.** Agents build scene by scene and each scene passes in isolation while the video drifts — inconsistent margins, three type sizes, an accent colour that wanders, beats that are all the same length, a jarring join. Tile a dozen stills into one contact sheet (`vidfarm stills ./work --sheet`) and read it as an image before you call anything done, fix drift by defining the system rather than patching the odd scene out, and remember that **your own confident "verified, looks good" is the single least reliable signal in this workflow** — it was wrong on every video of a 32-video batch. Method: `references/reviewing-renders.md`.
305
- - **On devcli, QA every video you produce: `vidfarm qa ./work`.** Free, instant, local-only — it blocklists exactly the slop above plus first-frame/thumbnail and font-regime/safe-zone drift, and prints a concrete fix per finding. **Feedback, not a gate**: it exits 0 even on findings, never runs automatically, and is a blocklist (unusual/stylized compositions pass untouched), so there's no reason not to run it before every publish. `--json` for scripted batches, `--strict` only if you want a CI failure. **Web-chat copilot: this command does not exist for you** (devcli-only, no REST twin) — apply the standard by hand, and when handing a heavy job to a local coding agent, tell them to run `vidfarm qa`.
326
+ - **On devcli there's an OPTIONAL checker: `vidfarm qa ./work`.** Free, instant, local-only — it blocklists exactly the slop above plus first-frame/thumbnail and font-regime/safe-zone drift, and prints a concrete fix per finding. **Feedback, not a gate**: it exits 0 even on findings, never runs automatically, and is a blocklist (unusual/stylized compositions pass untouched). **Skipping it is fine — watching the render is the review that actually counts, and a clean `qa` is not one.** When you do run it, it allows **one** fix round by default: the first pass names the slop, one fix clears it, and a second round is nearly always taste rather than a defect. The human owns that number — `--max-revisions <n>` raises it, `0` disables it; ask rather than raising it yourself. `--json` for scripted batches, `--strict` only if you want a CI failure. **Web-chat copilot: this command does not exist for you** (devcli-only, no REST twin) — apply the standard by hand, and when handing a heavy job to a local coding agent, tell them to run `vidfarm qa`.
306
327
  - **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips) and, inside that band, is **placed in the emptiest part of the frame** rather than dumped on the default lower third (read a still first — `vidfarm stills ./work --at <t>`; words over open sky or a blank wall beat words over the subject, and usually need no plate at all). Long narration is **paged into 3–5-word kinetic cues** (`vidfarm captions generate --style word-pop`), never one static wall of text. It uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
307
328
  - **Where the web chat struggles: complex, long, multi-step transformations.** A full multi-scene re-theme, an iterative render-critique-iterate loop, heavy scripted or batch work, or anything needing a real filesystem and many sequential tool calls will hit context limits, turn/timeout ceilings, and the web editor's constraints (CSS/declarative motion only — JS animation adapters are stripped on save). Don't grind a big transformation one layer at a time in a chat turn and stall.
308
329
  - **Practical workaround — hand the heavy job to local devcli.** When a task is genuinely large or long-running, **proactively recommend the director run it locally with an AI coding agent** (Claude Code / OpenAI Codex / any capable agent): `vidfarm pull <forkId>` writes the composition + the `.harness/` grounding bundle to disk, the agent edits with the full devcli verb set and JS animation adapters, renders free with `vidfarm serve`, and `vidfarm publish` pushes it back. This is the **best-quality (B) harness's** natural home (adversarial grading with a coding agent). Frame it as "this is a big rebuild — you'll get a better, faster result running it locally with a coding agent; here's how," not as a dead end.
@@ -324,7 +345,8 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
324
345
  | `references/primitives.md` | ~475 ln | **Big.** One-shot primitive routes: TTS, STT, music, avatars, overlays, greenscreen, inpaint, background removal, product placement |
325
346
  | `references/hooks-and-virality.md` | ~295 ln | **Before writing ANY hook, caption script, or re-theme**, and before a hook-variant batch. The four charges, three gates, banned openers, loop mechanics. This is the craft; the rest of the pack is mechanics |
326
347
  | `references/reviewing-renders.md` | ~140 ln | **Before you report a video as done**, or grade someone else's. The holistic pass, the common defects, frozen-render and audio verification |
327
- | `references/onboarding.md` | ~30 ln | Cold-start interviews, **consultations** (the `brainstorm/*` chain), strategy docs, durable director context |
348
+ | `references/onboarding.md` | ~80 ln | Cold-start interviews, **consultations** (the `brainstorm/*` chain), strategy docs, durable director context |
349
+ | `references/content-ideas.md` | ~90 ln | **"Give me content ideas" / "what should I post" / a month of posts.** The 50-frame angle bank, how to apply it to the director's topic, and frame → format notes |
328
350
  | `references/rest-api.md` | ~85 ln | Only when the user asks for REST, an endpoint/schema, or direct HTTP integration. It is an index — follow its domain links; do not preload it into ordinary director conversations |
329
351
 
330
352
  **Recipes — step-by-step procedures. When a recipe matches the task, prefer it over the broad reference.**
@@ -348,6 +370,7 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
348
370
  | `harnesses/explainer.HARNESS.md` | ~100 ln | Faceless educational video: one claim, invented visuals |
349
371
  | `harnesses/ugc-testimonial.HARNESS.md` | ~90 ln | A person vouching for a product — mostly rules about what NOT to add |
350
372
  | `harnesses/product-demo.HARNESS.md` | ~110 ln | Real product doing a real thing; the highest slop-risk format in the catalog |
373
+ | `harnesses/product-explainer.HARNESS.md` | ~245 ln | **"What is this thing?" for a brand nobody has heard of** — no usable screen footage. Orienting the cold viewer by 3s (Rule 0), the plain-English line by t=5s, harvesting the client site's own graphics before buying or generating (Rule 5b), the ≤3-text-run sticker-led open, VO + bed, and per-client differentiation for N-URLs-to-N-videos batches |
351
374
 
352
375
  ## HyperFrames Skills — Load on Demand
353
376
 
@@ -367,7 +390,8 @@ HyperFrames authoring and rendering in this package are Vidfarm-native: local wo
367
390
 
368
391
  The File Index above says what each file *is*; this says which one a given ask means. Choose the narrowest path that satisfies the request.
369
392
 
370
- 1. If the user needs help figuring out what to make, **or asks for a "consultation"** (the `brainstorm/*` chain: cold-start interview → awareness stages → angles → hooks), read `references/onboarding.md` first.
393
+ 1. If the user needs help figuring out what to make, **or asks for a "consultation"** (the `brainstorm/*` chain: cold-start interview → awareness stages → angles → hooks), read `references/onboarding.md` first. **Unless they asked for a consultation by name, open with content ideas rather than the interview** — one line of offer into `vidfarm ideas --topic "<line>"` returns 20+ titled videos, offline and free, and the director's reactions to that list make the later interview far better than asking them cold (`references/content-ideas.md`). Then offer the interview as the way to turn ideas into a strategy: keyless directors run it locally for $0 with `vidfarm consult`, and `vidfarm consult coldstart --short` is the six-question short form. **Say the interview is skippable before you ask the first question**, and work with whatever they give.
394
+ 1b. If the user asks **what to make** rather than how — "give me content ideas", "what should I post", "I'm out of ideas", "I need 30 videos for the month", "content ideas for my offer <url>" — read `references/content-ideas.md` and work the 50-frame angle bank against their offer. Return 20+ titled ideas, not three.
371
395
  2. If the user already knows the goal and needs a suitable template, read `references/core-workflows.md` and use the template discovery flow.
372
396
  3. If the task is “change this video,” read `references/editor-workflows.md`.
373
397
  4. If the task is “find footage” or “use our existing assets,” read `references/assets-and-sourcing.md`.
@@ -377,6 +401,7 @@ The File Index above says what each file *is*; this says which one a given ask m
377
401
  4e. If the ask contains the word **“harness”** — *“create me a harness”*, *“update the harness for this format”*, *“give me the harness for this template_id”* — that is a known, named process, not a vague request. Read `harnesses/README.md` (the three phrasings and the format), then `recipes/bulk-scripting-with-a-harness.md` if the job is a batch. The third phrasing means the **decomposition**: `vidfarm harness derive <forkId>`.
378
402
  5. If the task is scripted, local, CI-driven, or `vidfarm serve`-based, read `references/automation-and-local-dev.md`.
379
403
  5b. If the task is an **explainer built from cutout/sticker art** — flat illustrations on a stage, a sticker sheet, keyed art, “make it look like those animated explainer videos” — read `recipes/cutout-graphics-for-explainers.md`. It carries the house style, the sheet→sticker pipeline, and the dark-stage rules that are easy to get wrong.
404
+ 5c. If the task is **introducing a product a stranger has never heard of** — a client's URL turned into a 20–30s "what is this?" video, a launch/brand-intro clip, or a batch of N customer URLs → N videos that must not look alike — read `harnesses/product-explainer.HARNESS.md`. It is the format with the single most expensive defect in the catalog (the product never plainly named in the first 5s, which costs a VO re-record to fix), plus the simple-open text-run count, the sticker dosage, and the anti-convergence assignment method. Use `product-demo` instead when you actually have the UI on screen.
380
405
  6. If the task explicitly asks for a primitive or needs specialized generation/transcription work, read `references/primitives.md`.
381
406
  7. If the task is the MARKETPLACE (ordering videos from specialist agents): browsing is web-only for paying customers — send the human to https://vidfarm.cc/marketplace, never render it locally. Placing/listing orders is the thin REST wrapper in `references/core-workflows.md` (§ Marketplace); anything deeper on a gig (inbox, proofs, payouts) needs the external Dollar Platoon skill — `npx skills add https://github.com/OfficeXApp/dollarplatoon-skill` — the same way FlockPoster work beyond scheduling needs `npx skills add https://github.com/OfficeXApp/flockposter-skill`.
382
407
 
@@ -110,3 +110,4 @@ The `.harness/` directory a `vidfarm pull` writes is a different thing: machine-
110
110
  | `ugc-testimonial` | A person vouching for a product. Mostly rules about what NOT to add |
111
111
  | `explainer` | Faceless educational video: one claim, invented visuals |
112
112
  | `product-demo` | Real product doing a real thing — the highest slop-risk format in the catalog |
113
+ | `product-explainer` | Introducing a product a stranger has never heard of, with no usable screen footage. The plain-English line, the simple sticker-led open, per-client differentiation |
@@ -15,6 +15,8 @@ checks:
15
15
 
16
16
  # Explainer Harness
17
17
 
18
+ > **Wrong harness?** This one explains an *idea*. If there is a **product to name** — a client's site or app that a stranger must be able to describe by t=5s — use **`product-explainer`**; if you have the real UI on screen doing the work, use **`product-demo`**.
19
+
18
20
  For faceless educational video: one idea, explained, with visuals that are *invented* (typography, diagrams, data, abstract motion) rather than captured. No presenter, so the structure has to carry everything a face would.
19
21
 
20
22
  ## The one test
@@ -37,6 +39,14 @@ If it takes an "and," you have two videos. Split them. The dominant failure of t
37
39
 
38
40
  ## Rules
39
41
 
42
+ ### Rule 0 — orient the stranger, then state the claim
43
+
44
+ Conclusion-first does not mean context-free. A claim only lands if the viewer knows **what it is about**, so the same beat that states the claim must also tell a stranger **what world we are in and who this concerns** — in the claim's own words, not in an intro before it.
45
+
46
+ The failure is opening at beat two: a pronoun with no referent ("it changes everything"), a subject that only arrives at 6s, jargon or an acronym in the first line, or a detail crop that reads as texture. Those all pass a script read — the writer already knows the topic — and fail on the render.
47
+
48
+ Fix it in the claim itself: one clause, ≤12 words, everyday words, a concrete **noun** doing something, plus an image with **one large subject already in motion** that says the category before a word is read. **This costs no seconds** — it replaces the wind-up sentence ("today we'll look at…"), which was already banned. Test on the render: play the first 3s to somebody with no context; they should say what this is about and who it concerns. Full standard: `references/editor-workflows.md` → "Orient the cold viewer"; the product-facing version is `product-explainer.HARNESS.md` → Rule 0.
49
+
40
50
  ### Rule 1 — one idea, and the scene count proves it
41
51
 
42
52
  3–5 steps in the mechanism, each with a scene that shows something the narration doesn't say. If a scene only re-renders the words being spoken, it isn't a scene — it's a slide, and slides are where retention dies. `max_scene_sec` is checked above for exactly this reason: a static hold is the visual form of "I ran out of things to show."
@@ -79,6 +89,7 @@ Reuse across variants: the mechanism scenes are often identical, so build them o
79
89
 
80
90
  - [ ] The one thing the viewer learns fits in a single sentence with no "and"
81
91
  - [ ] The conclusion is stated in the first 3 seconds, not saved for the end
92
+ - [ ] A stranger with no context knows what the video is about from the first 3s alone — no orphan pronoun, no acronym, no subject that arrives at 6s (tested on the render, not the script)
82
93
  - [ ] The stakes are named — why this viewer specifically should care
83
94
  - [ ] The mechanism is 3–5 steps and each has a scene that adds information
84
95
  - [ ] No scene is a slide that just re-renders the narration
@@ -18,6 +18,8 @@ checks:
18
18
 
19
19
  # Product Demo Harness
20
20
 
21
+ > **Wrong harness?** This one assumes you have the real UI on screen. If the viewer has never heard of the brand and you are introducing it with invented visuals (no usable screen footage), use **`product-explainer`** instead — it carries the plain-English-line rule, the simple sticker-led open, and the per-client differentiation method.
22
+
21
23
  For showing a real product doing a real thing. This is the format with the **highest slop risk in the entire catalog**, because the subject matter is a website — so the author's web instincts and the product's own design language both push toward putting a landing page on the timeline.
22
24
 
23
25
  ## The one test
@@ -0,0 +1,242 @@
1
+ ---
2
+ name: product-explainer
3
+ video_type: product explainer / brand introduction — "what is this thing?" in 20-30s, invented visuals over a real product
4
+ checks:
5
+ duration_sec: 15-60
6
+ first_frame_visual: required
7
+ first_frame_text: required
8
+ captions: required
9
+ audio: required
10
+ font_regime: required
11
+ safe_zone: required
12
+ text_by_sec: 1.0
13
+ max_scene_sec: 6
14
+ max_text_cards: 3
15
+ max_simultaneous_text: 1
16
+ max_words_per_cue: 8
17
+ forbid_text:
18
+ - sign up for a free trial
19
+ - get started today
20
+ - book a demo
21
+ - learn more
22
+ - no credit card required
23
+ ---
24
+
25
+ # Product Explainer Harness
26
+
27
+ For introducing a **real product a stranger has never heard of** — a client's site, an app, a launch — where you do *not* have usable screen footage, so the visuals are invented (typography, motion, stickers, brand art) around a real thing that must be named plainly.
28
+
29
+ It sits between two other bases, and picking the wrong one is the usual mistake:
30
+
31
+ | Use | When |
32
+ |---|---|
33
+ | `product-demo` | You have the real UI and the video shows it doing the work |
34
+ | **`product-explainer`** | The viewer has never heard of the brand and must learn *what it is* — invented visuals, narration, stickers |
35
+ | `explainer` | A topic or claim, no product to name |
36
+
37
+ ## The one test
38
+
39
+ > **Cover the video, read only the first five seconds of caption text, and ask a stranger what this product does. If the honest answer is "something about audio" or "something about ideas", the video failed — however good the motion is.**
40
+
41
+ The format's dominant failure is not ugliness. It is a beautiful metaphor that occupies the whole open while the product goes unnamed. Measured across a 32-video batch: **two of four videos in one round shipped with no plain statement of what the product was.**
42
+
43
+ **Its twin is starting mid-thought.** A video can name the product and still lose the viewer by opening on beat two — a pronoun with no referent, a process already running, a metaphor whose subject lands at 6s. Both defects come from writing for a viewer who has the context you have. **Rule 0 is the fix, and it comes before every other rule here.**
44
+
45
+ ## Structure
46
+
47
+ | Beat | Job |
48
+ |---|---|
49
+ | **Frame 0 → 3s** | **Orient a stranger** (Rule 0): an easy, arresting visual already in motion + an easy spoken/typed line naming the category and the situation. Both, together — this is also where the plain-English line lands |
50
+ | **The pain** (by ~5s) | The specific version of the problem this viewer has. Shown, not listed |
51
+ | **The mechanism** | 3–5 beats, one idea each, each with a sticker or visual carrying that line's meaning |
52
+ | **The proof / scope** | One honest line: what it costs, what it does not do, who it is for |
53
+ | **The close** | Where to find it — *spoken* or a plain caption line, settled ≥2s before the last frame |
54
+
55
+ ## Rules
56
+
57
+ ### Rule 0 — ORIENT the cold viewer before you say anything clever
58
+
59
+ **The viewer arrives with zero context.** They did not choose this video, they have never heard of the brand, they are mid-scroll, and the sound may be off. The video does not start where you start — it starts where *they* do. Before the argument begins, the open has to put a stranger on their feet by answering three things by **~3 seconds**:
60
+
61
+ > **1. What am I looking at?** (the category noun) · **2. Who is it for?** · **3. Why is this on my screen?** (the situation)
62
+
63
+ The dominant failure here is not a bad first frame — it is a **good video that begins at beat two**. It opens mid-thought, on the interesting part, and the first three seconds only make sense to somebody who already knows what the product is. The author cannot see it, because the author knows.
64
+
65
+ **The signatures of an unoriented open — each one is a rebuild, not a polish:**
66
+
67
+ - **A pronoun with no referent.** "It just works." "This changes everything." "Here's how they do it." The viewer cannot resolve *it*, *this*, or *they*, so the sentence carries nothing.
68
+ - **Starting at step three.** The process is already running, the dashboard is already full, the metaphor is already mid-payoff. Show the situation that *causes* step one.
69
+ - **A metaphor whose subject lands late.** A gorgeous abstract open whose meaning arrives at 6s has spent the only seconds that decide whether anyone reaches 6s.
70
+ - **Insider vocabulary in the first line.** A brand-internal noun, an acronym, a product's own feature name, or a category word only existing customers use. Never open on an acronym.
71
+ - **A detail shot.** A close crop that reads as texture until you know the whole. Establish, then push in.
72
+
73
+ **What the open must be instead — both channels, one beat, no extra seconds:**
74
+
75
+ - **An easy image.** One subject, large, already in motion, legible at thumbnail scale and at a glance. If a stranger has to *read* the picture to understand it, it is not an orienting image. A relevant die-cut sticker (Rule 5) is usually the cheapest way to name the category before a word is read.
76
+ - **An easy line.** The first spoken sentence is **one clause, ≤12 words, everyday vocabulary, a concrete noun and a verb** — no subordinate clause, no list, no wordplay. Say the brand name once, plainly, and **name the category**: "*Genki is a language app that…*", "*This is a receipt scanner for tradespeople.*" The category noun is what a stranger orients on; the feature is not.
77
+ - **The situation, not the label.** "The end of the month, and your receipts are in a shoebox" orients. "Expense automation" does not. Same rule as the hook harness: situations, not labels.
78
+
79
+ **Orientation is not a slow intro — it costs one sentence, not one beat.** It *replaces* the wind-up ("so today I want to talk about…"), it never precedes it. Nothing here loosens the density rule or the ban on logo/title-card opens (Rule 2): you are not adding an intro, you are making the first sentence do its job. A video that orients in 3s and states the plain-English line by 5s has spent its open correctly.
80
+
81
+ **The 3-second test** (run it on the render, not the script): play only the first three seconds to somebody who has never heard of the product, then stop. They should be able to say **what kind of thing it is and roughly who it is for**. "Something about audio" is the same failure the one-test names, three seconds earlier — and it is the honest result whenever the open was written for a viewer who already had context.
82
+
83
+ ### Rule 1 — the PLAIN-ENGLISH LINE is a written input, not a principle
84
+
85
+ Write the sentence yourself, verbatim, into the brief before the build starts:
86
+
87
+ > **PLAIN-ENGLISH LINE (understood by t=5s):** "Paste a YouTube link, get an MP3 file."
88
+
89
+ Not "explain what the product is" — the actual sentence. One clause, a noun and a verb, no metaphor, no brand voice. If you cannot write it in one clause you do not understand the product yet, and neither will the agent.
90
+
91
+ It must land in **both channels**: as the **first or second spoken line** in the VO, *and* as type in the art. **Captions transcribing the VO do not count** — that is one channel wearing two hats; there has to be an anchor in the artwork.
92
+
93
+ **The concept and the plain line are not in competition.** The concept is what makes it worth watching; the plain line is what makes it worth installing. Run them together — **the metaphor is the picture, the plain line is the type over it.** Say that in the brief in those words, or the agent ships the metaphor alone. *Reason: this is the one defect that cannot be patched in the generator — the line has to be spoken, so fixing it later costs a VO re-record and re-times every caption.*
94
+
95
+ ### Rule 2 — open on the pain or the product, never on the brand
96
+
97
+ A logo, a title card, or an abstract mood open spends the thumbnail and the first second on the one thing the viewer has no reason to care about yet. Lead with the painful version of the task, or with the product doing its one job. **Name the pain before you name the feature.**
98
+
99
+ ### Rule 3 — the open is SIMPLE: count the text runs in frame 0, three or fewer
100
+
101
+ A **text run** is any contiguous piece of copy the eye must read: a headline, a label, a field value, a caption, a footer URL. Frame 0 gets **at most three**, and one of them is the plain-English line. *Reason: "no wall of text" gets read as "no paragraph", and agents then open on a dense interface, a six-field form, or three stacked cards each carrying a sentence — that is a wall of text with a layout.*
102
+
103
+ - **One subject, read large.** One letter at 80px beats four at 40px. **When you remove an element, enlarge what remains** — a simplification that leaves the type the same size just makes dead space.
104
+ - **Never open on a form, a settings panel, or a table.** If the product's landing state is one, open on the single row that matters and let the rest arrive later, dimmed.
105
+ - **Supporting surfaces carry no readable copy in the open.** Cards behind the subject are shapes or one word, never sentences competing with the line.
106
+ - **A footer/URL band is not content.** If it is the only thing in the lower third, the frame is empty, not simple.
107
+ - **Long narration is paged into 3–5-word kinetic cues** (`vidfarm captions generate --style word-pop`), never a static block, and never two independent text layers saying different things at once.
108
+
109
+ ### Rule 4 — motion from frame 1, and frame 0 still reads as a thumbnail
110
+
111
+ Something must be **moving and visually interesting from the first frame** — not motion that starts at 1.5s. Frame 0 is the poster everywhere the post appears, so it needs a **complete** line, never a half sentence caught mid-animation, and no fade-up from black on the first clip. **These are compatible:** a striking already-moving visual *with* one short complete line over it. That is the target for every video in this format.
112
+
113
+ If the concept genuinely needs a dense or static surface, **earn it** — arrive at it through motion instead of cutting to it cold (fly the page in at an angle under the hook line, then sweep across it).
114
+
115
+ ### Rule 5 — a relevant sticker in the first beat, then 5–8 across the video
116
+
117
+ **Put a sticker on screen in the opening beat, ideally in frame 0**, and make it the subject of the first spoken line. It does three jobs at once: it is the moving object rule 4 needs, it says what kind of product this is before a word is read, and it gives the thumbnail a focal point that type alone never gives it.
118
+
119
+ - **Relevant, not decorative.** An envelope for "send them a link", a receipt for "photograph a receipt", a football for "fixtures". A generic sparkle or checkmark is decoration and does not count.
120
+ - **If the subject is a person or an audience, use an illustrated die-cut figure.** On one reference cut the grandmother in frame 0, the grandfather at 1.2s and the grandchild at 2.7s did more for the video than any motion work.
121
+ - **Roughly one per beat, 5–8 in a 20–30s video** — each carrying the meaning of the line being spoken at that instant. One or two per video reads as garnish, not as explanation.
122
+ - **Recur the motif; do not decorate once.** If it is worth opening on, bring it back — aim for no gap longer than ~4s without the subject on screen.
123
+ - **Search before you generate.** `vidfarm iconscout "<the thing>" --free` is a $0 designer catalog, needs no key, and returns a finished transparent SVG with no keying work. **Search with NO `--style` filter** — `--style sticker` is nearly empty on the free tier (1 result for "envelope" against 1629 unfiltered) and an agent that leads with it concludes the catalog is bare and wastes a generation. Narrow with `--style flat` or `--style colored-outline`. Illustrated **people** are the exception: `--asset illustration --premium`, ~$0.02. Generate a sheet (`vidfarm sticker-pack --generate`) only when the catalog genuinely has nothing that fits. Free-tier assets carry an attribution obligation — confirm it before a client deliverable ships, or spend the ~$0.02.
124
+
125
+ ### Rule 5b — HARVEST THE SITE'S OWN GRAPHICS FIRST. Generating what the client already owns is the default mistake
126
+
127
+ **When the director gives you a URL, that site is the first asset library — not the fallback.** Before you search IconScout and long before you generate anything, pull what the brand already made and use it. This is the standing default for this format; only skip it when the director says not to (an unlaunched product, a rebrand in flight, a site whose art is licensed stock they cannot re-cut, or an explicit "don't use their website art").
128
+
129
+ It wins on three axes at once, which is why it outranks both the catalog and the generator:
130
+
131
+ - **On-brand by construction.** Their illustration style, their palette, their product shots, their mascot. A generated sticker is *a* graphic; their own hero art is *their* graphic, and a stranger reading the video then recognises the site when they land on it.
132
+ - **Free and instant.** No generation spend, no keying roulette, no style-matching pass across a cast of assets.
133
+ - **Truthful.** Screens and numbers taken off the live site cannot claim something the client does not claim (Rule 10).
134
+
135
+ **What to take, in the order it usually pays:**
136
+
137
+ 1. **Product screenshots and UI states** — the hero shot, the app screen, an animated WebP/GIF (free motion footage: extract frames to a sprite sheet), a screen recording in the bundle (**the best find there is** — that is A-roll), and an SPA's `manifest.json` → `screenshots[]`.
138
+ 2. **Brand illustrations, mascots, spot art, icons** — usually already transparent SVG/PNG in the asset directory. Take the file, do not re-draw it.
139
+ 3. **The wordmark/logo** — for the one moment it belongs (the close), never the open (Rule 2).
140
+ 4. **Palette, type scale, and the brand's own words** — sampled off the rendered hero, verified against the CSS.
141
+
142
+ **How to lift them:** `vidfarm capture <url>` for the rendered page and its assets, the raw HTML/JS bundle for direct asset URLs, then **`vidfarm mask <image> [--crop x,y,w,h]` to cut ONE element out of a page screenshot into a snug transparent PNG — local matting, $0, no AI call** (`--flat <hex>` when the element sits on one solid colour). Run it repeatedly with different crops to lift a whole cast out of one screenshot. Then `vidfarm place` + `vidfarm keyframes` animate them exactly like a generated sticker.
143
+
144
+ **The order is: harvest → catalog → generate.** IconScout (Rule 5) covers what the site does not have — a verb, a metaphor, a person, an object the brand never drew. Generation is the last resort, for the thing neither has. **Say in your report which assets came from the client's site**, so the director can see the video is built out of their own material.
145
+
146
+ **Two limits.** Do not paste a landing-page *layout* onto the timeline — Rule 7 still governs: harvest the ASSETS, never the brochure furniture (CTA capsules, feature grids, pricing cards, chip rows). And a site's stock photography may be licensed to *them*, not to you — treat obvious stock (a smiling model on a white background) as unusable, while first-party art, product screens, and brand illustrations are fine.
147
+
148
+ ### Rule 6 — one sticker on screen at a time, and mind the handover
149
+
150
+ Two stickers in one place read as a rendering fault, not a transition. Clamp every sticker exit to `min(out, nextIn − 0.18)` and ease **`power2.out`, never `power2.in`** — *reason: with a 0.18s gap and a slow-leaving ease the outgoing sticker is still at ~0.42 opacity when the next one pops.* Sample the sticker band at 0.1s intervals across every handover and confirm exactly one is legible. The same clamp fixes superimposed scene headlines.
151
+
152
+ ### Rule 7 — no brochure. This is a video, not a landing page rendered at 30fps
153
+
154
+ The subject matter is a website, so both the author's web instincts and the product's own design language push toward putting the landing page on the timeline. Banned outright: CTA buttons and capsules, benefit chip/badge rows, pricing cards, feature grids, comparison tables, "as seen in" logo strips, frosted/bordered cards holding a headline + URL, gradient text fills. **Nothing in a video is clickable.** The CTA is spoken, or it is a plain caption line.
155
+
156
+ The test is the **native-editor test**: could you have made this element with the tools inside TikTok's own editor — font, colour, stroke, shadow, tight text box, alignment, rotation, animation presets, stickers, emoji, drawn marks? If you reached past that set, cut it. A comparison is an animated before/after, not a two-column table.
157
+
158
+ ### Rule 8 — the TikTok font regime, not the web's
159
+
160
+ Captions and display type use the composition's bold font regime — **Montserrat (default) or TikTok Sans, weight 700–900, ~36–64px on a 1080-wide frame**, inside the **8%–85%** safe zone, placed in the emptiest part of the frame rather than dumped on the default lower third. **Web/Bootstrap type is the giveaway**: Inter / Roboto / Arial / system-ui at weight 400–600, thin light-grey subtitles, letter-spaced small caps. Matching the client's *brand* font is fine for a wordmark; it is not fine for the caption layer. Exactly one of four caption backgrounds: outline/stroke, plain + shadow, an active-word highlight pill, or a tight solid band (radius ≤8px). Caption colour and plate are **measured off the composited background**, one treatment for the whole video (`short-form.HARNESS.md` → "Caption styling is MEASURED off the background").
161
+
162
+ ### Rule 9 — narration voiceover AND a music bed, both, always
163
+
164
+ This format is a stranger being told what something is, so **a narrated VO is effectively required** — text alone makes the viewer do the work of reading and understanding at the same time. Under it, **a music bed is highly recommended**: silence behind a TTS voice sounds like a screen reader, and the bed is what makes the cut feel produced. Free path: local Kokoro TTS (`--engine local`, default voice `af_heart` — warm, even, unobtrusive) plus a CC0 bed, both $0.
165
+
166
+ - **Deviate from the default voice only for a reason you can state in one clause** — "the read should move because the product is sold on speed" is a reason; "for variety across the batch" is not.
167
+ - **Check the brand name's pronunciation** and respell it phonetically for the TTS if needed (Kokoro said "Jenki" for *Genki* until the input was written "Ghenki"). Confirm with a whisper round-trip — you are running whisper for word timings anyway.
168
+ - **For a calm, unhurried read, render line by line** — separate takes concatenated with measured silences — so the pauses are real instead of TTS filler. One continuous pass reads rushed however slow the copy is.
169
+ - **Never `adelay` the VO** — whisper timings and your captions are relative to raw `vo.wav`. Use `apad` + `atrim`.
170
+ - **Verify the mix by measurement, not by ear**: ~12–15 dB speech-over-bed, peak < 0 dBFS.
171
+
172
+ ### Rule 10 — production floor
173
+
174
+ Claim only what the client's own site claims · one honest scope line · frame 0 is a designed still · the end card settles ≥2s before the last frame · `vidfarm qa ./work` before you report.
175
+
176
+ ## Recon — do it in the main loop, before any agent starts
177
+
178
+ Cheap, deterministic, and it costs no creative capacity. **Recon is also the asset harvest (Rule 5b)** — everything below is collected so the video can be built out of the client's own graphics before a single one is bought or generated:
179
+
180
+ 1. **Headless DPR-2 capture** of the site in 1440×900 scroll slices, tiled into one 3×3 contact sheet — one image read instead of eight.
181
+ 2. **Pull the raw HTML** for verbatim copy, asset URLs and compiled CSS tokens.
182
+ 3. **Sample the palette off the RENDERED hero**, not the CSS (luminance < 110 → build the video dark). **Treat it as a hypothesis, not fact** — agents overturned it about a third of the time and were right every time; it had been sampled off a hero photograph, a seasonal login page, or a store's own chrome. Tell the agent to verify against the real CSS and say so if it overturns.
183
+
184
+ Special cases seen repeatedly: a **blank capture** means the landing state is a form — drive the app headlessly instead. A **tiny HTML shell** means an SPA — harvest the JS bundle for copy and assets, and **check `manifest.json` → `screenshots[]` first**, which repeatedly held ten real screenshots on apps written off as unreachable. **No website at all** (store listing only) — the listing is the asset base, but the store's chrome colour is not the brand. An **animated WebP/GIF on the site** is free motion footage — extract frames to a sprite sheet driven off the paused timeline. **A screen recording in the bundle is your A-roll**, and the best asset find there is.
185
+
186
+ ## Bulk-generation notes — differentiation is an INPUT, not a hope
187
+
188
+ N client URLs → N videos that must not read as N runs of one template. Agents told "make it good" converge; agents given assignments do not. Per brand, **assign**: format (9:16 / 16:9 / 4:5 / 1:1), theme (light/dark *and* the temperature — warm charcoal vs cold slate vs pure black), energy, type system, and voice. Then tell it to **derive its own concept from the brand's own words and assets**.
189
+
190
+ Maintain a `DIFFERENTIATION.md` across the whole engagement with two sections:
191
+
192
+ - **(a) concepts already used** — brand → concept → motion device, appended after every batch. No reuse, no reskin.
193
+ - **(b) forbidden generic shapes** — the abstractions those concepts occupy: *a progress rail advancing through N labelled stages · items dropping into a container · a lattice tessellating as a transition · text morphing into other text · a funnel narrowing many to one · a scan line travelling across a surface · cards flying apart and snapping back · one element pinned while the background swaps · concentric rings resolving into alignment.*
194
+
195
+ **Name each agent's own most predictable answer and forbid it** — no waveform for a music app, no node-graph for something called Nodebase, no word-flip for a language app. And point it at where good concepts come from: **something the brand already said or showed** ("no more FOMO", "two front doors", "split it down the middle") — the one idea that is *theirs*.
196
+
197
+ Split the labour: recon, the differentiation list, review, and watermark/delivery stay in the **main loop**; concept, layout, motion and script go to **one subagent per brand, in parallel** (5 at a time is comfortable, 10–12 works). Each agent owns its own slug's directories and never edits a shared script.
198
+
199
+ ## Pre-flight checklist
200
+
201
+ - [ ] **The open orients a stranger by ~3s**: the category noun, who it is for, and the situation are all available to someone who has never heard of the product
202
+ - [ ] The first spoken sentence is one clause, ≤12 words, everyday words — no acronym, no insider noun, no pronoun without a referent ("it", "this", "they")
203
+ - [ ] The opening image is legible at a glance and at thumbnail scale — one large subject, not a detail crop, a dense surface, or a metaphor whose meaning arrives later
204
+ - [ ] The video does not start at step three: the situation that causes step one is on screen first
205
+ - [ ] Orientation cost one sentence, not one beat — no wind-up sentence, no intro was added to make room for it
206
+ - [ ] The plain-English line was written verbatim into the brief before the build started
207
+ - [ ] It is the first or second spoken line, AND it appears as type in the art (not only in captions)
208
+ - [ ] Frame 0 has three or fewer text runs, one of them the plain line
209
+ - [ ] The open is not a form, a settings panel, a table, or three cards each carrying a sentence
210
+ - [ ] Something is moving from frame 1, and frame 0 carries a complete line that works as a thumbnail
211
+ - [ ] The video opens on the pain or the product, not on a logo, a title card, or a mood piece
212
+ - [ ] A relevant sticker is on screen in the first beat and is the subject of the first spoken line
213
+ - [ ] 5–8 stickers across the video, one per beat, each carrying that line's meaning; the motif recurs with no gap >4s
214
+ - [ ] **The client's own site was harvested FIRST** — screenshots, brand illustrations, icons, mascot, any WebP/GIF or screen recording — and the report names which assets came from it (Rule 5b)
215
+ - [ ] Page elements were lifted with `vidfarm mask` into snug transparent PNGs rather than re-drawn or re-generated
216
+ - [ ] Only what the site does not have went to IconScout, and only what neither has was generated
217
+ - [ ] IconScout was searched (`--free`, no `--style` filter) before anything was generated
218
+ - [ ] Exactly one sticker is legible through every handover (exits clamped, `power2.out`)
219
+ - [ ] Captions are 3–5-word kinetic cues; no static block, no two independent text layers at once
220
+ - [ ] Type is the bold TikTok regime in the safe zone — no Inter/Roboto/Arial at 400–600, no thin grey subtitles
221
+ - [ ] No CTA button, benefit chip row, pricing card, feature grid, comparison table, or logo strip
222
+ - [ ] There is a narration VO and a music bed, and the bed does not fight the read
223
+ - [ ] One honest scope line (price or limitation), and every claim is one the client's own site makes
224
+ - [ ] In a batch: assigned format/theme/energy/type/voice differ, and the concept is not on the used or forbidden list
225
+
226
+ **The stranger test — run it on the frames, never on the script.** The script always looks like it explains the product, because you already know what the product is.
227
+
228
+ - [ ] t=0,1,2,3,4,5 were extracted and read as images; from those six frames alone a stranger can say what the product does
229
+ - [ ] The harsher version passed too: cover the frames, read only the caption text of the first 5s — it still identifies the product
230
+ - [ ] **The 3-second test**: played the first 3s only to someone with no context — they can say what kind of thing this is and roughly who it is for (Rule 0). "Something about audio" is a fail, not a pass with notes
231
+
232
+ **Whole-video review** — on the render, not the plan
233
+ - [ ] A contact sheet of ~12 stills was read as an image: consistent margins, one type scale, one accent colour, one illustration style
234
+ - [ ] No large flat dead region, and no placeholder empty state that reads as a missing asset (sample at 0.2s through a transition — a transient empty wipe frame is fine, resting there >0.5s is not)
235
+ - [ ] No frame carries two contradictory numbers
236
+ - [ ] The CTA/end card is settled ≥2s before the last frame
237
+ - [ ] Frames from two different scenes were compared — a frozen render passes duration, frame-count and audio-hash checks
238
+ - [ ] Audio verified by measurement (~12–15 dB speech-over-bed, peak <0 dBFS), not by "it sounds fine"
239
+
240
+ **Client work** — when the product is someone else's, these are not optional. Use the `product-demo.HARNESS.md` "Client work" block verbatim: exact quotes on ratings/prices, no third party named negatively, humour aimed at the problem, no fear-selling, licence attributions on screen for the full runtime, personal data in screenshots a deliberate choice.
241
+
242
+ > **Bonus deliverable:** log what you find broken on their site while you work — across 32 brands this turned up contradictory stats, plaintext API keys in a JS bundle, a `vercel.app` preview URL in production marketing, and an app icon reading as the wrong country's flag. Clients sometimes value that more than the video.
@@ -26,7 +26,7 @@ Two ways in, depending on where the format came from:
26
26
 
27
27
  ```bash
28
28
  # (a) From a bundled base — when the format is one you're defining
29
- vidfarm harness list # short-form | hooks | ugc-testimonial | explainer | product-demo
29
+ vidfarm harness list # short-form | hooks | ugc-testimonial | explainer | product-demo | product-explainer
30
30
  vidfarm harness init hooks --out ./work/HARNESS.md
31
31
 
32
32
  # (b) From the template you're batching — when the format is one you're REPLICATING
@@ -22,6 +22,8 @@ The mechanical trio — **generate on a chroma plate → key it out → trim to
22
22
 
23
23
  **Illustrations default to simplicity — and to SOLID FILLS.** Whatever path you take to a sticker, aim for **flat vector, simple shapes, minimal detail, few colors, solid opaque fills, no background, no text baked in** — a friendly icon-grade illustration, not a rendered 3D scene or a detailed painting. Simple art keys cleanly, trims tight, scales without mush, animates readably at 9:16, and stays on-style across a whole cast. "Solid fills" is the load-bearing word: **outline-only art has its interior keyed away and comes back as a rim around a transparent hole** (see "Then make the ART key-safe too" below). When generating, say so in the prompt: `--generate "a coffee cup, simple flat vector illustration, minimal detail, 2-3 flat colors, solid filled shapes (not outline-only), no shadows"`.
24
24
 
25
+ **Before either path: check IconScout.** `vidfarm iconscout "<meaning>" --style sticker --free` searches a designer catalog of icons, stickers, illustrations, 3D props and Lottie for $0 — no key needed, and a free asset downloads as a clean transparent SVG/PNG that needs no keying at all. `vidfarm iconscout get <uuid> --format svg` gives you a durable URL. Only fall through to masking or generation when IconScout genuinely has nothing that fits.
26
+
25
27
  **In cost-saving mode, don't generate illustrations at all — mask them out of images the director already has.** If `vidfarm cost-mode` is `minimize` (or the director says "without burning credits"), the default for adding an illustration is `vidfarm mask <their-image> --crop …` — lifting art out of an infographic, poster, deck slide, brand sheet, or screenshot for **$0 and zero AI calls**. Ask for source art before you ask for a generation budget; the guided loop is **"Mask from an image you already have"** below. **If no source art exists and the graphic must be custom, you still don't have to spend** — hand the director a prompt for a **free** image generator (meta.ai / free ChatGPT / a Hugging Face Space) and cut the returned sheet into stickers locally: **"Free manual image-gen"** below.
26
28
 
27
29
  ### "A sticker pack" — what it means, and the one command for it
@@ -2,13 +2,15 @@
2
2
 
3
3
  Use this only when the director signals they do not know where to start.
4
4
 
5
- 1. Read `references/onboarding.md`.
6
- 2. Capture product context and save durable notes into the correct My Files folder.
7
- 3. Determine awareness stages, persuasive angles, and hooks with the brainstorm primitives.
8
- 4. Ask about brand assets, demos, and recurring characters; organize them in My Files.
9
- 5. Ask about budget and map it to the cost spectrum before recommending expensive generation.
10
- 6. Search for the best matching templates and fork one strong default.
11
- 7. Teach the house phrasing early: coach the director to ask for **"a vidfarm template that …"** rather than "a video," and explain the payoff — it's forkable forever and a teammate can **`vidfarm serve <template_id>`** to pull it onto their own machine (see SKILL.md, "Say 'create a vidfarm template that…'"). Getting this into their vocabulary on day one is the point.
12
- 8. Transition into the ordinary template-editing workflow.
5
+ 1. Read `references/onboarding.md`. Agree on one **working folder** for all of this director's Vidfarm work, and run every command from it.
6
+ 2. **Give them content ideas before you ask them anything.** Take the offer in one line (or a URL you read), run `vidfarm ideas --topic "<line>"`, sharpen the frames into 20+ titled videos, and save `content-ideas.md`. Offline, free, keyless. Their reactions to the list are the first real context you get.
7
+ 3. Offer the interview as the next step, not as a gate: `vidfarm consult coldstart --short` (six fixed questions) or the full `coldstart`. Say every question is skippable. Capture product context into `OFFER.md` and the durable answers into `CONTEXT.md`, in the working folder and mirrored to the right My Files folder.
8
+ 4. Determine awareness stages, persuasive angles, and hooks with the brainstorm primitives, then revisit `content-ideas.md` now that the awareness stage is known.
9
+ 5. Ask about brand assets, demos, and recurring characters; organize them in My Files.
10
+ 6. Ask about budget and map it to the cost spectrum before recommending expensive generation.
11
+ - Set the graphics default in the same breath: icons, stickers, illustrations, 3D props and Lottie come from `vidfarm iconscout`, never from an image model. No key, no setup, search is free, free assets are $0.
12
+ 7. Search for the best matching templates and fork one strong default.
13
+ 8. Teach the house phrasing early: coach the director to ask for **"a vidfarm template that …"** rather than "a video," and explain the payoff — it's forkable forever and a teammate can **`vidfarm serve <template_id>`** to pull it onto their own machine (see SKILL.md, "Say 'create a vidfarm template that…'"). Getting this into their vocabulary on day one is the point.
14
+ 9. Transition into the ordinary template-editing workflow.
13
15
 
14
16
  Do not force onboarding on users who already know what they want.