@officexapp/vidfarm-devcli 0.21.32 → 0.21.33

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -105,11 +105,12 @@ Full rationale + the devcli twins (`vidfarm stills --at 0`, `vidfarm qa`) in `vi
105
105
 
106
106
  You author into HTML, which makes it dangerously easy to build a **web page instead of a video**. This is the #1 way an AI-edited composition betrays itself. Apply on every text/graphic you place — and strip it when a decomposed fork or a pasted brand asset brings one in.
107
107
 
108
- **The test:** *would this element exist if the video were shot on a phone and captioned in CapCut?* If its whole job is to look **clickable**, cut it. Nothing in a video is clickable.
108
+ **The test — the native-editor test:** *could you have made this element with the tools inside TikTok's own editor?* Its text tool gives you a font, a color, a stroke/outline, a soft shadow, a tight text box, alignment, opacity, rotation, and animation presets plus stickers, emoji, drawn marks, and clips. It gives you **no padded capsule, no border, no gradient fill, no blur panel, no card**. If you reached past that toolset, you are decorating like a web designer and the frame will read as machine-made. Second half of the test: if the element's whole job is to look **clickable**, cut it. Nothing in a video is clickable.
109
109
 
110
110
  **BANNED (never `add_layer` / `replace_composition_html` these):**
111
111
  - **CTA buttons** — a filled/gradient capsule with action copy ("Sign Up for a Free Trial →", "Get Started", "Book a Call"), glow or drop shadow. A social CTA is *spoken* or a plain caption line.
112
- - **Badge / chip / pill rows** — "✓ ID-Verified · ✓ No Credit Card Needed · ✓ 30-Min Trial". Say them as three *timed caption lines* on the footage instead.
112
+ - **Badges / chips / pillsa row of them, and equally a SINGLE one.** The strip "✓ ID-Verified · ✓ No Credit Card Needed · ✓ 30-Min Trial" is obvious; the common miss is one lonely capsule around a stat or label — `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`, `( +40% )`. Alone ≠ native: a rounded, padded, filled tag on static text is a web badge. **The only legitimate pill is the active-word `spotlight`/`karaoke` highlight**, because it moves with the spoken word. Emphasize a stat the way the editor would — bigger, heavier, ALL-CAPS, an accent color, a drawn circle, or its own beat — and say benefits as *timed caption lines* on the footage, one at a time.
113
+ - **Rule of thumb:** on anything holding words, `border-radius` > ~8px **plus** a background fill **plus** padding = a badge. Drop the fill, or drop the radius until the band hugs the glyphs.
113
114
  - **Cards / panels** — a bordered, shadowed, or `backdrop-filter`-frosted rounded box holding a headline + subheading/URL. Text goes ON the footage, not in a floating panel.
114
115
  - **Gradient text fills, neon border glows, elevation shadows, glassmorphism**, navbars, hero sections, feature grids, `<ul>` bullet lists, tables, "as seen in" strips.
115
116
  - **Web-default type** — Inter/Roboto/system-ui/Arial/Helvetica at weight 400–600 and 16–24px.
@@ -121,7 +122,7 @@ You author into HTML, which makes it dangerously easy to build a **web page inst
121
122
  **Font + background regime (every caption/title, via `set_captions` / `set_layer_style` / `add_layer`):**
122
123
  - **Font:** Montserrat (default) or TikTok Sans / Abel / Source Code Pro / Yesteryear — a family the composition actually imports, or it silently falls back to the slop sans. Weight **700–900**. `font_size` in px of the render canvas: **~36–64px** on a 1080-wide frame; never <28, never 0 (invisible). ~2 lines, ~5 words per line; `line_height` 0.95–1.15.
123
124
  - **Position:** inside the **8%–85%** vertical safe zone (phone UI clips the edges) and clear of the right ~12% action rail — a centered box at `x:10 width:80` is safe. Lower-third ≈ `y:70`; a "POV:" top line ≈ `y:8`, never `y:0`.
124
- - **Background — exactly one of four:** `background_style:"outline"` (stroke, the default look) · `"plain"` (bare + soft shadow) · an **active-word highlight pill** via `set_captions caption_style:"spotlight"|"karaoke"` (the *only* legitimate pill — it tracks the spoken word) · `"highlight-solid"`/`"highlight-translucent"` as a band that **hugs** the text (radius ≤~8px, no border, no shadow, no gradient, no blur, one text run — never a heading+subheading+URL stacked inside it). Anything else is a web card.
125
+ - **Background — exactly one of four:** `background_style:"outline"` (stroke, the default look) · `"plain"` (bare + soft shadow) · an **active-word highlight pill** via `set_captions caption_style:"spotlight"|"karaoke"` (the *only* legitimate pill anywhere in the frame — it tracks the spoken word; a static label never gets one) · `"highlight-solid"`/`"highlight-translucent"` as a band that **hugs** the text (radius ≤~8px, no border, no shadow, no gradient, no blur, one text run — never a heading+subheading+URL stacked inside it). Anything else is a web card.
125
126
 
126
127
  **There is no QA tool for you.** The devcli ships `vidfarm qa <dir>` — a free local blocklist pass over exactly the rules above — but it is **devcli-only with no REST twin**, so in the web editor you enforce this by reading your own output. When you hand a heavy job off to a local coding agent, tell them to run `vidfarm qa ./work` before rendering.
127
128
 
@@ -269,13 +269,13 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
269
269
 
270
270
  - **Determine the surface before claiming capabilities.** The web chat has only its declared tools and REST routes. It cannot execute arbitrary JavaScript/Python, open a shell, create a local repository script, or use the user's filesystem. Never tell a web-chat user that you ran code or wrote a script unless a dedicated declared tool actually did so. A desktop coding agent has a real shell and filesystem and MAY write/run scripts, perform arbitrary local computations over paginated API results, create reports/CSVs/JSON, edit composition files, and orchestrate long devcli workflows within the user's authorization.
271
271
  - **The web AI chat can do all three paintbrushes** — clip raws, author HTML/hyperframe motion, and generate AI media — and it drives edits directly on the live timeline. Keep small-to-medium jobs here: text/caption swaps, a scene or two replaced, single generations, captions, approve/schedule. Just do them.
272
- - **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r`, stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
272
+ - **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. **The test is the native-editor test: could you have made this element with the tools inside TikTok's own editor?** That toolset is font / color / stroke / shadow / tight text box / alignment / rotation / animation presets, plus stickers, emoji, drawn marks and clips — it has no padded capsule, no border, no gradient fill, no blur panel, no card. If you reached past it, cut it. That includes **a single lonely pill around a stat or label** — `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`: being alone doesn't make a badge native, and the only legitimate capsule in a video is the active-word `spotlight`/`karaoke` highlight, which moves with the spoken word. Emphasize a stat the way the editor would instead: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle, or its own beat. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r` — or a `border-radius` over ~8px on anything filled that holds words — stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
273
273
  - **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm regime show hooks`.
274
274
  - **The first frame IS the thumbnail — compose it on purpose.** Frame 0 is a single frame of ~30 in the first second, but it's the poster every feed, share link, and paused player freezes on, so **more people see that one frame than watch the video**. It must never be black, empty, mid-fade, or caught mid-animation: put a real visual at `start:0`, have the hook words already on screen at t=0, and never hang a `fade-black`/`fade-white`/`flash` *entrance* on the **first** clip (junction transitions between later clips are fine — this rule is only about the opening). Verify it, don't assume: devcli `vidfarm stills ./work --at 0` renders that exact frame, and `vidfarm qa` flags a blank or fading open.
275
275
  - **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`QA_REGIME.md`** — because a loop of fifty videos has no human looking at every frame, and the regime is what replaces those eyes. Don't silently ship a one-off when they asked for volume, or drag someone into a harness when they wanted one clip.
276
276
  - **`QA_REGIME.md` is the director's own quality contract, and it's a first-class artifact.** `vidfarm qa`'s built-ins are universal (slop, fonts, the thumbnail frame); a regime is what makes *this* format good — audience, hook shape, banned vocabulary, pacing, compliance line. It lives next to the work, they own it, it stacks: `vidfarm regime init short-form --out ./work/QA_REGIME.md` (bundled bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo` — each a starting point to **edit**, never a house style), then `vidfarm qa ./work --regime hooks --regime ./brand/HOUSE.md`, and any user file anywhere is valid. Its `checks:` front matter is machine-settled; its `- [ ]` checklist comes back as **review items you answer honestly in your report** — never claim a video passed the half the CLI can't judge. When a batch teaches you something, **write it back into the regime**: that's the artifact that compounds. Details in `references/automation-and-local-dev.md` ("Scripting mode"), format in `regimes/README.md`.
277
277
  - **On devcli, QA every video you produce: `vidfarm qa ./work`.** Free, instant, local-only — it blocklists exactly the slop above plus first-frame/thumbnail and font-regime/safe-zone drift, and prints a concrete fix per finding. **Feedback, not a gate**: it exits 0 even on findings, never runs automatically, and is a blocklist (unusual/stylized compositions pass untouched), so there's no reason not to run it before every publish. `--json` for scripted batches, `--strict` only if you want a CI failure. **Web-chat copilot: this command does not exist for you** (devcli-only, no REST twin) — apply the standard by hand, and when handing a heavy job to a local coding agent, tell them to run `vidfarm qa`.
278
- - **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
278
+ - **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
279
279
  - **Where the web chat struggles: complex, long, multi-step transformations.** A full multi-scene re-theme, an iterative render-critique-iterate loop, heavy scripted or batch work, or anything needing a real filesystem and many sequential tool calls will hit context limits, turn/timeout ceilings, and the web editor's constraints (CSS/declarative motion only — JS animation adapters are stripped on save). Don't grind a big transformation one layer at a time in a chat turn and stall.
280
280
  - **Practical workaround — hand the heavy job to local devcli.** When a task is genuinely large or long-running, **proactively recommend the director run it locally with an AI coding agent** (Claude Code / OpenAI Codex / any capable agent): `vidfarm pull <forkId>` writes the composition + the `.harness/` grounding bundle to disk, the agent edits with the full devcli verb set and JS animation adapters, renders free with `vidfarm serve`, and `vidfarm publish` pushes it back. This is the **best-quality (B) harness's** natural home (adversarial grading with a coding agent). Frame it as "this is a big rebuild — you'll get a better, faster result running it locally with a coding agent; here's how," not as a dead end.
281
281
  - **Offer a handoff, do not impersonate the desktop agent.** When web chat reaches that boundary, offer to save a Markdown handoff in My Files containing the objective, selected template/fork IDs, asset paths, grounding, constraints, completed work, and suggested devcli commands. Create it only after the user agrees. The desktop agent should read that document, pull the referenced fork, and then use its actual code/shell capabilities.
@@ -7,7 +7,7 @@ Use this when a coding agent is doing the work locally or the user wants a repro
7
7
  3. Read `./work/.harness/agent-guide.md` and `./work/.harness/context.json` before editing.
8
8
  4. Make deterministic edits to `composition.html` and optionally `composition.json`.
9
9
  5. Validate with `vidfarm lint` or `vidfarm stills` when useful. **Always look at `vidfarm stills ./work --at 0`** — that frame becomes the thumbnail, so it must not be black, empty, or mid-fade.
10
- 6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, and flags a blank/fading first frame (the thumbnail). Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
10
+ 6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, a lone pill around a static stat/label, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, and flags a blank/fading first frame (the thumbnail). Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
11
11
  7. Render with `vidfarm render <forkId> --dir ./work --wait`.
12
12
  8. **Ask about deduplication before you approve** — "is this going out more than once (several accounts, another platform, a re-post later)?" If yes, run `vidfarm dedupe ./final.mp4 [--variants N]` on the **exported** MP4 (free, local ffmpeg, no re-render) and approve each variant separately. Asking here rather than after publication is what avoids paying for a second render. See `references/core-workflows.md` → *Deduplicate before you publish*.
13
13
  9. Approve the finished MP4 with `vidfarm approve --video <url|./final.mp4> --caption "..."`. This prints the shareable `share_url`.
@@ -320,6 +320,7 @@ What it flags:
320
320
  |---|---|---|
321
321
  | `cta-button` | error | Action copy ("Sign Up for a Free Trial →") **inside** a filled/gradient rounded capsule. Bare CTA copy in a caption is fine — "BUY NOW" is real social copy |
322
322
  | `badge-chip-row` | error | 2+ sibling small rounded filled tags — the "✓ No Credit Card Needed · ✓ 30-Min Trial" strip |
323
+ | `static-pill` | error | ONE filled, padded, ≥20px-radius capsule around static text — a stat/label badge like "10 hrs / week", "STEP 2", "EP.01". Skips active-word `spotlight`/`karaoke` highlights (the only legitimate pill) and mock social UI (chat bubbles, comment cards — mark yours `data-vf-mock-ui` if the heuristic misses it) |
323
324
  | `card-panel` / `glass-card` | error | A rounded box with a border/shadow/`backdrop-filter` wrapping 2+ elements. A tight legibility **band** (radius ≤8px, one text run, no border/shadow) stays legal |
324
325
  | `gradient-text` | error | `background-clip:text` gradient headline fills |
325
326
  | `clickable-element` | error/warn | `<button>`, `<form>`, `<input>`; `<a href>` warns |
@@ -496,12 +496,12 @@ Web copilot: same standard, applied by hand — check the opening layer's `start
496
496
 
497
497
  Compositions are authored in HTML, so the single most common way an AI-edited video betrays itself is by **looking like a web page**. A landing-page button, a row of feature chips, a frosted card with a gradient headline — these are everywhere on the web and **nowhere in a real social post**. A director scrolling TikTok/Reels/Shorts never sees them, so the moment one appears the video reads as machine-made and the retention dies. HTML and hyperframes are the right tool; **Bootstrap/Tailwind furniture is not.**
498
498
 
499
- **The one test:** *would this element exist if the video were shot on a phone and captioned in CapCut?* If it only makes sense inside a browser if its whole job is to look **clickable** cut it. **Nothing in a video is clickable.**
499
+ **The one test — the native-editor test:** *could you have made this element with the tools inside TikTok's (or CapCut's) own editor?* That editor's text tool gives you exactly this: a font, a color, a stroke/outline, a soft shadow, a tight text-box band, alignment, opacity, rotation, and animation presets plus stickers, emoji, drawn marks, and clips. It does **not** give you a padded capsule, a border, a gradient fill, a blur panel, or a card. If you had to reach past that toolset, you are decorating like a web designer, and the frame will read as machine-made no matter how good the copy is. The second half of the same test: if the element's whole job is to look **clickable**, cut it **nothing in a video is clickable.**
500
500
 
501
501
  **BANNED — never author, and strip on sight when a fork or a paste brings one in:**
502
502
 
503
503
  - **CTA / "button" shapes.** Any filled capsule or rounded rect containing action copy — "Sign Up for a Free Trial →", "Learn More", "Get Started", "Book a Call" — especially gradient-filled with a glow or drop shadow. A social CTA is *spoken*, or a plain caption line, or an arrow pointing at the real UI. Not a button.
504
- - **Badge / chip / pill rows.** The "✓ ID-Verified Tutors · ✓ No Credit Card Needed · ✓ 30-Min Trial" strip of small rounded tags. Nothing on TikTok is a `<span class="badge">`. Say the three benefits as three timed caption lines instead one at a time, on the footage.
504
+ - **Badges, chips, pills including a SINGLE one.** The "✓ ID-Verified Tutors · ✓ No Credit Card Needed · ✓ 30-Min Trial" strip is the obvious case, but the far more common one is **one lonely capsule holding a stat or a label**: `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`, `( BEGINNER )`, `( +40% )`. Being alone does not make it native — a rounded, padded, filled tag around static text is a `<span class="badge">` wearing a different hat, and it is one of the loudest web tells in the whole frame. **The only legitimate pill in a video is the active-word highlight** (`spotlight`/`karaoke`), because it tracks the spoken word and moves. Static text gets `outline`, `plain`, or a tight band that hugs the glyphs (radius ≤ ~8px). If a stat deserves emphasis, give it emphasis the *editor* can give: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle or underline around it, its own beat on screen.
505
505
  - **Cards, panels, containers.** A rounded-rect box (dark, bordered, shadowed, or `backdrop-filter` frosted) holding a heading plus a subheading/URL. Text belongs **on the footage**, not inside a floating panel with padding around it.
506
506
  - **Gradient text fills, neon border glows, elevation shadows, glassmorphism, hover-implying strokes.**
507
507
  - **Web page furniture of any kind:** navbars, hero sections, feature grids, two-column layouts, `<ul>` bullet lists with disc markers, tables, alerts, progress bars as decoration, logo-in-a-circle avatars, "as seen in" strips.
@@ -509,6 +509,8 @@ Compositions are authored in HTML, so the single most common way an AI-edited vi
509
509
 
510
510
  **Greppable smell test.** If you are typing `class="btn…"`, `badge`, `chip`, `tag`, `card`, `panel`, `container`, `row`/`col-`, `hero`, `cta`, `rounded-full`, `shadow-lg`, `backdrop-blur`, `bg-gradient-to-r`, `border border-…` — **stop.** You are building a web page, not a video. Rewrite as timed text on footage.
511
511
 
512
+ **The capsule rule of thumb.** On any element that holds words: `border-radius` over ~8px **combined with** a background fill and padding = a badge. Either take the fill away (bare text + outline/shadow) or take the radius and padding down until the band hugs the glyphs. There is no third option for static text.
513
+
512
514
  **ALLOWED and encouraged — these ARE social-native:**
513
515
 
514
516
  - **Arrows** (drawn, animated, hand-style), circles/scribbles/underline strokes highlighting part of the frame, hand-drawn marks, checkmarks *as glyphs inside a caption line* (not as chips).
@@ -535,7 +537,7 @@ Short-form is watched on a phone, and the phone's UI eats the frame's edges. **N
535
537
  | 3 | **Highlight pill behind the ACTIVE word only** | `set_captions caption_style:"spotlight"` / `"karaoke"` (+ `caption_highlight_color`) | Hormozi/CapCut word-by-word. **The only legitimate "pill" in a video** — it tracks the spoken word, so it isn't a badge |
536
538
  | 4 | **Solid band that tightly hugs the text lines** (CapCut "text box") | `background_style:"highlight-solid"` (or `"highlight-translucent"`) + a `background` color | Guaranteed legibility over noisy footage |
537
539
 
538
- Treatment 4 is a **band, not a card**: it hugs the glyphs with minimal padding, corner radius ≤ ~8px, **no border, no drop shadow, no gradient, no blur**, and it wraps *one* text run — never a heading + subheading + URL stacked inside one rounded box. The moment it grows padding, a stroke, or a second element inside it, it has become a web card. Fix it.
540
+ Treatment 4 is a **band, not a card**: it hugs the glyphs with minimal padding, corner radius ≤ ~8px, **no border, no drop shadow, no gradient, no blur**, and it wraps *one* text run — never a heading + subheading + URL stacked inside one rounded box. The moment it grows padding, a stroke, or a second element inside it, it has become a web card. Fix it. And the moment its radius goes fully round, it has become a **badge** — treatment 3 is the *only* capsule allowed, and only because it tracks the spoken word. A static "10 hrs / week" in a rounded pill is web furniture; the same words in treatment 1 or 2, bigger and heavier, are a beat.
539
541
 
540
542
  **A common trap: decomposed templates mirror the source's caption placement**, so a forked meme can arrive with its caption pinned at `top:0` in a non-regime font — and a re-theme prompt ("make this for my tutoring service") is exactly where an agent starts inventing landing-page CTAs and benefit chips because the *subject* is a SaaS product. **Fix to the standard, don't inherit it, and don't import the website's design language into the video.** When placing text yourself (`set_captions`, `set_layer_text`, `set_layer_style`, `add_layer`, devcli `place`/`captions`), set `y` / `font_family` / `font_weight` / `background_style` to the standard from the start.
541
543
 
package/SKILL.director.md CHANGED
@@ -269,13 +269,13 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
269
269
 
270
270
  - **Determine the surface before claiming capabilities.** The web chat has only its declared tools and REST routes. It cannot execute arbitrary JavaScript/Python, open a shell, create a local repository script, or use the user's filesystem. Never tell a web-chat user that you ran code or wrote a script unless a dedicated declared tool actually did so. A desktop coding agent has a real shell and filesystem and MAY write/run scripts, perform arbitrary local computations over paginated API results, create reports/CSVs/JSON, edit composition files, and orchestrate long devcli workflows within the user's authorization.
271
271
  - **The web AI chat can do all three paintbrushes** — clip raws, author HTML/hyperframe motion, and generate AI media — and it drives edits directly on the live timeline. Keep small-to-medium jobs here: text/caption swaps, a scene or two replaced, single generations, captions, approve/schedule. Just do them.
272
- - **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r`, stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
272
+ - **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. **The test is the native-editor test: could you have made this element with the tools inside TikTok's own editor?** That toolset is font / color / stroke / shadow / tight text box / alignment / rotation / animation presets, plus stickers, emoji, drawn marks and clips — it has no padded capsule, no border, no gradient fill, no blur panel, no card. If you reached past it, cut it. That includes **a single lonely pill around a stat or label** — `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`: being alone doesn't make a badge native, and the only legitimate capsule in a video is the active-word `spotlight`/`karaoke` highlight, which moves with the spoken word. Emphasize a stat the way the editor would instead: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle, or its own beat. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r` — or a `border-radius` over ~8px on anything filled that holds words — stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
273
273
  - **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm regime show hooks`.
274
274
  - **The first frame IS the thumbnail — compose it on purpose.** Frame 0 is a single frame of ~30 in the first second, but it's the poster every feed, share link, and paused player freezes on, so **more people see that one frame than watch the video**. It must never be black, empty, mid-fade, or caught mid-animation: put a real visual at `start:0`, have the hook words already on screen at t=0, and never hang a `fade-black`/`fade-white`/`flash` *entrance* on the **first** clip (junction transitions between later clips are fine — this rule is only about the opening). Verify it, don't assume: devcli `vidfarm stills ./work --at 0` renders that exact frame, and `vidfarm qa` flags a blank or fading open.
275
275
  - **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`QA_REGIME.md`** — because a loop of fifty videos has no human looking at every frame, and the regime is what replaces those eyes. Don't silently ship a one-off when they asked for volume, or drag someone into a harness when they wanted one clip.
276
276
  - **`QA_REGIME.md` is the director's own quality contract, and it's a first-class artifact.** `vidfarm qa`'s built-ins are universal (slop, fonts, the thumbnail frame); a regime is what makes *this* format good — audience, hook shape, banned vocabulary, pacing, compliance line. It lives next to the work, they own it, it stacks: `vidfarm regime init short-form --out ./work/QA_REGIME.md` (bundled bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo` — each a starting point to **edit**, never a house style), then `vidfarm qa ./work --regime hooks --regime ./brand/HOUSE.md`, and any user file anywhere is valid. Its `checks:` front matter is machine-settled; its `- [ ]` checklist comes back as **review items you answer honestly in your report** — never claim a video passed the half the CLI can't judge. When a batch teaches you something, **write it back into the regime**: that's the artifact that compounds. Details in `references/automation-and-local-dev.md` ("Scripting mode"), format in `regimes/README.md`.
277
277
  - **On devcli, QA every video you produce: `vidfarm qa ./work`.** Free, instant, local-only — it blocklists exactly the slop above plus first-frame/thumbnail and font-regime/safe-zone drift, and prints a concrete fix per finding. **Feedback, not a gate**: it exits 0 even on findings, never runs automatically, and is a blocklist (unusual/stylized compositions pass untouched), so there's no reason not to run it before every publish. `--json` for scripted batches, `--strict` only if you want a CI failure. **Web-chat copilot: this command does not exist for you** (devcli-only, no REST twin) — apply the standard by hand, and when handing a heavy job to a local coding agent, tell them to run `vidfarm qa`.
278
- - **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
278
+ - **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
279
279
  - **Where the web chat struggles: complex, long, multi-step transformations.** A full multi-scene re-theme, an iterative render-critique-iterate loop, heavy scripted or batch work, or anything needing a real filesystem and many sequential tool calls will hit context limits, turn/timeout ceilings, and the web editor's constraints (CSS/declarative motion only — JS animation adapters are stripped on save). Don't grind a big transformation one layer at a time in a chat turn and stall.
280
280
  - **Practical workaround — hand the heavy job to local devcli.** When a task is genuinely large or long-running, **proactively recommend the director run it locally with an AI coding agent** (Claude Code / OpenAI Codex / any capable agent): `vidfarm pull <forkId>` writes the composition + the `.harness/` grounding bundle to disk, the agent edits with the full devcli verb set and JS animation adapters, renders free with `vidfarm serve`, and `vidfarm publish` pushes it back. This is the **best-quality (B) harness's** natural home (adversarial grading with a coding agent). Frame it as "this is a big rebuild — you'll get a better, faster result running it locally with a coding agent; here's how," not as a dead end.
281
281
  - **Offer a handoff, do not impersonate the desktop agent.** When web chat reaches that boundary, offer to save a Markdown handoff in My Files containing the objective, selected template/fork IDs, asset paths, grounding, constraints, completed work, and suggested devcli commands. Create it only after the user agrees. The desktop agent should read that document, pull the referenced fork, and then use its actual code/shell capabilities.
@@ -1208,12 +1208,12 @@ Web copilot: same standard, applied by hand — check the opening layer's `start
1208
1208
 
1209
1209
  Compositions are authored in HTML, so the single most common way an AI-edited video betrays itself is by **looking like a web page**. A landing-page button, a row of feature chips, a frosted card with a gradient headline — these are everywhere on the web and **nowhere in a real social post**. A director scrolling TikTok/Reels/Shorts never sees them, so the moment one appears the video reads as machine-made and the retention dies. HTML and hyperframes are the right tool; **Bootstrap/Tailwind furniture is not.**
1210
1210
 
1211
- **The one test:** *would this element exist if the video were shot on a phone and captioned in CapCut?* If it only makes sense inside a browser if its whole job is to look **clickable** cut it. **Nothing in a video is clickable.**
1211
+ **The one test — the native-editor test:** *could you have made this element with the tools inside TikTok's (or CapCut's) own editor?* That editor's text tool gives you exactly this: a font, a color, a stroke/outline, a soft shadow, a tight text-box band, alignment, opacity, rotation, and animation presets plus stickers, emoji, drawn marks, and clips. It does **not** give you a padded capsule, a border, a gradient fill, a blur panel, or a card. If you had to reach past that toolset, you are decorating like a web designer, and the frame will read as machine-made no matter how good the copy is. The second half of the same test: if the element's whole job is to look **clickable**, cut it **nothing in a video is clickable.**
1212
1212
 
1213
1213
  **BANNED — never author, and strip on sight when a fork or a paste brings one in:**
1214
1214
 
1215
1215
  - **CTA / "button" shapes.** Any filled capsule or rounded rect containing action copy — "Sign Up for a Free Trial →", "Learn More", "Get Started", "Book a Call" — especially gradient-filled with a glow or drop shadow. A social CTA is *spoken*, or a plain caption line, or an arrow pointing at the real UI. Not a button.
1216
- - **Badge / chip / pill rows.** The "✓ ID-Verified Tutors · ✓ No Credit Card Needed · ✓ 30-Min Trial" strip of small rounded tags. Nothing on TikTok is a `<span class="badge">`. Say the three benefits as three timed caption lines instead one at a time, on the footage.
1216
+ - **Badges, chips, pills including a SINGLE one.** The "✓ ID-Verified Tutors · ✓ No Credit Card Needed · ✓ 30-Min Trial" strip is the obvious case, but the far more common one is **one lonely capsule holding a stat or a label**: `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`, `( BEGINNER )`, `( +40% )`. Being alone does not make it native — a rounded, padded, filled tag around static text is a `<span class="badge">` wearing a different hat, and it is one of the loudest web tells in the whole frame. **The only legitimate pill in a video is the active-word highlight** (`spotlight`/`karaoke`), because it tracks the spoken word and moves. Static text gets `outline`, `plain`, or a tight band that hugs the glyphs (radius ≤ ~8px). If a stat deserves emphasis, give it emphasis the *editor* can give: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle or underline around it, its own beat on screen.
1217
1217
  - **Cards, panels, containers.** A rounded-rect box (dark, bordered, shadowed, or `backdrop-filter` frosted) holding a heading plus a subheading/URL. Text belongs **on the footage**, not inside a floating panel with padding around it.
1218
1218
  - **Gradient text fills, neon border glows, elevation shadows, glassmorphism, hover-implying strokes.**
1219
1219
  - **Web page furniture of any kind:** navbars, hero sections, feature grids, two-column layouts, `<ul>` bullet lists with disc markers, tables, alerts, progress bars as decoration, logo-in-a-circle avatars, "as seen in" strips.
@@ -1221,6 +1221,8 @@ Compositions are authored in HTML, so the single most common way an AI-edited vi
1221
1221
 
1222
1222
  **Greppable smell test.** If you are typing `class="btn…"`, `badge`, `chip`, `tag`, `card`, `panel`, `container`, `row`/`col-`, `hero`, `cta`, `rounded-full`, `shadow-lg`, `backdrop-blur`, `bg-gradient-to-r`, `border border-…` — **stop.** You are building a web page, not a video. Rewrite as timed text on footage.
1223
1223
 
1224
+ **The capsule rule of thumb.** On any element that holds words: `border-radius` over ~8px **combined with** a background fill and padding = a badge. Either take the fill away (bare text + outline/shadow) or take the radius and padding down until the band hugs the glyphs. There is no third option for static text.
1225
+
1224
1226
  **ALLOWED and encouraged — these ARE social-native:**
1225
1227
 
1226
1228
  - **Arrows** (drawn, animated, hand-style), circles/scribbles/underline strokes highlighting part of the frame, hand-drawn marks, checkmarks *as glyphs inside a caption line* (not as chips).
@@ -1247,7 +1249,7 @@ Short-form is watched on a phone, and the phone's UI eats the frame's edges. **N
1247
1249
  | 3 | **Highlight pill behind the ACTIVE word only** | `set_captions caption_style:"spotlight"` / `"karaoke"` (+ `caption_highlight_color`) | Hormozi/CapCut word-by-word. **The only legitimate "pill" in a video** — it tracks the spoken word, so it isn't a badge |
1248
1250
  | 4 | **Solid band that tightly hugs the text lines** (CapCut "text box") | `background_style:"highlight-solid"` (or `"highlight-translucent"`) + a `background` color | Guaranteed legibility over noisy footage |
1249
1251
 
1250
- Treatment 4 is a **band, not a card**: it hugs the glyphs with minimal padding, corner radius ≤ ~8px, **no border, no drop shadow, no gradient, no blur**, and it wraps *one* text run — never a heading + subheading + URL stacked inside one rounded box. The moment it grows padding, a stroke, or a second element inside it, it has become a web card. Fix it.
1252
+ Treatment 4 is a **band, not a card**: it hugs the glyphs with minimal padding, corner radius ≤ ~8px, **no border, no drop shadow, no gradient, no blur**, and it wraps *one* text run — never a heading + subheading + URL stacked inside one rounded box. The moment it grows padding, a stroke, or a second element inside it, it has become a web card. Fix it. And the moment its radius goes fully round, it has become a **badge** — treatment 3 is the *only* capsule allowed, and only because it tracks the spoken word. A static "10 hrs / week" in a rounded pill is web furniture; the same words in treatment 1 or 2, bigger and heavier, are a beat.
1251
1253
 
1252
1254
  **A common trap: decomposed templates mirror the source's caption placement**, so a forked meme can arrive with its caption pinned at `top:0` in a non-regime font — and a re-theme prompt ("make this for my tutoring service") is exactly where an agent starts inventing landing-page CTAs and benefit chips because the *subject* is a SaaS product. **Fix to the standard, don't inherit it, and don't import the website's design language into the video.** When placing text yourself (`set_captions`, `set_layer_text`, `set_layer_style`, `add_layer`, devcli `place`/`captions`), set `y` / `font_family` / `font_weight` / `background_style` to the standard from the start.
1253
1255
 
@@ -2091,6 +2093,7 @@ What it flags:
2091
2093
  |---|---|---|
2092
2094
  | `cta-button` | error | Action copy ("Sign Up for a Free Trial →") **inside** a filled/gradient rounded capsule. Bare CTA copy in a caption is fine — "BUY NOW" is real social copy |
2093
2095
  | `badge-chip-row` | error | 2+ sibling small rounded filled tags — the "✓ No Credit Card Needed · ✓ 30-Min Trial" strip |
2096
+ | `static-pill` | error | ONE filled, padded, ≥20px-radius capsule around static text — a stat/label badge like "10 hrs / week", "STEP 2", "EP.01". Skips active-word `spotlight`/`karaoke` highlights (the only legitimate pill) and mock social UI (chat bubbles, comment cards — mark yours `data-vf-mock-ui` if the heuristic misses it) |
2094
2097
  | `card-panel` / `glass-card` | error | A rounded box with a border/shadow/`backdrop-filter` wrapping 2+ elements. A tight legibility **band** (radius ≤8px, one text run, no border/shadow) stays legal |
2095
2098
  | `gradient-text` | error | `background-clip:text` gradient headline fills |
2096
2099
  | `clickable-element` | error/warn | `<button>`, `<form>`, `<input>`; `<a href>` warns |
@@ -2798,7 +2801,7 @@ Use this when a coding agent is doing the work locally or the user wants a repro
2798
2801
  3. Read `./work/.harness/agent-guide.md` and `./work/.harness/context.json` before editing.
2799
2802
  4. Make deterministic edits to `composition.html` and optionally `composition.json`.
2800
2803
  5. Validate with `vidfarm lint` or `vidfarm stills` when useful. **Always look at `vidfarm stills ./work --at 0`** — that frame becomes the thumbnail, so it must not be black, empty, or mid-fade.
2801
- 6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, and flags a blank/fading first frame (the thumbnail). Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
2804
+ 6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, a lone pill around a static stat/label, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, and flags a blank/fading first frame (the thumbnail). Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
2802
2805
  7. Render with `vidfarm render <forkId> --dir ./work --wait`.
2803
2806
  8. **Ask about deduplication before you approve** — "is this going out more than once (several accounts, another platform, a re-post later)?" If yes, run `vidfarm dedupe ./final.mp4 [--variants N]` on the **exported** MP4 (free, local ffmpeg, no re-render) and approve each variant separately. Asking here rather than after publication is what avoids paying for a second render. See `references/core-workflows.md` → *Deduplicate before you publish*.
2804
2807
  9. Approve the finished MP4 with `vidfarm approve --video <url|./final.mp4> --caption "..."`. This prints the shareable `share_url`.
package/SKILL.md CHANGED
@@ -114,7 +114,7 @@ For composition *authoring* craft (motion, keyframes, scene design), Vidfarm shi
114
114
  - Never build composition HTML by string concatenation — parse, edit, re-serialize the DOM.
115
115
  - Render only through `POST /api/v1/compositions/:forkId/render`; never call the renderer directly.
116
116
  - Submissions are **not idempotent** — every render/primitive POST charges again. Check status before retrying.
117
- - **No HTML slop.** Compositions are HTML, but a video is not a web page: never author CTA "buttons", benefit chip/badge rows, frosted or bordered cards holding a headline + URL, gradient text, feature grids, or bullet lists none of that exists in a real TikTok, and nothing in a video is clickable. Say it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, color pops, emoji, cut-out stickers, and mock social UI are fine.
117
+ - **No HTML slop.** Compositions are HTML, but a video is not a web page. The test: *could you have made this element with the tools inside TikTok's own editor* (font, color, stroke, shadow, tight text box, rotation, animation presets, stickers, emoji, drawn marks)? If you reached past that — a padded capsule, a border, a gradient fill, a blur panel, a card — cut it. So: no CTA "buttons", no benefit chip/badge rows, **no single pill around a static stat or label** (`( 10 hrs / week )`, `( STEP 2 )` — the only legitimate capsule is the active-word `spotlight`/`karaoke` highlight), no frosted/bordered cards holding a headline + URL, no gradient text, feature grids, or bullet lists. Nothing in a video is clickable. Say it as timed text on the footage; emphasize with size, weight, ALL-CAPS, an accent color, or a drawn circle. Arrows, scribble/underline marks, italics, color pops, emoji, cut-out stickers, and mock social UI are fine.
118
118
  - **Structure before polish — the four charges, written before the timeline.** 🪝 **Hook**: the first line is a complete clause naming a *situation*, not a label; on screen at `start:0` (chunk 1 is read before any audio). 🔄 **Loop**: one open question by 0:10, stated on screen, **closing inside this video**, with an answer the viewer can't guess. 😍 **Payoff**: shown, not summarized, before the final beat. 🎣 **Bait**: one ask, final beat + post caption. Banned openers: throat-clearing, a logo, a title card, a fade from black. Full harness: `references/hooks-and-virality.md`; checkable form: `vidfarm regime show hooks`.
119
119
  - **The first frame is the thumbnail.** Frame 0 is one frame of ~30 in the first second, but it's the poster every feed and share link freezes on — so it's seen by more people than the video is. Never open on black, an empty frame, or a fade-up: a real visual at `start:0`, the hook words already up, and no *entrance* transition on the first clip (junction transitions between later clips are fine). Check it with `vidfarm stills <dir> --at 0`.
120
120
  - **On devcli, `vidfarm qa <dir>` before every render.** Free, instant, local-only blocklist for the slop above + the first frame + the font regime. Feedback, not a gate (exits 0, never automatic). No REST/web equivalent.
package/dist/src/cli.js CHANGED
@@ -2387,7 +2387,7 @@ Rules:
2387
2387
  - When swapping visuals, match both the literal scene DNA and the narrative purpose of the beat.
2388
2388
  - For replacement graphics, screenshots, or still-like scenes, prefer AI image generation plus Ken Burns before paying for AI video unless static_vs_pivot says motion footage is load-bearing.
2389
2389
  - If narration must be customized, default to premium ElevenLabs first, then the user's own ElevenLabs path, then BYOK OpenAI/Gemini/OpenRouter. If captions or scenes were timed to the old VO, retime them to the new narration.
2390
- - NO HTML SLOP. You are editing HTML, but the output is a social video, not a web page. Never author landing-page furniture: CTA "buttons" (a filled/gradient rounded capsule with action copy like "Sign Up for a Free Trial →"), benefit chip/badge rows ("✓ No Credit Card Needed"), bordered/shadowed/frosted cards holding a headline + URL, gradient text fills, feature grids, bulleted lists, or web-default fonts (Inter/Roboto/Arial/system-ui). None of that appears in a real TikTok, and nothing in a video is clickable — say it as timed text on the footage instead. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all fine. Captions use an imported family (Montserrat default / TikTok Sans / Abel / Source Code Pro / Yesteryear) at weight 700-900, ~36-64px on a 1080-wide frame, inside the 8%-85% safe zone, with exactly one of four backgrounds: outline, plain, an active-word spotlight/karaoke pill, or a tight-hugging solid band (radius <=8px, no border/shadow/gradient/blur).
2390
+ - NO HTML SLOP. You are editing HTML, but the output is a social video, not a web page. THE TEST IS THE NATIVE-EDITOR TEST: could you have made this element with the tools inside TikTok's own editor? That toolset is a font, a color, a stroke/outline, a soft shadow, a tight text box, alignment, opacity, rotation, animation presets — plus stickers, emoji, drawn marks and clips. It has NO padded capsule, NO border, NO gradient fill, NO blur panel, NO card. If you reached past it, cut it. Never author landing-page furniture: CTA "buttons" (a filled/gradient rounded capsule with action copy like "Sign Up for a Free Trial →"), benefit chip/badge rows ("✓ No Credit Card Needed"), bordered/shadowed/frosted cards holding a headline + URL, gradient text fills, feature grids, bulleted lists, or web-default fonts (Inter/Roboto/Arial/system-ui). AND NOT A SINGLE PILL EITHER: one lonely rounded, padded, filled capsule around a static stat or label — "10 hrs / week", "STEP 2", "EP.01", "+40%" — is a web badge, and being the only one on screen does not make it native. The ONLY legitimate capsule in a video is the active-word spotlight/karaoke caption highlight, because it moves with the spoken word. Emphasize a stat the way the editor would: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle or underline, or its own beat on screen. Rule of thumb on anything holding words: border-radius over ~8px PLUS a background fill PLUS padding = a badge; drop the fill or drop the radius until the band hugs the glyphs. None of this appears in a real TikTok, and nothing in a video is clickable — say it as timed text on the footage instead. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all fine. Captions use an imported family (Montserrat default / TikTok Sans / Abel / Source Code Pro / Yesteryear) at weight 700-900, ~36-64px on a 1080-wide frame, inside the 8%-85% safe zone, with exactly one of four backgrounds: outline, plain, an active-word spotlight/karaoke pill, or a tight-hugging solid band (radius <=8px, no border/shadow/gradient/blur).
2391
2391
  - STRUCTURE BEFORE POLISH — THE FOUR CHARGES, WRITTEN BEFORE YOU TOUCH THE TIMELINE. Most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) HOOK — write the opening line as text first: a complete clause (subject + verb), no jargon, naming a SITUATION ("I've quit six businesses") not a label ("anonymity"); it goes on screen at start:0, because caption chunk 1 is read before any audio and muted autoplay is the default. Banned openings: throat-clearing ("so I was thinking", "here's the thing"), a logo, a title card, a fade from black, context before the claim. (2) LOOP — one open question by 0:10, said ON SCREEN, closing INSIDE this video (state the timestamp it closes at; if you can't, there is no loop), and the withheld answer must be one the viewer CANNOT supply themselves — a formally-correct loop with a guessable answer passes every mechanical check and dies in the field. (3) PAYOFF — shown, not summarized, ≥5 uninterrupted seconds, landing BEFORE the final beat; the payoff is not the CTA. (4) BAIT — one ask in the final beat and in the post caption; never a DM funnel, "follow for part two", or ragebait. Then build the timeline. Re-theming a decomposed template: viral_dna already names the source's hook/retention/payoff — rebuild each charge for the new subject, never flatten the loop into a product statement. Full craft harness: the vidfarm skill's references/hooks-and-virality.md. Checkable form: \`vidfarm regime show hooks\`.
2392
2392
  - THE FIRST FRAME IS THE THUMBNAIL. Frame 0 is one frame of ~30 in the first second, but every feed card, share link, and paused player freezes on it — more people see that frame than watch the video. It must never be black, empty, mid-fade, or mid-animation: a real visual at start:0 (\`vidfarm retime . --layer <key> --start 0\`), the hook words already on screen at t=0, and NO entrance transition on the FIRST clip (\`vidfarm transitions set . --layer <key> --in none\`; junction transitions between later clips are fine). Look at the actual pixels before you render: \`vidfarm stills . --at 0\`.
2393
2393
  - ONE-TIME OR BULK? Ask before you build. If the director wants volume (daily posting, N variants, hook tests), that's SCRIPTING MODE: pin this fork as the base, vary exactly ONE thing per variant, and install a QA_REGIME.md — \`vidfarm regime init short-form --out ./QA_REGIME.md\` (bases: short-form, hooks, ugc-testimonial, explainer, product-demo), then EDIT it with them. It is their own written quality standard, and it exists because nobody watches variant #37 as carefully as #1. \`vidfarm qa .\` picks up ./QA_REGIME.md automatically; \`--regime <name|path>\` adds more (they stack, and any file of theirs anywhere is valid). Its \`checks:\` are machine-settled; its \`- [ ]\` items come back for YOU to answer honestly in your report — never claim a pass on the half the CLI can't judge. When a batch teaches you something, write it back into the regime.
@@ -103,6 +103,38 @@ function radiusPx(style) {
103
103
  const num = Number((raw.match(/(-?[\d.]+)px/) ?? [])[1] ?? 0);
104
104
  return Number.isFinite(num) ? num : 0;
105
105
  }
106
+ /** Walk self + ancestors up to the composition root. */
107
+ function selfAndAncestors(node) {
108
+ const chain = [];
109
+ let cursor = node;
110
+ while (cursor && typeof cursor.getAttribute === "function") {
111
+ chain.push(cursor);
112
+ if (cursor.getAttribute("data-composition-id") != null)
113
+ break;
114
+ cursor = cursor.parentNode;
115
+ }
116
+ return chain;
117
+ }
118
+ /**
119
+ * Is this node part of an animated caption run? The active-word highlight IS a
120
+ * pill, and it is the one legitimate pill in a video — it tracks the spoken
121
+ * word instead of sitting there like a badge.
122
+ */
123
+ function isAnimatedCaptionPart(node) {
124
+ return selfAndAncestors(node).some((n) => n.getAttribute("data-caption-animation") != null || n.getAttribute("data-cap-word") != null);
125
+ }
126
+ // Native platform artifacts that legitimately use rounded, filled bubbles:
127
+ // iMessage/DM threads, TikTok comment cards, fake chat UI. These are social
128
+ // furniture, not web furniture, so the badge rule must never fire on them.
129
+ const MOCK_UI_HINT = /(imessage|messenger|whatsapp|bubble|chat|sms|dm-|comment|reply|tweet|notification|caption-pill)/i;
130
+ function isMockSocialUi(node) {
131
+ return selfAndAncestors(node).some((n) => {
132
+ if (n.getAttribute("data-vf-mock-ui") != null)
133
+ return true;
134
+ const hay = `${n.getAttribute("class") ?? ""} ${n.getAttribute("data-label") ?? ""} ${n.getAttribute("data-slug") ?? ""} ${n.getAttribute("id") ?? ""}`;
135
+ return MOCK_UI_HINT.test(hay);
136
+ });
137
+ }
106
138
  function primaryFamily(raw) {
107
139
  return String(raw).split(",")[0].replace(/['"]/g, "").trim().toLowerCase();
108
140
  }
@@ -264,6 +296,42 @@ export function qaCompositionHtml(html) {
264
296
  fix: "Say each benefit as its OWN timed caption line on the footage, one at a time, in the font regime. One idea per beat reads far better than a chip strip."
265
297
  });
266
298
  }
299
+ // ── Rule: standalone badge pill ────────────────────────────────────────────
300
+ // The lonely capsule: ONE rounded, padded, filled tag holding a static stat or
301
+ // label — "10 hrs / week", "STEP 2", "EP.01", "+40%". badge-chip-row needs two
302
+ // siblings and cta-button needs action copy, so a single stat pill used to slip
303
+ // through both — and it is the most common surviving web tell in practice.
304
+ // Two signals: a capsule shape (radius well past a caption band) AND a fill.
305
+ // Excluded on purpose: active-word caption highlights (they MOVE with the
306
+ // spoken word — the one legitimate pill) and mock social UI, which is native.
307
+ for (const node of all) {
308
+ const text = textOf(node);
309
+ if (!text || text.length > 45)
310
+ continue;
311
+ const style = styleString(node);
312
+ const radius = radiusPx(style);
313
+ const filled = /background(?:-color|-image)?\s*:/.test(style) && !/background[^;]*:\s*(none|transparent)/.test(style);
314
+ // A caption band tops out around 8px; 20px+ on a text-sized box is a capsule.
315
+ if (!filled || radius < 20)
316
+ continue;
317
+ // Padding is what turns a hugging band into a tag. Either axis counts.
318
+ const padded = hasDecl(style, "padding") || hasDecl(style, "padding-left") || hasDecl(style, "padding-inline") ||
319
+ hasDecl(style, "padding-top") || hasDecl(style, "padding-block");
320
+ if (!padded)
321
+ continue;
322
+ if (isAnimatedCaptionPart(node) || isMockSocialUi(node))
323
+ continue;
324
+ // A capsule wrapping several elements is a card — card-panel's job, not ours.
325
+ if (Array.from(node.children ?? []).filter((c) => textOf(c)).length >= 2)
326
+ continue;
327
+ push({
328
+ rule: "static-pill",
329
+ severity: "error",
330
+ message: `Badge pill ("${text.slice(0, 30)}") — a filled ${radius >= 9999 ? "fully-rounded" : `${Math.round(radius)}px`} capsule with padding around static text. One is still a badge.`,
331
+ where: label(node, "pill"),
332
+ fix: "Drop the capsule and set the words themselves: bigger, heavier, ALL-CAPS, or an accent color — or circle/underline them. The only legitimate pill tracks the spoken word (set_captions spotlight/karaoke). Mock social UI is exempt: mark it data-vf-mock-ui."
333
+ });
334
+ }
267
335
  // ── Rule: card / panel / glassmorphism ─────────────────────────────────────
268
336
  // A rounded box with a border, shadow, or frosted blur, holding more than one
269
337
  // piece of content. A tight caption BAND is legal (small radius, one run) —
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@officexapp/vidfarm-devcli",
3
- "version": "0.21.32",
3
+ "version": "0.21.33",
4
4
  "description": "Local bridge for the Vidfarm Trackpad Editor. `vidfarm serve <template_id>` boots the FULL editor on localhost (disk-backed records/storage, free in-process render); edit composition.html on disk (Claude Code, Codex, etc.) and the browser live-morphs it.",
5
5
  "type": "module",
6
6
  "bin": {