@officexapp/vidfarm-devcli 0.21.31 → 0.21.33
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/editor-capabilities/SKILL.md +4 -3
- package/.agents/skills/vidfarm/SKILL.md +3 -2
- package/.agents/skills/vidfarm/recipes/bulk-scripting-with-a-regime.md +12 -0
- package/.agents/skills/vidfarm/recipes/local-edit-render-approve.md +3 -2
- package/.agents/skills/vidfarm/references/automation-and-local-dev.md +14 -1
- package/.agents/skills/vidfarm/references/core-workflows.md +62 -5
- package/.agents/skills/vidfarm/references/editor-workflows.md +5 -3
- package/.agents/skills/vidfarm/references/primitives.md +75 -32
- package/SKILL.director.md +174 -45
- package/SKILL.md +3 -1
- package/dist/src/cli.js +487 -4
- package/dist/src/devcli/dedupe-local.js +209 -0
- package/dist/src/devcli/qa-check.js +68 -0
- package/dist/src/lib/dedupe-recipe.js +420 -0
- package/package.json +3 -1
- package/public/assets/homepage-client-app.js +14 -14
package/SKILL.director.md
CHANGED
|
@@ -269,13 +269,13 @@ You may be running as the **in-web AI chat** (the /editor copilot, the chat dock
|
|
|
269
269
|
|
|
270
270
|
- **Determine the surface before claiming capabilities.** The web chat has only its declared tools and REST routes. It cannot execute arbitrary JavaScript/Python, open a shell, create a local repository script, or use the user's filesystem. Never tell a web-chat user that you ran code or wrote a script unless a dedicated declared tool actually did so. A desktop coding agent has a real shell and filesystem and MAY write/run scripts, perform arbitrary local computations over paginated API results, create reports/CSVs/JSON, edit composition files, and orchestrate long devcli workflows within the user's authorization.
|
|
271
271
|
- **The web AI chat can do all three paintbrushes** — clip raws, author HTML/hyperframe motion, and generate AI media — and it drives edits directly on the live timeline. Keep small-to-medium jobs here: text/caption swaps, a scene or two replaced, single generations, captions, approve/schedule. Just do them.
|
|
272
|
-
- **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r
|
|
272
|
+
- **No HTML slop — a video is not a web page.** Compositions are authored in HTML, so the #1 tell of an AI-edited video is landing-page furniture: gradient CTA "buttons" ("Sign Up for a Free Trial →"), rows of benefit chips/badges ("✓ No Credit Card Needed"), frosted/bordered cards holding a gradient headline + a URL, feature grids, bullet lists, web-default type (Inter/Roboto/Arial at weight 400-600). None of that exists in a real TikTok/Reel — **nothing in a video is clickable**. **The test is the native-editor test: could you have made this element with the tools inside TikTok's own editor?** That toolset is font / color / stroke / shadow / tight text box / alignment / rotation / animation presets, plus stickers, emoji, drawn marks and clips — it has no padded capsule, no border, no gradient fill, no blur panel, no card. If you reached past it, cut it. That includes **a single lonely pill around a stat or label** — `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`: being alone doesn't make a badge native, and the only legitimate capsule in a video is the active-word `spotlight`/`karaoke` highlight, which moves with the spoken word. Emphasize a stat the way the editor would instead: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle, or its own beat. If you are typing `btn` / `badge` / `chip` / `card` / `rounded-full` / `shadow-lg` / `backdrop-blur` / `bg-gradient-to-r` — or a `border-radius` over ~8px on anything filled that holds words — stop and rewrite it as timed text on the footage. Arrows, scribble/underline marks, italics, ALL-CAPS, single-word color pops, emoji, transparent cut-out stickers, and mock social UI (iMessage bubbles, comment cards) are all *fine* — they're native to the platform.
|
|
273
273
|
- **Every video gets the four charges — hook, loop, payoff, bait — and you write them BEFORE you touch the timeline.** This is the largest quality delta in the product and it costs nothing: most agent-made videos fail on structure, not polish, because the timeline is the fun part so it gets built first and the words get retrofitted. Invert it: (1) write the **hook line as text** — a complete clause, subject + verb, no jargon, naming a **situation** ("I've quit six businesses"), never a label ("anonymity") — and put it on screen at `start:0`; (2) name the **curiosity loop** and the timestamp it closes at, *inside this video* (if you can't state the timestamp, there is no loop, and the withheld answer must be one the viewer **can't guess**); (3) name the **payoff** — shown, not summarized, landing before the final beat; (4) write the **bait** — one ask in the final beat and in the post caption. Then build. Chunk 1 is read before any audio (muted autoplay is the default), so the text hook does more work than the spoken one. Full harness — the three gates, banned openers, loop mechanics, compliance, and diagnosis-by-charge — in `references/hooks-and-virality.md`; the checkable form is `vidfarm regime show hooks`.
|
|
274
274
|
- **The first frame IS the thumbnail — compose it on purpose.** Frame 0 is a single frame of ~30 in the first second, but it's the poster every feed, share link, and paused player freezes on, so **more people see that one frame than watch the video**. It must never be black, empty, mid-fade, or caught mid-animation: put a real visual at `start:0`, have the hook words already on screen at t=0, and never hang a `fade-black`/`fade-white`/`flash` *entrance* on the **first** clip (junction transitions between later clips are fine — this rule is only about the opening). Verify it, don't assume: devcli `vidfarm stills ./work --at 0` renders that exact frame, and `vidfarm qa` flags a blank or fading open.
|
|
275
275
|
- **Ask early: one-time video, or bulk?** "Make me a video about X" and "I need to post daily / give me 20 hook variants" are different jobs, and directors often don't know the second one has a name. Ask once, up front: *"One video, or should we set this up as a repeatable batch?"* Bulk = **scripting mode** (a pinned base fork + a loop that varies ONE thing per variant; a public-raws shelf is the cheapest source of the N), and every batch gets a **`QA_REGIME.md`** — because a loop of fifty videos has no human looking at every frame, and the regime is what replaces those eyes. Don't silently ship a one-off when they asked for volume, or drag someone into a harness when they wanted one clip.
|
|
276
276
|
- **`QA_REGIME.md` is the director's own quality contract, and it's a first-class artifact.** `vidfarm qa`'s built-ins are universal (slop, fonts, the thumbnail frame); a regime is what makes *this* format good — audience, hook shape, banned vocabulary, pacing, compliance line. It lives next to the work, they own it, it stacks: `vidfarm regime init short-form --out ./work/QA_REGIME.md` (bundled bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo` — each a starting point to **edit**, never a house style), then `vidfarm qa ./work --regime hooks --regime ./brand/HOUSE.md`, and any user file anywhere is valid. Its `checks:` front matter is machine-settled; its `- [ ]` checklist comes back as **review items you answer honestly in your report** — never claim a video passed the half the CLI can't judge. When a batch teaches you something, **write it back into the regime**: that's the artifact that compounds. Details in `references/automation-and-local-dev.md` ("Scripting mode"), format in `regimes/README.md`.
|
|
277
277
|
- **On devcli, QA every video you produce: `vidfarm qa ./work`.** Free, instant, local-only — it blocklists exactly the slop above plus first-frame/thumbnail and font-regime/safe-zone drift, and prints a concrete fix per finding. **Feedback, not a gate**: it exits 0 even on findings, never runs automatically, and is a blocklist (unusual/stylized compositions pass untouched), so there's no reason not to run it before every publish. `--json` for scripted batches, `--strict` only if you want a CI failure. **Web-chat copilot: this command does not exist for you** (devcli-only, no REST twin) — apply the standard by hand, and when handing a heavy job to a local coding agent, tell them to run `vidfarm qa`.
|
|
278
|
-
- **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
|
|
278
|
+
- **Every production adheres to the TikTok-native caption standard.** On-screen text lives inside the readable safe zone (**~8%–85%** of a 9:16 frame — never pinned to the top/bottom edges the phone UI clips), uses the composition's bold font regime (**Montserrat** default / TikTok Sans, weight **700–900**, ~36–64px on a 1080-wide frame), and uses **exactly one of four valid backgrounds**: outline/stroke (`background_style:"outline"`, the default), plain + shadow (`"plain"`), an active-word highlight pill (`set_captions` `spotlight`/`karaoke` — the only legitimate pill *in the whole frame*, static labels included), or a tight-hugging solid band (`"highlight-solid"`, radius ≤8px, no border/shadow/gradient/blur). Decomposed forks often inherit the source's edge-pinned caption in an off-regime font — fix it, don't inherit it. Local devcli renders auto-normalize position + font family only (never the slop), so author it correctly. Full rules in `references/editor-workflows.md` ("Social-native visual standard" + "TikTok-native caption standard").
|
|
279
279
|
- **Where the web chat struggles: complex, long, multi-step transformations.** A full multi-scene re-theme, an iterative render-critique-iterate loop, heavy scripted or batch work, or anything needing a real filesystem and many sequential tool calls will hit context limits, turn/timeout ceilings, and the web editor's constraints (CSS/declarative motion only — JS animation adapters are stripped on save). Don't grind a big transformation one layer at a time in a chat turn and stall.
|
|
280
280
|
- **Practical workaround — hand the heavy job to local devcli.** When a task is genuinely large or long-running, **proactively recommend the director run it locally with an AI coding agent** (Claude Code / OpenAI Codex / any capable agent): `vidfarm pull <forkId>` writes the composition + the `.harness/` grounding bundle to disk, the agent edits with the full devcli verb set and JS animation adapters, renders free with `vidfarm serve`, and `vidfarm publish` pushes it back. This is the **best-quality (B) harness's** natural home (adversarial grading with a coding agent). Frame it as "this is a big rebuild — you'll get a better, faster result running it locally with a coding agent; here's how," not as a dead end.
|
|
281
281
|
- **Offer a handoff, do not impersonate the desktop agent.** When web chat reaches that boundary, offer to save a Markdown handoff in My Files containing the objective, selected template/fork IDs, asset paths, grounding, constraints, completed work, and suggested devcli commands. Create it only after the user agrees. The desktop agent should read that document, pull the referenced fork, and then use its actual code/shell capabilities.
|
|
@@ -319,6 +319,7 @@ Choose the narrowest path that satisfies the request.
|
|
|
319
319
|
4. If the task is “find footage” or “use our existing assets,” read `references/assets-and-sourcing.md`.
|
|
320
320
|
4b. If the task is **“download this video/audio off a website”** (a pasted YouTube / TikTok / Instagram / X post URL the user wants the actual file from), Vidfarm does that for you on a **paid plan** — `POST /api/v1/primitives/videos/download` (or `/audio/download`), devcli `vidfarm download-video <url>` / `vidfarm download-audio <url>`. **Free plan → do not call it; walk the user through opening the URL in Chrome and downloading it from the page, then `vidfarm put-file` the local file in for $0.** Details in `references/assets-and-sourcing.md` → `references/primitives.md`.
|
|
321
321
|
4c. If the task is **“turn this Reddit/X thread, subreddit, or account into a video”** — “tweet to TikTok”, “Reddit to TikTok”, “make a video from this thread”, “what are the top comments saying” — run `vidfarm recycle <source>` (or `POST /api/v1/primitives/social/recycle`) with the URL. It **decomposes** the source into raw JSON (text, comment tree, media URLs, author pics, stats) and hands it back unranked so YOU pick what to remix. **Paid plan; `max_records` is the spend ceiling.** Brokers the reddit-lead-gen / x-lead-gen OfficeX apps, so it waits out their async job for you. Details in `references/assets-and-sourcing.md` → `references/primitives.md`.
|
|
322
|
+
4d. If the task is **“post this again / to several accounts / on another platform”**, or you are about to publish or bulk-produce at all — that is **deduplication**. Run `vidfarm dedupe <mp4> [--variants N]` on the **exported file** (free, local ffmpeg, no re-render), then approve/schedule each variant. **Ask the operator whether they want deduplicated copies, and how many, BEFORE the render/bulk run** — deciding after means paying for a second render. Details in `references/core-workflows.md` → *Deduplicate before you publish* and `references/primitives.md` → *Primitive: media_dedupe*.
|
|
322
323
|
5. If the task is scripted, local, CI-driven, or `vidfarm serve`-based, read `references/automation-and-local-dev.md`.
|
|
323
324
|
6. If the task explicitly asks for a primitive or needs specialized generation/transcription work, read `references/primitives.md`.
|
|
324
325
|
7. If the task is the MARKETPLACE (ordering videos from specialist agents): browsing is web-only for paying customers — send the human to https://vidfarm.cc/marketplace, never render it locally. Placing/listing orders is the thin REST wrapper in `references/core-workflows.md` (§ Marketplace); anything deeper on a gig (inbox, proofs, payouts) needs the external Dollar Platoon skill — `npx skills add https://github.com/OfficeXApp/dollarplatoon-skill` — the same way FlockPoster work beyond scheduling needs `npx skills add https://github.com/OfficeXApp/flockposter-skill`.
|
|
@@ -489,6 +490,25 @@ Every publish creates an immutable version snapshot at `versions/<N>/composition
|
|
|
489
490
|
|
|
490
491
|
The Web UI **Render** button and devcli render both use this same endpoint. The fast `202` response includes the deterministic `expectedOutputPublicUrl` so a caller can store or pass along the final public S3 URL before the render has completed, then poll by `renderId` until `status` settles.
|
|
491
492
|
|
|
493
|
+
## Deduplicate before you publish (ASK FIRST)
|
|
494
|
+
|
|
495
|
+
Social platforms fingerprint every upload. If a render is going out **more than once** — to several accounts, to a second platform, or again in a few weeks — the later copies get suppressed as duplicate/reused content unless each one carries a distinct fingerprint.
|
|
496
|
+
|
|
497
|
+
**Before you render for publication, and before any bulk run, ask the operator: "Do you want deduplicated copies for posting? How many?"** Ask *then*, not after — dedupe runs on the exported MP4, so the order is **render once → dedupe N times**. Getting the answer up front is what stops you paying for a second render later.
|
|
498
|
+
|
|
499
|
+
```bash
|
|
500
|
+
# free, offline, no wallet, no re-render — the default path
|
|
501
|
+
vidfarm dedupe ./out/final.mp4 # one distinct copy
|
|
502
|
+
vidfarm dedupe ./out/final.mp4 --variants 3 --out-dir ./out/posts
|
|
503
|
+
```
|
|
504
|
+
|
|
505
|
+
- Default preset `standard` = skew 2%, zoom 3%, rotate 2°, speed +2%, saturation +4%, plus contrast/brightness/hue/grain, a container-metadata strip and a per-variant CRF walk. Invisible to a viewer.
|
|
506
|
+
- `--variants N` mints N copies that differ from the original **and from each other** — one per account/slot. Two accounts posting the same variant defeats the point.
|
|
507
|
+
- `--preset light|standard|strong` for how hard to push; `--rotate 0` if the forced corner-hiding crop (~6.7% on a tall frame at 2°) matters more than fingerprint distance.
|
|
508
|
+
- Cloud equivalent: `POST /api/v1/primitives/media/dedupe` — same transforms, billed. See `references/primitives.md` → *Primitive: media_dedupe*.
|
|
509
|
+
|
|
510
|
+
Dedupe the **finished MP4**, then approve/schedule each variant separately. Do not dedupe the composition and re-render.
|
|
511
|
+
|
|
492
512
|
## Approve a finished post
|
|
493
513
|
|
|
494
514
|
A render produces a bare MP4 URL. **Approving** wraps that MP4 (plus caption, title, pinned comment, and any carousel slides) into a shareable preview page — the phone-mockup page a human opens to review and copy the post.
|
|
@@ -510,21 +530,59 @@ devcli: `vidfarm approve --video <mp4-url> --caption "..."` prints the `share_ur
|
|
|
510
530
|
|
|
511
531
|
## Schedule a post
|
|
512
532
|
|
|
513
|
-
|
|
533
|
+
A **destination** is one connected channel. There are exactly two kinds, and they are billed and owned differently:
|
|
534
|
+
|
|
535
|
+
- **`email`** — a verified email address on the vidfarm account. Delivered by vidfarm itself. Your vidfarm API key is the only credential involved.
|
|
536
|
+
- **`flockposter`** — a social account (TikTok/IG/X/…) connected through FlockPoster, a separate product. Needs the customer's FlockPoster key saved in vidfarm Settings → Channels. Vidfarm brokers the call; FlockPoster does the posting.
|
|
537
|
+
|
|
538
|
+
Every account always has at least one email destination: its own signup address, pre-verified, created automatically. So `destination_type: "email"` works on a fresh account with no setup at all.
|
|
539
|
+
|
|
540
|
+
**List destinations first** (or just send an address — see below):
|
|
541
|
+
|
|
542
|
+
```
|
|
543
|
+
GET /api/v1/user/me/channels
|
|
544
|
+
→ { "channels": [ { "destination_type": "email", "destination_id": "cus_…:default-email",
|
|
545
|
+
"handle": "operator@example.com", "status": "verified", "schedulable": true,
|
|
546
|
+
"accepts": ["cus_…:default-email", "operator@example.com", "operator"] } ],
|
|
547
|
+
"flockposter_connected": false, "flockposter_error": null }
|
|
548
|
+
```
|
|
549
|
+
|
|
550
|
+
Then schedule:
|
|
514
551
|
|
|
515
552
|
```
|
|
516
553
|
POST /api/v1/approved/posts/:postId/schedules
|
|
517
554
|
Content-Type: application/json
|
|
518
555
|
|
|
519
|
-
{ "destination_type": "flockposter" | "email", "destination_id": "<
|
|
556
|
+
{ "destination_type": "flockposter" | "email", "destination_id": "<see below>",
|
|
520
557
|
"scheduled_at": "2026-07-10T14:00:00Z", "timezone": "America/New_York", "additional_notes": "optional" }
|
|
521
558
|
```
|
|
522
559
|
|
|
523
|
-
|
|
560
|
+
**`destination_id` accepts whatever you have** — it is resolved server-side, so you do not need to look up an id:
|
|
524
561
|
|
|
525
|
-
|
|
562
|
+
| For `email` | For `flockposter` |
|
|
563
|
+
|---|---|
|
|
564
|
+
| the channel id (`cus_…:default-email`, or `email:<id>`) | the integration id |
|
|
565
|
+
| the address (`operator@example.com`) | the handle (`@brandname` or `brandname`) |
|
|
566
|
+
| the local part (`operator`) | the channel title |
|
|
567
|
+
| the channel title | the platform (`tiktok`) when exactly one is connected |
|
|
568
|
+
|
|
569
|
+
Matching is case-insensitive. When nothing matches, the `400` names what was searched for and lists the channels that do exist — read it rather than guessing another id.
|
|
570
|
+
|
|
571
|
+
Minimum 10-minute lead time. Response (`201`) is the schedule record.
|
|
572
|
+
|
|
573
|
+
**Managing a schedule** (all authed with the same vidfarm API key — an agent has full control):
|
|
574
|
+
|
|
575
|
+
- `GET /api/v1/approved/posts/:postId/schedules` — browse this post's schedules
|
|
576
|
+
- `PATCH /api/v1/approved/posts/:postId/schedules/:scheduleId` — reschedule; same body as POST. Cancels the queued send and re-queues it.
|
|
577
|
+
- `DELETE /api/v1/approved/posts/:postId/schedules/:scheduleId` — cancel the queued send
|
|
578
|
+
|
|
579
|
+
Cancel/reschedule reaches through to the provider (Resend for email, FlockPoster for social), so a `400` here means the send was **not** stopped — the email or post is still queued. Never report a failed cancel as cancelled.
|
|
580
|
+
|
|
581
|
+
Schedules created outside vidfarm (`managed_by: "external"`) are read-only here and return `409`; change those in FlockPoster.
|
|
582
|
+
|
|
583
|
+
devcli: `vidfarm channels` lists destinations, `vidfarm schedule <postId> --at <iso> --to <destination> [--type flockposter|email]` schedules, `vidfarm schedules <postId>` browses.
|
|
526
584
|
|
|
527
|
-
Deeper FlockPoster work (channel management, direct posting/analytics outside vidfarm's schedule wrapper) is FlockPoster's own API — grab its skill first: `npx skills add https://github.com/OfficeXApp/flockposter-skill` (mirrored as `vidfarm skills add flockposter`).
|
|
585
|
+
Deeper FlockPoster work (connecting accounts, channel management, direct posting/analytics outside vidfarm's schedule wrapper) is FlockPoster's own API, not vidfarm's — grab its skill first: `npx skills add https://github.com/OfficeXApp/flockposter-skill` (mirrored as `vidfarm skills add flockposter`).
|
|
528
586
|
|
|
529
587
|
## Marketplace — order videos from specialist agents
|
|
530
588
|
|
|
@@ -1150,12 +1208,12 @@ Web copilot: same standard, applied by hand — check the opening layer's `start
|
|
|
1150
1208
|
|
|
1151
1209
|
Compositions are authored in HTML, so the single most common way an AI-edited video betrays itself is by **looking like a web page**. A landing-page button, a row of feature chips, a frosted card with a gradient headline — these are everywhere on the web and **nowhere in a real social post**. A director scrolling TikTok/Reels/Shorts never sees them, so the moment one appears the video reads as machine-made and the retention dies. HTML and hyperframes are the right tool; **Bootstrap/Tailwind furniture is not.**
|
|
1152
1210
|
|
|
1153
|
-
**The one test:** *
|
|
1211
|
+
**The one test — the native-editor test:** *could you have made this element with the tools inside TikTok's (or CapCut's) own editor?* That editor's text tool gives you exactly this: a font, a color, a stroke/outline, a soft shadow, a tight text-box band, alignment, opacity, rotation, and animation presets — plus stickers, emoji, drawn marks, and clips. It does **not** give you a padded capsule, a border, a gradient fill, a blur panel, or a card. If you had to reach past that toolset, you are decorating like a web designer, and the frame will read as machine-made no matter how good the copy is. The second half of the same test: if the element's whole job is to look **clickable**, cut it — **nothing in a video is clickable.**
|
|
1154
1212
|
|
|
1155
1213
|
**BANNED — never author, and strip on sight when a fork or a paste brings one in:**
|
|
1156
1214
|
|
|
1157
1215
|
- **CTA / "button" shapes.** Any filled capsule or rounded rect containing action copy — "Sign Up for a Free Trial →", "Learn More", "Get Started", "Book a Call" — especially gradient-filled with a glow or drop shadow. A social CTA is *spoken*, or a plain caption line, or an arrow pointing at the real UI. Not a button.
|
|
1158
|
-
- **
|
|
1216
|
+
- **Badges, chips, pills — including a SINGLE one.** The "✓ ID-Verified Tutors · ✓ No Credit Card Needed · ✓ 30-Min Trial" strip is the obvious case, but the far more common one is **one lonely capsule holding a stat or a label**: `( 10 hrs / week )`, `( STEP 2 )`, `( EP.01 )`, `( BEGINNER )`, `( +40% )`. Being alone does not make it native — a rounded, padded, filled tag around static text is a `<span class="badge">` wearing a different hat, and it is one of the loudest web tells in the whole frame. **The only legitimate pill in a video is the active-word highlight** (`spotlight`/`karaoke`), because it tracks the spoken word and moves. Static text gets `outline`, `plain`, or a tight band that hugs the glyphs (radius ≤ ~8px). If a stat deserves emphasis, give it emphasis the *editor* can give: bigger, heavier, ALL-CAPS, an accent color, a hand-drawn circle or underline around it, its own beat on screen.
|
|
1159
1217
|
- **Cards, panels, containers.** A rounded-rect box (dark, bordered, shadowed, or `backdrop-filter` frosted) holding a heading plus a subheading/URL. Text belongs **on the footage**, not inside a floating panel with padding around it.
|
|
1160
1218
|
- **Gradient text fills, neon border glows, elevation shadows, glassmorphism, hover-implying strokes.**
|
|
1161
1219
|
- **Web page furniture of any kind:** navbars, hero sections, feature grids, two-column layouts, `<ul>` bullet lists with disc markers, tables, alerts, progress bars as decoration, logo-in-a-circle avatars, "as seen in" strips.
|
|
@@ -1163,6 +1221,8 @@ Compositions are authored in HTML, so the single most common way an AI-edited vi
|
|
|
1163
1221
|
|
|
1164
1222
|
**Greppable smell test.** If you are typing `class="btn…"`, `badge`, `chip`, `tag`, `card`, `panel`, `container`, `row`/`col-`, `hero`, `cta`, `rounded-full`, `shadow-lg`, `backdrop-blur`, `bg-gradient-to-r`, `border border-…` — **stop.** You are building a web page, not a video. Rewrite as timed text on footage.
|
|
1165
1223
|
|
|
1224
|
+
**The capsule rule of thumb.** On any element that holds words: `border-radius` over ~8px **combined with** a background fill and padding = a badge. Either take the fill away (bare text + outline/shadow) or take the radius and padding down until the band hugs the glyphs. There is no third option for static text.
|
|
1225
|
+
|
|
1166
1226
|
**ALLOWED and encouraged — these ARE social-native:**
|
|
1167
1227
|
|
|
1168
1228
|
- **Arrows** (drawn, animated, hand-style), circles/scribbles/underline strokes highlighting part of the frame, hand-drawn marks, checkmarks *as glyphs inside a caption line* (not as chips).
|
|
@@ -1189,7 +1249,7 @@ Short-form is watched on a phone, and the phone's UI eats the frame's edges. **N
|
|
|
1189
1249
|
| 3 | **Highlight pill behind the ACTIVE word only** | `set_captions caption_style:"spotlight"` / `"karaoke"` (+ `caption_highlight_color`) | Hormozi/CapCut word-by-word. **The only legitimate "pill" in a video** — it tracks the spoken word, so it isn't a badge |
|
|
1190
1250
|
| 4 | **Solid band that tightly hugs the text lines** (CapCut "text box") | `background_style:"highlight-solid"` (or `"highlight-translucent"`) + a `background` color | Guaranteed legibility over noisy footage |
|
|
1191
1251
|
|
|
1192
|
-
Treatment 4 is a **band, not a card**: it hugs the glyphs with minimal padding, corner radius ≤ ~8px, **no border, no drop shadow, no gradient, no blur**, and it wraps *one* text run — never a heading + subheading + URL stacked inside one rounded box. The moment it grows padding, a stroke, or a second element inside it, it has become a web card. Fix it.
|
|
1252
|
+
Treatment 4 is a **band, not a card**: it hugs the glyphs with minimal padding, corner radius ≤ ~8px, **no border, no drop shadow, no gradient, no blur**, and it wraps *one* text run — never a heading + subheading + URL stacked inside one rounded box. The moment it grows padding, a stroke, or a second element inside it, it has become a web card. Fix it. And the moment its radius goes fully round, it has become a **badge** — treatment 3 is the *only* capsule allowed, and only because it tracks the spoken word. A static "10 hrs / week" in a rounded pill is web furniture; the same words in treatment 1 or 2, bigger and heavier, are a beat.
|
|
1193
1253
|
|
|
1194
1254
|
**A common trap: decomposed templates mirror the source's caption placement**, so a forked meme can arrive with its caption pinned at `top:0` in a non-regime font — and a re-theme prompt ("make this for my tutoring service") is exactly where an agent starts inventing landing-page CTAs and benefit chips because the *subject* is a SaaS product. **Fix to the standard, don't inherit it, and don't import the website's design language into the video.** When placing text yourself (`set_captions`, `set_layer_text`, `set_layer_style`, `add_layer`, devcli `place`/`captions`), set `y` / `font_family` / `font_weight` / `background_style` to the standard from the start.
|
|
1195
1255
|
|
|
@@ -1791,6 +1851,17 @@ for VARIANT in "${VARIANTS[@]}"; do
|
|
|
1791
1851
|
done
|
|
1792
1852
|
```
|
|
1793
1853
|
|
|
1854
|
+
**Before you start a bulk run, ask the operator whether the output should be deduplicated, and for how many posting slots.** A bulk run's whole point is volume across accounts/platforms, which is exactly the shape platforms flag as duplicate content. Dedupe is a post-render ffmpeg pass, so asking up front is what keeps it at *render once → dedupe N* instead of a second render per slot:
|
|
1855
|
+
|
|
1856
|
+
```bash
|
|
1857
|
+
# after the loop: one distinct copy per posting slot, free and offline
|
|
1858
|
+
for MP4 in renders/*.mp4; do
|
|
1859
|
+
vidfarm dedupe "$MP4" --variants "$SLOTS" --seed "$(basename "$MP4" .mp4)" --out-dir ./posts
|
|
1860
|
+
done
|
|
1861
|
+
```
|
|
1862
|
+
|
|
1863
|
+
Reuse one `--seed` per source so a batch is reproducible, and post each variant to a **different** account — two accounts posting the same variant defeats the point. See `references/core-workflows.md` → *Deduplicate before you publish*.
|
|
1864
|
+
|
|
1794
1865
|
`vidfarm qa` still exits 0 on findings — the gate above is the *script's* choice, made explicit with `jq`, not a behavior change in the tool. Keep it that way: an agent that can't ship a deliberately weird variant will quietly stop trying weird variants.
|
|
1795
1866
|
|
|
1796
1867
|
This section is for a **desktop/local coding agent**, not the web copilot. A local Codex/Claude agent may use its shell and filesystem to write JavaScript/TypeScript/Python/shell scripts, fetch every API page, join and score catalog/library data, calculate statistics, emit CSV/JSON/Markdown reports, manipulate composition DOM files, and run iterative render/inspection loops. The web copilot cannot inherit those abilities from this document: it may only call its declared tools and bounded REST routes. If web chat prepares work for this flow, consume its My Files handoff document as input; do not claim the web chat itself executed the script.
|
|
@@ -1927,6 +1998,7 @@ The licensed harness also carries the **generative build workflow** guidance (ch
|
|
|
1927
1998
|
| `vidfarm inpaint <image> --mask <png> --prompt "…" [--region "label=…"] [--ref …] [--out <f>]` | `POST /api/v1/primitives/images/inpaint` (polls job) | masked image EDIT — replace ONLY the transparent-mask region, keep everything else (devcli twin of the /inpaint page) |
|
|
1928
1999
|
| `vidfarm create-overlay "<subject>" [--key-color #00FF00] [--aspect-ratio 1:1] [--place <dir>] [--out <f>]` | `POST /api/v1/primitives/images/create-overlay` (polls job) | **Vox-style** transparent OVERLAY — AI image on a forced key-color background, chroma-keyed out in one job → ready-to-composite transparent PNG |
|
|
1929
2000
|
| `vidfarm remove-greenscreen <image\|video> [--preset green\|blue\|white\|black\|digital-green\|magenta] [--key-color #00FF00] [--tolerance 0.3] [--local] [--gif] [--out <f>]` | `POST /api/v1/primitives/remove-greenscreen` (polls job) | chroma-key a FLAT solid background → transparent PNG/WebP (image) or WebM/VP9-alpha (video); auto-detects media kind. `--local` runs it FREE in-process (sharp/ffmpeg, no wallet); default cloud is billed at real compute × 1.2. **`--gif` writes a transparent GIF instead** (ANIMATED for a clip; `--gif-fps`/`--gif-width`/`--gif-alpha`) — local-only, 1-bit alpha, for GIF-only sticker surfaces; prefer PNG/WebP/WebM for compositing. Aliases: `greenscreen`, `remove-background-greenscreen`. |
|
|
2001
|
+
| `vidfarm dedupe <video\|image\|url> [--preset light\|standard\|strong] [--variants N] [--seed <s>] [--zoom/--rotate/--skew/--speed/--saturation/--hue/--noise/--flip] [--local\|--cloud] [--out <f>\|--out-dir <d>]` | **local, free, ffmpeg-only** by default (no job); `--cloud` = `POST /api/v1/primitives/media/dedupe` (polls job) | **DEDUPLICATION — the publish-safety pass.** Makes a finished render read as a NEW upload to a platform's duplicate-content detector, invisibly to a viewer. Default preset `standard` = skew 2%, zoom 3%, rotate 2°, speed +2%, saturation +4%, plus contrast/brightness/hue/grain, a container-metadata strip and a per-variant CRF walk. **Runs on the EXPORTED file — never re-render for this.** `--variants N` mints N copies that differ from the original AND from each other (jittered magnitudes, alternating signs), one per account/posting slot; `--seed` makes a batch reproducible. A rotate forces a bigger centre-crop to hide the black corners (~6.7% on a tall frame at 2°) and says so — pass `--rotate 0` when framing matters more. `--flip` is the strongest single knob but visibly reverses on-screen text. **Ask the operator whether they want this BEFORE publishing or bulk-producing.** Aliases: `dedup`, `deduplicate`, `uniquify`. |
|
|
1930
2002
|
| `vidfarm cutout <image\|url> [--generate "<prompt>"] [--preset green] [--pad <px>] [--alpha-threshold <n>] [--no-trim] [--output-format png\|webp] [--out <f>]` | **local, free, ffmpeg-only** (no job) — key + `alphaextract`/`cropdetect` trim | **The transparent explainer-STICKER maker.** Keys out the flat plate **and then shrinks the canvas to the cutout's true min width/height** (a 1024² mostly-empty plate → a snug sticker whose pixel size IS the subject) so you can scale/position it precisely. `--generate` AI-generates the graphic first on a matching chroma plate (that step is the billed image primitive), then keys+trims in one shot; without it, keys+trims a file/url you already have. **IMAGE-only** (a moving subject has no single bounding box — key a clip with `remove-greenscreen`). Prefer this over `create-overlay` locally: same idea, but free and auto-trimmed. `--pad` keeps transparent breathing room; `--json` reports final `width`/`height`/`area_reduced_pct`, plus `hole_pct`/`hollow` — the "the key ate the fill" check (outline-only art keys into a rim around a transparent hole; `--generate` prompts against it automatically, and the console prints a `Hollow:` warning with the fix). Alias: `sticker`. See recipe `cutout-graphics-for-explainers.md`. |
|
|
1931
2003
|
| `vidfarm mask <image\|url> [--crop x,y,w,h] [--flat <hex>] [--pad <px>] [--alpha-threshold <n>] [--no-trim] [--output-format png\|webp] [--keep-region <f>] [--out <f>]` | **local, free** (no job) — ffmpeg crop + ONNX matting (or ffmpeg chroma-key) + `cropdetect` trim | **Lift an illustration OUT of an image you already have** (infographic / poster / marketing graphic / brand sheet / screenshot) → snug transparent PNG, the same reusable explainer sticker `cutout` makes but with **$0 and zero AI generation** — the cost-saving move whenever source art exists. `--crop x,y,w,h` (pixels **or** %) isolates ONE element from a multi-illustration source before masking (re-run with different rects to grab each). Background removed by **local ONNX matting** (any/busy background) by default, or **`--flat <hexcolor>`** chroma-keys a solid fill for crisper edges (an infographic's cream/white paper); then trims to the subject's true min width/height. **IMAGE-only** (matte a clip with `remove-background`). Aliases: `isolate`, `extract`. See recipe `cutout-graphics-for-explainers.md` → "Mask from an image you already have". |
|
|
1932
2004
|
| `vidfarm sticker-pack [sheet\|url] [--generate "<theme>"] [--items "a,b,c"] [--count <n>] [--dry-run] [--gap <pct>] [--min-area <pct>] [--output-format png\|webp\|gif] [--out-dir <d>]` | **local, free, ffmpeg-only** (no job; only `--generate` bills, ONCE for the whole set) — key + alpha-channel segmentation + per-item trim | **The STICKER-PACK maker — the answer whenever a director asks for "a sticker pack" / prop set / icon set.** A pack is ONE greenscreen sheet holding every item, keyed once and then masked apart: 1/N the cost of N `cutout` calls, and the only way a cast stays on-style. Finds each item **automatically** by segmenting the keyed sheet's alpha into connected islands — no hand-measured `--crop` rects — and writes one snug transparent file per item (named from `--items`, reading order) plus a `stickers.json` manifest. `--dry-run` prints the detected boxes first; `--gap` merges (lower) or splits (raise) items that came out joined/broken; items have **no maximum size** — a full-frame landscape/backdrop is as valid a sticker as a 3% icon. **Plate color is chosen for you:** when generating it reads the subject and moves the plate off any hue the art uses (green → magenta → blue → black → white — a pack of leaves/frogs/money on GREEN would key holes through the art), and when splitting an existing sheet it DETECTS the plate from the sheet's four corners, so a red/purple sheet handed back from a web tool just works. Pin it with `--key-color`/`--preset`, or `--no-auto-key` for plain green. **The ART is made key-safe too:** the generation prompt is auto-appended with "closed, solidly filled shapes, no outline-only/hollow art, nothing in the plate hue or a near-shade, fully opaque, no glow/translucency" — the fix for stickers that come back as a rim around a transparent hole — and after keying each item reports `holes`/`hole_pct`/`hollow` (console `⚠ N% hollow` at ≥20%, plus `--json` and `stickers.json`). It **warns, never blocks** (a ring/frame/donut reads identically); re-generate with the fill clause, or lift that one item with `vidfarm mask --crop …`. `--output-format gif` emits 1-bit-alpha GIFs for GIF-only surfaces. IMAGE-only. Aliases: `stickers`, `sticker-sheet`. See recipe `cutout-graphics-for-explainers.md` → "A sticker pack". |
|
|
@@ -2021,6 +2093,7 @@ What it flags:
|
|
|
2021
2093
|
|---|---|---|
|
|
2022
2094
|
| `cta-button` | error | Action copy ("Sign Up for a Free Trial →") **inside** a filled/gradient rounded capsule. Bare CTA copy in a caption is fine — "BUY NOW" is real social copy |
|
|
2023
2095
|
| `badge-chip-row` | error | 2+ sibling small rounded filled tags — the "✓ No Credit Card Needed · ✓ 30-Min Trial" strip |
|
|
2096
|
+
| `static-pill` | error | ONE filled, padded, ≥20px-radius capsule around static text — a stat/label badge like "10 hrs / week", "STEP 2", "EP.01". Skips active-word `spotlight`/`karaoke` highlights (the only legitimate pill) and mock social UI (chat bubbles, comment cards — mark yours `data-vf-mock-ui` if the heuristic misses it) |
|
|
2024
2097
|
| `card-panel` / `glass-card` | error | A rounded box with a border/shadow/`backdrop-filter` wrapping 2+ elements. A tight legibility **band** (radius ≤8px, one text run, no border/shadow) stays legal |
|
|
2025
2098
|
| `gradient-text` | error | `background-clip:text` gradient headline fills |
|
|
2026
2099
|
| `clickable-element` | error/warn | `<button>`, `<form>`, `<input>`; `<a href>` warns |
|
|
@@ -2050,7 +2123,7 @@ The four modes, quoted as **cost per finished video**. The first two are spend p
|
|
|
2050
2123
|
|
|
2051
2124
|
**All of it bills to the user's own AI provider keys (BYOK)** — the keys saved with `vidfarm add-provider-key <provider> <key>` or at **Settings → Bring your own keys** (<https://vidfarm.cc/settings/developer>). The model providers charge those keys directly; Vidfarm wallet credits only come into play when the user deliberately runs on the platform key instead of their own. So `minimize` isn't "cheap", it's **zero**: nothing reaches a paid key at all.
|
|
2052
2125
|
|
|
2053
|
-
`vidfarm cost-mode <minimize|hybrid|rich-ai|pure-videogen>` records a single spend preference (in `~/.vidfarm/cost-mode.json`) that every **billed** command honors: `generate`, `music`, `decompose`, `inspiration-decompose`, `create`, `replicate`, `inpaint`, `create-overlay`, `cutout --generate` (only the generation step; a bare `cutout` on an existing graphic is free), and the cloud paths of `render --target cloud`, `tts --cloud`, `stt --cloud`, `remove-greenscreen` (non-`--local`)
|
|
2126
|
+
`vidfarm cost-mode <minimize|hybrid|rich-ai|pure-videogen>` records a single spend preference (in `~/.vidfarm/cost-mode.json`) that every **billed** command honors: `generate`, `music`, `decompose`, `inspiration-decompose`, `create`, `replicate`, `inpaint`, `create-overlay`, `cutout --generate` (only the generation step; a bare `cutout` on an existing graphic is free), and the cloud paths of `render --target cloud`, `tts --cloud`, `stt --cloud`, `remove-greenscreen` (non-`--local`), `dedupe --cloud`. FREE local engines never gate (`render` local default, `tts --engine local`, `stt --engine whisper`, `remove-greenscreen --local`, `dedupe --local`, `cutout` on a file/url, all the file-editing verbs). `vidfarm media search` (the free stock catalog — pixabay/openverse/iconify music, SFX, images, icons, stock video) is also free and never gates.
|
|
2054
2127
|
|
|
2055
2128
|
- **minimize ($0 videos)** — a billed op is **refused** unless you add `--yes`; the error names the free local alternative (which now includes the matching `vidfarm media search` for music/SFX/image/video). Use this to guarantee no surprise AI spend. Before paying to generate music, sound effects, or images, try `vidfarm media search "<meaning>" --type <bgm|sfx|image|vector|icon|video>` first — free royalty-free assets instead of a billed `music`/`generate` call. **Check the keyless sources first — Openverse (CC/CC0 music, SFX, images) and iconify (icons) need no account at all**, so they always work in `minimize`. Photos/vectors/stock-video need a **free Pixabay key** that **may already be saved** — check `vidfarm provider-keys` (or web **Settings → Bring your own keys** / <https://vidfarm.cc/settings/developer>) before assuming a short result means "no key." If absent, save one once: `vidfarm add-provider-key pixabay <key>` (free key from <https://pixabay.com/api/docs/>), the Settings surface, or hand it to the desktop AI agent to run that command.
|
|
2056
2129
|
- **minimize still gets CUSTOM images — via a free manual generator.** A refused `generate` is not the end of the road. Offer the user the manual loop (ask once, then make it the session default): **you write the prompt → they run it free in <https://meta.ai>, free-tier ChatGPT, or a free Hugging Face image Space (<https://huggingface.co/spaces>) → they hand the PNG back** via `vidfarm put-file ./sheet.png` or web **My Files**. Ask for **one sheet holding every graphic you need**, gridded on a **flat pure-green plate** (`#00FF00`), no text — one round trip instead of N, which saves the user's time and your tokens. Then split it locally for $0: `vidfarm mask ./sheet.png --crop x,y,w,h --flat "#00FF00" --out prop-a.png`, once per element (drop `--flat` and let local ONNX matting handle it if the tool ignored the green background). Full prompt template + loop: recipe `recipes/cutout-graphics-for-explainers.md` (“Free manual image-gen”).
|
|
@@ -2447,38 +2520,83 @@ curl -X POST "$VIDFARM_BASE/api/v1/primitives/videos/remove-captions" \
|
|
|
2447
2520
|
|
|
2448
2521
|
## Primitive: media_dedupe
|
|
2449
2522
|
|
|
2450
|
-
|
|
2523
|
+
**Deduplication** — make a finished image or video read as a **new upload** to a social platform's duplicate-content detector, while staying invisible to a viewer. Platforms fingerprint every upload; posting the same render twice (across accounts, or again next month) gets the later copy suppressed as duplicate/reused content. This primitive nudges geometry, color, timing and grain by a couple of percent, strips container metadata, and walks the encoder's CRF, so each copy carries a distinct fingerprint.
|
|
2524
|
+
|
|
2525
|
+
**Ask the operator before you publish or bulk-produce.** Dedupe runs on the EXPORTED file, so the correct order is *render once → dedupe N times*, never *render N times*. Deciding up front avoids paying for a second render later.
|
|
2451
2526
|
|
|
2452
2527
|
- `POST /api/v1/primitives/media/dedupe`
|
|
2453
2528
|
- Body: `{ "tracer": "...", "payload": { ...fields... }, "webhook_url"?: "..." }`
|
|
2454
2529
|
- Note: webhook delivery is not yet active — `webhook_url` is accepted and persisted on the job but never fired. Poll the job endpoints (`GET /api/v1/primitives/jobs/:jobId`) for completion.
|
|
2455
|
-
- Response: standard primitive job. Poll to completion, then read `primary_file_url` (also `video.file_url` for MP4 or `image.file_url` for stills)
|
|
2456
|
-
- Billing: metered
|
|
2530
|
+
- Response: standard primitive job. Poll to completion, then read `primary_file_url` (also `video.file_url` for MP4 or `image.file_url` for stills). `output.dedupe` carries the resolved effects, the effective zoom, the CRF and a human-readable `notes` list.
|
|
2531
|
+
- Billing: one ffmpeg pass, metered at real compute (much cheaper than a render). Free on a local serve box — and free ANYWHERE via `vidfarm dedupe --local`, which runs the identical filter graph on bundled ffmpeg.
|
|
2532
|
+
|
|
2533
|
+
### Presets
|
|
2534
|
+
|
|
2535
|
+
`preset` picks a calibrated transformation set. Default `standard`.
|
|
2536
|
+
|
|
2537
|
+
| preset | skew | zoom | rotate | speed | saturation | when |
|
|
2538
|
+
| --- | --- | --- | --- | --- | --- | --- |
|
|
2539
|
+
| `none` | — | — | — | — | — | no-op; re-encode only |
|
|
2540
|
+
| `light` | 1% | 2% | 0.75° | +1% | +2% | lightly reused footage, or tight framing you can't crop |
|
|
2541
|
+
| `standard` | 2% | 3% | 2° | +2% | +4% | **the default** — the house standard |
|
|
2542
|
+
| `strong` | 3.5% | 6% | 3° | +5% | +8% | an Nth re-post, or an account that already ran this clip |
|
|
2543
|
+
| `legacy` | — | 4% | 3° | +5% | +5% | the pre-ffmpeg composition-renderer defaults |
|
|
2457
2544
|
|
|
2458
|
-
|
|
2545
|
+
`light`/`standard`/`strong` also move contrast, brightness, hue and grain. Any individual knob in `effects` overrides the preset.
|
|
2546
|
+
|
|
2547
|
+
### Minting N distinct copies
|
|
2548
|
+
|
|
2549
|
+
`variant` (1-based) is what makes bulk posting work. Variant 1 is the preset as authored; later variants get deterministically jittered magnitudes and **alternating signs** (a sign flip moves a perceptual hash much further than a magnitude nudge), so N copies differ from the original *and from each other*. Reuse one `seed` across the batch, and post each variant to a different account/slot.
|
|
2550
|
+
|
|
2551
|
+
### Payload fields
|
|
2459
2552
|
|
|
2460
2553
|
- `source_media_url` (required, URL) — the image or video to transform
|
|
2461
2554
|
- `media_type` (`"image" | "video"`, optional) — auto-detected from URL extension if omitted (`.mp4/.mov/.webm/.m4v` → video, else image)
|
|
2462
|
-
- `
|
|
2463
|
-
|
|
2464
|
-
|
|
2465
|
-
|
|
2466
|
-
|
|
2467
|
-
- `
|
|
2468
|
-
- `
|
|
2469
|
-
- `
|
|
2470
|
-
- `
|
|
2471
|
-
- `
|
|
2472
|
-
- `
|
|
2473
|
-
- `
|
|
2474
|
-
- `
|
|
2475
|
-
- `
|
|
2476
|
-
- `
|
|
2477
|
-
- `
|
|
2478
|
-
- `
|
|
2555
|
+
- `preset` (`"none" | "light" | "standard" | "strong" | "legacy"`, default `"standard"`)
|
|
2556
|
+
- `engine` (`"ffmpeg" | "composition"`, default `"ffmpeg"`) — `ffmpeg` is a real pixel/timing transform on the source file (true shear, honest playback speed, grain, metadata strip) and is both cheaper and stronger. `composition` is the legacy HyperFrames-render path, kept only for callers that depend on its exact output.
|
|
2557
|
+
- `variant` (int ≥ 1, default `1`), `seed` (string, optional), `jitter` (bool, optional — defaults on for `variant > 1`)
|
|
2558
|
+
- `strip_metadata` (default `true`) — drop creation time / encoder / source handler. Several platforms compare that **before** they compare pixels.
|
|
2559
|
+
- `effects` (optional object). Every field optional; each one overrides the preset:
|
|
2560
|
+
- `zoom` — scale factor, centre-cropped back (`1.03` = 3% punch-in)
|
|
2561
|
+
- `skew` — horizontal shear as a **percent of frame width** (ffmpeg engine only)
|
|
2562
|
+
- `rotate` — degrees of 2D rotation
|
|
2563
|
+
- `tilt` — degrees of 3D X-axis tilt on the composition engine; folded into the shear budget on ffmpeg
|
|
2564
|
+
- `speed` — playback multiplier, video only; changes duration **and** pitch-preserved audio tempo
|
|
2565
|
+
- `saturation`, `contrast`, `brightness` — multipliers around `1`
|
|
2566
|
+
- `hue_rotate` — degrees
|
|
2567
|
+
- `noise` — film grain `0..100` (ffmpeg engine only). Cheap, invisible, moves a lot of hash.
|
|
2568
|
+
- `blur` — gaussian sigma in px. Usually `0` — blur is the one knob viewers notice.
|
|
2569
|
+
- `volume` — audio gain multiplier
|
|
2570
|
+
- `horizontal_flip` (default `false`) — the strongest single knob, but it visibly reverses on-screen text. Opt in deliberately.
|
|
2571
|
+
- `tint_color` (default `"#FF8C00"`), `tint_opacity` (default `0.08`) — flat color wash; set opacity `0` to skip
|
|
2572
|
+
- `width` / `height` — **optional on the ffmpeg engine**; omit to keep the source's own frame size (forcing 1080×1920 onto a 16:9 source would squash it). The composition engine falls back to 1080×1920.
|
|
2573
|
+
- `crf` — base x264 quality; jittered ±1 per variant so the coded bitstream differs too
|
|
2479
2574
|
- `output_format` (`"png" | "jpeg" | "webp"`, default `"png"`) — image mode only; video mode always outputs MP4
|
|
2575
|
+
- Composition-engine only: `duration_ms` / `fallback_duration_ms` (default `5000`), `object_fit`, `object_position`, `background_color`, `muted`
|
|
2576
|
+
|
|
2577
|
+
**A rotate forces a bigger crop than you asked for.** Black corners have to go somewhere, so the primitive raises `zoom` to the smallest value that covers the rotation and says so in `output.dedupe.notes`. On a tall 1080×1920 frame a 2° rotate costs ~6.7% of the frame. If framing matters more than fingerprint distance, pass `effects.rotate: 0`.
|
|
2480
2578
|
|
|
2481
|
-
Video example
|
|
2579
|
+
Video example — three copies of one render, one per account:
|
|
2580
|
+
|
|
2581
|
+
```bash
|
|
2582
|
+
for V in 1 2 3; do
|
|
2583
|
+
curl -X POST "$VIDFARM_BASE/api/v1/primitives/media/dedupe" \
|
|
2584
|
+
-H "vidfarm-api-key: $VIDFARM_API_KEY" \
|
|
2585
|
+
-H "content-type: application/json" \
|
|
2586
|
+
-d "{
|
|
2587
|
+
\"tracer\": \"dedupe-launch-reel-v$V\",
|
|
2588
|
+
\"payload\": {
|
|
2589
|
+
\"source_media_url\": \"https://cdn.example.com/reel.mp4\",
|
|
2590
|
+
\"media_type\": \"video\",
|
|
2591
|
+
\"preset\": \"standard\",
|
|
2592
|
+
\"variant\": $V,
|
|
2593
|
+
\"seed\": \"launch-reel\"
|
|
2594
|
+
}
|
|
2595
|
+
}"
|
|
2596
|
+
done
|
|
2597
|
+
```
|
|
2598
|
+
|
|
2599
|
+
Custom knobs (keep the framing, lean on color and timing instead):
|
|
2482
2600
|
|
|
2483
2601
|
```bash
|
|
2484
2602
|
curl -X POST "$VIDFARM_BASE/api/v1/primitives/media/dedupe" \
|
|
@@ -2489,15 +2607,7 @@ curl -X POST "$VIDFARM_BASE/api/v1/primitives/media/dedupe" \
|
|
|
2489
2607
|
"payload": {
|
|
2490
2608
|
"source_media_url": "https://cdn.example.com/reel.mp4",
|
|
2491
2609
|
"media_type": "video",
|
|
2492
|
-
"
|
|
2493
|
-
"effects": {
|
|
2494
|
-
"zoom": 1.05,
|
|
2495
|
-
"tilt": 2,
|
|
2496
|
-
"rotate": -2,
|
|
2497
|
-
"speed": 1.03,
|
|
2498
|
-
"hue_rotate": 4,
|
|
2499
|
-
"horizontal_flip": true
|
|
2500
|
-
},
|
|
2610
|
+
"effects": { "rotate": 0, "skew": 1.5, "zoom": 1.02, "speed": 1.03, "hue_rotate": 6, "noise": 2 },
|
|
2501
2611
|
"tint_color": "#00A3FF",
|
|
2502
2612
|
"tint_opacity": 0.06
|
|
2503
2613
|
}
|
|
@@ -2515,12 +2625,18 @@ curl -X POST "$VIDFARM_BASE/api/v1/primitives/media/dedupe" \
|
|
|
2515
2625
|
"payload": {
|
|
2516
2626
|
"source_media_url": "https://cdn.example.com/photo.jpg",
|
|
2517
2627
|
"media_type": "image",
|
|
2518
|
-
"
|
|
2628
|
+
"preset": "light",
|
|
2519
2629
|
"output_format": "webp"
|
|
2520
2630
|
}
|
|
2521
2631
|
}'
|
|
2522
2632
|
```
|
|
2523
2633
|
|
|
2634
|
+
Local equivalent (free, offline, identical transforms — prefer this):
|
|
2635
|
+
|
|
2636
|
+
```bash
|
|
2637
|
+
vidfarm dedupe ./out/final.mp4 --variants 3 --out-dir ./out/posts
|
|
2638
|
+
```
|
|
2639
|
+
|
|
2524
2640
|
## Primitive: music (text → music)
|
|
2525
2641
|
|
|
2526
2642
|
Generate music (instrumental, songs with lyrics, background beds, jingles, scores) via ElevenLabs. **Default this on freely — music is a core primitive.** `use_wallet_credits` defaults **true**: it runs on vidfarm's platform ElevenLabs key and bills the customer's wallet. Recommend keeping it on; set it false only to save wallet credits or to use the customer's OWN saved ElevenLabs key.
|
|
@@ -2685,9 +2801,10 @@ Use this when a coding agent is doing the work locally or the user wants a repro
|
|
|
2685
2801
|
3. Read `./work/.harness/agent-guide.md` and `./work/.harness/context.json` before editing.
|
|
2686
2802
|
4. Make deterministic edits to `composition.html` and optionally `composition.json`.
|
|
2687
2803
|
5. Validate with `vidfarm lint` or `vidfarm stills` when useful. **Always look at `vidfarm stills ./work --at 0`** — that frame becomes the thumbnail, so it must not be black, empty, or mid-fade.
|
|
2688
|
-
6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, and flags a blank/fading first frame (the thumbnail). Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
|
|
2804
|
+
6. **QA before you render: `vidfarm qa ./work`.** Free, instant, devcli-only. It blocklists HTML slop (CTA buttons, benefit chip rows, a lone pill around a static stat/label, frosted cards, gradient text, web-page classes/fonts), checks the caption font regime + safe zone, and flags a blank/fading first frame (the thumbnail). Feedback only — exit 0 even on findings, never automatic — but it catches the #1 tell of an agent-made video, so run it on every production. Fix what's real, ignore what's a deliberate style call, then render.
|
|
2689
2805
|
7. Render with `vidfarm render <forkId> --dir ./work --wait`.
|
|
2690
|
-
8.
|
|
2806
|
+
8. **Ask about deduplication before you approve** — "is this going out more than once (several accounts, another platform, a re-post later)?" If yes, run `vidfarm dedupe ./final.mp4 [--variants N]` on the **exported** MP4 (free, local ffmpeg, no re-render) and approve each variant separately. Asking here rather than after publication is what avoids paying for a second render. See `references/core-workflows.md` → *Deduplicate before you publish*.
|
|
2807
|
+
9. Approve the finished MP4 with `vidfarm approve --video <url|./final.mp4> --caption "..."`. This prints the shareable `share_url`.
|
|
2691
2808
|
|
|
2692
2809
|
**Approving a locally rendered file → cloud preview link.** Approve takes media by **URL**, not bytes, and an approved post is **permanent** — so the local MP4 must land in **durable My Files**, not the 30-day temp store (a temp video would 404 the share page after 30 days). The devcli presigns, PUTs the bytes **direct to S3**, finalizes, then approves with that durable URL — so `vidfarm approve --video ./final.mp4` handles files up to **200 MB**, bypasses the ~6 MB Lambda request-body limit, and the share link never breaks. By raw REST: `POST /api/v1/user/me/attachments/presign` → PUT to the returned S3 URL → `POST /api/v1/user/me/attachments` (finalize) → pass the returned `viewUrl` in the approve `media` array. Do not multipart-POST a big file to `.../attachments/upload` against the cloud host (Lambda-bound, ~6 MB cap). Add `vidfarm approve --temp` only when you want a disposable 30-day preview.
|
|
2693
2810
|
|
|
@@ -2747,6 +2864,18 @@ done
|
|
|
2747
2864
|
|
|
2748
2865
|
`vidfarm qa` still exits 0 on findings — the `jq -e` line is **your** gate, in your script, made explicit. Keep it that way; a hard gate inside the tool would quietly train the loop to stop trying anything unusual.
|
|
2749
2866
|
|
|
2867
|
+
### 5b. Deduplicate the renders (ask first)
|
|
2868
|
+
|
|
2869
|
+
A bulk run exists to put volume across accounts and platforms — which is exactly the shape a platform's duplicate-content detector flags. **Ask the director up front: "deduplicated copies for posting, and how many slots?"** Ask before the loop, not after: dedupe is a post-render ffmpeg pass, so answering early keeps it at *render once → dedupe N* rather than a second render per slot.
|
|
2870
|
+
|
|
2871
|
+
```bash
|
|
2872
|
+
for MP4 in renders/*.mp4; do
|
|
2873
|
+
vidfarm dedupe "$MP4" --variants "$SLOTS" --seed "$(basename "$MP4" .mp4)" --out-dir ./posts
|
|
2874
|
+
done
|
|
2875
|
+
```
|
|
2876
|
+
|
|
2877
|
+
Free, offline, no wallet. Variant 1 is the `standard` preset as authored (skew 2%, zoom 3%, rotate 2°, speed +2%, saturation +4%); later variants get jittered magnitudes and flipped signs, so they differ from the original **and from each other**. One variant per account — two accounts posting the same variant defeats the point. Reuse one `--seed` per source so the batch is reproducible.
|
|
2878
|
+
|
|
2750
2879
|
### 6. Answer the review items — don't skip this
|
|
2751
2880
|
|
|
2752
2881
|
The regime's `- [ ]` checklist comes back on every run because the CLI *can't* settle it. Machine checks catch a 13-word hook or a black first frame; only you can answer "is this variant genuinely different from its siblings?" or "can the viewer guess the withheld answer?" **Report both halves honestly**: what the machine checked, and what you judged. A batch report claiming a clean pass on the judgment half is worse than no report.
|
package/SKILL.md
CHANGED
|
@@ -103,6 +103,7 @@ For composition *authoring* craft (motion, keyframes, scene design), Vidfarm shi
|
|
|
103
103
|
4b. **"Download this video/audio from <a website URL>"** → Vidfarm fetches it for you on a **paid plan**: `POST /api/v1/primitives/videos/download` (or `/audio/download`); devcli `vidfarm download-video <url>` / `vidfarm download-audio <url>`. Works on YouTube, TikTok, Instagram, X, and other supported posts; returns a durable Vidfarm file (photo/carousel posts → an ordered slideshow). **Free plan gets a 402 — don't call it. Tell the user (or, with browser automation, do it yourself) to open the URL in Chrome and download it from the page, then `vidfarm put-file ./the-file.mp4` to bring it in for $0.** Never answer "I can't download that." Details: `references/assets-and-sourcing.md`.
|
|
104
104
|
4b-ii. **"Tweet to TikTok" / "Reddit to TikTok" / "make a video out of this thread / subreddit / account"** → `vidfarm recycle <source>` (devcli), or `POST /api/v1/primitives/social/recycle` with `{ tracer, payload: { source_url, max_records } }`. It **decomposes** a Reddit thread (post + comments), a subreddit (its threads), an X thread (tweet + replies), or an X profile (their posts) into raw JSON — text, comment tree, media URLs, author avatar, engagement stats — and hands it back **unranked and unsummarized** so you pick the hook, the punchline comment, the stat. Auto-paginates via `nextCursor`; `max_records` is the spend ceiling. Brokers the reddit-lead-gen / x-lead-gen OfficeX apps and waits out their async job for you. **Paid plan.** Details: `references/assets-and-sourcing.md`.
|
|
105
105
|
4c. **"Create an avatar"** (spokesperson / presenter / host / UGC creator / talking head) → always a **talking-head VIDEO with spoken audio**, generated on an exact-key-color **greenscreen** plate and keyed off it in the same job → a **transparent presenter** you layer over any background. `vidfarm avatar "<who they are>" --say "<their line>" [--ref headshot.png]` / `POST /api/v1/primitives/videos/create-avatar`. Details: `references/primitives.md` → "Primitive: talking_avatar".
|
|
106
|
+
4d. **"Post this again / to several accounts / on another platform"** → **deduplication**. Platforms fingerprint uploads; the second copy of the same render gets suppressed as duplicate/reused content. `vidfarm dedupe <mp4> [--variants N]` runs on the **exported file** — free, local ffmpeg, **no re-render** — nudging skew/zoom/rotate/speed/saturation/grain a couple of percent and stripping container metadata, invisibly to a viewer. `--variants N` mints N copies distinct from the original *and from each other*, one per account/slot. Cloud twin: `POST /api/v1/primitives/media/dedupe`. Details: `references/core-workflows.md` → *Deduplicate before you publish*.
|
|
106
107
|
5. "Script / batch / render loop" → `references/automation-and-local-dev.md`
|
|
107
108
|
6. "I need TTS / music / captions / background removal" → `references/primitives.md`
|
|
108
109
|
7. **"Update / upgrade vidfarm"** (or anything that smells like a stale install — a missing command, a 404 on a documented route, a version mismatch) → fetch <https://vidfarm.cc/update.md> and follow it. Update the skill pack and the devcli **together**; updating one alone is the usual cause of "the skill says to do X but it fails."
|
|
@@ -113,10 +114,11 @@ For composition *authoring* craft (motion, keyframes, scene design), Vidfarm shi
|
|
|
113
114
|
- Never build composition HTML by string concatenation — parse, edit, re-serialize the DOM.
|
|
114
115
|
- Render only through `POST /api/v1/compositions/:forkId/render`; never call the renderer directly.
|
|
115
116
|
- Submissions are **not idempotent** — every render/primitive POST charges again. Check status before retrying.
|
|
116
|
-
- **No HTML slop.** Compositions are HTML, but a video is not a web page:
|
|
117
|
+
- **No HTML slop.** Compositions are HTML, but a video is not a web page. The test: *could you have made this element with the tools inside TikTok's own editor* (font, color, stroke, shadow, tight text box, rotation, animation presets, stickers, emoji, drawn marks)? If you reached past that — a padded capsule, a border, a gradient fill, a blur panel, a card — cut it. So: no CTA "buttons", no benefit chip/badge rows, **no single pill around a static stat or label** (`( 10 hrs / week )`, `( STEP 2 )` — the only legitimate capsule is the active-word `spotlight`/`karaoke` highlight), no frosted/bordered cards holding a headline + URL, no gradient text, feature grids, or bullet lists. Nothing in a video is clickable. Say it as timed text on the footage; emphasize with size, weight, ALL-CAPS, an accent color, or a drawn circle. Arrows, scribble/underline marks, italics, color pops, emoji, cut-out stickers, and mock social UI are fine.
|
|
117
118
|
- **Structure before polish — the four charges, written before the timeline.** 🪝 **Hook**: the first line is a complete clause naming a *situation*, not a label; on screen at `start:0` (chunk 1 is read before any audio). 🔄 **Loop**: one open question by 0:10, stated on screen, **closing inside this video**, with an answer the viewer can't guess. 😍 **Payoff**: shown, not summarized, before the final beat. 🎣 **Bait**: one ask, final beat + post caption. Banned openers: throat-clearing, a logo, a title card, a fade from black. Full harness: `references/hooks-and-virality.md`; checkable form: `vidfarm regime show hooks`.
|
|
118
119
|
- **The first frame is the thumbnail.** Frame 0 is one frame of ~30 in the first second, but it's the poster every feed and share link freezes on — so it's seen by more people than the video is. Never open on black, an empty frame, or a fade-up: a real visual at `start:0`, the hook words already up, and no *entrance* transition on the first clip (junction transitions between later clips are fine). Check it with `vidfarm stills <dir> --at 0`.
|
|
119
120
|
- **On devcli, `vidfarm qa <dir>` before every render.** Free, instant, local-only blocklist for the slop above + the first frame + the font regime. Feedback, not a gate (exits 0, never automatic). No REST/web equivalent.
|
|
121
|
+
- **Ask about deduplication before you publish or bulk-produce.** "Is this going out more than once — several accounts, another platform, a re-post later? How many copies?" Ask *before* the render or the batch, not after: dedupe is a post-render ffmpeg pass, so answering early keeps it at **render once → dedupe N** instead of paying for a second render per slot. Then post each variant to a **different** account — two accounts posting the same variant defeats the point.
|
|
120
122
|
- **Ask one-time vs bulk before you build.** Volume = **scripting mode**: a pinned base fork, a loop varying ONE thing, and a **`QA_REGIME.md`** — the director's own written standard, because a fifty-video loop has no human watching every frame. `vidfarm regime init short-form --out ./work/QA_REGIME.md` (bases: `short-form`, `hooks`, `ugc-testimonial`, `explainer`, `product-demo`), then `vidfarm qa ./work --regime <name|path>` — stackable, any user file valid, auto-discovered from the work dir. Its `checks:` are machine-settled; its `- [ ]` items come back for **you** to answer honestly.
|
|
121
123
|
- **Caption regime is mandatory**: an imported display font (Montserrat default / TikTok Sans), weight 700–900, ~36–64px on a 1080-wide frame, inside the 8%–85% safe zone, and exactly one of four backgrounds — `outline`, `plain`, an active-word `spotlight`/`karaoke` pill, or a tight-hugging `highlight-solid` band (radius ≤8px, no border/shadow/gradient/blur).
|
|
122
124
|
- In the web editor, CSS/declarative motion only (JS animation adapters are stripped on save); locally via `vidfarm serve` the full JS adapters work.
|