@cueframe/skills 0.1.0 → 0.1.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +9 -104
- package/{skills → dist/skills}/add-music-bed/SKILL.md +12 -7
- package/{skills → dist/skills}/brand-reel/SKILL.md +19 -15
- package/{skills → dist/skills}/clip-a-talking-head/SKILL.md +15 -7
- package/dist/skills/composing-video/SKILL.md +447 -0
- package/{skills → dist/skills}/cueframe-brand-demo/SKILL.md +9 -2
- package/{skills → dist/skills}/cueframe-cli/SKILL.md +54 -49
- package/{skills → dist/skills}/cueframe-component-authoring/SKILL.md +14 -6
- package/{skills → dist/skills}/cueframe-compose-loop/SKILL.md +20 -6
- package/dist/skills/cueframe-compose-loop/references/builtin-insertion.md +49 -0
- package/dist/skills/cueframe-compose-loop/references/builtin-requests.json +111 -0
- package/dist/skills/cueframe-connect/SKILL.md +53 -0
- package/{skills → dist/skills}/cueframe-product-video/SKILL.md +15 -7
- package/{skills → dist/skills}/cueframe-scene-shot/SKILL.md +8 -0
- package/dist/skills/cueframe-storyboard/SKILL.md +76 -0
- package/{skills → dist/skills}/every-format-from-one-edit/SKILL.md +14 -8
- package/{skills → dist/skills}/extracting-brand-kits/SKILL.md +8 -0
- package/{skills → dist/skills}/launch-video/SKILL.md +17 -11
- package/{skills → dist/skills}/make-a-social-reel/SKILL.md +19 -15
- package/{skills → dist/skills}/rebrand-a-video/SKILL.md +18 -16
- package/dist/skills/video-craft-standards/SKILL.md +121 -0
- package/dist/skills/video-craft-standards/agents/openai.yaml +6 -0
- package/package.json +9 -28
- package/skills-dir.d.ts +1 -0
- package/skills-dir.js +2 -1
- package/.agents/plugins/marketplace.json +0 -12
- package/.claude-plugin/marketplace.json +0 -6
- package/.claude-plugin/plugin.json +0 -15
- package/.codex-plugin/plugin.json +0 -30
- package/.cursor-plugin/plugin.json +0 -1
- package/.mcp.json +0 -1
- package/AGENTS.md +0 -20
- package/assets/logo-400.png +0 -0
- package/gemini-extension.json +0 -1
- package/glama.json +0 -1
- package/hooks/hooks.json +0 -7
- package/hooks/session-inject.md +0 -15
- package/hooks/session-start.sh +0 -6
- package/llms-install.md +0 -47
- package/mcp.json +0 -1
- package/plugin.json +0 -46
- package/rules/cueframe.mdc +0 -19
- package/skills/composing-video/SKILL.md +0 -702
- package/skills/cueframe-connect/SKILL.md +0 -45
- package/skills/cueframe-storyboard/SKILL.md +0 -104
- package/skills/video-craft-standards/SKILL.md +0 -128
- package/skills/video-craft-standards/agents/openai.yaml +0 -6
- package/skills/video-craft-standards/assets/icon.svg +0 -16
- package/skills.sh.json +0 -1
- /package/{skills → dist/skills}/add-music-bed/agents/openai.yaml +0 -0
- /package/{assets → dist/skills/add-music-bed/assets}/icon.svg +0 -0
- /package/{skills → dist/skills}/brand-reel/agents/openai.yaml +0 -0
- /package/{skills/add-music-bed → dist/skills/brand-reel}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/clip-a-talking-head/agents/openai.yaml +0 -0
- /package/{skills/brand-reel → dist/skills/clip-a-talking-head}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/composing-video/agents/openai.yaml +0 -0
- /package/{skills/clip-a-talking-head → dist/skills/composing-video}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-brand-demo/agents/openai.yaml +0 -0
- /package/{skills/composing-video → dist/skills/cueframe-brand-demo}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-cli/agents/openai.yaml +0 -0
- /package/{skills/cueframe-brand-demo → dist/skills/cueframe-cli}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-component-authoring/agents/openai.yaml +0 -0
- /package/{skills/cueframe-cli → dist/skills/cueframe-component-authoring}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-compose-loop/agents/openai.yaml +0 -0
- /package/{skills/cueframe-component-authoring → dist/skills/cueframe-compose-loop}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-compose-loop/references/preview-workflow.md +0 -0
- /package/{skills → dist/skills}/cueframe-connect/agents/openai.yaml +0 -0
- /package/{skills/cueframe-compose-loop → dist/skills/cueframe-connect}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-product-video/agents/openai.yaml +0 -0
- /package/{skills/cueframe-connect → dist/skills/cueframe-product-video}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-scene-shot/agents/openai.yaml +0 -0
- /package/{skills/cueframe-product-video → dist/skills/cueframe-scene-shot}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-storyboard/agents/openai.yaml +0 -0
- /package/{skills/cueframe-scene-shot → dist/skills/cueframe-storyboard}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/every-format-from-one-edit/agents/openai.yaml +0 -0
- /package/{skills/cueframe-storyboard → dist/skills/every-format-from-one-edit}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/extracting-brand-kits/agents/openai.yaml +0 -0
- /package/{skills/every-format-from-one-edit → dist/skills/extracting-brand-kits}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/launch-video/agents/openai.yaml +0 -0
- /package/{skills/extracting-brand-kits → dist/skills/launch-video}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/make-a-social-reel/agents/openai.yaml +0 -0
- /package/{skills/launch-video → dist/skills/make-a-social-reel}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/rebrand-a-video/agents/openai.yaml +0 -0
- /package/{skills/make-a-social-reel → dist/skills/rebrand-a-video}/assets/icon.svg +0 -0
- /package/{skills/rebrand-a-video → dist/skills/video-craft-standards}/assets/icon.svg +0 -0
|
@@ -1,702 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: composing-video
|
|
3
|
-
description: Use when an agent is asked to make ANY video with CueFrame — a product launch, announcement, promo, teaser, social clip, reel, short, explainer, animated explainer, concept video, product/dev-tool demo, walkthrough, or founder/talking-head piece. Triggers on "make a video", "launch video", "promo", "teaser", "reel", "short", "explainer", "animated explainer", "demo video", "talking-head", or turning long-form footage into clips.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Composing Video: craft for CueFrame
|
|
7
|
-
|
|
8
|
-
For preview selection and small revisions, use `cueframe-compose-loop`'s
|
|
9
|
-
[preview workflow](../cueframe-compose-loop/references/preview-workflow.md). Retain the approved
|
|
10
|
-
brief/rubric for a local correction; check the affected behavior and relevant regressions instead
|
|
11
|
-
of restarting the new-film method below. A diagnostic or preview-only request does not require
|
|
12
|
-
a final export or a whole-film judge pass. Explicit user scope takes precedence over craft defaults.
|
|
13
|
-
|
|
14
|
-
> **The bar:** the `video-craft-standards` skill (GET /v1/skills/video-craft-standards) defines DONE for every artifact class — per-class MUST checklists and kill criteria. Read it before composing and score against it before rendering. When this skill and that document disagree, the standards win.
|
|
15
|
-
|
|
16
|
-
## The one idea
|
|
17
|
-
|
|
18
|
-
CueFrame composes **real footage** into finished video, and that is its edge: it detects faces and the active speaker and reframes/punches in on them; it renders titles behind the speaker's head; it times captions to the real transcript; it grades color and renders 4K. Real footage is the strongest grounding you can give a video and CueFrame's home turf — so it is the thing to reach for FIRST and usually your highest-leverage move. But the skill's job is **any good video**, footage or not: an animated explainer, an animated concept video, a data piece. These are first-class, not a lower tier.
|
|
19
|
-
|
|
20
|
-
What separates good from slop is not "footage vs. no footage" — it is **grounded specificity vs. the generic default**. The enemy is the interchangeable filler that could belong to any video: gradient/aurora backgrounds under kinetic text, stock "AI b-roll," faked UI, unmotivated motion. The antidote is a **controlling idea** plus **grounded, specific content** — real footage, OR real UI/screen captures, OR real data, OR a *purposeful* graphic/animation that DEPICTS THIS content specifically (an authored diagram of the actual mechanism, a data animation of the real numbers) — a data/mechanism visual, NOT animated copy. Animated/kinetic TEXT by itself is not grounding. A well-crafted animated explainer built on diagrams of the real mechanism is grounded and good; it is not slop.
|
|
21
|
-
|
|
22
|
-
**Generic default ⇒ slop. Grounded specificity ⇒ good.** Real footage is the strongest grounding and CueFrame's home turf — reach for it first — but purposeful, specific graphics count too. This is a mechanical prediction, not a taste call (see *Craft principles*). Your single most expensive mistake is walking past real footage that exists to build generic gradient slides instead — which is precisely what happened in the launch-video case below.
|
|
23
|
-
|
|
24
|
-
**The loophole to nail shut:** "purposeful graphics" means *specific to this content* — the real mechanism, the real data, a diagram of the actual thing. Animated/kinetic TEXT by itself is not that, no matter how deliberate — it is not grounding. A generic gradient or aurora is NEVER "purposeful"; it is always the slop default. The reframe broadens what counts as good; it does not wave decorative filler through.
|
|
25
|
-
|
|
26
|
-
**The boundary.** This is craft guidance you can override, not a gate. You own the **vision** (the one idea for this specific video) and you own **data correctness** (if you supply a wrong number, CueFrame renders it faithfully and the mistake is yours). CueFrame owns being faithful, legible, and teaching you the craft below. Reason from these principles; break a rule knowingly when your vision demands it. (Skip this whole flow only for a pure re-render of an already-approved composition, or a single trivial still — don't ceremony a one-liner.) Planning a new video from scratch? Do **`cueframe-storyboard`** first to get the beat flow, then apply this craft to each beat.
|
|
27
|
-
|
|
28
|
-
---
|
|
29
|
-
|
|
30
|
-
## Related skills & positioning
|
|
31
|
-
|
|
32
|
-
This skill is the **craft / quality layer** for an agent driving CueFrame over the **MCP tool surface** (`preview_frame` / `apply_composition` / `score_composition` / `create_render`) — *not* the `cueframe` CLI binary. It owns the decision layer: vision → rubric → grounding → iterate → verify. When a mechanic below is owned by a sibling skill, defer to that skill for the deep how-to:
|
|
33
|
-
|
|
34
|
-
- **`cueframe-storyboard`** — the intake/structure step BEFORE craft: interview the user for the brief and lay out the hook→body→cta beat flow. Start here for a new video from scratch; this skill then makes each beat grounded and good.
|
|
35
|
-
- **`cueframe-cli`** — the `cueframe` binary + wire mechanics (auth, projects, raw request/response shapes).
|
|
36
|
-
- **`cueframe-compose-loop`** — the full author→render→verify→fix **convergence loop**; it owns the loop budget and ship bar (this skill's step 6 mirrors it: 3–5 passes, ship at composite ≥ 7.5).
|
|
37
|
-
- **`cueframe-product-video`** — the click-points → reframe-segment **auto-zoom tiling** + picture-in-picture recipe for screen demos.
|
|
38
|
-
- **`cueframe-brand-demo`** — brand-token vocabulary + auto intro/outro/logo for branded launches.
|
|
39
|
-
- **`cueframe-component-authoring`** — the full fork → preview → push loop for custom components.
|
|
40
|
-
- **`cueframe-scene-shot`** — the 3D product shot authored as JSON (a tilted page screenshot, a lit laptop/phone showing a real app screenshot, camera move, focus). Reach for it *before* hand-writing a WebGL component for a device mockup.
|
|
41
|
-
|
|
42
|
-
Loading this skill and a sibling together: this one governs *what makes the video good*; the sibling governs *how that mechanic works*.
|
|
43
|
-
|
|
44
|
-
---
|
|
45
|
-
|
|
46
|
-
## If you read nothing else
|
|
47
|
-
|
|
48
|
-
Six non-negotiables. Everything below is the reasoning behind them.
|
|
49
|
-
|
|
50
|
-
1. **HUNT for real footage FIRST** — it's the strongest grounding and CueFrame's edge — before building without it; that failure was never looking.
|
|
51
|
-
2. **Write a harsh, specific rubric BEFORE you render a single frame** — including the three mandatory kill criteria (K1–K3), meaning preserved.
|
|
52
|
-
3. **Grounded, specific content must DOMINATE the runtime** — real footage, real UI/data, or purposeful content-specific graphics — never generic/decorative filler. (For footage-based formats, target real footage ≥60%.)
|
|
53
|
-
4. **Describe every rendered still in words before you score it** — if you can't say what's in the frame, you didn't look.
|
|
54
|
-
5. **Get an independent grade you can't game:** `score_composition` samples its own frames — you cannot cherry-pick flattering timestamps past it.
|
|
55
|
-
6. **Budget 3–5 verify passes, never set your ship bar below 7.5/10, then stop.**
|
|
56
|
-
|
|
57
|
-
---
|
|
58
|
-
|
|
59
|
-
## The method — run these steps in order, every time
|
|
60
|
-
|
|
61
|
-
The order is the forcing function. Steps 2, 3, and 5 are where agents cut corners and ship slop, so they are not optional.
|
|
62
|
-
|
|
63
|
-
### 1 — Decide the VISION, then the FORMAT
|
|
64
|
-
Write the controlling idea as one sentence with a value and a cause — "X wins *because* Y" (McKee, *Story*: a story is a charged value-shift, not a topic). Not "product launch" — a claim: "Two founders tell you, to your face, why their product kills the pipeline step you dread — and you watch it be faster." Then pick the format floor: product-launch / social-clip / explainer / demo / talking-head (playbooks below). The vision is the filter for what belongs; anything that doesn't express it gets cut.
|
|
65
|
-
|
|
66
|
-
### 2 — Write a HARSH, SPECIFIC rubric — BEFORE you render anything
|
|
67
|
-
Authored against the vision, before pixels exist, so it can't be rationalized after the fact. Use the template below: must-haves anchored to named frame timestamps, measurable thresholds, and kill criteria that fail the render outright.
|
|
68
|
-
|
|
69
|
-
**Hard gate:** do not call `preview_frame` until the rubric text exists in your context. No rubric → no pixels.
|
|
70
|
-
|
|
71
|
-
**These three kill criteria are MANDATORY — preserve the MEANING of K1–K3 in every rubric, adapting the phrasing to your video. You may add more; you may not remove or weaken these:**
|
|
72
|
-
- **K1.** Any sampled frame is a gradient/solid background with only text and NO grounded subject — no real footage, no real UI/data, no purposeful content-specific graphic, just decorative filler. Animated/kinetic TEXT over any background does NOT count as a grounded subject — only a visual of the real mechanism or real data does.
|
|
73
|
-
- **K2.** Any on-screen number you cannot trace to a written source, or any faked/mocked UI standing in for a real capture.
|
|
74
|
-
- **K3.** Generic/decorative filler dominates the runtime — grounded, specific content (real footage, real UI/data, or purposeful content-specific graphics) does not.
|
|
75
|
-
|
|
76
|
-
### 3 — Ground the video (hunt for footage first, then a grounding ladder — never straight to generic filler)
|
|
77
|
-
|
|
78
|
-
Real footage is the strongest grounding and CueFrame's edge, so hunt for it FIRST — but grounding is the goal, and purposeful, specific graphics are a legitimate way to reach it.
|
|
79
|
-
|
|
80
|
-
**3a. Hunt for footage FIRST, before building without it.** Before you build without footage, actively look, in order:
|
|
81
|
-
1. Did the user attach or name footage?
|
|
82
|
-
2. Does the subject (company/person/product) already have a launch video, founder clips, demo recordings, or a conference talk you can `import_media` from a public https URL?
|
|
83
|
-
3. Ask the user directly: *"Do you have a founder clip, screen capture, or event recording? CueFrame is built to compose real footage — it's the difference between a launch film and a slideshow."*
|
|
84
|
-
|
|
85
|
-
In the launch-video case, step 2 alone would have surfaced the founders-to-camera launch video and changed everything. Footage is the highest-leverage grounding; don't walk past it.
|
|
86
|
-
|
|
87
|
-
**3b. The stop fires only when the plan is GENERIC filler.** If your plan has no grounded content — no footage, no real UI/data, no purposeful content-specific graphic, just a gradient with kinetic text — STOP and say so, verbatim, before rendering anything:
|
|
88
|
-
> *"This has no grounded content — it's a gradient-and-text motion piece, the generic default, not a video composed from the real thing. I recommend we ground it in [a 30-second founder to-camera clip / a real screen capture / an authored diagram of the actual mechanism]. Want to proceed as generic motion-graphics, or ground it first?"*
|
|
89
|
-
|
|
90
|
-
Proceed only on an explicit "proceed." A *purposeful, specific* animation — an authored diagram of the real mechanism, a data animation of the real numbers — does NOT trip this stop; it is a legitimate grounded choice, not a consolation prize. The stop exists only for decorative, could-be-any-video filler. Silent generic slop is the failure; a stated, chosen tradeoff is fine.
|
|
91
|
-
|
|
92
|
-
**3c. The grounding ladder** (strongest first; every rung is legitimate except the last — degrade the *promise*, never the craft):
|
|
93
|
-
1. **Real footage** — `import_media` a founder clip, event recording, or demo. The strongest grounding and CueFrame's home turf.
|
|
94
|
-
2. **Real screen capture** — `import_media` (CueFrame cannot screen-record; that needs an external screen-capture client — e.g. a headed Playwright walkthrough recorded to MP4 — or the user's own recording).
|
|
95
|
-
3. **Real product screenshots / photos** as held stills with **motivated** motion.
|
|
96
|
-
4. **Crafted graphics OF THE ACTUAL THING** — an authored `{kind:'card', html, tokens}` clip diagramming the real mechanism (process/algorithm/data viz), or a data graphic of the real numbers specific to this content, shelled with `applyMotionPreset`. A **first-class grounded path**, not a fallback — a good explainer lives here. The grounding is that it depicts THIS thing — the real mechanism, the real data — not a decorative loop; animated/kinetic TEXT by itself does not qualify.
|
|
97
|
-
5. **Generic `generate_media` texture** as a *supporting* backdrop only (a stage, a texture) — never a fake-cinematic hero, never a re-created product UI (it hallucinates specifics).
|
|
98
|
-
6. **Generic gradient-with-text** is the **last** resort and is exactly what triggers the 3b stop.
|
|
99
|
-
|
|
100
|
-
Do not fake the missing footage, and do not dress up generic filler as grounding.
|
|
101
|
-
|
|
102
|
-
### 4 — Author against `get_media_context`
|
|
103
|
-
When you have imported media, pull the detected **faces** (subject boxes per frame, with `faceId`s) and **transcript** before composing. Align reframe/punch-in to the faces, caption words to the transcript timing, and cut points to the speech. Discover primitives with `list_catalog` before hand-placing hex and coordinates — a built-in almost always exists.
|
|
104
|
-
|
|
105
|
-
If `faces.status` is `not_detected` and the edit uses `active-speaker`/`face` framing, call `prepare_media` with `kind:"subjectTrack"` and the clip's exact SOURCE trim window, then poll `get_media_facts` with that same kind/window until `exact.state:"ready"`. Do the same with `kind:"matte"` before sighting a `zPlane:"behind-subject"` graphic. `preview_frame` and `score_composition` fail loud when required perception is unresolved; never treat a flattened title or a static wide crop as acceptable preview evidence.
|
|
106
|
-
|
|
107
|
-
### 5 — Iterate on REAL PIXELS (and prove you looked)
|
|
108
|
-
`preview_frame` PNG stills at your rubric's named timestamps — **always including t=0 and hook-end** — and **look at them**. It queues durable capture and returns a `jobId`; call `wait_job(kind:"preview_frame")` until the evidence is terminal.
|
|
109
|
-
|
|
110
|
-
For every still, before you score it, write one literal line of what is visible: *Is there a grounded subject (real footage, real UI/data, or a purposeful content-specific graphic)? Real UI, or a styled div? Is the background mostly empty decoration? Is the text legible over what's behind it? Does any on-screen number match its written source?* If you can't describe the pixels, you did not look — render again. Score only from these descriptions, never from what you intended to author. Grade the first frame hardest: if t=0 is a bare gradient with the product name, K1 trips and you're done.
|
|
111
|
-
|
|
112
|
-
`validate_composition` only checks shape (membership, params, track-family, reframe coverage) — a pass there means nothing about quality. Judge the pixels, not the JSON.
|
|
113
|
-
|
|
114
|
-
### 6 — Get an independent second opinion, then stop
|
|
115
|
-
You grading your own render is the weak link. Call `score_composition` with your VISION line as `editorialIntent` and brand facts as `brandContext`, then call `wait_job(kind:"verify")`. It renders and grades **four fixed server-side criteria** — editorial, spatial, brand, caption (0–10 each) — plus a weighted composite and a worst-first critique, **on frames it samples itself.** Your `editorialIntent` steers *what "good" means for this video*; it does **not** replace the criteria, and you cannot feed it your own rubric or pick its frames — that un-gameable sampling *is* the independence. Treat a low score on a frame you didn't choose as ground truth, not noise.
|
|
116
|
-
|
|
117
|
-
Fix the worst thing via `apply_composition`, preview the affected frames and motion, and re-score when useful. For a whole-film review, budget up to 3–5 verify passes; two no-improvement passes are a reason to stop spending and report remaining gaps (`cueframe-compose-loop` owns the convergence loop). The rubric's acceptance criteria still apply; a score around 7.5 is guidance, not permission to ship a known mismatch. Hosted preview and scoring are stateless durable jobs; there is no hosted session to close.
|
|
118
|
-
|
|
119
|
-
---
|
|
120
|
-
|
|
121
|
-
## The film spine — how graphics-led video coheres (narration first, always)
|
|
122
|
-
|
|
123
|
-
Benchmarked against a top studio launch film: a scene detector reads theirs as **one continuous 44s shot**, because nothing ever changes the whole frame at once — and its spine is a continuous voiceover whose words run word-synced along the bottom of *every* frame. A film coheres when one thing never stops; visuals are then free to change every 1–2 seconds. Build in this order:
|
|
124
|
-
|
|
125
|
-
1. **Write the narration script FIRST — it is the film.** 10–14 short sentences, one idea each (~2.2 words/sec). Do not author a single visual before the script exists. Then `generate_media { generator: "text-to-speech" }` for one continuous take, transcribe it (or measure sentence boundaries), and **cut every beat boundary to the voice** — never the reverse. A silent text-driven film forces every card to hold long enough to be *read*; that dead time is the slideshow texture. With a voice carrying the through-line, each visual is a 2–3s illustration of the phrase being spoken.
|
|
126
|
-
2. **Author the caption spine with `setCaptions`** — segments of word timings matching the VO, styled small/bottom/`plate:"none"`. The persistent word-synced line is the strongest continuity device the format has; the `caption-spine` cohesion floor flags a vo-role clip without it.
|
|
127
|
-
3. **One world.** A single background system (color, texture, light) across every beat. Chapters are tiny inline corner labels (`01 — INPUTS`) inside the flowing world — **never full-screen interstitial cards**, which hard-stop the film. No beat gets its own room (a black hook, a dark grid finale): a full-frame background change reads as a scene change, and enough of them make a slide deck.
|
|
128
|
-
4. **One part vocabulary + one recurring actor.** Everything on screen is the same part (one card/chip grammar: same radius, shadow, label style) at different densities. Pick one small element (a dot, a mark) that appears in every beat and *hands off* across cuts — end a beat with it where the next beat starts it (an exit that launches it up off-frame matched by an entrance that drops it in is a match cut).
|
|
129
|
-
5. **Exit phases, no dead holds.** Every beat's elements leave the frame during its last ~0.4s; no component ends at rest, so motion carries across every cut. No visual state holds longer than ~3s (the `static-hold` floor fires on cards/still images past 4s). Cards can carry this motion themselves now — see the motion runtime below; components (`create_component`) remain the medium for choreography beyond it.
|
|
130
|
-
6. **The falsifiable bar:** run scene detection (or imagine it) over your cut — if it finds a hard boundary at every beat, you built a deck, not a film. Sparse frames are a tell too: elements sized for a fraction of the canvas leave dead space the reference never has.
|
|
131
|
-
|
|
132
|
-
### The card token contract — `tokens` key `x` is read as `var(--cf-x)`
|
|
133
|
-
|
|
134
|
-
A `{kind:'card', html, tokens}` clip binds brand-swappable values through `tokens`. Each
|
|
135
|
-
**key** is projected onto the card's wrapper as a CSS custom property with a `--cf-`
|
|
136
|
-
prefix, so the html reads it back **prefixed**:
|
|
137
|
-
|
|
138
|
-
```jsonc
|
|
139
|
-
{ "kind": "card", "tokens": { "plate": "#0af", "fg": "$brand:colors.text" },
|
|
140
|
-
"html": "<div style=\"background:var(--cf-plate);color:var(--cf-fg)\">…</div>" }
|
|
141
|
-
```
|
|
142
|
-
|
|
143
|
-
Values may be literals or whole-string `$brand:<path>` refs (resolved against the
|
|
144
|
-
project's brand kit at render).
|
|
145
|
-
|
|
146
|
-
**Do not put dashes in the key.** `tokens: {"--plate": …}` projects to `--cf---plate`, so
|
|
147
|
-
the card's `var(--plate)` matches nothing — and CSS treats an undefined custom property as
|
|
148
|
-
*invalid at computed-value time*, which means it does not error, it **degrades**:
|
|
149
|
-
`background` falls back to transparent and `color` inherits (usually black). The render
|
|
150
|
-
succeeds and the graphic is unreadable. `apply_composition` returns `warnings[]` when it
|
|
151
|
-
catches this, but the fix is to never write the dashes.
|
|
152
|
-
|
|
153
|
-
Two `--cf-*` families share the wrapper: **your token vars** (above) and the **motion
|
|
154
|
-
vars** below, which are always injected and are not yours to declare.
|
|
155
|
-
|
|
156
|
-
### The motion runtime — three ways anything you author can move
|
|
157
|
-
|
|
158
|
-
CSS `@keyframes`/`transition` **never animate in a render** (frames are independent
|
|
159
|
-
Chromium captures; the clock never advances). Use these instead:
|
|
160
|
-
|
|
161
|
-
1. **Card motion vars (`--cf-*`)** — every card's wrapper carries frame-driven, unitless
|
|
162
|
-
custom properties: `--cf-progress` (0→1 across the clip), `--cf-ease` (eased progress),
|
|
163
|
-
`--cf-enter` (0→1 over the first ~0.6s), `--cf-exit` (1→0 over the last ~0.6s), plus
|
|
164
|
-
`--cf-frame` / `--cf-t` / `--cf-fps` / `--cf-duration`. Author motion directly in card
|
|
165
|
-
CSS: `opacity: var(--cf-enter)`, `transform: translateY(calc((1 - var(--cf-enter)) * 40px))`,
|
|
166
|
-
an exit slide via `--cf-exit`. This satisfies rule 5's exit phases without leaving the
|
|
167
|
-
card medium — and a `preview_frame` still at time *t* shows the exact *t*-frozen state.
|
|
168
|
-
2. **`transitionIn` on any overlay clip** — `{ transitionIn: { type: "fade", duration: 0.4 } }`
|
|
169
|
-
now applies to overlay-family clips (cards, components, primitives), not just video.
|
|
170
|
-
Use it for entrances; author exits in-card via `--cf-exit`.
|
|
171
|
-
3. **`applyMotionPreset` op** — shell a static card into a frame-driven component in one op:
|
|
172
|
-
`{ type: "applyMotionPreset", clipId, preset }` (unknown presets are rejected with the
|
|
173
|
-
full preset list). The layout stays verbatim; the preset adds entrance/idle/exit
|
|
174
|
-
choreography. The result is a component you can `get_component_source` and refine.
|
|
175
|
-
|
|
176
|
-
---
|
|
177
|
-
|
|
178
|
-
## SFX pack — the curated CC0 sound palette
|
|
179
|
-
|
|
180
|
-
CueFrame ships a **closed, versioned pack of CC0 sound effects** (40 sounds, all
|
|
181
|
-
peak-normalized). There are two ways to place them. **Automatically:** the compose
|
|
182
|
-
SFX pass binds them **deterministically** — a hit on every eligible graphic
|
|
183
|
-
entrance (family from the primitive: `impact` for hero/title, `pop` for
|
|
184
|
-
stat/value, `tick` for enumeration items), a `whoosh` on video transitions, and a
|
|
185
|
-
`riser` ending at the hero entrance — timed to the same frames as the overlay,
|
|
186
|
-
ducked under speech. **By hand:** a client author places any pack sound directly —
|
|
187
|
-
`import_resource { kind:"sfx-pack", id:"<soundId>" }` returns the org's media
|
|
188
|
-
receipt (`{ id, status:"ready" }`) for that sound, which you then drop in as a
|
|
189
|
-
normal sfx clip (`mediaId`). Idempotent per (org, soundId); an unknown soundId is
|
|
190
|
-
a loud 422. Same curated bytes either path.
|
|
191
|
-
Restraint (density caps, min spacing, speech collision) comes from the AUDIO_MIX
|
|
192
|
-
SSOT the eval floors grade with, so placement and grading can't disagree.
|
|
193
|
-
|
|
194
|
-
**Variety is a floor, not a nicety.** The `sfx-same-sound-reuse` craft floor fires
|
|
195
|
-
when a dense mix (≥6 SFX hits) draws from too small a palette — distinct sounds ÷
|
|
196
|
-
hits below **0.4**. The pack carries **≥4 distinct sounds per transient family**
|
|
197
|
-
precisely so a memoir-class film satisfies it from the pack alone; leaning every
|
|
198
|
-
hit on one sound is the tell it guards. Browse the full palette structurally on
|
|
199
|
-
`list_catalog` (the `sfx` array — `soundId`, `family`, `envelopeClass`,
|
|
200
|
-
`description`, `durationSec`).
|
|
201
|
-
|
|
202
|
-
**One palette, org sounds first.** `list_catalog` also returns an `audio` array —
|
|
203
|
-
the ONE unified sound palette: the curated packs (SFX one-shots **and** CC0 music
|
|
204
|
-
beds) unioned with THIS org's own uploads/imports (`source:"org"`) and generated
|
|
205
|
-
audio (`source:"generated"`), each entry a uniform `{ref, role, envelopeClass,
|
|
206
|
-
description, durationSec, evidence, source, license, provenance}` (music entries
|
|
207
|
-
carry `bpm`). Provenance rides every entry — pack attribution, `org-owned`, or the
|
|
208
|
-
generator + prompt. **Prefer an org sound when one fits the family/envelope class:**
|
|
209
|
-
an org that uploads its product's real UI sounds hears ITS product, not a generic
|
|
210
|
-
pack tick — the compose SFX pass does this automatically (org sound of the matching
|
|
211
|
-
envelope class beats the pack default, exactly like a brand kit overrides the
|
|
212
|
-
default font), and by hand you place an org `ref` (a `mediaId`) directly. **The
|
|
213
|
-
music tier is curated beds + a generate escape hatch:** browse the `source:"pack"`
|
|
214
|
-
`role:"music"` beds (each with `bpm` + mood, spanning slow/mid/fast) and place one
|
|
215
|
-
via `import_resource { kind:"music-pack", id:"<bedId>" }` (FREE, like sfx-pack); for
|
|
216
|
-
a bed the pack lacks, `generate_media { generator:"text-to-music" }`. Buy/curate SFX,
|
|
217
|
-
curate-or-generate music.
|
|
218
|
-
|
|
219
|
-
**Transient one-shots** (entrance hits, per-element ticks):
|
|
220
|
-
|
|
221
|
-
| family | soundIds (subtle→strong) |
|
|
222
|
-
|---|---|
|
|
223
|
-
| `tick` | `tick-soft` · `tick-toggle` · `tick-scroll` · `tick-select` · `tick-click` · `tick-mouse` · `tick-shutter` · `tick-switch` |
|
|
224
|
-
| `pop` | `pop-drop` · `pop-glass` · `pop-pluck` · `pop-drop-low` · `pop-bong` · `pop-confirm` |
|
|
225
|
-
| `impact` | `impact-generic` · `impact-soft` · `impact-tin` · `impact-metal` · `impact-wood` · `impact-firm` · `impact-punch` · `impact-bell` |
|
|
226
|
-
| `confirm` | `confirm-soft` · `confirm-chime` · `confirm-question` · `confirm-bright` |
|
|
227
|
-
| `error` | `error-soft` · `error-buzz` · `error-tone` · `error-hard` |
|
|
228
|
-
|
|
229
|
-
**Sweeps** (transitions + hero lead-in, ~0.15–1.2s):
|
|
230
|
-
|
|
231
|
-
| family | soundIds |
|
|
232
|
-
|---|---|
|
|
233
|
-
| `whoosh` | `whoosh-air` · `whoosh-whip` · `whoosh-page` · `whoosh-sweep` · `whoosh-phaser` · `whoosh-low` |
|
|
234
|
-
| `riser` | `riser-jump` · `riser-phaser` · `riser-sweep` · `riser-power` |
|
|
235
|
-
|
|
236
|
-
(The SFX pack is transient/sweep one-shots; a sustained bed is a MUSIC-role clip —
|
|
237
|
-
a curated `music-pack` bed or a `text-to-music` generation, not a pack hit; see the
|
|
238
|
-
`audio` palette above.) For a sound the pack lacks, search
|
|
239
|
-
CC0 stock at curation time via `search_resources { kind:"sfx" }` (Freesound,
|
|
240
|
-
duration-bounded) and `import_resource` it — the pack is the fast default, not the
|
|
241
|
-
ceiling. For a sound that doesn't exist anywhere, generate it:
|
|
242
|
-
`generate_media { generator: "text-to-sfx", prompt: "short bright metallic tick, fast decay, no tail" }`
|
|
243
|
-
(0.5–22s; the delivered file is envelope-checked before it goes ready).
|
|
244
|
-
|
|
245
|
-
### Audio design — how to USE sounds (canonical doctrine + our calibration)
|
|
246
|
-
|
|
247
|
-
This is established film/UI sound-design craft — Chion, Murch, Thom, and the
|
|
248
|
-
Material sound guidelines — with the numbers calibrated on our own measured films.
|
|
249
|
-
The stakes are empirical, not aesthetic: films with sound effects measure >3x higher
|
|
250
|
-
perceived immersion than without (Kock & Louven), and audio-aware models consistently
|
|
251
|
-
beat vision-only models at predicting short-video engagement (VQualA 2025) — the
|
|
252
|
-
soundtrack is a measurable share of whether anyone keeps watching.
|
|
253
|
-
|
|
254
|
-
0. **Who owns time — pick the rhythm authority FIRST.** Before any hit or bed,
|
|
255
|
-
decide what the cut serves. A film has ONE rhythm authority (word timings, beat
|
|
256
|
-
maps, and shot boundaries are the same TYPE — a temporal grid; pick which one
|
|
257
|
-
drives). **VO-driven** (a narration film): cut every beat boundary to the
|
|
258
|
-
*voice* — the film-spine rule (author the script first, then read the word grid
|
|
259
|
-
off `get_media_context` — `wordGrid.wordTimesMs`, the word onsets in ms — and
|
|
260
|
-
land each cut/caption on a word). **Music-driven** (a montage/no-VO piece): cut
|
|
261
|
-
on the *beat* — read `beatGrid` off `get_media_context` (bpm + `beatTimesMs`),
|
|
262
|
-
land cut points and transition MIDPOINTS on beats, and resolve risers exactly on
|
|
263
|
-
a downbeat/strong beat (every 4th beat from the grid start). A grid is EVIDENCE,
|
|
264
|
-
not autopilot — you choose which beats carry cuts; on-beat reads as craft,
|
|
265
|
-
off-beat as accident (synchresis, Chion). Trust `beatGrid` for cutting when its
|
|
266
|
-
`confidence` is high; a beatless bed reports low and carries no reliable grid.
|
|
267
|
-
Once you have picked the authority, **materialize its grid onto the timeline
|
|
268
|
-
with the `materializeGrid` op** (`{clipId, grid:"beat"|"word", maxMarkers?}`):
|
|
269
|
-
the server reads that clip's source grid and lands `kind:"beat"|"word"` markers
|
|
270
|
-
at the correct TIMELINE times (projected through the clip's trim/playbackRate/
|
|
271
|
-
excludedRanges — you never do the media→timeline math yourself). Then place cuts
|
|
272
|
-
and transition **midpoints** on the markers — a transition's midpoint IS the cut
|
|
273
|
-
point, so start a 0.5s transition 0.25s *before* the marker so its center lands
|
|
274
|
-
on the beat. Materialize ONCE per authority clip; re-running the op refreshes the
|
|
275
|
-
markers in place (it replaces the prior set from the same clip+grid), so re-run
|
|
276
|
-
after you retrim or restretch the clip. (`grid:"shot"` is reserved — not exposed
|
|
277
|
-
yet.)
|
|
278
|
-
|
|
279
|
-
0b. **Declare the audio's ROLE at ingest — evidence follows the declaration.**
|
|
280
|
-
Uploads/imports accept an optional `audioRole: vo|music|sfx|ambience` (finalize
|
|
281
|
-
and import bodies). Declaring it routes enrichment: `vo` → transcript + word
|
|
282
|
-
grid; `music` → `beatGrid` (detected, no prompt needed); `sfx`/`ambience` →
|
|
283
|
-
`envelope` evidence on `get_media_context` (`envelopeClass:
|
|
284
|
-
transient|sweep|ambient` + attackMs/crest stats — pick pack-style hits by
|
|
285
|
-
class, not by listening). Undeclared audio is transcribed as before, and a
|
|
286
|
-
no-speech file still comes back envelope-classed. **`ambience` is the fourth
|
|
287
|
-
role**: a bed that grounds the SPACE — it loops, sits at the bottom of the mix,
|
|
288
|
-
and HOLDS under VO (it is never ducked; only `music` ducks under speech). Use
|
|
289
|
-
ambience for room tone/atmosphere continuity across cuts (the room-changes
|
|
290
|
-
floor's natural fix), never for anything that must be *noticed*.
|
|
291
|
-
|
|
292
|
-
0c. **The engine owns the default mix — author levels only to deviate.** A clip
|
|
293
|
-
tagged with an `audioRole` but no authored `volumeDb`/`volume` is auto-staged
|
|
294
|
-
to the clarity hierarchy (VO 0 dB, sfx −6, music −16 under VO else −8,
|
|
295
|
-
ambience −22 — keeping VO ≥6 dB over the bed). An authored level always wins,
|
|
296
|
-
a role-less clip stays at unity, and authored levels that invert the hierarchy
|
|
297
|
-
(e.g. music over VO) raise a non-blocking `mix_hierarchy_inverted` warning on
|
|
298
|
-
`create_render` naming the offending clips. Tag roles and let the engine mix;
|
|
299
|
-
reach for `volumeDb` only when you intend to break the hierarchy.
|
|
300
|
-
|
|
301
|
-
0d. **A brand kit can carry a SOUND identity — set it once, it rides every
|
|
302
|
-
compose.** `create_brand_kit`'s `audio` block (echoed on `get_brand_kit` +
|
|
303
|
-
`get_profile`) holds the brand's sonic taste; every ref is a palette reference
|
|
304
|
-
(a pack `soundId` | `bedId` | an org `mediaId` from `list_catalog`'s `audio[]`).
|
|
305
|
-
Four optional fields, resolution precedence **brand > org > pack** (the audio
|
|
306
|
-
analog of how a brand FONT sits above org/pack defaults):
|
|
307
|
-
- **`sonicLogo` `{ ref, placement: intro|outro|both, offsetMs? }`** — the brand
|
|
308
|
-
signature. When the kit ALSO carries the matching bookend (`intro`/`outro`
|
|
309
|
-
bumper), compose binds an sfx-role clip at the body edge that bookend abuts
|
|
310
|
-
(intro → body start; outro → body end − logo length; `offsetMs` shifts
|
|
311
|
-
within). No matching bumper ⇒ a non-blocking `brand_sonic_logo_no_bookend`
|
|
312
|
-
advisory and nothing placed (the sonic logo is the AUDIO companion to the
|
|
313
|
-
VISUAL bookend). A `ref` that no longer resolves fails loud
|
|
314
|
-
`brand_audio_unresolvable` — re-point it or clear the sonic logo.
|
|
315
|
-
- **`uiSoundSet` `{ family → ref }`** — per-family SFX overrides (the brand's
|
|
316
|
-
own tick/pop/impact/…). Consumed by the compose SFX draw as the HIGHEST tier:
|
|
317
|
-
a family the brand overrides plays the brand's sound, everything else falls to
|
|
318
|
-
the org palette then the pack default.
|
|
319
|
-
- **`bedStyle` `{ bedId? | genre?, bpmRange? }`** and **`ambiencePreference`
|
|
320
|
-
`{ ref } | null`** — SURFACED-ONLY today: compose auto-selects neither a music
|
|
321
|
-
bed nor ambience, so these carry the brand's preference for a DRIVING agent to
|
|
322
|
-
honor when it pins/generates a bed (`brief.audio.music`) or places ambience.
|
|
323
|
-
Extraction can't hear: `extract_brand_kit` (website → colors/fonts) never fills
|
|
324
|
-
the `audio` block — set the sound identity explicitly on `create_brand_kit`.
|
|
325
|
-
|
|
326
|
-
1. **Every hit is a SYNC POINT — sound welded to a visible event.** Chion
|
|
327
|
-
(*Audio-Vision*) calls this **synchresis**: the immediate, involuntary weld the
|
|
328
|
-
brain makes between a sound and the image it lands on — that weld is where
|
|
329
|
-
"added value" comes from, and a hit with no visible cause produces none (it's
|
|
330
|
-
the audio version of unmotivated motion, slop tell #6). Bind hits to their
|
|
331
|
-
events via `anchor: { clipId, offsetMs }` so the weld survives re-edits. And
|
|
332
|
-
per Randy Thom (*Designing a Movie for Sound*): sound is a storyteller, not
|
|
333
|
-
decoration — score the moments that matter, let ordinary cuts breathe.
|
|
334
|
-
2. **Class → moment.** `transient` (tick/pop/impact/confirm/error) = entrances,
|
|
335
|
-
cuts, UI semantics — Material's sound guidelines are the reference for UI
|
|
336
|
-
semantics (each sound expresses its place in the hierarchy; decorative sound
|
|
337
|
-
used sparingly). `sweep`: whoosh = motion carrying ACROSS a cut (pair with the
|
|
338
|
-
exit→entrance handoff); riser = tension INTO a reveal — resolve exactly on the
|
|
339
|
-
downbeat, never into nothing (an unresolved riser is a broken promise).
|
|
340
|
-
3. **Density has a perceptual ceiling.** Murch (*Dense Clarity — Clear Density*):
|
|
341
|
-
the brain tracks only ~two-and-a-half simultaneous sound streams; past that,
|
|
342
|
-
individual sounds stop reading and merge into texture — his mixes convey
|
|
343
|
-
complex scenes with a FEW carefully chosen elements. Our calibration: the
|
|
344
|
-
41.6s film that beat its reference carried 9 hits; the rejected cut carried 15
|
|
345
|
-
from 3 sounds. At 6+ hits the variety floor applies (`sfx-same-sound-reuse`
|
|
346
|
-
fires below 0.4 distinct/hits).
|
|
347
|
-
4. **Hierarchy within a family** (Murch's balanced-spectrum principle + Material's
|
|
348
|
-
sound hierarchy): the pack orders soundIds subtle→strong; the biggest beat gets
|
|
349
|
-
the strong impact ONCE, everything else sits a tier down. Fifteen strong hits
|
|
350
|
-
= zero strong hits.
|
|
351
|
-
5. **Gain staging.** SFX under a music bed must READ: target the hit ~≥6 dB above
|
|
352
|
-
the bed at its transient (our measured failure: ticks 1.5 dB over bed =
|
|
353
|
-
inaudible mush). Working precedent: sfx `volumeDb: -12`, bed lower still.
|
|
354
|
-
Author sane levels and STOP — the engine masters to −14 LUFS (the streaming
|
|
355
|
-
loudness standard); never pre-compress or push peaks to compensate.
|
|
356
|
-
6. **The bed is tempo, not wallpaper.** Match bed BPM/energy to the cut rate (our
|
|
357
|
-
fast recut earned a 124 BPM driving bed; a calm film wants sparser pulse), arc
|
|
358
|
-
the energy with the narrative, and always `duck` under VO
|
|
359
|
-
(`duck: { duckDb, attackMs, releaseMs }`) — dialogue sits atop Murch's
|
|
360
|
-
encoded–embodied spectrum; sfx punctuate BETWEEN phrases, never fight words.
|
|
361
|
-
7. **Silence is a device** (Thom: quiet is the most underrated tool in the
|
|
362
|
-
soundtrack). Clean air before the first hit; a dropout before the biggest
|
|
363
|
-
reveal is worth more than any riser. If the whole film is scored, nothing is.
|
|
364
|
-
|
|
365
|
-
---
|
|
366
|
-
|
|
367
|
-
## Recognizing slop (the mechanism — so you can predict it, not just avoid it)
|
|
368
|
-
|
|
369
|
-
Slop is the model's most common answer to an under-specified prompt: what the path of least resistance produces once the human decision layers are removed. It has a recognizable signature *because* each missing decision compounds:
|
|
370
|
-
|
|
371
|
-
- **No juxtaposition, no meaning.** Meaning in video is manufactured by adjacency — the same neutral shot reads as grief or desire depending only on the shot next to it (Kuleshov effect; Eisenstein's montage of collision). A gradient next to a gradient creates nothing. Grounded, specific shots in sequence (problem-shot → product-shot → relief-shot) make a claim the individual shots never state. Compose in **pairs and triads**, ordered so adjacency does the arguing.
|
|
372
|
-
- **No specificity, no belief.** Images are encoded more deeply than abstract words (picture-superiority effect, Paivio), and — a separate, stronger mechanism — a real human face is among the most attention-binding things a screen can show. "10x faster" on a card is an abstraction the viewer can't feel; real people talking in a real room, or a diagram of the *actual* mechanism, is embodied and credible — it *shows*.
|
|
373
|
-
- **No subject, no eye-trace, no peak.** With nothing specific on screen there's no focus of interest to cut on (Murch's eye-trace) and no candidate for the one emotional peak the viewer remembers (Kahneman's Peak-End Rule). A flat decorated slideshow has no peak and a whimper ending.
|
|
374
|
-
|
|
375
|
-
So grounding the video (step 3) is load-bearing, not a preference. Missing it guarantees the rest.
|
|
376
|
-
|
|
377
|
-
### Slop tells — you are making slop if a frame has any of these
|
|
378
|
-
1. **Gradient/aurora background under centered kinetic text**, no grounded subject. The literal default of every text-to-video toy. Zero information per frame — animating the text, or calling it "purposeful kinetic typography," does NOT ground it; only a visual of the real mechanism or real data does.
|
|
379
|
-
2. **Generic AI b-roll** — glowing brains, blue circuit boards, particle networks, rotating server racks, holographic dashboards. "Technology" as a vibe; interchangeable across a thousand companies.
|
|
380
|
-
3. **Faked or mocked UI/terminals** — a styled `<div>` pretending to be the product. The cursor doesn't blink, nothing scrolls, the output is too clean. A real 4-second capture beats it every time.
|
|
381
|
-
4. **On-screen numbers with no real source** — "10x", "5x", "99.9%" over a gradient. Unsourced is the content-farm tell; *wrong* is a credibility bomb.
|
|
382
|
-
5. **All generic filler, nothing grounded** — 100% decorative generated backgrounds and kinetic text, with no real footage, no real UI/data, and no purposeful content-specific graphic. The master failure that produces all the others. (A *purposeful* animated explainer — one built on diagrams of the real mechanism or animations of the real data, NOT animated copy — is grounded and does NOT count as this tell.)
|
|
383
|
-
6. **Unmotivated motion** — everything drifts, floats, Ken-Burns-pans, swoosh-transitions for no reason. The screensaver aesthetic: soft, well-oiled, no harsh edges, regardless of content.
|
|
384
|
-
7. **Template rhythm** — title card → 3 bullet slides → CTA, every beat the same length and entrance. The eye recognizes the mold.
|
|
385
|
-
8. **Metronomic pacing** — every clip the same length, every caption the same duration, no held beat before the payoff. The tell of an algorithm chunking on a timer.
|
|
386
|
-
9. **Off-brand default type / over-captioning** — Inter/Arial where a brand typeface exists; karaoke captions on a slideshow with no speech to caption.
|
|
387
|
-
10. **Dead-center symmetry, one z-plane** — subject and text fighting for the same pixels, throwing away the depth CueFrame gives you (an overlay title at `zPlane:"behind-subject"`).
|
|
388
|
-
|
|
389
|
-
### The positive inverse — author toward these
|
|
390
|
-
One controlling idea every clip serves · grounded content of the real thing (real footage, real UI/data, or a purposeful diagram of the actual mechanism) · motivated motion (the frame moves because the subject moved or the idea turned) · restraint (fewer elements, held longer) · typographic discipline (brand type, ≤2 weights, one hierarchy) · honest data with visible provenance ("8.5%, with the source on screen" beats "10x") · varied pace that builds to one memorable beat · real depth (a title behind the head, layered z-planes).
|
|
391
|
-
|
|
392
|
-
### The launch-video case — the exact contrast to burn in
|
|
393
|
-
**Brief:** a product-launch video for a developer-tools product. (Composite example.)
|
|
394
|
-
**What the agent shipped (slop):** ~37s of purple-gradient slides with centered kinetic text; a static mocked terminal card pretending to be the product running; a fabricated "10x faster" where the sourced result was **8.5%**; zero grounded content. A clean sweep of tells #1, #3, #4, #5, #7, #10.
|
|
395
|
-
**What was sitting right there, ignored:** the team's own launch footage — people talking to camera in a nice office. *Precisely* CueFrame's home turf: `all-faces` on a two-shot / `active-speaker` on the singles, transcript-timed captions, a lower-third per speaker, a color grade, a 4K render.
|
|
396
|
-
**Craft looks like (same tool, same 37s):** `import_media` the interview footage → `get_media_context` for faces + transcript → reframe to whoever is speaking → transcript-tracked captions with a legibility plate → a title behind the head → a `lower-third` naming each speaker → the team's real brand type → if a benchmark appears, the real **8.5%** with its source visible, ideally over a real capture of the product running. One idea, grounded content, motivated cuts, honest data, a payoff. (Even the no-footage version was fixable: an authored diagram of the product's *actual* pipeline step beats a gradient — that's grounded too.)
|
|
397
|
-
|
|
398
|
-
---
|
|
399
|
-
|
|
400
|
-
## Operating floors (fold into every format)
|
|
401
|
-
|
|
402
|
-
Directional, not laws — reason above them:
|
|
403
|
-
|
|
404
|
-
- **Hook = first 3 seconds.** Viewers decide to stay or scroll almost immediately; the first frame is the highest-leverage frame in the piece. Budget craft disproportionately to 0–3s. Put a real face or real motion there — never a logo fade.
|
|
405
|
-
- **Caption speech, always.** Most social video is watched sound-off, and captions lift retention and completion — treat them as mandatory, not decoration. Attach `captions` with word-level timing from `get_media_context`; set legibility via `setCaptionStyle` (`position` + `emphasisPlate`). Caption **speech** — do not karaoke a slideshow.
|
|
406
|
-
- **Length: earn it.** Engagement decays with length; short pieces complete far better than long. A sub-1-min teaser is watched to the end at rates a 5-min film never sees; instructional pieces are the exception *only* when they genuinely teach. Never pad — every second must carry the idea forward.
|
|
407
|
-
- **Cut cadence (motivated, never for its own sake):** a visual change every ~1.5–2s for sub-60s pieces; B-roll shots ~1.5–3s (>3s stalls); cinematic ~4–6s. For talking-head/demo the two operations differ: **reframe/punch-in every ~3–4s *within* a held speaker shot**, and insert a B-roll **cutaway every ~5–8s** of continuous talking. Drive with clip `trim`/`duration` and overlay entrance timing — but every change of state must be *motivated* (a reframe to the speaker, a lower-third entrance, a caption emphasis).
|
|
408
|
-
- **Word budget:** ~130–150 wpm; 60s ≈ ~130 words; a 12–15-word hook ≈ 5–6s of speech. Past ~160 wpm reads rushed.
|
|
409
|
-
- **Sound sets perceived pace.** CueFrame composes to an existing audio bed — use it: land caption entrances and cut points on stressed words / beats via the transcript's word timing (a cut on the beat reads as craft, off it reads as accident — synchresis, Chion); duck music under speech; hold one beat of silence before the payoff line. Silence before the peak is a tool, not an absence.
|
|
410
|
-
- **The saturated default.** Gradient-text-over-music synced to a beat is the *standard* software-launch look, not a differentiator. When you reach for that generic default instead of grounding the piece, you land in exactly that interchangeable tier — which is the whole reason step 3 exists.
|
|
411
|
-
|
|
412
|
-
---
|
|
413
|
-
|
|
414
|
-
## Format playbooks
|
|
415
|
-
|
|
416
|
-
Each playbook is a **floor**, not a template — reason above it. Every beat maps to a real CueFrame capability. Default to **16:9 landscape** for product/demo/explainer/founder-web; **9:16 vertical** only for social cutdowns (a wide source crammed vertical crops the subject wrong and wastes the frame).
|
|
417
|
-
|
|
418
|
-
### 1 — Product launch / announcement
|
|
419
|
-
**When:** launching or announcing a product; teaser or full launch film.
|
|
420
|
-
**Beat sheet (45s, 16:9):**
|
|
421
|
-
| Beat | Time | Content | Words |
|
|
422
|
-
|---|---|---|---|
|
|
423
|
-
| Cold-open hook | 0–3s | One wow frame or the single pain moment. No logo. | ≤10 |
|
|
424
|
-
| Problem | 3–12s | The specific status-quo pain, shown not told. | ~20 |
|
|
425
|
-
| Reveal | 12–30s | The product doing the thing on REAL footage/screen. | ~35 |
|
|
426
|
-
| Proof | 30–40s | One concrete result or credible demo moment. | ~20 |
|
|
427
|
-
| CTA + date | 40–45s | Exact action + launch date. Logo last. | ≤12 |
|
|
428
|
-
|
|
429
|
-
**Frameworks:** 3-act teaser (Tease → Reveal → CTA) or StoryBrand SB7 (customer is hero, you are guide). One pain or one wow — not five diluted features.
|
|
430
|
-
**Reach for:** `import_media` a real demo capture or founder clip · `list_catalog` → the `product-launch-trailer` and `cinematic-title` scenes (use `cinematic-title` for the *closing* wordmark, never the open) · `color-grade` for the hero look · date + CTA on the final held frame.
|
|
431
|
-
**Anti-patterns:** opening on a logo animation; stacking five features; narrating specs; a fabricated/rounded stat; a mocked terminal standing in for a real screen.
|
|
432
|
-
**Quality bar:** the first frame survives with zero text on it. If your open is a gradient with the product name, you've already lost.
|
|
433
|
-
|
|
434
|
-
### 2 — Short-form social clip (TikTok / Reels / Shorts)
|
|
435
|
-
**When:** a vertical clip for social, usually cut from existing long-form.
|
|
436
|
-
**Beat sheet (22s, 9:16):**
|
|
437
|
-
| Beat | Time | Content |
|
|
438
|
-
|---|---|---|
|
|
439
|
-
| Visual + verbal hook | 0–3s | Bold claim / question / pattern-interrupt, on screen AND spoken. Answer teased, not given. |
|
|
440
|
-
| Setup | 3–8s | Frame the stakes; open loop #1. |
|
|
441
|
-
| Escalation | 8–16s | Value in ≤2 beats, each opening the next loop. Cut every ~1.5–2s. |
|
|
442
|
-
| Payoff | 16–20s | Close the main loop; put the payoff word on screen. |
|
|
443
|
-
| CTA / loop-back | 20–22s | Soft CTA or a re-hook that rewards replay. |
|
|
444
|
-
|
|
445
|
-
**Frameworks:** Hook → Retention → Payoff with curiosity stacking. Keep cutdowns ≤30s — shorter clips complete far more often.
|
|
446
|
-
**Reach for:** `suggest_briefs` on the long-form media (the AI clip-finder authors trim + speaker framing + captions from the transcript for you), then `compose({ suggestionId })` (Director ensemble) or author yourself via `apply_composition` · reframe `focus: {mode:"active-speaker"}` to keep the speaker centered in the vertical crop · word-level captions with `emphasis` on the payoff word and `entrance:"word-pop"` · `derive_composition` to auto-reframe a 16:9 master to 9:16 content-aware (not letterboxed).
|
|
447
|
-
**Anti-patterns:** a wide source letterboxed into 9:16 with black bars instead of subject-tracked reframe; no captions (dies on mute); a hook that describes instead of provokes ("In this video I'll…").
|
|
448
|
-
**Quality bar:** the hook sentence starts at t ≤ 1.5s; the face box is a real fraction of the frame with the top of the head in-frame on every sampled still.
|
|
449
|
-
|
|
450
|
-
### 3 — Explainer
|
|
451
|
-
**When:** explaining a concept, mechanism, or product so a viewer *gets* it.
|
|
452
|
-
**Beat sheet (75s, 16:9, ~160 words @ ~130 wpm):**
|
|
453
|
-
| Beat | Time | Words | Content |
|
|
454
|
-
|---|---|---|---|
|
|
455
|
-
| Hook | 0–5s | ~12–15 | The "why watch" — land inside 3s. |
|
|
456
|
-
| Problem | 5–20s | ~30 | Name the exact pain; they must feel understood. |
|
|
457
|
-
| Solution | 20–45s | ~55 | Your answer and how it differs. Show it working. |
|
|
458
|
-
| Proof | 45–62s | ~35 | One result / mechanism / credible demo. |
|
|
459
|
-
| CTA | 62–75s | ~20 | One clear next step. |
|
|
460
|
-
|
|
461
|
-
**Frameworks:** Hook → Problem → Solution → Proof → CTA. Decide the CTA first, script it last; write for the ear.
|
|
462
|
-
**A fully animated explainer is a FIRST-CLASS grounded format**, not a fallback for missing footage. An explainer built on authored diagrams of the *actual* mechanism (its real pipeline, the real data) is grounded — the grounding is that each visual depicts THIS thing specifically, not a decorative loop. Reach for it deliberately, not as a consolation for "no footage."
|
|
463
|
-
**Reach for:** an authored `card` clip (your own HTML/typography, shelled with `applyMotionPreset`) to anchor an abstract mechanism with a real diagram — process/algorithm/data viz, not cinematic footage · `lower-third` for term labels · one idea per beat · when using real footage, narrate to the detected transcript timing.
|
|
464
|
-
**Anti-patterns:** a wall of narration over stock gradient loops with no visual keyed to the words (the generic-filler tell — a purposeful diagram of the real mechanism is the opposite); explaining features before establishing the problem; using `generate_media` text-to-video to fake a "product" that doesn't match the real UI (it hallucinates specifics).
|
|
465
|
-
**Quality bar:** every beat has a visual keyed to what's being said — grounded and specific to this content, not decoration; the hook lands in 3s.
|
|
466
|
-
|
|
467
|
-
### 4 — SaaS / dev-tool demo
|
|
468
|
-
**When:** showing a product/CLI/app actually doing a real task.
|
|
469
|
-
**Beat sheet (~2 min, 16:9, real screen capture):**
|
|
470
|
-
| Beat | Time | Content |
|
|
471
|
-
|---|---|---|
|
|
472
|
-
| Hook + problem | 0–15s | The workflow that hurts today. Real UI or founder to-camera. |
|
|
473
|
-
| Setup | 15–30s | The task we'll accomplish; the promised outcome. |
|
|
474
|
-
| Walkthrough | 30–95s | Do the real thing on the real screen. Punch-in on the exact UI region per step; ~3–4s cadence; caption each action. |
|
|
475
|
-
| Outcome | 95–110s | Result achieved; the payoff vs the opening pain. |
|
|
476
|
-
| CTA | 110–120s | One action: start free / book / docs. |
|
|
477
|
-
|
|
478
|
-
**Frameworks:** Problem → Product-in-action → Outcome, wrapped in one completed task, not a feature tour. ~2 min is a good ceiling; chapter beyond that.
|
|
479
|
-
**Reach for:** `import_media` a REAL screen recording (captured by the user, or by an external screen-capture client such as a headed Playwright walkthrough recorded to MP4 — CueFrame does not screen-record) · reframe punch-in to the active UI element each step (`focus:{mode:"point",x,y}` on the click target, or a `face` focus for a founder inset) — the single biggest amateur→pro delta in demos, because the click target is ~30px in a 3840px 4K frame and nobody sees it unzoomed (**see `cueframe-product-video`** for the full click-point→reframe-segment auto-zoom tiling + PiP recipe) · a `lower-third` label per step.
|
|
480
|
-
**Anti-patterns:** the cardinal demo sin — a mocked/fake terminal or static UI card instead of the real app; showing the full 4K screen unzoomed; narrating every menu.
|
|
481
|
-
**Quality bar:** every UI/terminal on screen is a real capture; the actual action is visibly punched-in, not lost in a wide frame.
|
|
482
|
-
|
|
483
|
-
### 5 — Founder / talking-head (CueFrame's home turf)
|
|
484
|
-
**When:** a real person speaking to camera — founder update, announcement, testimonial. The format CueFrame composes best, and the one that agent walked past.
|
|
485
|
-
**Beat sheet (75s, 16:9 for web; 9:16 cutdown for social):**
|
|
486
|
-
| Beat | Time | Content |
|
|
487
|
-
|---|---|---|
|
|
488
|
-
| Attention / Problem | 0–5s | Founder states the pain or bold claim to camera. Eyes to lens. |
|
|
489
|
-
| Interest / Agitate | 5–25s | Why it matters now; personal stakes. B-roll cutaway. |
|
|
490
|
-
| Desire / Solution | 25–55s | What they built and the shift it creates. Product B-roll. |
|
|
491
|
-
| Proof | 55–68s | One credible result. Lower-third for name/title/metric. |
|
|
492
|
-
| Action | 68–75s | One CTA. |
|
|
493
|
-
|
|
494
|
-
**Frameworks:** PAS (Problem → Agitate → Solution) or AIDA. Framing: eyes on the upper-third line, small headroom, subject looks into the lens; break every ~5–8s of pure talking head with a cutaway.
|
|
495
|
-
**Reach for:** THE canonical CueFrame job. `get_media_context` for faces + transcript → reframe `focus:{mode:"active-speaker"}` for a solo speaker, `{mode:"all-faces"}` to hold a two-shot, or `{mode:"face",faceId,shotScale:"close"}` to punch in on emphasis → transcript-timed `captions` with a legibility plate (`setCaptionStyle` `emphasisPlate`, `position:"bottom"`) → a text overlay at `zPlane:"behind-subject"` when you want a title/word to sit *behind* the head (the signature look; render bakes the person matte) → a `lower-third` for name/title, shown once → `color-grade`. Zero generation needed.
|
|
496
|
-
**Anti-patterns:** a static wide talking head running 30s uncut (reads as a raw webcam upload); captions slapped over the chin with no plate; head dead-center or too much headroom; missing the lower-third identity.
|
|
497
|
-
**Quality bar:** a face is on screen within 3s; the person framed is the person speaking; captions never cover the mouth or eyes.
|
|
498
|
-
|
|
499
|
-
---
|
|
500
|
-
|
|
501
|
-
## Map craft to CueFrame tools (never invent a tool, primitive, or param)
|
|
502
|
-
|
|
503
|
-
| Craft goal | Do this in CueFrame |
|
|
504
|
-
|---|---|
|
|
505
|
-
| Bring in the real thing | `import_media` — requires `url` (public https) + `filename` + `contentType` (from the fixed enum: `video/mp4`, `video/quicktime`, `image/png`, `audio/mpeg`, …). Poll `list_media` or `create_webhook` on `media.completed` (status inside: complete|failed). |
|
|
506
|
-
| Know where subject + speech are | `get_media_context` → face boxes per frame (with `faceId`) + transcript. If faces are not detected, `prepare_media(kind:"subjectTrack", intent:{startSec,endSec})` → poll `get_media_facts` for the same SOURCE window until ready. |
|
|
507
|
-
| Keep the speaker framed / punch-in | `setCropIntents` op (the PREFERRED authoring path) or inline `reframe.segments` on a `clip.add` source. Each segment needs `startSec`/`endSec`/`focus`/`zoom`; **segments must TILE the clip** (an uncovered span → `reframe_coverage_gap` at validate). `focus.mode`: `frame-center` \| `point{x,y}` \| `face{faceId, shotScale?: close\|standard\|wide}` \| `all-faces` \| `active-speaker`. `zoom<1` punches in, `=1` full frame, `>1` rejected (min 0.1). Times in **seconds**. |
|
|
508
|
-
| Captions that track real speech | `composition.captions.segments[].words[]` = `{text, startMs, endMs, emphasis?, annotate?: circle\|underline}` — word times in **milliseconds**. Author the words directly with the `setCaptions` op (segments wholesale — how a generated-VO film gets its word-synced spine without a source transcript); style via `setCaptionStyle`: `position` (top\|center\|bottom), `entrance` (fade\|word-pop\|stagger-up), `emphasisPlate` + plate colors for legibility. Captions default to front; set top-level `composition.captions.zPlane:"behind-subject"` only when the user wants the whole caption layer behind the speaker. It resolves person cutouts for every overlapping source window and degrades visibly to front where the legibility gate refuses the effect. |
|
|
509
|
-
| Text behind the speaker's head | An `overlay` clip (a title/label primitive) with `zPlane:"behind-subject"`. This is the focused hero treatment; caption-layer depth is a separate whole-layer choice. Before preview/judgment, `prepare_media(kind:"matte", intent:{startSec,endSec})` for the overlapping clip's SOURCE trim and poll `get_media_facts` until ready. Render also resolves the person cutout on miss. |
|
|
510
|
-
| Name someone / label a step | `lower-third` scene from `list_catalog` |
|
|
511
|
-
| Title / closing wordmark | `cinematic-title` scene (closing, not opening) |
|
|
512
|
-
| Grade the look | `color-grade` scene/effect |
|
|
513
|
-
| Anchor an abstract mechanism | An authored `card` clip + `applyMotionPreset` (diagrams of the real mechanism — NOT cinematic; a first-class grounded path) |
|
|
514
|
-
| Long-form → clips | `suggest_briefs` → `wait_job(kind:"suggestions")` → `compose({ suggestionId })` |
|
|
515
|
-
| Auto-reframe to another aspect | `derive_composition` (content-aware 9:16 / 1:1 / 4:5 from a 16:9 master) |
|
|
516
|
-
| Discover what to build with | `list_catalog` — over MCP it takes **NO input params**, so the FULL live catalog + your installed components returns, and you filter the RESULT **client-side** by `category`/`kind`/`tier`/`useCase`/`mood`/`intent`. Each entry carries `placement` (how to author it), `tier`/`useCase`, and `fixedCopy` (baked text no param changes). Read `cueframe://component/{id}` for the full **prop schema**. |
|
|
517
|
-
| Own / fork a custom graphic | `get_component_source` / `create_component` (self-contained tsxSource) / `update_component` — **see `cueframe-component-authoring`** for the full fork→preview→push loop |
|
|
518
|
-
| Several cells that share space and must reflow together (feature grid, rail + fluid content, cast + type) | An authored component on `FlexLayout` from `@cueframe/animate` — tracks are **weights** (`{ weight: 0, basis: px }` = content-sized, pushes neighbours), `tracksEnd` animates the reflow per frame, `fits` per cell (`stretch`/`contain`/`cover`/`position`/`scale`/`matte`). Static PiP / split is clip `region` + `fit` instead. Never CSS-scale the layout. **See `cueframe-component-authoring`** |
|
|
519
|
-
| Seed + edit the timeline | `apply_composition` (atomic batch ops; `dry_run:true` to check, `if_match` for OCC) — this is where clips, reframe, captions, and format are authored |
|
|
520
|
-
| Set the brand | `create_brand_kit` (402-gated) → set `brandKitId` on the **project** (`create_project` / `update_project`). It lives on the project, NOT on a composition op — trying to set it elsewhere is silently dropped. |
|
|
521
|
-
| Shape-check (NOT quality) | `validate_composition` → `{ valid, errors[] }` (`unknown_primitive`, `invalid_primitive_params`, `clip_source_track_mismatch`, `reframe_coverage_gap`, …); does **not** check trim bounds or image-on-video-track (those fail at save/render) |
|
|
522
|
-
| Look at real pixels | `preview_frame` at timestamps → PNG stills → LOOK |
|
|
523
|
-
| Judge motion or audio timing | `preview_clip` for the affected `[fromSec, toSec)` window → `wait_job(kind:"render")` → watch/listen; use matching reference timestamps |
|
|
524
|
-
| Check an isolated card | `preview_card` with its HTML and tokens; no scratch composition needed |
|
|
525
|
-
| Preview one graphic | `preview_component` (202 + jobId → `wait_job(kind:"preview")`; returns the verified artifact from the same immutable bundle used by production; compile/runtime/capability errors land on the job's failed terminal with the author-fixable message) |
|
|
526
|
-
| Independent grade (steered, un-gameable) | `score_composition` (sessionless; `editorialIntent`=your vision, `brandContext`=brand facts) → four fixed criteria + composite + worst-first critique, on frames it samples itself |
|
|
527
|
-
| Director ensemble compose | `compose` — set EXACTLY ONE of `fromComposition:true` (author from the seeded composition) or `suggestionId` (clipping path) → `wait_job(kind:"compose")` |
|
|
528
|
-
| Final MP4 | `create_render` → `wait_job(kind:"render")` or `create_webhook` on `render.completed`; `retry_render` (transient only) / `cancel_render` |
|
|
529
|
-
| Orient before spending | `get_account` FIRST — read `balance`, NOT `included`. Rendering / generation / previews / judge / Director share ONE credit wallet (`fundedBy`); each `balance` restates that same wallet in that feature's unit. Balance > 0 = proceed — never tell the user a capability is missing, never send them to checkout. New orgs get a free credit grant, so an empty usage history is not a blocker. 402 `billing_required` fires once the balance is spent; `generate_media` also 402s over its USD ceiling |
|
|
530
|
-
|
|
531
|
-
Everything needs a `projectId` — `create_project` is step 0.
|
|
532
|
-
|
|
533
|
-
**Hard limits — respect them, don't fake around them.** No cinematic AI generation: `generate_media` is limited text-to-video/text-to-image/image-to-video, plus audio: `text-to-speech` (narration VO) and `text-to-music` (instrumental bed); treat generated *visual* assets as *supporting graphics* (a diagram, a texture), never a fake-cinematic hero. No screen capture: that needs an external screen-capture client (e.g. a headed Playwright walkthrough recorded to MP4). When the right move needs footage CueFrame can't generate, `import_media` real footage — do not synthesize a substitute. **Units trap:** `startTime`/`duration`/`trim`/reframe `startSec`/`endSec`/`ease` are **seconds**; caption word times are **milliseconds**. **Render-time fail-loud errors to expect** (render never ships a wrong frame silently): `content_anchor` (a clip's `trim` frames *different* footage than its captions show — align the trim to the captioned moment), `behind_split_unsupported` (one media used under two *different* behind-subject windows — one behind-subject window per MEDIA; to render two different behind windows of one source, import it as two separate MEDIA ITEMS (distinct mediaIds) — duplicating the same mediaId into two clips still throws), `BehindMatteUnresolved`/`TrajectoryUnresolved` (the behind/face bake found no detectable person — drop the intent or fix the source).
|
|
534
|
-
|
|
535
|
-
**Worked alignment example (the error-prone mechanic).** You want the caption emphasis to land on the founder's stressed word "faster," and the frame to punch in on the same beat. From `get_media_context`, the transcript word "faster" runs `startMs:8120, endMs:8560` (transcript/source time). **For these round numbers, assume an untrimmed clip placed at timeline 0** — then source time, timeline time, and clip-relative time all coincide. So: put that word in `captions.segments[].words[]` with `emphasis:true` at `startMs:8120`; and author reframe with `setCropIntents` as a **contiguous segment list that tiles the whole clip** — e.g. a wide segment `{startSec:0, endSec:8.12, focus:{mode:"active-speaker"}, zoom:1}` then the punch `{startSec:8.12, endSec:<clipEnd>, focus:{mode:"face",faceId:"…",shotScale:"close"}, zoom:0.8, ease:{in:0.4,out:0}}` (ms → s: 8120ms = 8.12s). A lone mid-clip segment that leaves `[0..8.12]` uncovered fails validation with `reframe_coverage_gap`.
|
|
536
|
-
|
|
537
|
-
**Watch the timebase: reframe segment times are CLIP-RELATIVE; caption/transcript times are TIMELINE-based — convert when the clip is trimmed or not at t=0.** They coincide only for an untrimmed clip placed at timeline 0 (the case above). For a clip trimmed to start at source `trim.start` and sitting at `clip.startTime` on the timeline, the reframe boundary is `startSec = transcriptSec − trim.start`, and the caption time is `clip.startTime + (transcriptSec − trim.start)`. (The simple case just has `trim.start = 0` and `clip.startTime = 0`, so both reduce to `transcriptSec`.) Land the caption entrance *on* the stressed word, not a round frame like 8.0s — on-beat reads as craft. And if the clip's `trim` window doesn't actually contain source-second 8.12, the captioned word shows different footage → `content_anchor` at render.
|
|
538
|
-
|
|
539
|
-
---
|
|
540
|
-
|
|
541
|
-
## Craft principles (reason FROM these — the rules above are their consequences)
|
|
542
|
-
|
|
543
|
-
- **Curiosity is an information gap** (Loewenstein, 1994). A slide that *states* a fact satiates before it primes; a cold open that *poses* ("…and that's the number that made us kill the old pipeline") opens a gap the viewer must watch to close. Order clips to pose, not answer. Never lead with a definitional title card.
|
|
544
|
-
- **Emotion outweighs everything** (Murch's Rule of Six: emotion ~51%, then story, rhythm, eye-trace, planarity, spatial continuity). The slop instinct optimizes the bottom of the list — aligned text, smooth easing — while ignoring the top. When choosing where to cut, ask "does this land the emotion / advance the point?" first, "is it geometrically neat?" last.
|
|
545
|
-
- **Attention habituates; pattern interrupts reset it** (orienting response). Constant motion is itself a flat pattern the brain adapts to. Plan a *motivated* change of state every few seconds; vary shot scale between clips.
|
|
546
|
-
- **Sound is half the picture** (Chion, *Audio-Vision*; synchresis). Cut on stressed words / beats via transcript word-timing; hold a silence beat before the peak line. On-beat reads as craft, off-beat as accident.
|
|
547
|
-
- **Story is a value-shift** (McKee). Even a 30s piece must move from pain (real) → product (real) → win. A launch with no reversal is a list, not a story. (Ira Glass's storytelling model says it well another way: build a *sequence of raised-and-answered questions* — which is exactly "pose, don't answer.")
|
|
548
|
-
- **Premium = legible intent, not polish.** Audiences forgive imperfect video with a point of view; they don't forgive adequate video with no soul. Restraint, typographic discipline, and one consistent color/type/motion system across the whole piece are the premium signal. If you can't say *why* something moves, it shouldn't.
|
|
549
|
-
|
|
550
|
-
### Deconstructed grammars (steal the structure, map to the tools)
|
|
551
|
-
These are *archetypes*, not documented cuts of specific films — reason from the grammar, don't cargo-cult a company.
|
|
552
|
-
|
|
553
|
-
- **The hardware-reveal grammar → "the product is the opening shot."** Open on black; the product materializes — no logo splash, no title card. One capability per shot. Numbers rationed, each on its own held frame. Negative space around one object = confidence. *CueFrame:* `import_media` real product footage → `preview_frame` the opening frame and ask "would this survive with zero text?" · `cinematic-title` for the *closing* wordmark only · `color-grade` for the hero look.
|
|
554
|
-
- **The founder-to-camera grammar → CueFrame's home field (the shape that agent walked past).** Founder, medium shot, direct address: "Hey, I want to show you something." No sting, just a face and eye contact. Screen-share woven in as real capture under the continuing VO. Captions burned in, plated. Lower-third names the founder once. Imperfect on purpose — that's the credibility. *CueFrame:* `get_media_context` → `active-speaker`/`all-faces` reframe → transcript-timed captions with an `emphasisPlate` → a `zPlane:"behind-subject"` title → `lower-third` → `color-grade`. Zero generation.
|
|
555
|
-
- **The dark-kinetic-product grammar → "a gradient is a stage, not a subject."** Hard cut in on the UI in motion. Feature montage ~2–4s each, every beat a *real* UI interaction, cut on the music. The UI floats on a near-black stage with a subtle gradient — but the gradient is the *stage*; the real UI is always on top. One hero feature gets 8–10s to breathe. *CueFrame:* `import_media` UI captures → tight per-clip `trim`, cut on beat via clip durations → a dark `color-grade` + a `backgrounds` primitive as the *stage*, real footage on top. `preview_frame` a mid-montage frame: if there's no product on it, it's slop.
|
|
556
|
-
|
|
557
|
-
*The compressed lessons:* the product/human is the opening shot, not a title card · a real screen recording beats a beautiful fake · a gradient is a stage, not a subject · open on your single most satisfying real moment in the first 5s · founder-to-camera + real footage is home field · a purposeful diagram of the real mechanism is grounded too · show one task completed end-to-end, not a feature grid.
|
|
558
|
-
|
|
559
|
-
---
|
|
560
|
-
|
|
561
|
-
## The rubric template (write your video-specific version BEFORE rendering)
|
|
562
|
-
|
|
563
|
-
```
|
|
564
|
-
VISION (one line): the ONE controlling idea for THIS video — a value + a cause.
|
|
565
|
-
FORMAT: [product-launch | social-clip | explainer | demo | talking-head]
|
|
566
|
-
SOURCE REALITY CHECK: what GROUNDS this video (after HUNTING for footage — 3a)? Real
|
|
567
|
-
footage, real UI/data, or a purposeful content-specific graphic (e.g. an authored diagram
|
|
568
|
-
of the actual mechanism). If the only plan is generic filler, the 3b stop fired and
|
|
569
|
-
the user has explicitly said "proceed."
|
|
570
|
-
|
|
571
|
-
MUST-HAVES (each scoreable YES/NO on a NAMED frame, or 0–10):
|
|
572
|
-
H1. [t=…] …
|
|
573
|
-
... (5–8 lines; anchor every one to a timestamp a PNG can answer)
|
|
574
|
-
|
|
575
|
-
KILL CRITERIA (any TRUE on any sampled frame = FAIL, regardless of score):
|
|
576
|
-
K1. [MANDATORY — preserve the meaning] Any sampled frame is a gradient/solid background with only
|
|
577
|
-
text and NO grounded subject — no real footage, no real UI/data, no purposeful
|
|
578
|
-
content-specific graphic, just decorative filler.
|
|
579
|
-
K2. [MANDATORY — preserve the meaning] Any on-screen number you cannot trace to a written
|
|
580
|
-
source, or any faked/mocked UI standing in for a real capture.
|
|
581
|
-
K3. [MANDATORY — preserve the meaning] Generic/decorative filler dominates the runtime — grounded,
|
|
582
|
-
specific content (real footage, real UI/data, or purposeful content-specific
|
|
583
|
-
graphics) does not.
|
|
584
|
-
K4+. … (add your own; encode this video's specific slop modes as disqualifiers)
|
|
585
|
-
|
|
586
|
-
PROVENANCE: for every on-screen number, write the source next to it here
|
|
587
|
-
(e.g. "8.5% — benchmark table, <link to the writeup>"). No written source → it does not go on screen.
|
|
588
|
-
|
|
589
|
-
PER-BEAT CHECKS: B1 (0:00–0:0X) … B2 … (each beat exists and lands, with a duration)
|
|
590
|
-
|
|
591
|
-
MEASURABLE THRESHOLDS:
|
|
592
|
-
T1. First grounded frame (real footage / real UI/data / purposeful graphic) at t ≤ __s.
|
|
593
|
-
T2. Grounded, content-specific visuals (real footage, real UI/data, or a visual of the
|
|
594
|
-
real mechanism/data) ≥ 60% of runtime; a beat whose only visual is animated text
|
|
595
|
-
over a background counts as filler, not grounded. (Footage-based formats: real
|
|
596
|
-
footage ≥ 60%.)
|
|
597
|
-
T3. ≤ __ on-screen stats, each held ≥ __s, each source-traceable.
|
|
598
|
-
T4. No single frame > __% empty decorative background.
|
|
599
|
-
T5. Caption legible (contrast/size, plate) on named frames.
|
|
600
|
-
T6. Cut/beat count in range __–__ (not a static hold, not a strobe).
|
|
601
|
-
|
|
602
|
-
SCORING: (a) your own frame-by-frame descriptions pass every MUST-HAVE with ZERO kill
|
|
603
|
-
criteria; AND (b) score_composition (editorialIntent = the VISION line;
|
|
604
|
-
brandContext = brand facts) composite ≥ 7.5/10 on its OWN sampled frames
|
|
605
|
-
(never set your bar below 7.5). Sample frames at [list incl t=0]. Budget 3–5 passes.
|
|
606
|
-
```
|
|
607
|
-
|
|
608
|
-
### Filled example — product launch (the anti-slop rubric)
|
|
609
|
-
```
|
|
610
|
-
VISION: "Two founders who clearly believe in the product tell you, to your face, why it's
|
|
611
|
-
faster — and you watch it BE faster."
|
|
612
|
-
FORMAT: talking-head + real demo (NOT slides).
|
|
613
|
-
SOURCE REALITY CHECK: HUNT (3a) → the team's real launch footage = people to camera → import
|
|
614
|
-
it. If no footage: a real capture of the product running, or a first-class authored diagram of
|
|
615
|
-
its actual pipeline (grounded, not a consolation) — NOT a gradient deck. Generic
|
|
616
|
-
filler = fail.
|
|
617
|
-
|
|
618
|
-
MUST-HAVES:
|
|
619
|
-
H1. [t=0–3s] A human FACE on screen within 3s (subject present, headroom).
|
|
620
|
-
H2. [t=0–3s] NO title card / logo splash before the face.
|
|
621
|
-
H3. [any speaking frame] The person framed is the person talking (reframe to
|
|
622
|
-
get_media_context faces; all-faces on the two-shot, active-speaker on singles).
|
|
623
|
-
H4. [speaking frames] Legible transcript captions, plated, never over the face.
|
|
624
|
-
H5. [one frame] Speaker name in a lower-third, shown once, spelled correctly.
|
|
625
|
-
H6. [demo beat] The real product running (captured pixels), not a re-created UI card.
|
|
626
|
-
H7. [stat beat] Any number matches the source exactly (source says 8.5% → the frame
|
|
627
|
-
says 8.5%, not 10x).
|
|
628
|
-
H8. [final] Wordmark + one line lands LAST, held ≥1.5s.
|
|
629
|
-
|
|
630
|
-
KILL CRITERIA:
|
|
631
|
-
K1. [mandatory] Gradient/solid + only text, no grounded subject (no real footage,
|
|
632
|
-
real UI/data, or purposeful content-specific graphic).
|
|
633
|
-
K2. [mandatory] Untraceable number (e.g. "10x") or a fake/static terminal.
|
|
634
|
-
K3. [mandatory] Generic/decorative filler dominates; grounded content does not.
|
|
635
|
-
K4. On >2 consecutive sampled beats, the frame's largest element is text on a
|
|
636
|
-
flat/gradient field with no grounded subject behind or beside it.
|
|
637
|
-
|
|
638
|
-
PROVENANCE: "8.5%" — the benchmark table in the sourced writeup. (No other on-screen numbers.)
|
|
639
|
-
|
|
640
|
-
PER-BEAT: B1(0:00–0:05) cold open on a speaker, direct address, caption running ·
|
|
641
|
-
B2(0:05–0:20) the "why," problem in one sentence, on the face · B3(0:20–0:35) real
|
|
642
|
-
product demo, the one satisfying moment, captioned · B4(0:35–0:45) the honest 8.5% on a
|
|
643
|
-
clean held frame · B5(0:45–0:50) face again + CTA + wordmark.
|
|
644
|
-
|
|
645
|
-
THRESHOLDS: T1 first grounded frame ≤ 3s · T2 real footage ≥ 60% of runtime · T3 ≤ 2
|
|
646
|
-
stats, each held ≥ 1.5s, source-traceable · T4 no frame > 30% empty background ·
|
|
647
|
-
T5 caption plate legible on B1,B3 · T6 ≥ 1 active-speaker/all-faces reframe.
|
|
648
|
-
|
|
649
|
-
SCORING: own descriptions pass H1–H8, zero kills; score_composition (editorialIntent
|
|
650
|
-
= VISION) composite ≥ 7.5/10 on its sampled frames. Sample at 0,3,10,20,35,45s.
|
|
651
|
-
Budget 3–5 passes.
|
|
652
|
-
```
|
|
653
|
-
|
|
654
|
-
### Filled example (compact) — social clip (9:16 from long-form)
|
|
655
|
-
```
|
|
656
|
-
VISION: "In 20s, one surprising claim from the talk makes you stop scrolling."
|
|
657
|
-
FORMAT: social-clip (9:16, sound-off).
|
|
658
|
-
SOURCE: a real talking-head/podcast clip (the whole point of the clipping path).
|
|
659
|
-
MUST-HAVES: H1[0–1.5s] the single most surprising sentence is the FIRST thing said
|
|
660
|
-
(align to transcript) · H2[every frame] face in the vertical safe area, top of head
|
|
661
|
-
in-frame · H3[speaking] word-level animated captions, high-contrast, plated · H4
|
|
662
|
-
captions never cover mouth/eyes · H5[last 2s] soft loop or one-line handle.
|
|
663
|
-
KILL: K1[mandatory] gradient+text, no grounded subject · K2[mandatory] untraceable
|
|
664
|
-
number/fake UI · K3[mandatory] generic filler dominates, grounded content doesn't ·
|
|
665
|
-
K4 face cropped at forehead/chin OR vertical frame mostly empty · K5 opens on filler
|
|
666
|
-
("um, so…") not the hook.
|
|
667
|
-
THRESHOLDS: 15–30s · hook ≤ 1.5s · caption height ≥ ~6% of frame · face box ≥ 25% of
|
|
668
|
-
frame area · real footage ≥ 60% · nothing critical in bottom ~12% or top ~10%.
|
|
669
|
-
SCORING: own descriptions pass, zero kills; score_composition (editorialIntent =
|
|
670
|
-
VISION) composite ≥ 7.5/10 on its sampled frames. Sample 0,1.5,5,10,18s. Budget 3–5 passes.
|
|
671
|
-
```
|
|
672
|
-
|
|
673
|
-
Every line names a frame and asks a question a PNG can answer — is a grounded subject present, is text legible, is the background mostly empty decoration, does the number match its written source, is the hook first. None require reading the composition JSON. The kill criteria *are* the enforcement: K1–K3 encode that trial's failure modes as automatic disqualifiers a render can't be rationalized past.
|
|
674
|
-
|
|
675
|
-
---
|
|
676
|
-
|
|
677
|
-
## Red flags — stop and fix if any is true
|
|
678
|
-
|
|
679
|
-
- **You concluded "no footage" without hunting** → run step 3a (user-attached? subject's existing footage? ask). That agent's whole failure was skipping this — footage is the strongest grounding.
|
|
680
|
-
- **Generic/decorative filler dominates the runtime** → grounded, specific content (real footage, real UI/data, or a purposeful content-specific graphic) doesn't. K3 trips. (For footage-based formats, real footage under ~60% is the usual symptom.)
|
|
681
|
-
- **An on-screen number the user didn't give you** → you fabricated it. CueFrame renders your number faithfully, so a wrong one is on you — write every stat's source in the rubric; no written source → it doesn't go on screen.
|
|
682
|
-
- **A terminal/UI that's a styled component, not a real capture** → it's a mock. Import the real capture or drop it.
|
|
683
|
-
- **A gradient/aurora you're calling "purposeful"** → it isn't. Purposeful means specific to THIS content (the real mechanism, the real data). A decorative gradient is always the slop default, never grounding.
|
|
684
|
-
- **You rendered before writing the rubric** → stop, write the rubric (step 2), then look at pixels.
|
|
685
|
-
- **You scored without describing the pixels** → you didn't look. Write one line of what's visible per still, then score.
|
|
686
|
-
- **You're the sole judge** → run `score_composition` (editorialIntent = vision); it samples its own frames — treat a low score on a frame you didn't pick as ground truth.
|
|
687
|
-
|
|
688
|
-
---
|
|
689
|
-
|
|
690
|
-
## Pre-render checklist (before `create_render`)
|
|
691
|
-
|
|
692
|
-
0. `create_project` done; `get_account` shows enough `balance` on the wallet funding render/compose/generate (and the generate cost ceiling clears). Judge inclusion by balance, not by `included`.
|
|
693
|
-
1. Vision written as one sentence (value + cause); format chosen.
|
|
694
|
-
2. Footage HUNTED (3a); if the plan is generic filler, the 3b tradeoff stated to the user and an explicit "proceed" received (a purposeful, specific graphic — e.g. an authored diagram of the real mechanism — does NOT require the stop).
|
|
695
|
-
3. Rubric written *before* any render — named-frame must-haves, measurable thresholds (grounded content dominates the runtime; footage-based formats real-footage ≥60%), the mandatory K1–K3 (meaning preserved, phrasing adapted to your video), and written provenance for every number.
|
|
696
|
-
4. Real source imported (`list_media` shows it ready); `get_media_context` faces + transcript pulled and used for reframe + caption timing (when the piece uses footage).
|
|
697
|
-
5. Components discovered via `list_catalog` (no hand-placed hex where a built-in exists); `fixedCopy` checked so no baked text lies. Brand set via `brandKitId` on the project if a brand kit exists.
|
|
698
|
-
6. `validate_composition` is `{ valid: [] }` — shape sound, no `reframe_coverage_gap` (shape ≠ quality).
|
|
699
|
-
7. `preview_frame` stills taken at every rubric timestamp incl t=0, each **described in words** then scored against the rubric.
|
|
700
|
-
8. `score_composition` run with `editorialIntent`=vision, `brandContext`=brand; worst-first critique addressed; composite ≥ 7.5 on its own sampled frames.
|
|
701
|
-
9. Every kill criterion false on every sampled frame; every on-screen number traces to its written source and matches it.
|
|
702
|
-
10. Verify passes budgeted (3–5) and the rubric genuinely passes — then `create_render` → `wait_job(kind:"render")`.
|