@cueframe/skills 0.1.0 → 0.1.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +9 -104
- package/{skills → dist/skills}/add-music-bed/SKILL.md +12 -7
- package/{skills → dist/skills}/brand-reel/SKILL.md +19 -15
- package/{skills → dist/skills}/clip-a-talking-head/SKILL.md +15 -7
- package/dist/skills/composing-video/SKILL.md +447 -0
- package/{skills → dist/skills}/cueframe-brand-demo/SKILL.md +9 -2
- package/{skills → dist/skills}/cueframe-cli/SKILL.md +54 -49
- package/{skills → dist/skills}/cueframe-component-authoring/SKILL.md +14 -6
- package/{skills → dist/skills}/cueframe-compose-loop/SKILL.md +20 -6
- package/dist/skills/cueframe-compose-loop/references/builtin-insertion.md +49 -0
- package/dist/skills/cueframe-compose-loop/references/builtin-requests.json +111 -0
- package/dist/skills/cueframe-connect/SKILL.md +53 -0
- package/{skills → dist/skills}/cueframe-product-video/SKILL.md +15 -7
- package/{skills → dist/skills}/cueframe-scene-shot/SKILL.md +8 -0
- package/dist/skills/cueframe-storyboard/SKILL.md +76 -0
- package/{skills → dist/skills}/every-format-from-one-edit/SKILL.md +14 -8
- package/{skills → dist/skills}/extracting-brand-kits/SKILL.md +8 -0
- package/{skills → dist/skills}/launch-video/SKILL.md +17 -11
- package/{skills → dist/skills}/make-a-social-reel/SKILL.md +19 -15
- package/{skills → dist/skills}/rebrand-a-video/SKILL.md +18 -16
- package/dist/skills/video-craft-standards/SKILL.md +121 -0
- package/dist/skills/video-craft-standards/agents/openai.yaml +6 -0
- package/package.json +9 -28
- package/skills-dir.d.ts +1 -0
- package/skills-dir.js +2 -1
- package/.agents/plugins/marketplace.json +0 -12
- package/.claude-plugin/marketplace.json +0 -6
- package/.claude-plugin/plugin.json +0 -15
- package/.codex-plugin/plugin.json +0 -30
- package/.cursor-plugin/plugin.json +0 -1
- package/.mcp.json +0 -1
- package/AGENTS.md +0 -20
- package/assets/logo-400.png +0 -0
- package/gemini-extension.json +0 -1
- package/glama.json +0 -1
- package/hooks/hooks.json +0 -7
- package/hooks/session-inject.md +0 -15
- package/hooks/session-start.sh +0 -6
- package/llms-install.md +0 -47
- package/mcp.json +0 -1
- package/plugin.json +0 -46
- package/rules/cueframe.mdc +0 -19
- package/skills/composing-video/SKILL.md +0 -702
- package/skills/cueframe-connect/SKILL.md +0 -45
- package/skills/cueframe-storyboard/SKILL.md +0 -104
- package/skills/video-craft-standards/SKILL.md +0 -128
- package/skills/video-craft-standards/agents/openai.yaml +0 -6
- package/skills/video-craft-standards/assets/icon.svg +0 -16
- package/skills.sh.json +0 -1
- /package/{skills → dist/skills}/add-music-bed/agents/openai.yaml +0 -0
- /package/{assets → dist/skills/add-music-bed/assets}/icon.svg +0 -0
- /package/{skills → dist/skills}/brand-reel/agents/openai.yaml +0 -0
- /package/{skills/add-music-bed → dist/skills/brand-reel}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/clip-a-talking-head/agents/openai.yaml +0 -0
- /package/{skills/brand-reel → dist/skills/clip-a-talking-head}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/composing-video/agents/openai.yaml +0 -0
- /package/{skills/clip-a-talking-head → dist/skills/composing-video}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-brand-demo/agents/openai.yaml +0 -0
- /package/{skills/composing-video → dist/skills/cueframe-brand-demo}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-cli/agents/openai.yaml +0 -0
- /package/{skills/cueframe-brand-demo → dist/skills/cueframe-cli}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-component-authoring/agents/openai.yaml +0 -0
- /package/{skills/cueframe-cli → dist/skills/cueframe-component-authoring}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-compose-loop/agents/openai.yaml +0 -0
- /package/{skills/cueframe-component-authoring → dist/skills/cueframe-compose-loop}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-compose-loop/references/preview-workflow.md +0 -0
- /package/{skills → dist/skills}/cueframe-connect/agents/openai.yaml +0 -0
- /package/{skills/cueframe-compose-loop → dist/skills/cueframe-connect}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-product-video/agents/openai.yaml +0 -0
- /package/{skills/cueframe-connect → dist/skills/cueframe-product-video}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-scene-shot/agents/openai.yaml +0 -0
- /package/{skills/cueframe-product-video → dist/skills/cueframe-scene-shot}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/cueframe-storyboard/agents/openai.yaml +0 -0
- /package/{skills/cueframe-scene-shot → dist/skills/cueframe-storyboard}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/every-format-from-one-edit/agents/openai.yaml +0 -0
- /package/{skills/cueframe-storyboard → dist/skills/every-format-from-one-edit}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/extracting-brand-kits/agents/openai.yaml +0 -0
- /package/{skills/every-format-from-one-edit → dist/skills/extracting-brand-kits}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/launch-video/agents/openai.yaml +0 -0
- /package/{skills/extracting-brand-kits → dist/skills/launch-video}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/make-a-social-reel/agents/openai.yaml +0 -0
- /package/{skills/launch-video → dist/skills/make-a-social-reel}/assets/icon.svg +0 -0
- /package/{skills → dist/skills}/rebrand-a-video/agents/openai.yaml +0 -0
- /package/{skills/make-a-social-reel → dist/skills/rebrand-a-video}/assets/icon.svg +0 -0
- /package/{skills/rebrand-a-video → dist/skills/video-craft-standards}/assets/icon.svg +0 -0
|
@@ -0,0 +1,447 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: composing-video
|
|
3
|
+
description: Use when an agent is asked to make ANY video with CueFrame — a product launch, announcement, promo, teaser, social clip, reel, short, explainer, animated explainer, concept video, product/dev-tool demo, walkthrough, or founder/talking-head piece. Triggers on "make a video", "launch video", "promo", "teaser", "reel", "short", "explainer", "animated explainer", "demo video", "talking-head", or turning long-form footage into clips.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Composing Video: craft for CueFrame
|
|
7
|
+
|
|
8
|
+
## Scope and capability policy
|
|
9
|
+
|
|
10
|
+
User intent and acceptance criteria override this recipe's aesthetic defaults. Silent, typography-only, uncaptioned and slow work are valid; no universal narration, footage, CTA, beat count or score target. Contracts, factual fidelity, media bindings and current source revisions remain mandatory.
|
|
11
|
+
|
|
12
|
+
Read available tools and current project; preserve the chosen local/hosted lane. Supported local work needs no hosted registration. Do not silently install, authenticate, switch projects or upload local content. Before metered hosted work, check `get_account`, entitlements, balance and quotes. Desktop setup: https://docs.cueframe.ai.
|
|
13
|
+
|
|
14
|
+
For scoped revisions, preserve unrelated timing, media, audio and shared brand-kit state. Edit affected content; inspect relevant frames, continuous motion and affected audio. No restarted intake, new approval pause, paid judge or final export by default. Judge only when useful; scores never override acceptance failures. Stop/report after two no-improvement passes. See [preview workflow](../cueframe-compose-loop/references/preview-workflow.md).
|
|
15
|
+
|
|
16
|
+
For preview selection and small revisions, use `cueframe-compose-loop`'s
|
|
17
|
+
[preview workflow](../cueframe-compose-loop/references/preview-workflow.md). Retain the approved
|
|
18
|
+
brief/rubric for a local correction; check the affected behavior and relevant regressions instead
|
|
19
|
+
of restarting the new-film method below. A diagnostic or preview-only request does not require
|
|
20
|
+
a final export or a whole-film judge pass. Explicit user scope takes precedence over craft defaults.
|
|
21
|
+
|
|
22
|
+
## Plan and author to the request
|
|
23
|
+
|
|
24
|
+
Use `cueframe-storyboard` only when new work lacks structure. For an existing title correction, retain the approved brief and edit the title directly; no whole-film rubric, paid judge or final export is required.
|
|
25
|
+
|
|
26
|
+
Choose content that serves the requested idea. Real footage, actual product captures, sourced data, diagrams, abstract motion and typography are all valid. Do not fabricate facts or present simulated UI as a real product capture. Search for footage when the request needs it; an explicit typography-only request needs no footage hunt or permission to remain typography-only.
|
|
27
|
+
|
|
28
|
+
For new work, state concise acceptance criteria from the brief before authoring. For broader review, reuse the approved criteria. Preview affected frames, inspect motion in the relevant window and listen when sound changes. Validate contracts and source/media bindings independently of aesthetic review. For face/active-speaker framing or behind-subject graphics, prepare the required subjectTrack/matte for the exact source trim window and wait for ready evidence; unresolved perception is not acceptable preview proof.
|
|
29
|
+
|
|
30
|
+
Use `score_composition` only when its broader editorial critique is useful, available and within budget. The hosted judge has editorial, spatial, brand and caption criteria; `editorialIntent` does not remove fixed criteria. A caption penalty for intentionally uncaptioned or silent work is a judge limitation, not a reason to add captions or declare user intent failed. Skill text does not repair the judge. Stop after two no-improvement passes and report gaps.
|
|
31
|
+
|
|
32
|
+
## Choose the film spine from the leading material
|
|
33
|
+
|
|
34
|
+
- Narrated explainer: draft the script, align visuals to its phrases, verify intelligibility and any requested captions against actual word timings.
|
|
35
|
+
- Recorded interview: preserve meaning and source order where required; cut and frame from the transcript and actual speaker evidence, with no invented narration.
|
|
36
|
+
- Product action: plan around real captured interactions and results; add voice or music only when requested or appropriate.
|
|
37
|
+
- Silent brand study or typography: plan visual rhythm, reading time and transitions directly. Silence is already decided by the request; no audio generation or footage requirement.
|
|
38
|
+
- Visual/music motion or observational sequence: use the chosen motion, music, activity or deliberate holds as the timing authority. No universal beat count, cut interval or CTA.
|
|
39
|
+
|
|
40
|
+
Use `cueframe-component-authoring` for custom 2D motion. Load `cueframe-scene-shot` for supported device, camera or 3D tasks. `cueframe-compose-loop` owns scoped preview and optional review; `video-craft-standards` supplies aesthetic defaults under the scope policy above.
|
|
41
|
+
|
|
42
|
+
### The card token contract — `tokens` key `x` is read as `var(--cf-x)`
|
|
43
|
+
|
|
44
|
+
A `{kind:'card', html, tokens}` clip binds brand-swappable values through `tokens`. Each
|
|
45
|
+
**key** is projected onto the card's wrapper as a CSS custom property with a `--cf-`
|
|
46
|
+
prefix, so the html reads it back **prefixed**:
|
|
47
|
+
|
|
48
|
+
```jsonc
|
|
49
|
+
{ "kind": "card", "tokens": { "plate": "#0af", "fg": "$brand:colors.text" },
|
|
50
|
+
"html": "<div style=\"background:var(--cf-plate);color:var(--cf-fg)\">…</div>" }
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Values may be literals or whole-string `$brand:<path>` refs (resolved against the
|
|
54
|
+
project's brand kit at render).
|
|
55
|
+
|
|
56
|
+
**Do not put dashes in the key.** `tokens: {"--plate": …}` projects to `--cf---plate`, so
|
|
57
|
+
the card's `var(--plate)` matches nothing — and CSS treats an undefined custom property as
|
|
58
|
+
*invalid at computed-value time*, which means it does not error, it **degrades**:
|
|
59
|
+
`background` falls back to transparent and `color` inherits (usually black). The render
|
|
60
|
+
succeeds and the graphic is unreadable. `apply_composition` returns `warnings[]` when it
|
|
61
|
+
catches this, but the fix is to never write the dashes.
|
|
62
|
+
|
|
63
|
+
Two `--cf-*` families share the wrapper: **your token vars** (above) and the **motion
|
|
64
|
+
vars** below, which are always injected and are not yours to declare.
|
|
65
|
+
|
|
66
|
+
### The motion runtime — three ways anything you author can move
|
|
67
|
+
|
|
68
|
+
CSS `@keyframes`/`transition` **never animate in a render** (frames are independent
|
|
69
|
+
Chromium captures; the clock never advances). Use these instead:
|
|
70
|
+
|
|
71
|
+
1. **Card motion vars (`--cf-*`)** — every card's wrapper carries frame-driven, unitless
|
|
72
|
+
custom properties: `--cf-progress` (0→1 across the clip), `--cf-ease` (eased progress),
|
|
73
|
+
`--cf-enter` (0→1 over the first ~0.6s), `--cf-exit` (1→0 over the last ~0.6s), plus
|
|
74
|
+
`--cf-frame` / `--cf-t` / `--cf-fps` / `--cf-duration`. Author motion directly in card
|
|
75
|
+
CSS: `opacity: var(--cf-enter)`, `transform: translateY(calc((1 - var(--cf-enter)) * 40px))`,
|
|
76
|
+
an exit slide via `--cf-exit`. This supports entrance/exit choreography within the
|
|
77
|
+
card medium — and a `preview_frame` still at time *t* shows the exact *t*-frozen state.
|
|
78
|
+
2. **`transitionIn` on any overlay clip** — `{ transitionIn: { type: "fade", duration: 0.4 } }`
|
|
79
|
+
now applies to overlay-family clips (cards, components, primitives), not just video.
|
|
80
|
+
Use it for entrances; author exits in-card via `--cf-exit`.
|
|
81
|
+
3. **`applyMotionPreset` op** — shell a static card into a frame-driven component in one op:
|
|
82
|
+
`{ type: "applyMotionPreset", clipId, preset }` (unknown presets are rejected with the
|
|
83
|
+
full preset list). The layout stays verbatim; the preset adds entrance/idle/exit
|
|
84
|
+
choreography. The result is a component you can `get_component_source` and refine.
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
## SFX pack — the curated CC0 sound palette
|
|
89
|
+
|
|
90
|
+
CueFrame ships a **closed, versioned pack of CC0 sound effects** (40 sounds, all
|
|
91
|
+
peak-normalized). There are two ways to place them. **Automatically:** the compose
|
|
92
|
+
SFX pass binds them **deterministically** — a hit on every eligible graphic
|
|
93
|
+
entrance (family from the primitive: `impact` for hero/title, `pop` for
|
|
94
|
+
stat/value, `tick` for enumeration items), a `whoosh` on video transitions, and a
|
|
95
|
+
`riser` ending at the hero entrance — timed to the same frames as the overlay,
|
|
96
|
+
ducked under speech. **By hand:** a client author places any pack sound directly —
|
|
97
|
+
`import_resource { kind:"sfx-pack", id:"<soundId>" }` returns the org's media
|
|
98
|
+
receipt (`{ id, status:"ready" }`) for that sound, which you then drop in as a
|
|
99
|
+
normal sfx clip (`mediaId`). Idempotent per (org, soundId); an unknown soundId is
|
|
100
|
+
a loud 422. Same curated bytes either path.
|
|
101
|
+
Restraint (density caps, min spacing, speech collision) comes from the AUDIO_MIX
|
|
102
|
+
SSOT the eval floors grade with, so placement and grading can't disagree.
|
|
103
|
+
|
|
104
|
+
**Optional audio review:** the `sfx-same-sound-reuse` craft floor can report
|
|
105
|
+
when a dense mix (≥6 SFX hits) draws from too small a palette — distinct sounds ÷
|
|
106
|
+
hits below **0.4**. The pack carries **≥4 distinct sounds per transient family**
|
|
107
|
+
to support a varied palette. This is reported judge behavior, not a universal creative requirement: preserve intentional repetition or silence and assess it against the brief. Browse the full palette structurally on
|
|
108
|
+
`list_catalog` (the `sfx` array — `soundId`, `family`, `envelopeClass`,
|
|
109
|
+
`description`, `durationSec`).
|
|
110
|
+
|
|
111
|
+
**One palette, org sounds first.** `list_catalog` also returns an `audio` array —
|
|
112
|
+
the ONE unified sound palette: the curated packs (SFX one-shots **and** CC0 music
|
|
113
|
+
beds) unioned with THIS org's own uploads/imports (`source:"org"`) and generated
|
|
114
|
+
audio (`source:"generated"`), each entry a uniform `{ref, role, envelopeClass,
|
|
115
|
+
description, durationSec, evidence, source, license, provenance}` (music entries
|
|
116
|
+
carry `bpm`). Provenance rides every entry — pack attribution, `org-owned`, or the
|
|
117
|
+
generator + prompt. **Prefer an org sound when one fits the family/envelope class:**
|
|
118
|
+
an org that uploads its product's real UI sounds hears ITS product, not a generic
|
|
119
|
+
pack tick — the compose SFX pass does this automatically (org sound of the matching
|
|
120
|
+
envelope class beats the pack default, exactly like a brand kit overrides the
|
|
121
|
+
default font), and by hand you place an org `ref` (a `mediaId`) directly. **The
|
|
122
|
+
music tier is curated beds + a generate escape hatch:** browse the `source:"pack"`
|
|
123
|
+
`role:"music"` beds (each with `bpm` + mood, spanning slow/mid/fast) and place one
|
|
124
|
+
via `import_resource { kind:"music-pack", id:"<bedId>" }` (FREE, like sfx-pack); for
|
|
125
|
+
a bed the pack lacks, `generate_media { generator:"text-to-music" }`. Buy/curate SFX,
|
|
126
|
+
curate-or-generate music.
|
|
127
|
+
|
|
128
|
+
**Transient one-shots** (entrance hits, per-element ticks):
|
|
129
|
+
|
|
130
|
+
| family | soundIds (subtle→strong) |
|
|
131
|
+
|---|---|
|
|
132
|
+
| `tick` | `tick-soft` · `tick-toggle` · `tick-scroll` · `tick-select` · `tick-click` · `tick-mouse` · `tick-shutter` · `tick-switch` |
|
|
133
|
+
| `pop` | `pop-drop` · `pop-glass` · `pop-pluck` · `pop-drop-low` · `pop-bong` · `pop-confirm` |
|
|
134
|
+
| `impact` | `impact-generic` · `impact-soft` · `impact-tin` · `impact-metal` · `impact-wood` · `impact-firm` · `impact-punch` · `impact-bell` |
|
|
135
|
+
| `confirm` | `confirm-soft` · `confirm-chime` · `confirm-question` · `confirm-bright` |
|
|
136
|
+
| `error` | `error-soft` · `error-buzz` · `error-tone` · `error-hard` |
|
|
137
|
+
|
|
138
|
+
**Sweeps** (transitions + hero lead-in, ~0.15–1.2s):
|
|
139
|
+
|
|
140
|
+
| family | soundIds |
|
|
141
|
+
|---|---|
|
|
142
|
+
| `whoosh` | `whoosh-air` · `whoosh-whip` · `whoosh-page` · `whoosh-sweep` · `whoosh-phaser` · `whoosh-low` |
|
|
143
|
+
| `riser` | `riser-jump` · `riser-phaser` · `riser-sweep` · `riser-power` |
|
|
144
|
+
|
|
145
|
+
(The SFX pack is transient/sweep one-shots; a sustained bed is a MUSIC-role clip —
|
|
146
|
+
a curated `music-pack` bed or a `text-to-music` generation, not a pack hit; see the
|
|
147
|
+
`audio` palette above.) For a sound the pack lacks, search
|
|
148
|
+
CC0 stock at curation time via `search_resources { kind:"sfx" }` (Freesound,
|
|
149
|
+
duration-bounded) and `import_resource` it — the pack is the fast default, not the
|
|
150
|
+
ceiling. For a sound that doesn't exist anywhere, generate it:
|
|
151
|
+
`generate_media { generator: "text-to-sfx", prompt: "short bright metallic tick, fast decay, no tail" }`
|
|
152
|
+
(0.5–22s; the delivered file is envelope-checked before it goes ready).
|
|
153
|
+
|
|
154
|
+
### Audio design — how to USE sounds (canonical doctrine + our calibration)
|
|
155
|
+
|
|
156
|
+
This is established film/UI sound-design craft — Chion, Murch, Thom, and the
|
|
157
|
+
Material sound guidelines — with the numbers calibrated on our own measured films.
|
|
158
|
+
The stakes are empirical, not aesthetic: films with sound effects measure >3x higher
|
|
159
|
+
perceived immersion than without (Kock & Louven), and audio-aware models consistently
|
|
160
|
+
beat vision-only models at predicting short-video engagement (VQualA 2025) — the
|
|
161
|
+
soundtrack is a measurable share of whether anyone keeps watching.
|
|
162
|
+
|
|
163
|
+
0. **Who owns time — pick the rhythm authority FIRST.** Before any hit or bed,
|
|
164
|
+
decide what the cut serves. A film has ONE rhythm authority (word timings, beat
|
|
165
|
+
maps, and shot boundaries are the same TYPE — a temporal grid; pick which one
|
|
166
|
+
drives). **VO-driven** (a narration film): cut every beat boundary to the
|
|
167
|
+
*voice* (author the script first for a narrated piece, then read the word grid
|
|
168
|
+
off `get_media_context` — `wordGrid.wordTimesMs`, the word onsets in ms — and
|
|
169
|
+
land each cut/caption on a word). **Music-driven** (a montage/no-VO piece): cut
|
|
170
|
+
on the *beat* — read `beatGrid` off `get_media_context` (bpm + `beatTimesMs`),
|
|
171
|
+
land cut points and transition MIDPOINTS on beats, and resolve risers exactly on
|
|
172
|
+
a downbeat/strong beat (every 4th beat from the grid start). A grid is EVIDENCE,
|
|
173
|
+
not autopilot — you choose which beats carry cuts; on-beat reads as craft,
|
|
174
|
+
off-beat as accident (synchresis, Chion). Trust `beatGrid` for cutting when its
|
|
175
|
+
`confidence` is high; a beatless bed reports low and carries no reliable grid.
|
|
176
|
+
Once you have picked the authority, **materialize its grid onto the timeline
|
|
177
|
+
with the `materializeGrid` op** (`{clipId, grid:"beat"|"word", maxMarkers?}`):
|
|
178
|
+
the server reads that clip's source grid and lands `kind:"beat"|"word"` markers
|
|
179
|
+
at the correct TIMELINE times (projected through the clip's trim/playbackRate/
|
|
180
|
+
excludedRanges — you never do the media→timeline math yourself). Then place cuts
|
|
181
|
+
and transition **midpoints** on the markers — a transition's midpoint IS the cut
|
|
182
|
+
point, so start a 0.5s transition 0.25s *before* the marker so its center lands
|
|
183
|
+
on the beat. Materialize ONCE per authority clip; re-running the op refreshes the
|
|
184
|
+
markers in place (it replaces the prior set from the same clip+grid), so re-run
|
|
185
|
+
after you retrim or restretch the clip. (`grid:"shot"` is reserved — not exposed
|
|
186
|
+
yet.)
|
|
187
|
+
|
|
188
|
+
0b. **Declare the audio's ROLE at ingest — evidence follows the declaration.**
|
|
189
|
+
Uploads/imports accept an optional `audioRole: vo|music|sfx|ambience` (finalize
|
|
190
|
+
and import bodies). Declaring it routes enrichment: `vo` → transcript + word
|
|
191
|
+
grid; `music` → `beatGrid` (detected, no prompt needed); `sfx`/`ambience` →
|
|
192
|
+
`envelope` evidence on `get_media_context` (`envelopeClass:
|
|
193
|
+
transient|sweep|ambient` + attackMs/crest stats — pick pack-style hits by
|
|
194
|
+
class, not by listening). Undeclared audio is transcribed as before, and a
|
|
195
|
+
no-speech file still comes back envelope-classed. **`ambience` is the fourth
|
|
196
|
+
role**: a bed that grounds the SPACE — it loops, sits at the bottom of the mix,
|
|
197
|
+
and HOLDS under VO (it is never ducked; only `music` ducks under speech). Use
|
|
198
|
+
ambience for room tone/atmosphere continuity across cuts (the room-changes
|
|
199
|
+
floor's natural fix), never for anything that must be *noticed*.
|
|
200
|
+
|
|
201
|
+
0c. **The engine owns the default mix — author levels only to deviate.** A clip
|
|
202
|
+
tagged with an `audioRole` but no authored `volumeDb`/`volume` is auto-staged
|
|
203
|
+
to the clarity hierarchy (VO 0 dB, sfx −6, music −16 under VO else −8,
|
|
204
|
+
ambience −22 — keeping VO ≥6 dB over the bed). An authored level always wins,
|
|
205
|
+
a role-less clip stays at unity, and authored levels that invert the hierarchy
|
|
206
|
+
(e.g. music over VO) raise a non-blocking `mix_hierarchy_inverted` warning on
|
|
207
|
+
`create_render` naming the offending clips. Tag roles and let the engine mix;
|
|
208
|
+
reach for `volumeDb` only when you intend to break the hierarchy.
|
|
209
|
+
|
|
210
|
+
0d. **A brand kit can carry a SOUND identity — set it once, it rides every
|
|
211
|
+
compose.** `create_brand_kit`'s `audio` block (echoed on `get_brand_kit` +
|
|
212
|
+
`get_profile`) holds the brand's sonic taste; every ref is a palette reference
|
|
213
|
+
(a pack `soundId` | `bedId` | an org `mediaId` from `list_catalog`'s `audio[]`).
|
|
214
|
+
Four optional fields, resolution precedence **brand > org > pack** (the audio
|
|
215
|
+
analog of how a brand FONT sits above org/pack defaults):
|
|
216
|
+
- **`sonicLogo` `{ ref, placement: intro|outro|both, offsetMs? }`** — the brand
|
|
217
|
+
signature. When the kit ALSO carries the matching bookend (`intro`/`outro`
|
|
218
|
+
bumper), compose binds an sfx-role clip at the body edge that bookend abuts
|
|
219
|
+
(intro → body start; outro → body end − logo length; `offsetMs` shifts
|
|
220
|
+
within). No matching bumper ⇒ a non-blocking `brand_sonic_logo_no_bookend`
|
|
221
|
+
advisory and nothing placed (the sonic logo is the AUDIO companion to the
|
|
222
|
+
VISUAL bookend). A `ref` that no longer resolves fails loud
|
|
223
|
+
`brand_audio_unresolvable` — re-point it or clear the sonic logo.
|
|
224
|
+
- **`uiSoundSet` `{ family → ref }`** — per-family SFX overrides (the brand's
|
|
225
|
+
own tick/pop/impact/…). Consumed by the compose SFX draw as the HIGHEST tier:
|
|
226
|
+
a family the brand overrides plays the brand's sound, everything else falls to
|
|
227
|
+
the org palette then the pack default.
|
|
228
|
+
- **`bedStyle` `{ bedId? | genre?, bpmRange? }`** and **`ambiencePreference`
|
|
229
|
+
`{ ref } | null`** — SURFACED-ONLY today: compose auto-selects neither a music
|
|
230
|
+
bed nor ambience, so these carry the brand's preference for a DRIVING agent to
|
|
231
|
+
honor when it pins/generates a bed (`brief.audio.music`) or places ambience.
|
|
232
|
+
Extraction can't hear: `extract_brand_kit` (website → colors/fonts) never fills
|
|
233
|
+
the `audio` block — set the sound identity explicitly on `create_brand_kit`.
|
|
234
|
+
|
|
235
|
+
1. **Every hit is a SYNC POINT — sound welded to a visible event.** Chion
|
|
236
|
+
(*Audio-Vision*) calls this **synchresis**: the immediate, involuntary weld the
|
|
237
|
+
brain makes between a sound and the image it lands on — that weld is where
|
|
238
|
+
"added value" comes from, and a hit with no visible cause produces none (it's
|
|
239
|
+
the audio version of unmotivated motion, slop tell #6). Bind hits to their
|
|
240
|
+
events via `anchor: { clipId, offsetMs }` so the weld survives re-edits. And
|
|
241
|
+
per Randy Thom (*Designing a Movie for Sound*): sound is a storyteller, not
|
|
242
|
+
decoration — score the moments that matter, let ordinary cuts breathe.
|
|
243
|
+
2. **Class → moment.** `transient` (tick/pop/impact/confirm/error) = entrances,
|
|
244
|
+
cuts, UI semantics — Material's sound guidelines are the reference for UI
|
|
245
|
+
semantics (each sound expresses its place in the hierarchy; decorative sound
|
|
246
|
+
used sparingly). `sweep`: whoosh = motion carrying ACROSS a cut (pair with the
|
|
247
|
+
exit→entrance handoff); riser = tension INTO a reveal — resolve exactly on the
|
|
248
|
+
downbeat, never into nothing (an unresolved riser is a broken promise).
|
|
249
|
+
3. **Density has a perceptual ceiling.** Murch (*Dense Clarity — Clear Density*):
|
|
250
|
+
the brain tracks only ~two-and-a-half simultaneous sound streams; past that,
|
|
251
|
+
individual sounds stop reading and merge into texture — his mixes convey
|
|
252
|
+
complex scenes with a FEW carefully chosen elements. Our calibration: the
|
|
253
|
+
41.6s film that beat its reference carried 9 hits; the rejected cut carried 15
|
|
254
|
+
from 3 sounds. At 6+ hits optional audio review may report `sfx-same-sound-reuse`
|
|
255
|
+
below 0.4 distinct/hits.
|
|
256
|
+
4. **Hierarchy within a family** (Murch's balanced-spectrum principle + Material's
|
|
257
|
+
sound hierarchy): the pack orders soundIds subtle→strong; the biggest beat gets
|
|
258
|
+
the strong impact ONCE, everything else sits a tier down. Fifteen strong hits
|
|
259
|
+
= zero strong hits.
|
|
260
|
+
5. **Gain staging.** SFX under a music bed must READ: target the hit ~≥6 dB above
|
|
261
|
+
the bed at its transient (our measured failure: ticks 1.5 dB over bed =
|
|
262
|
+
inaudible mush). Working precedent: sfx `volumeDb: -12`, bed lower still.
|
|
263
|
+
Author sane levels and STOP — the engine masters to −14 LUFS (the streaming
|
|
264
|
+
loudness standard); never pre-compress or push peaks to compensate.
|
|
265
|
+
6. **The bed is tempo, not wallpaper.** Match bed BPM/energy to the cut rate (our
|
|
266
|
+
fast recut earned a 124 BPM driving bed; a calm film wants sparser pulse), arc
|
|
267
|
+
the energy with the narrative, and always `duck` under VO
|
|
268
|
+
(`duck: { duckDb, attackMs, releaseMs }`) — dialogue sits atop Murch's
|
|
269
|
+
encoded–embodied spectrum; sfx punctuate BETWEEN phrases, never fight words.
|
|
270
|
+
7. **Silence is a device** (Thom: quiet is the most underrated tool in the
|
|
271
|
+
soundtrack). Clean air before the first hit; a dropout before the biggest
|
|
272
|
+
reveal is worth more than any riser. If the whole film is scored, nothing is.
|
|
273
|
+
|
|
274
|
+
---
|
|
275
|
+
|
|
276
|
+
## Craft defaults
|
|
277
|
+
|
|
278
|
+
Make deliberate choices about hierarchy, type, framing, juxtaposition and motion. For a product demonstration, use real captures and sourced claims; never present a mock UI as documentary evidence or invent a performance number. For an explicit abstract or typography study, judge its reading time, rhythm and visual intent directly: text and gradients are valid material, not automatic failures.
|
|
279
|
+
|
|
280
|
+
Platform conventions can suggest hooks, captions, short cuts and clear closing actions. Apply them only when they serve the requested format. An observational long hold, slow logo-first opening, silent piece or no-CTA ending is not a failure merely because another format prefers something else. Caption speech when requested or appropriate; verify actual word timing. Use audio or visual rhythm as the timing authority selected in the brief.
|
|
281
|
+
|
|
282
|
+
## Format playbooks (examples, not fixed structures)
|
|
283
|
+
|
|
284
|
+
Each playbook is an illustrative format default; explicit intent and scope govern. Every beat maps to a real CueFrame capability. Default to **16:9 landscape** for product/demo/explainer/founder-web; **9:16 vertical** for social cutdowns or other explicitly vertical work (a wide source crammed vertical crops the subject wrong and wastes the frame).
|
|
285
|
+
|
|
286
|
+
### 1 — Product launch / announcement
|
|
287
|
+
**When:** launching or announcing a product; teaser or full launch film.
|
|
288
|
+
**Beat sheet (45s, 16:9):**
|
|
289
|
+
| Beat | Time | Content | Words |
|
|
290
|
+
|---|---|---|---|
|
|
291
|
+
| Cold-open hook | 0–3s | One wow frame or the single pain moment. No logo. | ≤10 |
|
|
292
|
+
| Problem | 3–12s | The specific status-quo pain, shown not told. | ~20 |
|
|
293
|
+
| Reveal | 12–30s | The product doing the thing on REAL footage/screen. | ~35 |
|
|
294
|
+
| Proof | 30–40s | One concrete result or credible demo moment. | ~20 |
|
|
295
|
+
| CTA + date | 40–45s | Exact action + launch date. Logo last. | ≤12 |
|
|
296
|
+
|
|
297
|
+
**Frameworks:** 3-act teaser (Tease → Reveal → CTA) or StoryBrand SB7 (customer is hero, you are guide). One pain or one wow — not five diluted features.
|
|
298
|
+
**Reach for:** `import_media` a real demo capture or founder clip · `list_catalog` → the `product-launch-trailer` and `cinematic-title` scenes (use `cinematic-title` for a wordmark where the brief calls for it) · `color-grade` for the hero look · date + CTA on the final held frame.
|
|
299
|
+
**Anti-patterns:** opening on a logo animation; stacking five features; narrating specs; a fabricated/rounded stat; a mocked terminal standing in for a real screen.
|
|
300
|
+
**Quality bar:** the first frame survives with zero text on it. If your open is a gradient with the product name, you've already lost.
|
|
301
|
+
|
|
302
|
+
### 2 — Short-form social clip (TikTok / Reels / Shorts)
|
|
303
|
+
**When:** a vertical clip for social, usually cut from existing long-form.
|
|
304
|
+
**Beat sheet (22s, 9:16):**
|
|
305
|
+
| Beat | Time | Content |
|
|
306
|
+
|---|---|---|
|
|
307
|
+
| Visual + verbal hook | 0–3s | Bold claim / question / pattern-interrupt, on screen AND spoken. Answer teased, not given. |
|
|
308
|
+
| Setup | 3–8s | Frame the stakes; open loop #1. |
|
|
309
|
+
| Escalation | 8–16s | Value in ≤2 beats, each opening the next loop. Cut every ~1.5–2s. |
|
|
310
|
+
| Payoff | 16–20s | Close the main loop; put the payoff word on screen. |
|
|
311
|
+
| CTA / loop-back | 20–22s | Soft CTA or a re-hook that rewards replay. |
|
|
312
|
+
|
|
313
|
+
**Frameworks:** Hook → Retention → Payoff with curiosity stacking. Keep cutdowns ≤30s — shorter clips complete far more often.
|
|
314
|
+
**Reach for:** `suggest_briefs` on the long-form media (the AI clip-finder authors trim + speaker framing + captions from the transcript for you), then `compose({ suggestionId })` (Director ensemble) or author yourself via `apply_composition` · reframe `focus: {mode:"active-speaker"}` to keep the speaker centered in the vertical crop · word-level captions with `emphasis` on the payoff word and `entrance:"word-pop"` · `derive_composition` to auto-reframe a 16:9 master to 9:16 content-aware (not letterboxed).
|
|
315
|
+
**Anti-patterns:** a wide source letterboxed into 9:16 with black bars instead of subject-tracked reframe; no captions (dies on mute); a hook that describes instead of provokes ("In this video I'll…").
|
|
316
|
+
**Quality bar:** the hook sentence starts at t ≤ 1.5s; the face box is a real fraction of the frame with the top of the head in-frame on every sampled still.
|
|
317
|
+
|
|
318
|
+
### 3 — Explainer
|
|
319
|
+
**When:** explaining a concept, mechanism, or product so a viewer *gets* it.
|
|
320
|
+
**Beat sheet (75s, 16:9, ~160 words @ ~130 wpm):**
|
|
321
|
+
| Beat | Time | Words | Content |
|
|
322
|
+
|---|---|---|---|
|
|
323
|
+
| Hook | 0–5s | ~12–15 | The "why watch" — land inside 3s. |
|
|
324
|
+
| Problem | 5–20s | ~30 | Name the exact pain; they must feel understood. |
|
|
325
|
+
| Solution | 20–45s | ~55 | Your answer and how it differs. Show it working. |
|
|
326
|
+
| Proof | 45–62s | ~35 | One result / mechanism / credible demo. |
|
|
327
|
+
| CTA | 62–75s | ~20 | One clear next step. |
|
|
328
|
+
|
|
329
|
+
**Frameworks:** Hook → Problem → Solution → Proof → CTA. Decide the CTA first, script it last; write for the ear.
|
|
330
|
+
**A fully animated explainer is a FIRST-CLASS grounded format**, not a fallback for missing footage. An explainer built on authored diagrams of the *actual* mechanism (its real pipeline, the real data) is grounded — the grounding is that each visual depicts THIS thing specifically, not a decorative loop. Reach for it deliberately, not as a consolation for "no footage."
|
|
331
|
+
**Reach for:** an authored `card` clip (your own HTML/typography, shelled with `applyMotionPreset`) to anchor an abstract mechanism with a real diagram — process/algorithm/data viz, not cinematic footage · `lower-third` for term labels · one idea per beat · when using real footage, narrate to the detected transcript timing.
|
|
332
|
+
**Anti-patterns:** a wall of narration over stock gradient loops with no visual keyed to the words (the generic-filler tell — a purposeful diagram of the real mechanism is the opposite); explaining features before establishing the problem; using `generate_media` text-to-video to fake a "product" that doesn't match the real UI (it hallucinates specifics).
|
|
333
|
+
**Quality bar:** every beat has a visual keyed to what's being said — grounded and specific to this content, not decoration; the hook lands in 3s.
|
|
334
|
+
|
|
335
|
+
### 4 — SaaS / dev-tool demo
|
|
336
|
+
**When:** showing a product/CLI/app actually doing a real task.
|
|
337
|
+
**Beat sheet (~2 min, 16:9, real screen capture):**
|
|
338
|
+
| Beat | Time | Content |
|
|
339
|
+
|---|---|---|
|
|
340
|
+
| Hook + problem | 0–15s | The workflow that hurts today. Real UI or founder to-camera. |
|
|
341
|
+
| Setup | 15–30s | The task we'll accomplish; the promised outcome. |
|
|
342
|
+
| Walkthrough | 30–95s | Do the real thing on the real screen. Punch-in on the exact UI region per step; ~3–4s cadence; caption each action. |
|
|
343
|
+
| Outcome | 95–110s | Result achieved; the payoff vs the opening pain. |
|
|
344
|
+
| CTA | 110–120s | One action: start free / book / docs. |
|
|
345
|
+
|
|
346
|
+
**Frameworks:** Problem → Product-in-action → Outcome, wrapped in one completed task, not a feature tour. ~2 min is a good ceiling; chapter beyond that.
|
|
347
|
+
**Reach for:** `import_media` a REAL screen recording (captured by the user, or by an external screen-capture client such as a headed Playwright walkthrough recorded to MP4 — CueFrame does not screen-record) · reframe punch-in to the active UI element each step (`focus:{mode:"point",x,y}` on the click target, or a `face` focus for a founder inset) — the single biggest amateur→pro delta in demos, because the click target is ~30px in a 3840px 4K frame and nobody sees it unzoomed (**see `cueframe-product-video`** for the full click-point→reframe-segment auto-zoom tiling + PiP recipe) · a `lower-third` label per step.
|
|
348
|
+
**Anti-patterns:** the cardinal demo sin — a mocked/fake terminal or static UI card instead of the real app; showing the full 4K screen unzoomed; narrating every menu.
|
|
349
|
+
**Quality bar:** every UI/terminal on screen is a real capture; the actual action is visibly punched-in, not lost in a wide frame.
|
|
350
|
+
|
|
351
|
+
### 5 — Founder / talking-head (CueFrame's home turf)
|
|
352
|
+
**When:** a real person speaking to camera — founder update, announcement, testimonial. The format CueFrame composes best, and the one that agent walked past.
|
|
353
|
+
**Beat sheet (75s, 16:9 for web; 9:16 cutdown for social):**
|
|
354
|
+
| Beat | Time | Content |
|
|
355
|
+
|---|---|---|
|
|
356
|
+
| Attention / Problem | 0–5s | Founder states the pain or bold claim to camera. Eyes to lens. |
|
|
357
|
+
| Interest / Agitate | 5–25s | Why it matters now; personal stakes. B-roll cutaway. |
|
|
358
|
+
| Desire / Solution | 25–55s | What they built and the shift it creates. Product B-roll. |
|
|
359
|
+
| Proof | 55–68s | One credible result. Lower-third for name/title/metric. |
|
|
360
|
+
| Action | 68–75s | One CTA. |
|
|
361
|
+
|
|
362
|
+
**Frameworks:** PAS (Problem → Agitate → Solution) or AIDA. Framing: eyes on the upper-third line, small headroom, subject looks into the lens; break every ~5–8s of pure talking head with a cutaway.
|
|
363
|
+
**Reach for:** THE canonical CueFrame job. `get_media_context` for faces + transcript → reframe `focus:{mode:"active-speaker"}` for a solo speaker, `{mode:"all-faces"}` to hold a two-shot, or `{mode:"face",faceId,shotScale:"close"}` to punch in on emphasis → transcript-timed `captions` with a legibility plate (`setCaptionStyle` `emphasisPlate`, `position:"bottom"`) → a text overlay at `zPlane:"behind-subject"` when you want a title/word to sit *behind* the head (the signature look; render bakes the person matte) → a `lower-third` for name/title, shown once → `color-grade`. Zero generation needed.
|
|
364
|
+
**Anti-patterns:** a static wide talking head running 30s uncut (reads as a raw webcam upload); captions slapped over the chin with no plate; head dead-center or too much headroom; missing the lower-third identity.
|
|
365
|
+
**Quality bar:** a face is on screen within 3s; the person framed is the person speaking; captions never cover the mouth or eyes.
|
|
366
|
+
|
|
367
|
+
---
|
|
368
|
+
|
|
369
|
+
## Map craft to CueFrame tools (never invent a tool, primitive, or param)
|
|
370
|
+
|
|
371
|
+
| Craft goal | Do this in CueFrame |
|
|
372
|
+
|---|---|
|
|
373
|
+
| Bring in the real thing | `import_media` — requires `url` (public https) + `filename` + `contentType` (from the fixed enum: `video/mp4`, `video/quicktime`, `image/png`, `audio/mpeg`, …). Poll `list_media` or `create_webhook` on `media.completed` (status inside: complete|failed). |
|
|
374
|
+
| Know where subject + speech are | `get_media_context` → face boxes per frame (with `faceId`) + transcript. If faces are not detected, `prepare_media(kind:"subjectTrack", intent:{startSec,endSec})` → poll `get_media_facts` for the same SOURCE window until ready. |
|
|
375
|
+
| Keep the speaker framed / punch-in | `setCropIntents` op (the PREFERRED authoring path) or inline `reframe.segments` on a `clip.add` source. Each segment needs `startSec`/`endSec`/`focus`/`zoom`; **segments must TILE the clip** (an uncovered span → `reframe_coverage_gap` at validate). `focus.mode`: `frame-center` \| `point{x,y}` \| `face{faceId, shotScale?: close\|standard\|wide}` \| `all-faces` \| `active-speaker`. `zoom<1` punches in, `=1` full frame, `>1` rejected (min 0.1). Times in **seconds**. |
|
|
376
|
+
| Captions that track real speech | `composition.captions.segments[].words[]` = `{text, startMs, endMs, emphasis?, annotate?: circle\|underline}` — word times in **milliseconds**. Author the words directly with the `setCaptions` op (segments wholesale — how a generated-VO film gets its word-synced spine without a source transcript); style via `setCaptionStyle`: `position` (top\|center\|bottom), `entrance` (fade\|word-pop\|stagger-up), `emphasisPlate` + plate colors for legibility. Captions default to front; set top-level `composition.captions.zPlane:"behind-subject"` only when the user wants the whole caption layer behind the speaker. It resolves person cutouts for every overlapping source window and degrades visibly to front where the legibility gate refuses the effect. |
|
|
377
|
+
| Text behind the speaker's head | An `overlay` clip (a title/label primitive) with `zPlane:"behind-subject"`. This is the focused hero treatment; caption-layer depth is a separate whole-layer choice. Before preview/judgment, `prepare_media(kind:"matte", intent:{startSec,endSec})` for the overlapping clip's SOURCE trim and poll `get_media_facts` until ready. Render also resolves the person cutout on miss. |
|
|
378
|
+
| Name someone / label a step | `lower-third` scene from `list_catalog` |
|
|
379
|
+
| Title / closing wordmark | `cinematic-title` scene (closing, not opening) |
|
|
380
|
+
| Grade the look | `color-grade` scene/effect |
|
|
381
|
+
| Anchor an abstract mechanism | An authored `card` clip + `applyMotionPreset` (diagrams of the real mechanism — NOT cinematic; a first-class grounded path) |
|
|
382
|
+
| Long-form → clips | `suggest_briefs` → `wait_job(kind:"suggestions")` → `compose({ suggestionId })` |
|
|
383
|
+
| Auto-reframe to another aspect | `derive_composition` (content-aware 9:16 / 1:1 / 4:5 from a 16:9 master) |
|
|
384
|
+
| Discover what to build with | `list_catalog` — pass filters such as `category`/`kind`/`tier`/`useCase`/`mood`/`intent`, plus `limit` and the returned pagination cursor. For builtin packages, request `projection:"component-library"` and `origin:"builtin"`, then resolve the selected exact identity before insertion. Each entry carries `placement` (how to author it), `tier`/`useCase`, and `fixedCopy` (baked text no param changes). Read `cueframe://component/{id}` for the full **prop schema**. |
|
|
385
|
+
| Own / fork a custom graphic | `get_component_source` / `create_component` (self-contained tsxSource) / `update_component` — **see `cueframe-component-authoring`** for the full fork→preview→push loop |
|
|
386
|
+
| Several cells that share space and must reflow together (feature grid, rail + fluid content, cast + type) | An authored component on `FlexLayout` from `@cueframe/animate` — tracks are **weights** (`{ weight: 0, basis: px }` = content-sized, pushes neighbours), `tracksEnd` animates the reflow per frame, `fits` per cell (`stretch`/`contain`/`cover`/`position`/`scale`/`matte`). Static PiP / split is clip `region` + `fit` instead. Never CSS-scale the layout. **See `cueframe-component-authoring`** |
|
|
387
|
+
| Seed + edit the timeline | `apply_composition` (atomic batch ops; `dry_run:true` to check, `if_match` for OCC) — this is where clips, reframe, captions, and format are authored |
|
|
388
|
+
| Set the brand | `create_brand_kit` (402-gated) → set `brandKitId` on the **project** (`create_project` / `update_project`). It lives on the project, NOT on a composition op — trying to set it elsewhere is silently dropped. |
|
|
389
|
+
| Shape-check (NOT quality) | `validate_composition` → `{ valid, errors[] }` (`unknown_primitive`, `invalid_primitive_params`, `clip_source_track_mismatch`, `reframe_coverage_gap`, …); does **not** check trim bounds or image-on-video-track (those fail at save/render) |
|
|
390
|
+
| Look at real pixels | `preview_frame` at timestamps → PNG stills → LOOK |
|
|
391
|
+
| Judge motion or audio timing | `preview_clip` for the affected `[fromSec, toSec)` window → `wait_job(kind:"render")` → watch/listen; use matching reference timestamps |
|
|
392
|
+
| Check an isolated card | `preview_card` with its HTML and tokens; no scratch composition needed |
|
|
393
|
+
| Preview one graphic | `preview_component` (202 + jobId → `wait_job(kind:"preview")`; returns the verified artifact from the same immutable bundle used by production; compile/runtime/capability errors land on the job's failed terminal with the author-fixable message) |
|
|
394
|
+
| Independent grade (steered, un-gameable) | `score_composition` (sessionless; `editorialIntent`=your vision, `brandContext`=brand facts) → four fixed criteria + composite + worst-first critique, on frames it samples itself |
|
|
395
|
+
| Director ensemble compose | `compose` — set EXACTLY ONE of `fromComposition:true` (author from the seeded composition) or `suggestionId` (clipping path) → `wait_job(kind:"compose")` |
|
|
396
|
+
| Final MP4 | `create_render` → `wait_job(kind:"render")` or `create_webhook` on `render.completed`; `retry_render` (transient only) / `cancel_render` |
|
|
397
|
+
| Orient before spending | `get_account` FIRST — read `balance`, NOT `included`. Rendering / generation / previews / judge / Director share ONE credit wallet (`fundedBy`); each `balance` restates that same wallet in that feature's unit. Positive balance is not proof of capability or entitlement: check the exposed contract, entitlements and quote before proceeding. New orgs get a free credit grant, so an empty usage history is not a blocker. 402 `billing_required` fires once the balance is spent; `generate_media` also 402s over its USD ceiling |
|
|
398
|
+
|
|
399
|
+
Everything needs a `projectId` — `create_project` is step 0.
|
|
400
|
+
|
|
401
|
+
**Hard limits — respect them, don't fake around them.** No cinematic AI generation: `generate_media` is limited text-to-video/text-to-image/image-to-video, plus audio: `text-to-speech` (narration VO) and `text-to-music` (instrumental bed); use generated visuals where their actual capabilities fit the brief, and do not present synthetic footage as documentary evidence. No screen capture: that needs an external screen-capture client (e.g. a headed Playwright walkthrough recorded to MP4). When the right move needs footage CueFrame can't generate, `import_media` real footage — do not synthesize a substitute. **Units trap:** `startTime`/`duration`/`trim`/reframe `startSec`/`endSec`/`ease` are **seconds**; caption word times are **milliseconds**. **Render-time fail-loud errors to expect** (render never ships a wrong frame silently): `content_anchor` (a clip's `trim` frames *different* footage than its captions show — align the trim to the captioned moment), `behind_split_unsupported` (one media used under two *different* behind-subject windows — one behind-subject window per MEDIA; to render two different behind windows of one source, import it as two separate MEDIA ITEMS (distinct mediaIds) — duplicating the same mediaId into two clips still throws), `BehindMatteUnresolved`/`TrajectoryUnresolved` (the behind/face bake found no detectable person — drop the intent or fix the source).
|
|
402
|
+
|
|
403
|
+
**Worked alignment example (the error-prone mechanic).** You want the caption emphasis to land on the founder's stressed word "faster," and the frame to punch in on the same beat. From `get_media_context`, the transcript word "faster" runs `startMs:8120, endMs:8560` (transcript/source time). **For these round numbers, assume an untrimmed clip placed at timeline 0** — then source time, timeline time, and clip-relative time all coincide. So: put that word in `captions.segments[].words[]` with `emphasis:true` at `startMs:8120`; and author reframe with `setCropIntents` as a **contiguous segment list that tiles the whole clip** — e.g. a wide segment `{startSec:0, endSec:8.12, focus:{mode:"active-speaker"}, zoom:1}` then the punch `{startSec:8.12, endSec:<clipEnd>, focus:{mode:"face",faceId:"…",shotScale:"close"}, zoom:0.8, ease:{in:0.4,out:0}}` (ms → s: 8120ms = 8.12s). A lone mid-clip segment that leaves `[0..8.12]` uncovered fails validation with `reframe_coverage_gap`.
|
|
404
|
+
|
|
405
|
+
**Watch the timebase: reframe segment times are CLIP-RELATIVE; caption/transcript times are TIMELINE-based — convert when the clip is trimmed or not at t=0.** They coincide only for an untrimmed clip placed at timeline 0 (the case above). For a clip trimmed to start at source `trim.start` and sitting at `clip.startTime` on the timeline, the reframe boundary is `startSec = transcriptSec − trim.start`, and the caption time is `clip.startTime + (transcriptSec − trim.start)`. (The simple case just has `trim.start = 0` and `clip.startTime = 0`, so both reduce to `transcriptSec`.) Land the caption entrance *on* the stressed word, not a round frame like 8.0s — on-beat reads as craft. And if the clip's `trim` window doesn't actually contain source-second 8.12, the captioned word shows different footage → `content_anchor` at render.
|
|
406
|
+
|
|
407
|
+
---
|
|
408
|
+
|
|
409
|
+
## Craft principles (aesthetic defaults) (reason FROM these — the rules above are their consequences)
|
|
410
|
+
|
|
411
|
+
- **Curiosity is an information gap** (Loewenstein, 1994). A slide that *states* a fact satiates before it primes; a cold open that *poses* ("…and that's the number that made us kill the old pipeline") opens a gap the viewer must watch to close. Order clips to pose, not answer. A definitional title card is appropriate when requested or useful for this format.
|
|
412
|
+
- **Emotion outweighs everything** (Murch's Rule of Six: emotion ~51%, then story, rhythm, eye-trace, planarity, spatial continuity). The slop instinct optimizes the bottom of the list — aligned text, smooth easing — while ignoring the top. When choosing where to cut, ask "does this land the emotion / advance the point?" first, "is it geometrically neat?" last.
|
|
413
|
+
- **Attention and rhythm.** Consider motivated changes of state or shot scale when attention is the goal. Deliberate repetition, continuous motion and long holds can serve other requested formats.
|
|
414
|
+
- **Sound is half the picture** (Chion, *Audio-Vision*; synchresis). Cut on stressed words / beats via transcript word-timing; hold a silence beat before the peak line. On-beat reads as craft, off-beat as accident.
|
|
415
|
+
- **Story can use a value-shift** (McKee). Pain → product → win is one launch structure; a sequence of raised-and-answered questions is another. Neither is required for a brand study, observation or a different requested arc.
|
|
416
|
+
- **Premium = legible intent, not polish.** Audiences forgive imperfect video with a point of view; they don't forgive adequate video with no soul. Restraint, typographic discipline, and one consistent color/type/motion system across the whole piece are the premium signal. Consider why motion serves the requested visual experience.
|
|
417
|
+
|
|
418
|
+
### Deconstructed grammars (steal the structure, map to the tools)
|
|
419
|
+
These are *archetypes*, not documented cuts of specific films — reason from the grammar, don't cargo-cult a company.
|
|
420
|
+
|
|
421
|
+
- **Hardware reveal example:** a product materializing against negative space can carry a reveal, with one capability per shot and sourced numbers held long enough to read. This is one option; a requested logo or title opening is equally valid. Use actual product footage when claiming to show the real hardware.
|
|
422
|
+
|
|
423
|
+
- **Human testimony example:** founder-to-camera footage can give a launch specificity. Reframe from actual speaker facts and use transcript timing for requested captions. Do not add a person or voice to a piece that calls for another medium.
|
|
424
|
+
|
|
425
|
+
- **Product-action example:** real UI interactions on a restrained stage can make a software demo readable. Choose holds, cuts, music and graphic-only intervals from the brief; a frame without the product is not an automatic failure. A typography or gradient study can use those elements as its subject.
|
|
426
|
+
|
|
427
|
+
These examples suggest choices, not a required opening shot, visual medium or film structure. Real captures matter when presented as product evidence; authored diagrams and typography are valid requested media.
|
|
428
|
+
|
|
429
|
+
---
|
|
430
|
+
|
|
431
|
+
## Acceptance template
|
|
432
|
+
|
|
433
|
+
Use only what the request needs:
|
|
434
|
+
|
|
435
|
+
```
|
|
436
|
+
INTENT: the requested outcome or visual experience.
|
|
437
|
+
SCOPE: new work | affected revision | broader review | final delivery.
|
|
438
|
+
PROJECT/LANE: current project and capabilities actually exposed.
|
|
439
|
+
ACCEPTANCE: observable requirements at relevant timestamps/windows.
|
|
440
|
+
PRESERVE: unrelated timing, media, audio, source identity and shared kit state.
|
|
441
|
+
FACTS: sources for factual claims; identify simulated UI honestly.
|
|
442
|
+
EVIDENCE: matching frames plus motion/audio checks where relevant; source revision and job identity.
|
|
443
|
+
REVIEW: optional judge critique within budget; acceptance failures take precedence over scores.
|
|
444
|
+
DELIVERY: preview evidence or verified final export as requested; report remaining gaps.
|
|
445
|
+
```
|
|
446
|
+
|
|
447
|
+
Before final export, verify the requested acceptance criteria, current source revision and media bindings, valid composition shape, relevant motion/audio, and applicable hosted account/quote safeguards. Return the artifact location and distinguish preview proof from final-export proof.
|
|
@@ -5,8 +5,15 @@ description: Use when turning a local app into a brand-consistent product demo w
|
|
|
5
5
|
|
|
6
6
|
# CueFrame — Brand demo recipe
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
## Scope and capability policy
|
|
9
|
+
|
|
10
|
+
User intent and acceptance criteria override this recipe's aesthetic defaults. Silent, typography-only, uncaptioned and slow work are valid; no universal narration, footage, CTA, beat count or score target. Contracts, factual fidelity, media bindings and current source revisions remain mandatory.
|
|
11
|
+
|
|
12
|
+
Read available tools and current project; preserve the chosen local/hosted lane. Supported local work needs no hosted registration. Do not silently install, authenticate, switch projects or upload local content. Before metered hosted work, check `get_account`, entitlements, balance and quotes. Desktop setup: https://docs.cueframe.ai.
|
|
9
13
|
|
|
14
|
+
For scoped revisions, preserve unrelated timing, media, audio and shared brand-kit state. Edit affected content; inspect relevant frames, continuous motion and affected audio. No restarted intake, new approval pause, paid judge or final export by default. Judge only when useful; scores never override acceptance failures. Stop/report after two no-improvement passes. See [preview workflow](../cueframe-compose-loop/references/preview-workflow.md).
|
|
15
|
+
|
|
16
|
+
> **Kit truth over MCP (current limits):** kit CONTENTS are not yet inspectable over MCP (`get_profile` lists names only), so `$brand:` tokens resolve blind — verify resolved fonts/colors with a `preview_frame` still, and when the kit conflicts with an explicit client spec, the spec wins (hardcode it and note the conflict). Bind kits by the SERVER id from `get_profile` (slug binding is unreliable — it can silently bind nothing); always confirm via the project response `brandKitId`.
|
|
10
17
|
|
|
11
18
|
A **use-case recipe** that runs end-to-end from inside the user's own project
|
|
12
19
|
folder. You (Claude) already see their code and their running app — that's the
|
|
@@ -254,4 +261,4 @@ frames) before declaring done.
|
|
|
254
261
|
which is the engine's job now.
|
|
255
262
|
- This recipe lives in a skill on purpose: the engine stays use-case-free; the
|
|
256
263
|
product opinions (what to extract, which overlays, how to brand) live here and
|
|
257
|
-
ship via `npx -y cueframe@0.
|
|
264
|
+
ship via `npx -y cueframe@0.6 install`.
|