creator-editing-studio 1.1.1 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "creator-editing-studio",
3
- "version": "1.1.1",
3
+ "version": "1.2.0",
4
4
  "description": "Scaffolds a Creator Editing OS studio — components, scripts and rules for editing video from a plain-English brief.",
5
5
  "license": "UNLICENSED",
6
6
  "type": "module",
@@ -48,257 +48,264 @@ Everything else in `src/design/` is the system and should be left alone. If a vi
48
48
  a colour or a size that is not in tokens.ts, the answer is to add it to tokens.ts, never to
49
49
  type it into a component.
50
50
 
51
- ## Non-negotiables
52
-
53
- 1. **No magic numbers.** Durations, curves, springs, staggers, blur come from `src/design/motion.ts`.
54
- Colours, fonts, sizes, spacing come from `src/design/tokens.ts`. If a needed token does not
55
- exist, add it to the token file (with a comment) instead of inlining a value.
56
- 2. **Timing derives from the spine.** Animation start frames come from word timestamps, beats,
57
- or other elements' timings via helpers — never hand-typed frame numbers.
58
- 3. **`npm run check` and `motion-check --scale=2` must pass** before any render is shown to the user.
59
- 4. **Verify visually before claiming done** (see Verification). Never say something "looks good"
60
- without having rendered and viewed the frames.
61
-
51
+ ## Non-negotiables
52
+
53
+ 1. **No magic numbers.** Durations, curves, springs, staggers, blur come from `src/design/motion.ts`.
54
+ Colours, fonts, sizes, spacing come from `src/design/tokens.ts`. If a needed token does not
55
+ exist, add it to the token file (with a comment) instead of inlining a value.
56
+ 2. **Timing derives from the spine.** Animation start frames come from word timestamps, beats,
57
+ or other elements' timings via helpers — never hand-typed frame numbers.
58
+ 3. **`npm run check` and `motion-check --scale=2` must pass** before any render is shown to the user.
59
+ 4. **Verify visually before claiming done** (see Verification). Never say something "looks good"
60
+ without having rendered and viewed the frames.
61
+
62
62
  ---
63
63
 
64
- ## The 12 motion laws
65
-
66
- 1. **No linear easing** — except constant-velocity loops/scrolls (mark with `// motion-ok`).
67
- 2. **Entrances ease-out (`EASE.out`), exits ease-in (`EASE.in`), on-screen A-to-B moves ease-in-out (`EASE.inOut`).**
68
- 3. **Exits are faster than entrances** — `EXIT_RATIO` (0.65). Use `exitFrames()`.
69
- 4. **Max 2–3 animated properties per element.** No fade + scale + rotate + blur together.
70
- 5. **Never opacity alone.** Pair with 8–24px travel or scale >= 0.94. Opacity completes at
71
- `OPACITY_LEAD` (60%) of the transform.
72
- 6. **Never scale from below 0.9.** Text scaling from zero is the #1 AI-edit tell.
73
- 7. **Blur-in capped at 12px**, fully resolved before the transform ends.
74
- 8. **Groups always stagger** (`STAGGER.tight/normal/dramatic`). Never simultaneous.
75
- 9. **Distance couples to duration sub-linearly** (`travelFrames()`), never linearly.
76
- 10. **Anything that moves fast is motion-blurred** (see Smoothness system). No hard edge may jump
77
- across the frame unblurred.
78
- 11. **Readable hold:** text stays fully visible >= `readFrames(text)` (min 0.9s, +55ms/word).
79
- 12. **Continuity:** every transition inherits the outgoing motion (see below).
80
-
81
- ### Continuity (what makes it feel connected, not cut together)
82
-
83
- - **Shared elements never re-enter.** If an element exists across two beats, `morph()` its
84
- position/size/radius between layouts; never exit and re-animate it. (TransitionTest: the
85
- eyebrow dot stays on screen and grows into the next card's icon tile.)
86
- - **Velocity handoff.** An exit toward a direction is answered by an entrance continuing that
87
- direction — `handoff(exitTo)`.
88
- - **Overlap, never gap.** The next beat starts before the previous one settles —
89
- `nextBeatAt(prevSettled)`.
90
- - **One camera over cuts.** Lay scenes out on one canvas and glide a camera between them
91
- (`src/lib/camera.ts`) instead of hard-cutting between unrelated layouts. Content that should
92
- stay put during a move (shared elements) lives in screen space, outside the camera.
93
- - **Nothing is cut off.** Compute exit start times backward from the end so the slowest,
94
- most-staggered exit completes on or before the last frame.
95
- - **Sound glues cuts.** Whooshes start `TRANSITION.sfxLeadMs` before a visual change and tail
96
- `sfxTailMs` after.
97
-
64
+ ## The 12 motion laws
65
+
66
+ 1. **No linear easing** — except constant-velocity loops/scrolls (mark with `// motion-ok`).
67
+ 2. **Entrances ease-out (`EASE.out`), exits ease-in (`EASE.in`), on-screen A-to-B moves ease-in-out (`EASE.inOut`).**
68
+ 3. **Exits are faster than entrances** — `EXIT_RATIO` (0.65). Use `exitFrames()`.
69
+ 4. **Max 2–3 animated properties per element.** No fade + scale + rotate + blur together.
70
+ 5. **Never opacity alone.** Pair with 8–24px travel or scale >= 0.94. Opacity completes at
71
+ `OPACITY_LEAD` (60%) of the transform.
72
+ 6. **Never scale from below 0.9.** Text scaling from zero is the #1 AI-edit tell.
73
+ 7. **Blur-in capped at 12px**, fully resolved before the transform ends.
74
+ 8. **Groups always stagger** (`STAGGER.tight/normal/dramatic`). Never simultaneous.
75
+ 9. **Distance couples to duration sub-linearly** (`travelFrames()`), never linearly.
76
+ 10. **Anything that moves fast is motion-blurred** (see Smoothness system). No hard edge may jump
77
+ across the frame unblurred.
78
+ 11. **Readable hold:** text stays fully visible >= `readFrames(text)` (min 0.9s, +55ms/word).
79
+ 12. **Continuity:** every transition inherits the outgoing motion (see below).
80
+
81
+ ### Continuity (what makes it feel connected, not cut together)
82
+
83
+ - **Shared elements never re-enter.** If an element exists across two beats, `morph()` its
84
+ position/size/radius between layouts; never exit and re-animate it. (TransitionTest: the
85
+ eyebrow dot stays on screen and grows into the next card's icon tile.)
86
+ - **Velocity handoff.** An exit toward a direction is answered by an entrance continuing that
87
+ direction — `handoff(exitTo)`.
88
+ - **Overlap, never gap.** The next beat starts before the previous one settles —
89
+ `nextBeatAt(prevSettled)`.
90
+ - **One camera over cuts.** Lay scenes out on one canvas and glide a camera between them
91
+ (`src/lib/camera.ts`) instead of hard-cutting between unrelated layouts. Content that should
92
+ stay put during a move (shared elements) lives in screen space, outside the camera.
93
+ - **Nothing is cut off.** Compute exit start times backward from the end so the slowest,
94
+ most-staggered exit completes on or before the last frame.
95
+ - **Sound glues cuts.** Whooshes start `TRANSITION.sfxLeadMs` before a visual change and tail
96
+ `sfxTailMs` after.
97
+
98
98
  ---
99
99
 
100
- ## Smoothness system (every one of these was a real defect found by motion-check)
101
-
102
- - **Use the components, not raw transforms.** `<Presence>`, `<MaskReveal>`, `<HighlightMark>`,
103
- `<RollingNumber>` add velocity motion blur automatically. Anything else that moves wraps its
104
- moving element in `<VelocityBlur vx vy>` with `perFrame()` velocities.
105
- - **Velocity motion blur** (`VelocityBlur`): directional blur sized to this frame's movement with
106
- a 180-degree shutter; fades to zero as the element settles. Near-free to render.
107
- - **Wipes** use `wipeMask(p, perFrame(...))`: the leading edge softens with speed. Never wipe with
108
- a hard `clipPath` edge.
109
- - **Camera moves**: wrap the camera canvas in a full-frame `VelocityBlur` with `bleed={0}` and
110
- `overflow: hidden`, velocity from `cameraShot` (see TransitionTest).
111
- - **Zooms / rotations** (not translations): use `<MotionBlur active>` (temporal sampling, centred
112
- on the frame so toggling it never shifts timing). It costs 10x render time while active —
113
- enable only during the move. Never use `@remotion/motion-blur`'s `CameraMotionBlur`: it samples
114
- ahead of the frame, so switching it on/off causes a timing hitch.
115
- - **Counters roll, never tick.** Use `<RollingNumber>`. A counting number changes glyphs in single
116
- frames as it slows, which motion-check flags as pops.
117
- - **Never `will-change`.** It makes each parallel render tab cache text at a different sub-pixel
118
- position, so settled elements flicker between 4 versions. (Lint enforces this.)
119
- - **Screen recordings used as b-roll usually scroll in steps** (they judder). Use a still frame with a
120
- smooth `ScreenView` pan instead, and blur names on the still.
121
- - **Presence `dur` sets entrance AND exit length.** `dur="instant"` makes exits last 3 frames (a snap) —
122
- for containers that only fade in, use at least `"base"`.
123
- - **Final renders are supersampled at `--scale=2`**, then downscaled to 1080x1920. At 1x, Chrome
124
- snaps text to whole pixels, so the slow tail of every ease-out steps ("move 1px, hold, move 1px")
125
- and motion-check reports stutters. Run motion-check with the same `--scale`.
126
- - **Frame rate:** match the footage. If footage is 60fps, build at 60 — per-frame jumps halve.
127
- All tokens are in ms, so compositions work at either rate.
128
-
100
+ ## Smoothness system (every one of these was a real defect found by motion-check)
101
+
102
+ - **Use the components, not raw transforms.** `<Presence>`, `<MaskReveal>`, `<HighlightMark>`,
103
+ `<RollingNumber>` add velocity motion blur automatically. Anything else that moves wraps its
104
+ moving element in `<VelocityBlur vx vy>` with `perFrame()` velocities.
105
+ - **Velocity motion blur** (`VelocityBlur`): directional blur sized to this frame's movement with
106
+ a 180-degree shutter; fades to zero as the element settles. Near-free to render.
107
+ - **Wipes** use `wipeMask(p, perFrame(...))`: the leading edge softens with speed. Never wipe with
108
+ a hard `clipPath` edge.
109
+ - **Camera moves**: wrap the camera canvas in a full-frame `VelocityBlur` with `bleed={0}` and
110
+ `overflow: hidden`, velocity from `cameraShot` (see TransitionTest).
111
+ - **Zooms / rotations** (not translations): use `<MotionBlur active>` (temporal sampling, centred
112
+ on the frame so toggling it never shifts timing). It costs 10x render time while active —
113
+ enable only during the move. Never use `@remotion/motion-blur`'s `CameraMotionBlur`: it samples
114
+ ahead of the frame, so switching it on/off causes a timing hitch.
115
+ - **Counters roll, never tick.** Use `<RollingNumber>`. A counting number changes glyphs in single
116
+ frames as it slows, which motion-check flags as pops.
117
+ - **Never `will-change`.** It makes each parallel render tab cache text at a different sub-pixel
118
+ position, so settled elements flicker between 4 versions. (Lint enforces this.)
119
+ - **Screen recordings used as b-roll usually scroll in steps** (they judder). Use a still frame with a
120
+ smooth `ScreenView` pan instead, and blur names on the still.
121
+ - **Presence `dur` sets entrance AND exit length.** `dur="instant"` makes exits last 3 frames (a snap) —
122
+ for containers that only fade in, use at least `"base"`.
123
+ - **Final renders are supersampled at `--scale=2`**, then downscaled to 1080x1920. At 1x, Chrome
124
+ snaps text to whole pixels, so the slow tail of every ease-out steps ("move 1px, hold, move 1px")
125
+ and motion-check reports stutters. Run motion-check with the same `--scale`.
126
+ - **Frame rate:** match the footage. If footage is 60fps, build at 60 — per-frame jumps halve.
127
+ All tokens are in ms, so compositions work at either rate.
128
+
129
129
  ---
130
130
 
131
- ## Design rules
132
-
133
- **Read `STYLE.md` before any design decision.** It is the house style (type scale, layout, motion,
134
- sound, approved and rejected beats) and wins over anything below or in a reference. Update it
135
- whenever the user approves or rejects something.
136
-
137
- - Everything visual comes from `src/design/tokens.ts`, which is yours to define.
138
- - **Brand essentials:** your display typeface only, at every weight (headlines/stats 800). One accent: accent
139
- `your accent` on ink and cream/white. No blue, no purple, no serif. Signature emphasis is the accent
140
- marker behind the key phrase — use `<HighlightMark>`; don't invent other emphasis styles.
141
- Headlines are short two-part fragments with the payoff highlighted. Numbers are shown raw and
142
- bold ("$1.3M", "11x") and roll in. Eyebrows are uppercase pills with a accent-ringed dot.
143
- Radii are generous, shadows soft, buttons/badges pills.
144
- - Secondary text uses `COLOR.fgSecondary`; `COLOR.fgMuted` fails contrast on cream — decoration only.
145
- - **User style decisions (first project):**
146
- - Captions are **crisp white** (`CAPTION` token, tight + soft dark shadow). Never ink text with a light glow.
147
- Over the cream background white is invisible, so in card layouts captions go **inside the video
148
- card's bottom edge** on the card's dark scrim (`StageState.scrim`).
149
- - Big headline text is **accent with a glow** (`HeroBlock` / `HeroTag`, `heroGlow()`).
150
- - Pointer is a **solid accent arrow** like the reference (`Cursor`), ~68px.
151
- - Brand logos are **real full-colour logos** (`LogoIcon`, SVG Logos CC0 via `src/design/logos.json` —
152
- extract only the icons used). Never approximate a logo that isn't in the set; use a neutral
153
- icon and ask for the official file (Klaviyo is missing).
154
- - A corner tag (e.g. "11 MINUTES" pill) is removed when the video returns full screen.
155
- - **Text behind the subject must stay readable:** the head covers only the lower ~20–30% of the
156
- headline. Head height changes shot to shot (leaning forward lifts it ~80px) — check a still of
157
- EVERY hero moment and raise the block (`top`; put labels above as `eyebrow`) where needed.
158
- - Keep accent highlights away from the accent background glow so they don't disappear into it.
159
- - Stats use proportional figures, not tabular (tabular makes "11" look gappy in your display typeface).
160
- - **Web design systems do not map 1:1 to video.** Keep colour, font families/weights, type-scale
161
- ratios, radii, spacing rhythm, iconography and imagery style — but rescale sizes for a
162
- 1080-wide frame viewed on a phone (body roughly 40–48px, not 16px). Web UI components mostly
163
- do not apply.
164
- - Respect safe zones: `LAYOUT.safeTop` / `LAYOUT.safeBottom` for Reels/Shorts UI.
165
- - Text contrast >= 4.5:1 against whatever is behind it, including over footage.
166
- - Draw graphics in code (SVG/CSS) so every part can animate. Use flat images only for things
167
- that must be real (logos, screenshots, photos). Avoid AI-generated imagery for anything the
168
- viewer is meant to look at.
169
- - Icons: SVG drawn in code or an installed icon library, never image files.
170
-
131
+ ## Design rules
132
+
133
+ **Read `STYLE.md` before any design decision.** It is the house style (type scale, layout, motion,
134
+ sound, approved and rejected beats) and wins over anything below or in a reference. Update it
135
+ whenever the user approves or rejects something.
136
+
137
+ - Everything visual comes from `src/design/tokens.ts`, which is yours to define.
138
+ - **Brand essentials:** your display typeface only, at every weight (headlines/stats 800). One accent: accent
139
+ `your accent` on ink and cream/white. No blue, no purple, no serif. Signature emphasis is the accent
140
+ marker behind the key phrase — use `<HighlightMark>`; don't invent other emphasis styles.
141
+ Headlines are short two-part fragments with the payoff highlighted. Numbers are shown raw and
142
+ bold ("$1.3M", "11x") and roll in. Eyebrows are uppercase pills with a accent-ringed dot.
143
+ Radii are generous, shadows soft, buttons/badges pills.
144
+ - Secondary text uses `COLOR.fgSecondary`; `COLOR.fgMuted` fails contrast on cream — decoration only.
145
+ - **Your style decisions:** *(empty on purpose — this fills in as you make them)*
146
+
147
+ When something gets settled about how your videos look, write it here as one line. Where
148
+ captions sit in a card layout. Whether headlines take the accent colour or the ink. What a
149
+ corner tag does when the video returns to full screen. This file is read first every
150
+ session, so a decision recorded here is one you never have to make twice, and the studio
151
+ stops asking.
152
+
153
+ Do not fill this with defaults copied from somewhere else. An empty section produces
154
+ questions; a borrowed one produces someone else's videos.
155
+
156
+ The one rule here that is not a preference: brand logos are **real full-colour logos**
157
+ (`LogoIcon`, SVG Logos CC0 via `src/design/logos.json`, extract only the icons used). Never
158
+ approximate a logo that is not in the set. Use a neutral icon and ask for the official file.
159
+ - **Text behind the subject must stay readable:** the head covers only the lower ~20–30% of the
160
+ headline. Head height changes shot to shot (leaning forward lifts it ~80px) — check a still of
161
+ EVERY hero moment and raise the block (`top`; put labels above as `eyebrow`) where needed.
162
+ - Keep accent highlights away from the accent background glow so they don't disappear into it.
163
+ - Stats use proportional figures, not tabular (tabular makes "11" look gappy in your display typeface).
164
+ - **Web design systems do not map 1:1 to video.** Keep colour, font families/weights, type-scale
165
+ ratios, radii, spacing rhythm, iconography and imagery style — but rescale sizes for a
166
+ 1080-wide frame viewed on a phone (body roughly 40–48px, not 16px). Web UI components mostly
167
+ do not apply.
168
+ - Respect safe zones: `LAYOUT.safeTop` / `LAYOUT.safeBottom` for Reels/Shorts UI.
169
+ - Text contrast >= 4.5:1 against whatever is behind it, including over footage.
170
+ - Draw graphics in code (SVG/CSS) so every part can animate. Use flat images only for things
171
+ that must be real (logos, screenshots, photos). Avoid AI-generated imagery for anything the
172
+ viewer is meant to look at.
173
+ - Icons: SVG drawn in code or an installed icon library, never image files.
174
+
171
175
  ---
172
176
 
173
- ## Compositing: text behind the subject
174
-
175
- Layer order, bottom to top:
176
-
177
- ```
178
- background plate (designed, or the original footage)
179
- text / graphics <- 1–3px blur to sit on the background's focal plane
180
- subject cutout <- same clip, frame-aligned
181
- light wrap <- blurred background screened onto the subject's inner edge, 15–30%
182
- grain + grade <- applied over everything so layers share one look
183
- ```
184
-
185
- - **The cutout must come from the same clip as its background plate**, frame-aligned. Never
186
- combine a cutout from one take with footage from another.
187
- - **Green screen:** key once to a transparent intermediate in `work/` (not per render). Clean up
188
- with spill suppression, ~0.5px choke, 0.5–1px feather. Temporal smoothing on the alpha
189
- prevents edge flicker.
190
- - **Paired clips (the user's standard delivery):** Video 1 = original; Video 2 = the same clip with the
191
- background replaced by flat green. Use Video 2 **only to build the matte (alpha)** and take the
192
- subject's colour from Video 1 — this removes green fringe/spill entirely, because the subject pixels
193
- never touched green. Before building, verify the pair: identical duration, fps, resolution and
194
- frame count, and that the subject lines up on the same frame (a 1-frame offset makes a halo or ghost).
195
- Any VFR conform must be applied identically to both. If the background-removal tool can export real
196
- transparency, that is even better than green.
197
- - **Supplied cutouts** must have real transparency (ProRes 4444, WebM with alpha, or PNG sequence).
198
- An MP4 cannot hold transparency — if one arrives, stop and tell the user.
199
- - Play transparent video with `<OffthreadVideo transparent src={staticFile(...)} />`.
200
- - Text behind a moving subject gets 0.85–0.95x counter-parallax; static text reads as a sticker.
201
- - Text should animate in while already partly occluded, not appear fully and then get covered.
202
- - Pull text colour toward the footage's black/white points; pure white over graded footage looks
203
- pasted on.
204
-
177
+ ## Compositing: text behind the subject
178
+
179
+ Layer order, bottom to top:
180
+
181
+ ```
182
+ background plate (designed, or the original footage)
183
+ text / graphics <- 1–3px blur to sit on the background's focal plane
184
+ subject cutout <- same clip, frame-aligned
185
+ light wrap <- blurred background screened onto the subject's inner edge, 15–30%
186
+ grain + grade <- applied over everything so layers share one look
187
+ ```
188
+
189
+ - **The cutout must come from the same clip as its background plate**, frame-aligned. Never
190
+ combine a cutout from one take with footage from another.
191
+ - **The matte is built automatically from ordinary footage.** No green screen, and never ask the
192
+ user for one — most people have never shot against green and do not know the term. Segment the
193
+ subject per frame, then **temporally smooth the alpha**, which is not optional: a per-frame matte
194
+ shimmers along the edge, and type sitting against a head is exactly where that shows. Finish with
195
+ spill-free cleanup: ~0.5px choke, 0.5–1px feather.
196
+ - **Quality is worst where it is most visible:** flyaway hair, motion blur, and a hand crossing in
197
+ front of the text. The light wrap below hides a good deal of it. If a shot simply will not hold,
198
+ say so and move the text rather than shipping a halo.
199
+ - **If the user happens to have a background-free or green version, prefer it** — chroma keying
200
+ still beats automatic matting on hair. Use it **only for the alpha** and take the subject's colour
201
+ from the original, so no green ever reaches the edges. Verify the pair first: identical duration,
202
+ fps, resolution and frame count, subject aligned on the same frame (one frame out makes a ghost).
203
+ This is an upgrade for people who already work that way, never a requirement.
204
+ - **Supplied cutouts** must have real transparency (ProRes 4444, WebM with alpha, or PNG sequence).
205
+ An MP4 cannot hold transparency — if one arrives, stop and tell the user.
206
+ - Play transparent video with `<OffthreadVideo transparent src={staticFile(...)} />`.
207
+ - Text behind a moving subject gets 0.85–0.95x counter-parallax; static text reads as a sticker.
208
+ - Text should animate in while already partly occluded, not appear fully and then get covered.
209
+ - Pull text colour toward the footage's black/white points; pure white over graded footage looks
210
+ pasted on.
211
+
205
212
  ---
206
213
 
207
- ## Per-video pipeline
208
-
209
- 1. **Ingest:** probe every clip (fps, resolution, colour, alpha).
210
- - **Phone footage is usually variable frame rate (VFR).** Conform it to constant frame rate
211
- with ffmpeg before use, or it stutters and drifts out of sync in Remotion.
212
- - Conform everything to one fps. Never mix frame rates.
213
- 2. **Spine:** word-level timings + music beats -> `work/<video>/spine.json`.
214
- Cut pauses/filler from word gaps. All animation timing reads from the spine.
215
- - If the user supplies a script or SRT, it is the **text truth** (spelling, names, brand words):
216
- force-align that exact text to the audio with WhisperX's aligner. Never use SRT timings
217
- directly — they are line-level and padded for readability, not word-accurate.
218
- - With no script, transcribe with WhisperX, then show the transcript for correction before building.
219
- 3. **Cutouts:** key green screen or validate supplied cutouts (see above).
220
- 4. **Beat sheet:** from the script and the user's marked punch lines, write a short plan — which
221
- moment gets which treatment — and confirm with the user before building.
222
- - **Read their saved references first** (`editing_os_references`). They are the whole reason
223
- the edit should look like theirs rather than like the defaults, and a plan written without
224
- them is a plan written for anybody. Say which reference each decision came from, so the
225
- user can see their own taste being applied and correct it when it is being misread.
226
- - Where references are silent, the defaults in STYLE.md apply. Where references and defaults
227
- disagree, the reference wins — it is a real preference, and the default is only a sensible
228
- starting point.
229
- 5. **Build** in `src/compositions/<Video>.tsx` from components + helpers + tokens only.
230
- 6. **Verify** (below), then show the user.
231
- 7. **Sound:** the user adds SFX and music themselves. Deliver a cue list (timestamps of every
232
- transition and hit, with suggested whoosh start = `TRANSITION.sfxLeadMs` before the change) so
233
- they can place sounds quickly. Render without music unless asked.
234
- 8. **Render** at `--scale=2`, downscaled to 1080x1920 unless told otherwise.
235
-
214
+ ## Per-video pipeline
215
+
216
+ 1. **Ingest:** probe every clip (fps, resolution, colour, alpha).
217
+ - **Phone footage is usually variable frame rate (VFR).** Conform it to constant frame rate
218
+ with ffmpeg before use, or it stutters and drifts out of sync in Remotion.
219
+ - Conform everything to one fps. Never mix frame rates.
220
+ 2. **Spine:** word-level timings + music beats -> `work/<video>/spine.json`.
221
+ Cut pauses/filler from word gaps. All animation timing reads from the spine.
222
+ - If the user supplies a script or SRT, it is the **text truth** (spelling, names, brand words):
223
+ force-align that exact text to the audio with WhisperX's aligner. Never use SRT timings
224
+ directly — they are line-level and padded for readability, not word-accurate.
225
+ - With no script, transcribe with WhisperX, then show the transcript for correction before building.
226
+ 3. **Cutouts:** build the matte from the footage automatically, or validate a supplied cutout (see above). Only needed for shots that want text behind the subject.
227
+ 4. **Beat sheet:** from the script and the user's marked punch lines, write a short plan — which
228
+ moment gets which treatment — and confirm with the user before building.
229
+ - **Read their saved references first** (`editing_os_references`). They are the whole reason
230
+ the edit should look like theirs rather than like the defaults, and a plan written without
231
+ them is a plan written for anybody. Say which reference each decision came from, so the
232
+ user can see their own taste being applied and correct it when it is being misread.
233
+ - Where references are silent, the defaults in STYLE.md apply. Where references and defaults
234
+ disagree, the reference wins — it is a real preference, and the default is only a sensible
235
+ starting point.
236
+ 5. **Build** in `src/compositions/<Video>.tsx` from components + helpers + tokens only.
237
+ 6. **Verify** (below), then show the user.
238
+ 7. **Sound:** the user adds SFX and music themselves. Deliver a cue list (timestamps of every
239
+ transition and hit, with suggested whoosh start = `TRANSITION.sfxLeadMs` before the change) so
240
+ they can place sounds quickly. Render without music unless asked.
241
+ 8. **Render** at `--scale=2`, downscaled to 1080x1920 unless told otherwise.
242
+
236
243
  ---
237
244
 
238
- ## Audio standard (audio is half the reel — the user will not accept a bad mix)
239
-
240
- What went wrong on an early video: the voice itself was bit-for-bit intact (measured against the source),
241
- but 2.3–3.5s whoosh files and a 2.4s riser kept playing UNDER the speech, which made the voice sound
242
- wrong. Claude cannot hear, so the mix must be engineered and measured, never guessed.
243
-
244
- 1. **The voice is never touched.** Render the picture `--muted`. The voice in every deliverable is the
245
- original recording's audio (stream-copied, or at most one encode), plus only a single gain change.
246
- No limiter, no dynamic normaliser (ffmpeg `loudnorm` single-pass pumps), no repeated AAC encodes.
247
- 2. **Sound effects are short and placed in gaps.** Trim every SFX to what the motion needs (whoosh
248
- 0.4–0.8s, hits 0.3–0.6s) with a 60–120ms fade-out. Nothing longer than ~0.8s may sit under a spoken
249
- word unless it is at least 24 dB under the voice. Risers go in pauses, not under speech.
250
- 3. **Duck SFX under the voice**: mix with ffmpeg `sidechaincompress` keyed from the voice (SFX drop
251
- ~6 dB while a word is spoken), and high-pass whooshes/impacts around 120–150 Hz so they don't muddy it.
252
- 4. **Measure before delivering** (the window scan in the project folder used `f32le` PCM from ffmpeg):
253
- voice-vs-mix difference per 0.5s window — any window where SFX energy is within 12 dB of the
254
- voice while a word is spoken gets fixed. Final: -14 LUFS integrated, true peak <= -1 dBTP, one
255
- AAC encode at 256k.
256
- 5. **Always also deliver stems + a cue sheet**: `voice.wav`, `sfx.wav` (the placed, trimmed SFX bus),
257
- and `SFX-CUES.txt` with numbered, timecoded SFX files — so the user can rebalance in Premiere in
258
- minutes. Picture-only MP4 alongside.
259
- 6. The user picks and judges sounds by ear; Claude says plainly that it cannot hear and asks for
260
- timecoded audio notes on the first draft.
261
-
245
+ ## Audio standard (audio is half the reel — the user will not accept a bad mix)
246
+
247
+ What went wrong on an early video: the voice itself was bit-for-bit intact (measured against the source),
248
+ but 2.3–3.5s whoosh files and a 2.4s riser kept playing UNDER the speech, which made the voice sound
249
+ wrong. Claude cannot hear, so the mix must be engineered and measured, never guessed.
250
+
251
+ 1. **The voice is never touched.** Render the picture `--muted`. The voice in every deliverable is the
252
+ original recording's audio (stream-copied, or at most one encode), plus only a single gain change.
253
+ No limiter, no dynamic normaliser (ffmpeg `loudnorm` single-pass pumps), no repeated AAC encodes.
254
+ 2. **Sound effects are short and placed in gaps.** Trim every SFX to what the motion needs (whoosh
255
+ 0.4–0.8s, hits 0.3–0.6s) with a 60–120ms fade-out. Nothing longer than ~0.8s may sit under a spoken
256
+ word unless it is at least 24 dB under the voice. Risers go in pauses, not under speech.
257
+ 3. **Duck SFX under the voice**: mix with ffmpeg `sidechaincompress` keyed from the voice (SFX drop
258
+ ~6 dB while a word is spoken), and high-pass whooshes/impacts around 120–150 Hz so they don't muddy it.
259
+ 4. **Measure before delivering** (the window scan in the project folder used `f32le` PCM from ffmpeg):
260
+ voice-vs-mix difference per 0.5s window — any window where SFX energy is within 12 dB of the
261
+ voice while a word is spoken gets fixed. Final: -14 LUFS integrated, true peak <= -1 dBTP, one
262
+ AAC encode at 256k.
263
+ 5. **Always also deliver stems + a cue sheet**: `voice.wav`, `sfx.wav` (the placed, trimmed SFX bus),
264
+ and `SFX-CUES.txt` with numbered, timecoded SFX files — so the user can rebalance in Premiere in
265
+ minutes. Picture-only MP4 alongside.
266
+ 6. The user picks and judges sounds by ear; Claude says plainly that it cannot hear and asks for
267
+ timecoded audio notes on the first draft.
268
+
262
269
  ---
263
270
 
264
- ## Verification (required before saying anything is done)
265
-
266
- Claude cannot watch video. Quality is checked with frames and maths, and the user makes the
267
- final call at playback speed. Be explicit about that when reporting.
268
-
269
- 1. `npm run check` passes.
270
- 2. `npm run motion-check -- <Comp> --scale=2` reports **no POP, CUT or STUTTER**. Allow intentional
271
- cuts with `--allow=<frame>`. For each FAST warning, render that frame and confirm the moving
272
- edge is blurred. Read the motion timeline for dead air and rhythm.
273
- 3. Render stills and **look at them**: first frame, each element mid-entrance, peak-speed frames
274
- (blur visible), fully settled (crisp), mid-exit, and the **last frame** (clean).
275
- 4. Check: no text collides with the subject's face, safe zones respected, contrast holds,
276
- no clipped descenders (g, p, y, commas) in masked text.
277
- 5. Report what was verified and what still needs the user's eyes at full speed.
278
-
271
+ ## Verification (required before saying anything is done)
272
+
273
+ Claude cannot watch video. Quality is checked with frames and maths, and the user makes the
274
+ final call at playback speed. Be explicit about that when reporting.
275
+
276
+ 1. `npm run check` passes.
277
+ 2. `npm run motion-check -- <Comp> --scale=2` reports **no POP, CUT or STUTTER**. Allow intentional
278
+ cuts with `--allow=<frame>`. For each FAST warning, render that frame and confirm the moving
279
+ edge is blurred. Read the motion timeline for dead air and rhythm.
280
+ 3. Render stills and **look at them**: first frame, each element mid-entrance, peak-speed frames
281
+ (blur visible), fully settled (crisp), mid-exit, and the **last frame** (clean).
282
+ 4. Check: no text collides with the subject's face, safe zones respected, contrast holds,
283
+ no clipped descenders (g, p, y, commas) in masked text.
284
+ 5. Report what was verified and what still needs the user's eyes at full speed.
285
+
279
286
  ---
280
287
 
281
- ## Remotion gotchas
282
-
283
- - Never use CSS `transition`/`animation` — drive everything from `useCurrentFrame()`.
284
- - Never use `Math.random()` — use `random(seed)` from `remotion`.
285
- - Never use `will-change` (render-tab flicker, see Smoothness system).
286
- - **`<Presence>` owns its element's `transform` and `opacity`.** Any extra transform (centring with
287
- `translateX(-50%)`) or opacity (fading content) passed in `style` is silently overwritten — put it
288
- on an inner wrapper. This bug appeared twice in the first project (off-centre tag, labels that
289
- never faded).
290
- - em-based spacing (gaps, padding) resolves against the element's own font-size: set `fontSize`
291
- on the container, or word gaps collapse.
292
- - Masked text needs travel > 100% and container padding, or descenders stay visible.
293
- - Load fonts through `@remotion/fonts` (brand your display typeface is local in `assets/fonts/`) so renders wait for them.
294
-
288
+ ## Remotion gotchas
289
+
290
+ - Never use CSS `transition`/`animation` — drive everything from `useCurrentFrame()`.
291
+ - Never use `Math.random()` — use `random(seed)` from `remotion`.
292
+ - Never use `will-change` (render-tab flicker, see Smoothness system).
293
+ - **`<Presence>` owns its element's `transform` and `opacity`.** Any extra transform (centring with
294
+ `translateX(-50%)`) or opacity (fading content) passed in `style` is silently overwritten — put it
295
+ on an inner wrapper. This bug appeared twice in the first project (off-centre tag, labels that
296
+ never faded).
297
+ - em-based spacing (gaps, padding) resolves against the element's own font-size: set `fontSize`
298
+ on the container, or word gaps collapse.
299
+ - Masked text needs travel > 100% and container padding, or descenders stay visible.
300
+ - Load fonts through `@remotion/fonts` (brand your display typeface is local in `assets/fonts/`) so renders wait for them.
301
+
295
302
  ---
296
303
 
297
- ## Feedback format from the user
298
-
299
- Notes arrive as timecode + specific issue ("0:14 title lands 3 frames late, bounce too strong").
300
- Translate each note into token or timing changes; if a note conflicts with a law, say so and ask.
301
-
304
+ ## Feedback format from the user
305
+
306
+ Notes arrive as timecode + specific issue ("0:14 title lands 3 frames late, bounce too strong").
307
+ Translate each note into token or timing changes; if a note conflicts with a law, say so and ask.
308
+
302
309
  ---
303
310
 
304
311
  ## What is already installed
@@ -1,17 +1,23 @@
1
- # The Python half of the studio: transcription, word alignment, cut detection,
2
- # green-screen cutouts and colour grading.
3
- #
4
- # python -m venv .venv
5
- # .venv/bin/pip install -r requirements.txt (Windows: .venv\Scripts\pip)
6
- #
7
- # Python 3.11 or 3.12. Not 3.13+ — WhisperX's dependencies do not support it yet, and the
8
- # failure is a wall of build errors rather than a clear message.
9
- #
10
- # certifi is not optional on macOS: torch downloads the alignment model with urllib, which
11
- # on a python.org build has no root certificates at all.
12
-
13
- whisperx==3.8.6
14
- torch>=2.2
15
- numpy>=1.26
16
- Pillow>=10.0
17
- certifi>=2024.2.2
1
+ # The Python half of the studio: transcription, word alignment, cut detection,
2
+ # green-screen cutouts and colour grading.
3
+ #
4
+ # python -m venv .venv
5
+ # .venv/bin/pip install -r requirements.txt (Windows: .venv\Scripts\pip)
6
+ #
7
+ # Python 3.11 or 3.12. Not 3.13+ — WhisperX's dependencies do not support it yet, and the
8
+ # failure is a wall of build errors rather than a clear message.
9
+ #
10
+ # certifi is not optional on macOS: torch downloads the alignment model with urllib, which
11
+ # on a python.org build has no root certificates at all.
12
+
13
+ whisperx==3.8.6
14
+ # BiRefNet, for cutting the subject out of ordinary footage so text can sit behind them.
15
+ # MIT licensed and self-hosted: nothing uploads, and there is no per-video cost.
16
+ transformers>=4.44
17
+ timm>=1.0
18
+ einops>=0.8
19
+ kornia>=0.7
20
+ torch>=2.2
21
+ numpy>=1.26
22
+ Pillow>=10.0
23
+ certifi>=2024.2.2
@@ -1,68 +1,181 @@
1
- """Build a transparent cutout of the subject for a frame range of a paired clip.
2
-
3
- Usage (from the project root):
4
- python scripts/make_cutout.py <original> <greenscreen> <out_dir> <first_frame> <last_frame>
5
-
6
- The green-screen clip is used ONLY for the matte (alpha). Colour comes from the original clip, so no
7
- green ever reaches the subject's edges. Writes RGBA PNGs named f<frame>.png (absolute source frame
8
- numbers) so compositions can map composition frames to cutout frames directly.
9
- """
10
-
11
- import subprocess
12
- import sys
13
- from pathlib import Path
14
-
15
- import numpy as np
16
- from PIL import Image
17
- import sys as _sys, pathlib as _pl
18
- _sys.path.insert(0, str(_pl.Path(__file__).resolve().parent / 'lib'))
19
- from media import FFMPEG, FFPROBE # resolved per platform - never hardcode a tool path
20
-
21
-
22
-
23
- W, H, FPS = 1920, 1080, 30
24
-
25
-
26
- def read_frames(path: str, first: int, count: int):
27
- proc = subprocess.Popen(
28
- [FFMPEG, "-v", "error", "-ss", f"{first / FPS:.6f}", "-i", path, "-frames:v", str(count),
29
- "-f", "rawvideo", "-pix_fmt", "rgb24", "-"],
30
- stdout=subprocess.PIPE,
31
- )
32
- size = W * H * 3
33
- for _ in range(count):
34
- buf = proc.stdout.read(size)
35
- if len(buf) < size:
36
- break
37
- yield np.frombuffer(buf, np.uint8).reshape(H, W, 3)
38
- proc.stdout.close()
39
- proc.wait()
40
-
41
-
42
- def matte(green: np.ndarray) -> np.ndarray:
43
- """Alpha 0..1 from how green each pixel is, relative to the flat background green."""
44
- g = green.astype(np.float32)
45
- greenness = g[..., 1] - np.maximum(g[..., 0], g[..., 2])
46
- background = np.median(greenness[greenness > 60]) if (greenness > 60).any() else 150.0
47
- lo, hi = 0.2 * background, 0.7 * background
48
- t = np.clip((greenness - lo) / (hi - lo), 0, 1)
49
- alpha = 1 - t * t * (3 - 2 * t) # smoothstep: soft, stable edge
50
- # Choke by 1px so the background never leaks in, then feather lightly.
51
- choked = np.minimum.reduce([alpha, np.roll(alpha, 1, 0), np.roll(alpha, -1, 0), np.roll(alpha, 1, 1), np.roll(alpha, -1, 1)])
52
- feathered = (choked * 4 + np.roll(choked, 1, 0) + np.roll(choked, -1, 0) + np.roll(choked, 1, 1) + np.roll(choked, -1, 1)) / 8
53
- return feathered
54
-
55
-
56
- def main() -> None:
57
- original, greenscreen, out_dir, first, last = sys.argv[1], sys.argv[2], Path(sys.argv[3]), int(sys.argv[4]), int(sys.argv[5])
58
- out_dir.mkdir(parents=True, exist_ok=True)
59
- count = last - first + 1
60
- for offset, (rgb, green) in enumerate(zip(read_frames(original, first, count), read_frames(greenscreen, first, count))):
61
- alpha = (matte(green) * 255).round().astype(np.uint8)
62
- rgba = np.dstack([rgb, alpha])
63
- Image.fromarray(rgba, "RGBA").save(out_dir / f"f{first + offset}.png", compress_level=1)
64
- print(f"wrote {count} frames to {out_dir}")
65
-
66
-
67
- if __name__ == "__main__":
68
- main()
1
+ """Build a transparent cutout of the subject for a frame range, from ordinary footage.
2
+
3
+ Usage (from the project root):
4
+ python scripts/make_cutout.py <video> <out_dir> <first_frame> <last_frame>
5
+ python scripts/make_cutout.py <video> <out_dir> <first> <last> --alpha-from <clip>
6
+ python scripts/make_cutout.py <video> <out_dir> <first> <last> --model full
7
+
8
+ Writes RGBA PNGs named f<frame>.png (absolute source frame numbers), so a composition can map
9
+ composition frames to cutout frames directly.
10
+
11
+ No green screen. The matte is segmented from the footage itself, because asking a customer to
12
+ shoot against green makes the feature unavailable to almost all of them - most have never done
13
+ it and do not know the term.
14
+
15
+ --alpha-from stays for the case where a background-free or green copy already exists, which is
16
+ faster and still sharper on hair. It is used ONLY for the alpha; colour always comes from the
17
+ original clip, so no green ever reaches the subject's edges.
18
+
19
+ Two things decide whether this reads as real rather than pasted on:
20
+
21
+ * The alpha is smoothed ACROSS frames. A per-frame matte shimmers along the edge, and type
22
+ sitting against someone's head is precisely where the eye catches it. A 3-frame median kills
23
+ single-frame flicker without the lag an average introduces on fast motion.
24
+ * The edge is choked before it is feathered, so the old background cannot leak into the seam.
25
+
26
+ It is worst where it is most visible: flyaway hair, motion blur, a hand crossing the text. If a
27
+ shot will not hold, move the text rather than ship a halo.
28
+ """
29
+
30
+ import argparse
31
+ import json
32
+ import subprocess
33
+ import sys
34
+ from pathlib import Path
35
+
36
+ import numpy as np
37
+ from PIL import Image
38
+ import sys as _sys, pathlib as _pl
39
+ _sys.path.insert(0, str(_pl.Path(__file__).resolve().parent / 'lib'))
40
+ from media import FFMPEG, FFPROBE # resolved per platform - never hardcode a tool path
41
+
42
+ MODELS = {"lite": "ZhengPeng7/BiRefNet_lite", "full": "ZhengPeng7/BiRefNet"}
43
+ SEG_SIZE = 1024
44
+
45
+
46
+ def probe(path: str):
47
+ """Real dimensions and frame rate. Never assume 1920x1080 at 30 - vertical phone footage is normal."""
48
+ out = subprocess.run(
49
+ [FFPROBE, "-v", "quiet", "-print_format", "json", "-show_streams", "-select_streams", "v:0", path],
50
+ capture_output=True, text=True, check=True,
51
+ ).stdout
52
+ s = json.loads(out)["streams"][0]
53
+ num, den = (s.get("r_frame_rate") or "30/1").split("/")
54
+ return int(s["width"]), int(s["height"]), float(num) / float(den or 1)
55
+
56
+
57
+ def read_frames(path: str, first: int, count: int, w: int, h: int, fps: float):
58
+ proc = subprocess.Popen(
59
+ [FFMPEG, "-v", "error", "-ss", f"{first / fps:.6f}", "-i", path, "-frames:v", str(count),
60
+ "-f", "rawvideo", "-pix_fmt", "rgb24", "-"],
61
+ stdout=subprocess.PIPE,
62
+ )
63
+ size = w * h * 3
64
+ for _ in range(count):
65
+ buf = proc.stdout.read(size)
66
+ if len(buf) < size:
67
+ break
68
+ yield np.frombuffer(buf, np.uint8).reshape(h, w, 3)
69
+ proc.stdout.close()
70
+ proc.wait()
71
+
72
+
73
+ def green_matte(green: np.ndarray) -> np.ndarray:
74
+ """Alpha 0..1 from how green each pixel is, relative to the flat background green."""
75
+ g = green.astype(np.float32)
76
+ greenness = g[..., 1] - np.maximum(g[..., 0], g[..., 2])
77
+ background = np.median(greenness[greenness > 60]) if (greenness > 60).any() else 150.0
78
+ lo, hi = 0.2 * background, 0.7 * background
79
+ t = np.clip((greenness - lo) / (hi - lo), 0, 1)
80
+ return 1 - t * t * (3 - 2 * t) # smoothstep: soft, stable edge
81
+
82
+
83
+ def load_segmenter(which: str):
84
+ """BiRefNet, lazily. Importing torch costs seconds; do not pay it on the --alpha-from path."""
85
+ try:
86
+ import torch
87
+ from transformers import AutoModelForImageSegmentation
88
+ except ModuleNotFoundError as exc:
89
+ print(f"Cannot import {exc.name}. The matting model needs the studio's Python environment.", file=sys.stderr)
90
+ print(" see what is installed with: npm run doctor", file=sys.stderr)
91
+ raise SystemExit(2)
92
+
93
+ # trust_remote_code: these repos ship their own model class, which transformers executes.
94
+ model = AutoModelForImageSegmentation.from_pretrained(MODELS[which], trust_remote_code=True)
95
+ model.eval()
96
+ # float32 on CPU: half precision is slower there rather than faster, and can produce NaNs.
97
+ device = "cuda" if torch.cuda.is_available() else "cpu"
98
+ model.to(device)
99
+ return torch, model, device
100
+
101
+
102
+ def segment(torch, model, device, rgb: np.ndarray) -> np.ndarray:
103
+ """Alpha 0..1 for one frame, returned at the frame's own resolution."""
104
+ h, w = rgb.shape[:2]
105
+ img = Image.fromarray(rgb).resize((SEG_SIZE, SEG_SIZE), Image.BILINEAR)
106
+ x = np.asarray(img, dtype=np.float32) / 255.0
107
+ x = (x - np.array([0.485, 0.456, 0.406], np.float32)) / np.array([0.229, 0.224, 0.225], np.float32)
108
+ t = torch.from_numpy(x).permute(2, 0, 1).unsqueeze(0).to(device)
109
+ with torch.no_grad():
110
+ pred = model(t)[-1].sigmoid()[0, 0].cpu().numpy()
111
+ return np.asarray(Image.fromarray((pred * 255).astype(np.uint8)).resize((w, h), Image.BILINEAR), np.float32) / 255.0
112
+
113
+
114
+ def smooth_in_time(alphas):
115
+ """3-frame median. Removes single-frame flicker without the lag an average would add."""
116
+ if len(alphas) < 3:
117
+ return alphas
118
+ out = [alphas[0]]
119
+ for i in range(1, len(alphas) - 1):
120
+ out.append(np.median(np.stack([alphas[i - 1], alphas[i], alphas[i + 1]]), axis=0))
121
+ out.append(alphas[-1])
122
+ return out
123
+
124
+
125
+ def clean_edge(alpha: np.ndarray) -> np.ndarray:
126
+ """Choke by a pixel so the old background cannot leak into the seam, then feather lightly."""
127
+ choked = np.minimum.reduce([alpha, np.roll(alpha, 1, 0), np.roll(alpha, -1, 0), np.roll(alpha, 1, 1), np.roll(alpha, -1, 1)])
128
+ return (choked * 4 + np.roll(choked, 1, 0) + np.roll(choked, -1, 0) + np.roll(choked, 1, 1) + np.roll(choked, -1, 1)) / 8
129
+
130
+
131
+ def main() -> None:
132
+ ap = argparse.ArgumentParser(description="Transparent cutout of the subject, from ordinary footage.")
133
+ ap.add_argument("video")
134
+ ap.add_argument("out_dir", type=Path)
135
+ ap.add_argument("first", type=int)
136
+ ap.add_argument("last", type=int)
137
+ ap.add_argument("--alpha-from", default=None,
138
+ help="Optional background-free or green copy of the SAME take, used only for the alpha.")
139
+ ap.add_argument("--model", choices=sorted(MODELS), default="lite",
140
+ help="lite is several times faster and good enough for most shots; full is sharper on hair.")
141
+ args = ap.parse_args()
142
+
143
+ w, h, fps = probe(args.video)
144
+ count = args.last - args.first + 1
145
+ args.out_dir.mkdir(parents=True, exist_ok=True)
146
+
147
+ frames = list(read_frames(args.video, args.first, count, w, h, fps))
148
+ if not frames:
149
+ print("No frames read - check the frame range against the clip's length.", file=sys.stderr)
150
+ raise SystemExit(1)
151
+
152
+ if args.alpha_from:
153
+ aw, ah, afps = probe(args.alpha_from)
154
+ if (aw, ah) != (w, h) or abs(afps - fps) > 0.01:
155
+ print(f"The alpha clip does not match: {aw}x{ah}@{afps:.2f} against {w}x{h}@{fps:.2f}.", file=sys.stderr)
156
+ print("A mismatched pair makes a halo or a ghost. Re-export it to match, or drop --alpha-from.", file=sys.stderr)
157
+ raise SystemExit(1)
158
+ print(f"Alpha from the supplied clip, colour from the original - {count} frames.")
159
+ alphas = [green_matte(g) for g in read_frames(args.alpha_from, args.first, count, w, h, fps)]
160
+ else:
161
+ print(f"Segmenting {count} frames at {w}x{h} with BiRefNet ({args.model}).")
162
+ print("The first run downloads the model. On a CPU expect a second or two per frame.")
163
+ torch, model, device = load_segmenter(args.model)
164
+ print(f"Running on {device}.")
165
+ alphas = []
166
+ for i, rgb in enumerate(frames):
167
+ alphas.append(segment(torch, model, device, rgb))
168
+ if (i + 1) % 10 == 0 or i + 1 == count:
169
+ print(f" {i + 1}/{count}", flush=True)
170
+
171
+ for i, a in enumerate(smooth_in_time(alphas)):
172
+ alpha = (clean_edge(a).clip(0, 1) * 255).round().astype(np.uint8)
173
+ Image.fromarray(np.dstack([frames[i], alpha]), "RGBA").save(
174
+ args.out_dir / f"f{args.first + i}.png", compress_level=1
175
+ )
176
+
177
+ print(f"wrote {len(alphas)} frames to {args.out_dir}")
178
+
179
+
180
+ if __name__ == "__main__":
181
+ main()
@@ -16,7 +16,12 @@ import {spawnSync} from 'node:child_process';
16
16
  import {existsSync, mkdirSync, readdirSync, readFileSync, writeFileSync} from 'node:fs';
17
17
  import {basename, extname, join} from 'node:path';
18
18
 
19
- const FFMPEG = 'D:\\Tools\\ffmpeg\\bin\\ffmpeg.exe';
19
+ // Never hardcode a tool path. media.mjs finds the ffmpeg that ships with the studio
20
+ // (ffmpeg-static, pulled in by npm install) before it looks at PATH, which is what lets this
21
+ // run on a machine that has never had ffmpeg installed on it. A hardcoded path failed
22
+ // silently here: the sheets were simply never generated and nothing said why.
23
+ import {FFMPEG} from './lib/media.mjs';
24
+
20
25
  const REF_DIR = 'references';
21
26
  const OUT = join('work', 'refs');
22
27
  const all = process.argv.includes('--all');
@@ -13,7 +13,10 @@ import {spawnSync} from 'node:child_process';
13
13
  import {mkdirSync, rmSync, writeFileSync} from 'node:fs';
14
14
  import {join, resolve} from 'node:path';
15
15
 
16
- const FFMPEG = 'D:\\Tools\\ffmpeg\\bin\\ffmpeg.exe';
16
+ // Never hardcode a tool path. media.mjs finds the ffmpeg that ships with the studio
17
+ // (ffmpeg-static, pulled in by npm install) before it looks at PATH, which is what lets this
18
+ // run on a machine that has never had ffmpeg installed on it.
19
+ import {FFMPEG} from './lib/media.mjs';
17
20
 
18
21
  const args = process.argv.slice(2);
19
22
  const id = args.find((a) => !a.startsWith('--'));
@@ -1,88 +1,90 @@
1
- # Brief: {{name}}
2
-
3
- Fill in what you can. Plain words are fine. Leave a line blank if you don't care.
4
- When done: drop your files in this folder's `footage/` (and `screenshots/`), then run
5
- `npm run prep -- {{name}}` and start the session with "start {{name}}".
6
-
7
- ## Basics
8
- Format: 16:9 <!-- 16:9, 9:16, or both -->
9
- Language: en <!-- en, hi, ar ... the language you speak in the video -->
10
- Talking head file: <!-- file name in footage/, e.g. —.mp4. Leave blank if there's only one clip. -->
11
- Green-screen twin: <!-- e.g. —-GS.mp4 — only if you want text behind you -->
12
- Length target: <!-- e.g. 25s, 90s -->
13
-
14
- ## Script
15
- <!-- Paste the script exactly as you read it. Numbers as digits: $99, 11 minutes, 3 days. -->
16
-
17
-
18
- ## The one thing viewers must remember
19
- <!-- One sentence. Everything in the edit serves this. -->
20
-
21
-
22
- ## Must-have moments
23
- <!-- By the words you say, not by time. Plain words are fine.
24
- - When I say "look at this" → my clip shrinks and other creators' clips appear
25
- - "millions of views" → a counter racing up to 1M+ -->
26
-
27
-
28
- ## Moments from references
29
- <!-- You DON'T need the name of the effect. Clip + time + what you like about it.
30
- Format: - clip-file.mp4 mm:ss-mm:ss — what you like (feel words are fine: snappy, floaty, punchy)
31
- - a-reel-you-like.mp4 00:02-00:04 — the colour screens flipping with each phrase, feels punchy -->
32
-
33
-
34
- ## Effects from the library
35
- <!-- Optional: pick any by name (see the list at the bottom), and say where.
36
- - colour flip → "different person, different story" -->
37
-
38
-
39
- ## B-roll
40
- <!-- Which clip where, by the words you say. Mark clips you DON'T want used.
41
- - creators' clips → "all of these people" -->
42
-
43
-
44
- ## Screenshots / things to blur
45
- <!-- Names, emails, prices, logos that must not be readable. -->
46
-
47
-
48
- ## CTA / ending
49
- <!-- What the viewer should do, and what the last frame shows. -->
50
-
51
-
52
- ## Sound
53
- SFX library: audio/ <!-- put sound effects in this folder's audio/, or say "none" -->
54
- Music: added by me in Premiere / Instagram <!-- or a file name in audio/ -->
55
-
56
- ## Don't want
57
- <!-- Anything you dislike: e.g. boxes over my footage, huge text, bouncy motion. -->
58
-
59
-
60
- ---
61
-
62
- ### Effects library (names you can use)
63
- | Name | What it looks like |
64
- |---|---|
65
- | Text behind me | Big headline sits behind your head (needs the green-screen twin) |
66
- | Card swap | You shrink into a rounded card, turn into another clip mid-shrink, a row of cards slides in and drifts |
67
- | Colour flip | Hard cuts to solid accent / ink screens on each phrase; a word rolls into the next and the colours swap |
68
- | Counter race | A number races up from 0 and lands on a value (e.g. 1,000,000+) |
69
- | Zoom-through cut | A jump cut hidden by a quick push-in that eases back — used on every cut by default |
70
- | Gradient CTA | Dark gradient rises from the bottom, small label + big word typed on the VO, then ticked pointers |
71
- | Browser mockup | Screenshots inside a clean browser window that pans and zooms to the detail |
72
- | Pointer click | An arrow glides to a spot and clicks, with a ripple |
73
- | Rolling number | Digits roll from one value to another (e.g. $26.50 → $48.94) |
74
- | Highlight marker | Accent marker wipes behind a key phrase |
75
- | Side card | You move into a card beside the graphics (16:9) or below them (9:16) |
76
- | Headline to pill | A big headline shrinks into a small corner tag that stays on screen |
77
- | Word-by-word captions | On by default (looks: clean, block, highlight, pop) |
78
- | Whip pan | Fast blurred pan between two unrelated shots |
79
- | Push | The next scene pushes the last one out |
80
- | Mask wipe | A soft-edged wipe reveals the next scene |
81
- | Match cut | Cut hidden inside a shared shape or movement |
82
- | Dip | Quick dip through the background colour |
83
- | Speed ramp | Glide into slow-mo (or speed-up) and back out |
84
- | Beat cut | Cuts land on the beat of your music |
85
- | Tighten it | Silence and filler words removed, every join hidden as a punch-in |
86
- | Cut to the music | Whole edit built off the track: slow intro, faster build, hardest on the drop |
87
- | Film look | Grain, halation, light leak, gate weave (looks: clean, film, super8, cinema) |
88
- | Sound pass | Whooshes, impacts and risers placed off the cut list, ducked under the voice |
1
+ # Brief: {{name}}
2
+
3
+ Fill in what you can. Plain words are fine. Leave a line blank if you don't care.
4
+ When done: drop your files in this folder's `footage/` (and `screenshots/`), then run
5
+ `npm run prep -- {{name}}` and start the session with "start {{name}}".
6
+
7
+ ## Basics
8
+ Format: 16:9 <!-- 16:9, 9:16, or both -->
9
+ Language: en <!-- en, hi, ar ... the language you speak in the video -->
10
+ Talking head file: <!-- file name in footage/, e.g. —.mp4. Leave blank if there's only one clip. -->
11
+ Background-free copy: <!-- Optional. Leave blank. Only fill this in if you already export a
12
+ version with the background removed or replaced by flat green; the
13
+ cutout is built automatically from your normal footage otherwise. -->
14
+ Length target: <!-- e.g. 25s, 90s -->
15
+
16
+ ## Script
17
+ <!-- Paste the script exactly as you read it. Numbers as digits: $99, 11 minutes, 3 days. -->
18
+
19
+
20
+ ## The one thing viewers must remember
21
+ <!-- One sentence. Everything in the edit serves this. -->
22
+
23
+
24
+ ## Must-have moments
25
+ <!-- By the words you say, not by time. Plain words are fine.
26
+ - When I say "look at this" → my clip shrinks and other creators' clips appear
27
+ - "millions of views" → a counter racing up to 1M+ -->
28
+
29
+
30
+ ## Moments from references
31
+ <!-- You DON'T need the name of the effect. Clip + time + what you like about it.
32
+ Format: - clip-file.mp4 mm:ss-mm:ss — what you like (feel words are fine: snappy, floaty, punchy)
33
+ - a-reel-you-like.mp4 00:02-00:04 — the colour screens flipping with each phrase, feels punchy -->
34
+
35
+
36
+ ## Effects from the library
37
+ <!-- Optional: pick any by name (see the list at the bottom), and say where.
38
+ - colour flip → "different person, different story" -->
39
+
40
+
41
+ ## B-roll
42
+ <!-- Which clip where, by the words you say. Mark clips you DON'T want used.
43
+ - creators' clips → "all of these people" -->
44
+
45
+
46
+ ## Screenshots / things to blur
47
+ <!-- Names, emails, prices, logos that must not be readable. -->
48
+
49
+
50
+ ## CTA / ending
51
+ <!-- What the viewer should do, and what the last frame shows. -->
52
+
53
+
54
+ ## Sound
55
+ SFX library: audio/ <!-- put sound effects in this folder's audio/, or say "none" -->
56
+ Music: added by me in Premiere / Instagram <!-- or a file name in audio/ -->
57
+
58
+ ## Don't want
59
+ <!-- Anything you dislike: e.g. boxes over my footage, huge text, bouncy motion. -->
60
+
61
+
62
+ ---
63
+
64
+ ### Effects library (names you can use)
65
+ | Name | What it looks like |
66
+ |---|---|
67
+ | Text behind me | Big headline sits behind your head (works from normal footage) |
68
+ | Card swap | You shrink into a rounded card, turn into another clip mid-shrink, a row of cards slides in and drifts |
69
+ | Colour flip | Hard cuts to solid accent / ink screens on each phrase; a word rolls into the next and the colours swap |
70
+ | Counter race | A number races up from 0 and lands on a value (e.g. 1,000,000+) |
71
+ | Zoom-through cut | A jump cut hidden by a quick push-in that eases back — used on every cut by default |
72
+ | Gradient CTA | Dark gradient rises from the bottom, small label + big word typed on the VO, then ticked pointers |
73
+ | Browser mockup | Screenshots inside a clean browser window that pans and zooms to the detail |
74
+ | Pointer click | An arrow glides to a spot and clicks, with a ripple |
75
+ | Rolling number | Digits roll from one value to another (e.g. $26.50 → $48.94) |
76
+ | Highlight marker | Accent marker wipes behind a key phrase |
77
+ | Side card | You move into a card beside the graphics (16:9) or below them (9:16) |
78
+ | Headline to pill | A big headline shrinks into a small corner tag that stays on screen |
79
+ | Word-by-word captions | On by default (looks: clean, block, highlight, pop) |
80
+ | Whip pan | Fast blurred pan between two unrelated shots |
81
+ | Push | The next scene pushes the last one out |
82
+ | Mask wipe | A soft-edged wipe reveals the next scene |
83
+ | Match cut | Cut hidden inside a shared shape or movement |
84
+ | Dip | Quick dip through the background colour |
85
+ | Speed ramp | Glide into slow-mo (or speed-up) and back out |
86
+ | Beat cut | Cuts land on the beat of your music |
87
+ | Tighten it | Silence and filler words removed, every join hidden as a punch-in |
88
+ | Cut to the music | Whole edit built off the track: slow intro, faster build, hardest on the drop |
89
+ | Film look | Grain, halation, light leak, gate weave (looks: clean, film, super8, cinema) |
90
+ | Sound pass | Whooshes, impacts and risers placed off the cut list, ducked under the voice |