creator-editing-studio 1.1.1 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/template/CLAUDE.md +240 -233
- package/template/requirements.txt +23 -17
- package/template/scripts/make_cutout.py +181 -68
- package/template/scripts/refs.mjs +6 -1
- package/template/scripts/stills.mjs +4 -1
- package/template/templates/BRIEF.md +90 -88
package/package.json
CHANGED
package/template/CLAUDE.md
CHANGED
|
@@ -48,257 +48,264 @@ Everything else in `src/design/` is the system and should be left alone. If a vi
|
|
|
48
48
|
a colour or a size that is not in tokens.ts, the answer is to add it to tokens.ts, never to
|
|
49
49
|
type it into a component.
|
|
50
50
|
|
|
51
|
-
## Non-negotiables
|
|
52
|
-
|
|
53
|
-
1. **No magic numbers.** Durations, curves, springs, staggers, blur come from `src/design/motion.ts`.
|
|
54
|
-
Colours, fonts, sizes, spacing come from `src/design/tokens.ts`. If a needed token does not
|
|
55
|
-
exist, add it to the token file (with a comment) instead of inlining a value.
|
|
56
|
-
2. **Timing derives from the spine.** Animation start frames come from word timestamps, beats,
|
|
57
|
-
or other elements' timings via helpers — never hand-typed frame numbers.
|
|
58
|
-
3. **`npm run check` and `motion-check --scale=2` must pass** before any render is shown to the user.
|
|
59
|
-
4. **Verify visually before claiming done** (see Verification). Never say something "looks good"
|
|
60
|
-
without having rendered and viewed the frames.
|
|
61
|
-
|
|
51
|
+
## Non-negotiables
|
|
52
|
+
|
|
53
|
+
1. **No magic numbers.** Durations, curves, springs, staggers, blur come from `src/design/motion.ts`.
|
|
54
|
+
Colours, fonts, sizes, spacing come from `src/design/tokens.ts`. If a needed token does not
|
|
55
|
+
exist, add it to the token file (with a comment) instead of inlining a value.
|
|
56
|
+
2. **Timing derives from the spine.** Animation start frames come from word timestamps, beats,
|
|
57
|
+
or other elements' timings via helpers — never hand-typed frame numbers.
|
|
58
|
+
3. **`npm run check` and `motion-check --scale=2` must pass** before any render is shown to the user.
|
|
59
|
+
4. **Verify visually before claiming done** (see Verification). Never say something "looks good"
|
|
60
|
+
without having rendered and viewed the frames.
|
|
61
|
+
|
|
62
62
|
---
|
|
63
63
|
|
|
64
|
-
## The 12 motion laws
|
|
65
|
-
|
|
66
|
-
1. **No linear easing** — except constant-velocity loops/scrolls (mark with `// motion-ok`).
|
|
67
|
-
2. **Entrances ease-out (`EASE.out`), exits ease-in (`EASE.in`), on-screen A-to-B moves ease-in-out (`EASE.inOut`).**
|
|
68
|
-
3. **Exits are faster than entrances** — `EXIT_RATIO` (0.65). Use `exitFrames()`.
|
|
69
|
-
4. **Max 2–3 animated properties per element.** No fade + scale + rotate + blur together.
|
|
70
|
-
5. **Never opacity alone.** Pair with 8–24px travel or scale >= 0.94. Opacity completes at
|
|
71
|
-
`OPACITY_LEAD` (60%) of the transform.
|
|
72
|
-
6. **Never scale from below 0.9.** Text scaling from zero is the #1 AI-edit tell.
|
|
73
|
-
7. **Blur-in capped at 12px**, fully resolved before the transform ends.
|
|
74
|
-
8. **Groups always stagger** (`STAGGER.tight/normal/dramatic`). Never simultaneous.
|
|
75
|
-
9. **Distance couples to duration sub-linearly** (`travelFrames()`), never linearly.
|
|
76
|
-
10. **Anything that moves fast is motion-blurred** (see Smoothness system). No hard edge may jump
|
|
77
|
-
across the frame unblurred.
|
|
78
|
-
11. **Readable hold:** text stays fully visible >= `readFrames(text)` (min 0.9s, +55ms/word).
|
|
79
|
-
12. **Continuity:** every transition inherits the outgoing motion (see below).
|
|
80
|
-
|
|
81
|
-
### Continuity (what makes it feel connected, not cut together)
|
|
82
|
-
|
|
83
|
-
- **Shared elements never re-enter.** If an element exists across two beats, `morph()` its
|
|
84
|
-
position/size/radius between layouts; never exit and re-animate it. (TransitionTest: the
|
|
85
|
-
eyebrow dot stays on screen and grows into the next card's icon tile.)
|
|
86
|
-
- **Velocity handoff.** An exit toward a direction is answered by an entrance continuing that
|
|
87
|
-
direction — `handoff(exitTo)`.
|
|
88
|
-
- **Overlap, never gap.** The next beat starts before the previous one settles —
|
|
89
|
-
`nextBeatAt(prevSettled)`.
|
|
90
|
-
- **One camera over cuts.** Lay scenes out on one canvas and glide a camera between them
|
|
91
|
-
(`src/lib/camera.ts`) instead of hard-cutting between unrelated layouts. Content that should
|
|
92
|
-
stay put during a move (shared elements) lives in screen space, outside the camera.
|
|
93
|
-
- **Nothing is cut off.** Compute exit start times backward from the end so the slowest,
|
|
94
|
-
most-staggered exit completes on or before the last frame.
|
|
95
|
-
- **Sound glues cuts.** Whooshes start `TRANSITION.sfxLeadMs` before a visual change and tail
|
|
96
|
-
`sfxTailMs` after.
|
|
97
|
-
|
|
64
|
+
## The 12 motion laws
|
|
65
|
+
|
|
66
|
+
1. **No linear easing** — except constant-velocity loops/scrolls (mark with `// motion-ok`).
|
|
67
|
+
2. **Entrances ease-out (`EASE.out`), exits ease-in (`EASE.in`), on-screen A-to-B moves ease-in-out (`EASE.inOut`).**
|
|
68
|
+
3. **Exits are faster than entrances** — `EXIT_RATIO` (0.65). Use `exitFrames()`.
|
|
69
|
+
4. **Max 2–3 animated properties per element.** No fade + scale + rotate + blur together.
|
|
70
|
+
5. **Never opacity alone.** Pair with 8–24px travel or scale >= 0.94. Opacity completes at
|
|
71
|
+
`OPACITY_LEAD` (60%) of the transform.
|
|
72
|
+
6. **Never scale from below 0.9.** Text scaling from zero is the #1 AI-edit tell.
|
|
73
|
+
7. **Blur-in capped at 12px**, fully resolved before the transform ends.
|
|
74
|
+
8. **Groups always stagger** (`STAGGER.tight/normal/dramatic`). Never simultaneous.
|
|
75
|
+
9. **Distance couples to duration sub-linearly** (`travelFrames()`), never linearly.
|
|
76
|
+
10. **Anything that moves fast is motion-blurred** (see Smoothness system). No hard edge may jump
|
|
77
|
+
across the frame unblurred.
|
|
78
|
+
11. **Readable hold:** text stays fully visible >= `readFrames(text)` (min 0.9s, +55ms/word).
|
|
79
|
+
12. **Continuity:** every transition inherits the outgoing motion (see below).
|
|
80
|
+
|
|
81
|
+
### Continuity (what makes it feel connected, not cut together)
|
|
82
|
+
|
|
83
|
+
- **Shared elements never re-enter.** If an element exists across two beats, `morph()` its
|
|
84
|
+
position/size/radius between layouts; never exit and re-animate it. (TransitionTest: the
|
|
85
|
+
eyebrow dot stays on screen and grows into the next card's icon tile.)
|
|
86
|
+
- **Velocity handoff.** An exit toward a direction is answered by an entrance continuing that
|
|
87
|
+
direction — `handoff(exitTo)`.
|
|
88
|
+
- **Overlap, never gap.** The next beat starts before the previous one settles —
|
|
89
|
+
`nextBeatAt(prevSettled)`.
|
|
90
|
+
- **One camera over cuts.** Lay scenes out on one canvas and glide a camera between them
|
|
91
|
+
(`src/lib/camera.ts`) instead of hard-cutting between unrelated layouts. Content that should
|
|
92
|
+
stay put during a move (shared elements) lives in screen space, outside the camera.
|
|
93
|
+
- **Nothing is cut off.** Compute exit start times backward from the end so the slowest,
|
|
94
|
+
most-staggered exit completes on or before the last frame.
|
|
95
|
+
- **Sound glues cuts.** Whooshes start `TRANSITION.sfxLeadMs` before a visual change and tail
|
|
96
|
+
`sfxTailMs` after.
|
|
97
|
+
|
|
98
98
|
---
|
|
99
99
|
|
|
100
|
-
## Smoothness system (every one of these was a real defect found by motion-check)
|
|
101
|
-
|
|
102
|
-
- **Use the components, not raw transforms.** `<Presence>`, `<MaskReveal>`, `<HighlightMark>`,
|
|
103
|
-
`<RollingNumber>` add velocity motion blur automatically. Anything else that moves wraps its
|
|
104
|
-
moving element in `<VelocityBlur vx vy>` with `perFrame()` velocities.
|
|
105
|
-
- **Velocity motion blur** (`VelocityBlur`): directional blur sized to this frame's movement with
|
|
106
|
-
a 180-degree shutter; fades to zero as the element settles. Near-free to render.
|
|
107
|
-
- **Wipes** use `wipeMask(p, perFrame(...))`: the leading edge softens with speed. Never wipe with
|
|
108
|
-
a hard `clipPath` edge.
|
|
109
|
-
- **Camera moves**: wrap the camera canvas in a full-frame `VelocityBlur` with `bleed={0}` and
|
|
110
|
-
`overflow: hidden`, velocity from `cameraShot` (see TransitionTest).
|
|
111
|
-
- **Zooms / rotations** (not translations): use `<MotionBlur active>` (temporal sampling, centred
|
|
112
|
-
on the frame so toggling it never shifts timing). It costs 10x render time while active —
|
|
113
|
-
enable only during the move. Never use `@remotion/motion-blur`'s `CameraMotionBlur`: it samples
|
|
114
|
-
ahead of the frame, so switching it on/off causes a timing hitch.
|
|
115
|
-
- **Counters roll, never tick.** Use `<RollingNumber>`. A counting number changes glyphs in single
|
|
116
|
-
frames as it slows, which motion-check flags as pops.
|
|
117
|
-
- **Never `will-change`.** It makes each parallel render tab cache text at a different sub-pixel
|
|
118
|
-
position, so settled elements flicker between 4 versions. (Lint enforces this.)
|
|
119
|
-
- **Screen recordings used as b-roll usually scroll in steps** (they judder). Use a still frame with a
|
|
120
|
-
smooth `ScreenView` pan instead, and blur names on the still.
|
|
121
|
-
- **Presence `dur` sets entrance AND exit length.** `dur="instant"` makes exits last 3 frames (a snap) —
|
|
122
|
-
for containers that only fade in, use at least `"base"`.
|
|
123
|
-
- **Final renders are supersampled at `--scale=2`**, then downscaled to 1080x1920. At 1x, Chrome
|
|
124
|
-
snaps text to whole pixels, so the slow tail of every ease-out steps ("move 1px, hold, move 1px")
|
|
125
|
-
and motion-check reports stutters. Run motion-check with the same `--scale`.
|
|
126
|
-
- **Frame rate:** match the footage. If footage is 60fps, build at 60 — per-frame jumps halve.
|
|
127
|
-
All tokens are in ms, so compositions work at either rate.
|
|
128
|
-
|
|
100
|
+
## Smoothness system (every one of these was a real defect found by motion-check)
|
|
101
|
+
|
|
102
|
+
- **Use the components, not raw transforms.** `<Presence>`, `<MaskReveal>`, `<HighlightMark>`,
|
|
103
|
+
`<RollingNumber>` add velocity motion blur automatically. Anything else that moves wraps its
|
|
104
|
+
moving element in `<VelocityBlur vx vy>` with `perFrame()` velocities.
|
|
105
|
+
- **Velocity motion blur** (`VelocityBlur`): directional blur sized to this frame's movement with
|
|
106
|
+
a 180-degree shutter; fades to zero as the element settles. Near-free to render.
|
|
107
|
+
- **Wipes** use `wipeMask(p, perFrame(...))`: the leading edge softens with speed. Never wipe with
|
|
108
|
+
a hard `clipPath` edge.
|
|
109
|
+
- **Camera moves**: wrap the camera canvas in a full-frame `VelocityBlur` with `bleed={0}` and
|
|
110
|
+
`overflow: hidden`, velocity from `cameraShot` (see TransitionTest).
|
|
111
|
+
- **Zooms / rotations** (not translations): use `<MotionBlur active>` (temporal sampling, centred
|
|
112
|
+
on the frame so toggling it never shifts timing). It costs 10x render time while active —
|
|
113
|
+
enable only during the move. Never use `@remotion/motion-blur`'s `CameraMotionBlur`: it samples
|
|
114
|
+
ahead of the frame, so switching it on/off causes a timing hitch.
|
|
115
|
+
- **Counters roll, never tick.** Use `<RollingNumber>`. A counting number changes glyphs in single
|
|
116
|
+
frames as it slows, which motion-check flags as pops.
|
|
117
|
+
- **Never `will-change`.** It makes each parallel render tab cache text at a different sub-pixel
|
|
118
|
+
position, so settled elements flicker between 4 versions. (Lint enforces this.)
|
|
119
|
+
- **Screen recordings used as b-roll usually scroll in steps** (they judder). Use a still frame with a
|
|
120
|
+
smooth `ScreenView` pan instead, and blur names on the still.
|
|
121
|
+
- **Presence `dur` sets entrance AND exit length.** `dur="instant"` makes exits last 3 frames (a snap) —
|
|
122
|
+
for containers that only fade in, use at least `"base"`.
|
|
123
|
+
- **Final renders are supersampled at `--scale=2`**, then downscaled to 1080x1920. At 1x, Chrome
|
|
124
|
+
snaps text to whole pixels, so the slow tail of every ease-out steps ("move 1px, hold, move 1px")
|
|
125
|
+
and motion-check reports stutters. Run motion-check with the same `--scale`.
|
|
126
|
+
- **Frame rate:** match the footage. If footage is 60fps, build at 60 — per-frame jumps halve.
|
|
127
|
+
All tokens are in ms, so compositions work at either rate.
|
|
128
|
+
|
|
129
129
|
---
|
|
130
130
|
|
|
131
|
-
## Design rules
|
|
132
|
-
|
|
133
|
-
**Read `STYLE.md` before any design decision.** It is the house style (type scale, layout, motion,
|
|
134
|
-
sound, approved and rejected beats) and wins over anything below or in a reference. Update it
|
|
135
|
-
whenever the user approves or rejects something.
|
|
136
|
-
|
|
137
|
-
- Everything visual comes from `src/design/tokens.ts`, which is yours to define.
|
|
138
|
-
- **Brand essentials:** your display typeface only, at every weight (headlines/stats 800). One accent: accent
|
|
139
|
-
`your accent` on ink and cream/white. No blue, no purple, no serif. Signature emphasis is the accent
|
|
140
|
-
marker behind the key phrase — use `<HighlightMark>`; don't invent other emphasis styles.
|
|
141
|
-
Headlines are short two-part fragments with the payoff highlighted. Numbers are shown raw and
|
|
142
|
-
bold ("$1.3M", "11x") and roll in. Eyebrows are uppercase pills with a accent-ringed dot.
|
|
143
|
-
Radii are generous, shadows soft, buttons/badges pills.
|
|
144
|
-
- Secondary text uses `COLOR.fgSecondary`; `COLOR.fgMuted` fails contrast on cream — decoration only.
|
|
145
|
-
- **
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
-
|
|
165
|
-
|
|
166
|
-
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
-
|
|
170
|
-
|
|
131
|
+
## Design rules
|
|
132
|
+
|
|
133
|
+
**Read `STYLE.md` before any design decision.** It is the house style (type scale, layout, motion,
|
|
134
|
+
sound, approved and rejected beats) and wins over anything below or in a reference. Update it
|
|
135
|
+
whenever the user approves or rejects something.
|
|
136
|
+
|
|
137
|
+
- Everything visual comes from `src/design/tokens.ts`, which is yours to define.
|
|
138
|
+
- **Brand essentials:** your display typeface only, at every weight (headlines/stats 800). One accent: accent
|
|
139
|
+
`your accent` on ink and cream/white. No blue, no purple, no serif. Signature emphasis is the accent
|
|
140
|
+
marker behind the key phrase — use `<HighlightMark>`; don't invent other emphasis styles.
|
|
141
|
+
Headlines are short two-part fragments with the payoff highlighted. Numbers are shown raw and
|
|
142
|
+
bold ("$1.3M", "11x") and roll in. Eyebrows are uppercase pills with a accent-ringed dot.
|
|
143
|
+
Radii are generous, shadows soft, buttons/badges pills.
|
|
144
|
+
- Secondary text uses `COLOR.fgSecondary`; `COLOR.fgMuted` fails contrast on cream — decoration only.
|
|
145
|
+
- **Your style decisions:** *(empty on purpose — this fills in as you make them)*
|
|
146
|
+
|
|
147
|
+
When something gets settled about how your videos look, write it here as one line. Where
|
|
148
|
+
captions sit in a card layout. Whether headlines take the accent colour or the ink. What a
|
|
149
|
+
corner tag does when the video returns to full screen. This file is read first every
|
|
150
|
+
session, so a decision recorded here is one you never have to make twice, and the studio
|
|
151
|
+
stops asking.
|
|
152
|
+
|
|
153
|
+
Do not fill this with defaults copied from somewhere else. An empty section produces
|
|
154
|
+
questions; a borrowed one produces someone else's videos.
|
|
155
|
+
|
|
156
|
+
The one rule here that is not a preference: brand logos are **real full-colour logos**
|
|
157
|
+
(`LogoIcon`, SVG Logos CC0 via `src/design/logos.json`, extract only the icons used). Never
|
|
158
|
+
approximate a logo that is not in the set. Use a neutral icon and ask for the official file.
|
|
159
|
+
- **Text behind the subject must stay readable:** the head covers only the lower ~20–30% of the
|
|
160
|
+
headline. Head height changes shot to shot (leaning forward lifts it ~80px) — check a still of
|
|
161
|
+
EVERY hero moment and raise the block (`top`; put labels above as `eyebrow`) where needed.
|
|
162
|
+
- Keep accent highlights away from the accent background glow so they don't disappear into it.
|
|
163
|
+
- Stats use proportional figures, not tabular (tabular makes "11" look gappy in your display typeface).
|
|
164
|
+
- **Web design systems do not map 1:1 to video.** Keep colour, font families/weights, type-scale
|
|
165
|
+
ratios, radii, spacing rhythm, iconography and imagery style — but rescale sizes for a
|
|
166
|
+
1080-wide frame viewed on a phone (body roughly 40–48px, not 16px). Web UI components mostly
|
|
167
|
+
do not apply.
|
|
168
|
+
- Respect safe zones: `LAYOUT.safeTop` / `LAYOUT.safeBottom` for Reels/Shorts UI.
|
|
169
|
+
- Text contrast >= 4.5:1 against whatever is behind it, including over footage.
|
|
170
|
+
- Draw graphics in code (SVG/CSS) so every part can animate. Use flat images only for things
|
|
171
|
+
that must be real (logos, screenshots, photos). Avoid AI-generated imagery for anything the
|
|
172
|
+
viewer is meant to look at.
|
|
173
|
+
- Icons: SVG drawn in code or an installed icon library, never image files.
|
|
174
|
+
|
|
171
175
|
---
|
|
172
176
|
|
|
173
|
-
## Compositing: text behind the subject
|
|
174
|
-
|
|
175
|
-
Layer order, bottom to top:
|
|
176
|
-
|
|
177
|
-
```
|
|
178
|
-
background plate (designed, or the original footage)
|
|
179
|
-
text / graphics <- 1–3px blur to sit on the background's focal plane
|
|
180
|
-
subject cutout <- same clip, frame-aligned
|
|
181
|
-
light wrap <- blurred background screened onto the subject's inner edge, 15–30%
|
|
182
|
-
grain + grade <- applied over everything so layers share one look
|
|
183
|
-
```
|
|
184
|
-
|
|
185
|
-
- **The cutout must come from the same clip as its background plate**, frame-aligned. Never
|
|
186
|
-
combine a cutout from one take with footage from another.
|
|
187
|
-
- **
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
-
|
|
201
|
-
|
|
202
|
-
-
|
|
203
|
-
|
|
204
|
-
|
|
177
|
+
## Compositing: text behind the subject
|
|
178
|
+
|
|
179
|
+
Layer order, bottom to top:
|
|
180
|
+
|
|
181
|
+
```
|
|
182
|
+
background plate (designed, or the original footage)
|
|
183
|
+
text / graphics <- 1–3px blur to sit on the background's focal plane
|
|
184
|
+
subject cutout <- same clip, frame-aligned
|
|
185
|
+
light wrap <- blurred background screened onto the subject's inner edge, 15–30%
|
|
186
|
+
grain + grade <- applied over everything so layers share one look
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
- **The cutout must come from the same clip as its background plate**, frame-aligned. Never
|
|
190
|
+
combine a cutout from one take with footage from another.
|
|
191
|
+
- **The matte is built automatically from ordinary footage.** No green screen, and never ask the
|
|
192
|
+
user for one — most people have never shot against green and do not know the term. Segment the
|
|
193
|
+
subject per frame, then **temporally smooth the alpha**, which is not optional: a per-frame matte
|
|
194
|
+
shimmers along the edge, and type sitting against a head is exactly where that shows. Finish with
|
|
195
|
+
spill-free cleanup: ~0.5px choke, 0.5–1px feather.
|
|
196
|
+
- **Quality is worst where it is most visible:** flyaway hair, motion blur, and a hand crossing in
|
|
197
|
+
front of the text. The light wrap below hides a good deal of it. If a shot simply will not hold,
|
|
198
|
+
say so and move the text rather than shipping a halo.
|
|
199
|
+
- **If the user happens to have a background-free or green version, prefer it** — chroma keying
|
|
200
|
+
still beats automatic matting on hair. Use it **only for the alpha** and take the subject's colour
|
|
201
|
+
from the original, so no green ever reaches the edges. Verify the pair first: identical duration,
|
|
202
|
+
fps, resolution and frame count, subject aligned on the same frame (one frame out makes a ghost).
|
|
203
|
+
This is an upgrade for people who already work that way, never a requirement.
|
|
204
|
+
- **Supplied cutouts** must have real transparency (ProRes 4444, WebM with alpha, or PNG sequence).
|
|
205
|
+
An MP4 cannot hold transparency — if one arrives, stop and tell the user.
|
|
206
|
+
- Play transparent video with `<OffthreadVideo transparent src={staticFile(...)} />`.
|
|
207
|
+
- Text behind a moving subject gets 0.85–0.95x counter-parallax; static text reads as a sticker.
|
|
208
|
+
- Text should animate in while already partly occluded, not appear fully and then get covered.
|
|
209
|
+
- Pull text colour toward the footage's black/white points; pure white over graded footage looks
|
|
210
|
+
pasted on.
|
|
211
|
+
|
|
205
212
|
---
|
|
206
213
|
|
|
207
|
-
## Per-video pipeline
|
|
208
|
-
|
|
209
|
-
1. **Ingest:** probe every clip (fps, resolution, colour, alpha).
|
|
210
|
-
- **Phone footage is usually variable frame rate (VFR).** Conform it to constant frame rate
|
|
211
|
-
with ffmpeg before use, or it stutters and drifts out of sync in Remotion.
|
|
212
|
-
- Conform everything to one fps. Never mix frame rates.
|
|
213
|
-
2. **Spine:** word-level timings + music beats -> `work/<video>/spine.json`.
|
|
214
|
-
Cut pauses/filler from word gaps. All animation timing reads from the spine.
|
|
215
|
-
- If the user supplies a script or SRT, it is the **text truth** (spelling, names, brand words):
|
|
216
|
-
force-align that exact text to the audio with WhisperX's aligner. Never use SRT timings
|
|
217
|
-
directly — they are line-level and padded for readability, not word-accurate.
|
|
218
|
-
- With no script, transcribe with WhisperX, then show the transcript for correction before building.
|
|
219
|
-
3. **Cutouts:**
|
|
220
|
-
4. **Beat sheet:** from the script and the user's marked punch lines, write a short plan — which
|
|
221
|
-
moment gets which treatment — and confirm with the user before building.
|
|
222
|
-
- **Read their saved references first** (`editing_os_references`). They are the whole reason
|
|
223
|
-
the edit should look like theirs rather than like the defaults, and a plan written without
|
|
224
|
-
them is a plan written for anybody. Say which reference each decision came from, so the
|
|
225
|
-
user can see their own taste being applied and correct it when it is being misread.
|
|
226
|
-
- Where references are silent, the defaults in STYLE.md apply. Where references and defaults
|
|
227
|
-
disagree, the reference wins — it is a real preference, and the default is only a sensible
|
|
228
|
-
starting point.
|
|
229
|
-
5. **Build** in `src/compositions/<Video>.tsx` from components + helpers + tokens only.
|
|
230
|
-
6. **Verify** (below), then show the user.
|
|
231
|
-
7. **Sound:** the user adds SFX and music themselves. Deliver a cue list (timestamps of every
|
|
232
|
-
transition and hit, with suggested whoosh start = `TRANSITION.sfxLeadMs` before the change) so
|
|
233
|
-
they can place sounds quickly. Render without music unless asked.
|
|
234
|
-
8. **Render** at `--scale=2`, downscaled to 1080x1920 unless told otherwise.
|
|
235
|
-
|
|
214
|
+
## Per-video pipeline
|
|
215
|
+
|
|
216
|
+
1. **Ingest:** probe every clip (fps, resolution, colour, alpha).
|
|
217
|
+
- **Phone footage is usually variable frame rate (VFR).** Conform it to constant frame rate
|
|
218
|
+
with ffmpeg before use, or it stutters and drifts out of sync in Remotion.
|
|
219
|
+
- Conform everything to one fps. Never mix frame rates.
|
|
220
|
+
2. **Spine:** word-level timings + music beats -> `work/<video>/spine.json`.
|
|
221
|
+
Cut pauses/filler from word gaps. All animation timing reads from the spine.
|
|
222
|
+
- If the user supplies a script or SRT, it is the **text truth** (spelling, names, brand words):
|
|
223
|
+
force-align that exact text to the audio with WhisperX's aligner. Never use SRT timings
|
|
224
|
+
directly — they are line-level and padded for readability, not word-accurate.
|
|
225
|
+
- With no script, transcribe with WhisperX, then show the transcript for correction before building.
|
|
226
|
+
3. **Cutouts:** build the matte from the footage automatically, or validate a supplied cutout (see above). Only needed for shots that want text behind the subject.
|
|
227
|
+
4. **Beat sheet:** from the script and the user's marked punch lines, write a short plan — which
|
|
228
|
+
moment gets which treatment — and confirm with the user before building.
|
|
229
|
+
- **Read their saved references first** (`editing_os_references`). They are the whole reason
|
|
230
|
+
the edit should look like theirs rather than like the defaults, and a plan written without
|
|
231
|
+
them is a plan written for anybody. Say which reference each decision came from, so the
|
|
232
|
+
user can see their own taste being applied and correct it when it is being misread.
|
|
233
|
+
- Where references are silent, the defaults in STYLE.md apply. Where references and defaults
|
|
234
|
+
disagree, the reference wins — it is a real preference, and the default is only a sensible
|
|
235
|
+
starting point.
|
|
236
|
+
5. **Build** in `src/compositions/<Video>.tsx` from components + helpers + tokens only.
|
|
237
|
+
6. **Verify** (below), then show the user.
|
|
238
|
+
7. **Sound:** the user adds SFX and music themselves. Deliver a cue list (timestamps of every
|
|
239
|
+
transition and hit, with suggested whoosh start = `TRANSITION.sfxLeadMs` before the change) so
|
|
240
|
+
they can place sounds quickly. Render without music unless asked.
|
|
241
|
+
8. **Render** at `--scale=2`, downscaled to 1080x1920 unless told otherwise.
|
|
242
|
+
|
|
236
243
|
---
|
|
237
244
|
|
|
238
|
-
## Audio standard (audio is half the reel — the user will not accept a bad mix)
|
|
239
|
-
|
|
240
|
-
What went wrong on an early video: the voice itself was bit-for-bit intact (measured against the source),
|
|
241
|
-
but 2.3–3.5s whoosh files and a 2.4s riser kept playing UNDER the speech, which made the voice sound
|
|
242
|
-
wrong. Claude cannot hear, so the mix must be engineered and measured, never guessed.
|
|
243
|
-
|
|
244
|
-
1. **The voice is never touched.** Render the picture `--muted`. The voice in every deliverable is the
|
|
245
|
-
original recording's audio (stream-copied, or at most one encode), plus only a single gain change.
|
|
246
|
-
No limiter, no dynamic normaliser (ffmpeg `loudnorm` single-pass pumps), no repeated AAC encodes.
|
|
247
|
-
2. **Sound effects are short and placed in gaps.** Trim every SFX to what the motion needs (whoosh
|
|
248
|
-
0.4–0.8s, hits 0.3–0.6s) with a 60–120ms fade-out. Nothing longer than ~0.8s may sit under a spoken
|
|
249
|
-
word unless it is at least 24 dB under the voice. Risers go in pauses, not under speech.
|
|
250
|
-
3. **Duck SFX under the voice**: mix with ffmpeg `sidechaincompress` keyed from the voice (SFX drop
|
|
251
|
-
~6 dB while a word is spoken), and high-pass whooshes/impacts around 120–150 Hz so they don't muddy it.
|
|
252
|
-
4. **Measure before delivering** (the window scan in the project folder used `f32le` PCM from ffmpeg):
|
|
253
|
-
voice-vs-mix difference per 0.5s window — any window where SFX energy is within 12 dB of the
|
|
254
|
-
voice while a word is spoken gets fixed. Final: -14 LUFS integrated, true peak <= -1 dBTP, one
|
|
255
|
-
AAC encode at 256k.
|
|
256
|
-
5. **Always also deliver stems + a cue sheet**: `voice.wav`, `sfx.wav` (the placed, trimmed SFX bus),
|
|
257
|
-
and `SFX-CUES.txt` with numbered, timecoded SFX files — so the user can rebalance in Premiere in
|
|
258
|
-
minutes. Picture-only MP4 alongside.
|
|
259
|
-
6. The user picks and judges sounds by ear; Claude says plainly that it cannot hear and asks for
|
|
260
|
-
timecoded audio notes on the first draft.
|
|
261
|
-
|
|
245
|
+
## Audio standard (audio is half the reel — the user will not accept a bad mix)
|
|
246
|
+
|
|
247
|
+
What went wrong on an early video: the voice itself was bit-for-bit intact (measured against the source),
|
|
248
|
+
but 2.3–3.5s whoosh files and a 2.4s riser kept playing UNDER the speech, which made the voice sound
|
|
249
|
+
wrong. Claude cannot hear, so the mix must be engineered and measured, never guessed.
|
|
250
|
+
|
|
251
|
+
1. **The voice is never touched.** Render the picture `--muted`. The voice in every deliverable is the
|
|
252
|
+
original recording's audio (stream-copied, or at most one encode), plus only a single gain change.
|
|
253
|
+
No limiter, no dynamic normaliser (ffmpeg `loudnorm` single-pass pumps), no repeated AAC encodes.
|
|
254
|
+
2. **Sound effects are short and placed in gaps.** Trim every SFX to what the motion needs (whoosh
|
|
255
|
+
0.4–0.8s, hits 0.3–0.6s) with a 60–120ms fade-out. Nothing longer than ~0.8s may sit under a spoken
|
|
256
|
+
word unless it is at least 24 dB under the voice. Risers go in pauses, not under speech.
|
|
257
|
+
3. **Duck SFX under the voice**: mix with ffmpeg `sidechaincompress` keyed from the voice (SFX drop
|
|
258
|
+
~6 dB while a word is spoken), and high-pass whooshes/impacts around 120–150 Hz so they don't muddy it.
|
|
259
|
+
4. **Measure before delivering** (the window scan in the project folder used `f32le` PCM from ffmpeg):
|
|
260
|
+
voice-vs-mix difference per 0.5s window — any window where SFX energy is within 12 dB of the
|
|
261
|
+
voice while a word is spoken gets fixed. Final: -14 LUFS integrated, true peak <= -1 dBTP, one
|
|
262
|
+
AAC encode at 256k.
|
|
263
|
+
5. **Always also deliver stems + a cue sheet**: `voice.wav`, `sfx.wav` (the placed, trimmed SFX bus),
|
|
264
|
+
and `SFX-CUES.txt` with numbered, timecoded SFX files — so the user can rebalance in Premiere in
|
|
265
|
+
minutes. Picture-only MP4 alongside.
|
|
266
|
+
6. The user picks and judges sounds by ear; Claude says plainly that it cannot hear and asks for
|
|
267
|
+
timecoded audio notes on the first draft.
|
|
268
|
+
|
|
262
269
|
---
|
|
263
270
|
|
|
264
|
-
## Verification (required before saying anything is done)
|
|
265
|
-
|
|
266
|
-
Claude cannot watch video. Quality is checked with frames and maths, and the user makes the
|
|
267
|
-
final call at playback speed. Be explicit about that when reporting.
|
|
268
|
-
|
|
269
|
-
1. `npm run check` passes.
|
|
270
|
-
2. `npm run motion-check -- <Comp> --scale=2` reports **no POP, CUT or STUTTER**. Allow intentional
|
|
271
|
-
cuts with `--allow=<frame>`. For each FAST warning, render that frame and confirm the moving
|
|
272
|
-
edge is blurred. Read the motion timeline for dead air and rhythm.
|
|
273
|
-
3. Render stills and **look at them**: first frame, each element mid-entrance, peak-speed frames
|
|
274
|
-
(blur visible), fully settled (crisp), mid-exit, and the **last frame** (clean).
|
|
275
|
-
4. Check: no text collides with the subject's face, safe zones respected, contrast holds,
|
|
276
|
-
no clipped descenders (g, p, y, commas) in masked text.
|
|
277
|
-
5. Report what was verified and what still needs the user's eyes at full speed.
|
|
278
|
-
|
|
271
|
+
## Verification (required before saying anything is done)
|
|
272
|
+
|
|
273
|
+
Claude cannot watch video. Quality is checked with frames and maths, and the user makes the
|
|
274
|
+
final call at playback speed. Be explicit about that when reporting.
|
|
275
|
+
|
|
276
|
+
1. `npm run check` passes.
|
|
277
|
+
2. `npm run motion-check -- <Comp> --scale=2` reports **no POP, CUT or STUTTER**. Allow intentional
|
|
278
|
+
cuts with `--allow=<frame>`. For each FAST warning, render that frame and confirm the moving
|
|
279
|
+
edge is blurred. Read the motion timeline for dead air and rhythm.
|
|
280
|
+
3. Render stills and **look at them**: first frame, each element mid-entrance, peak-speed frames
|
|
281
|
+
(blur visible), fully settled (crisp), mid-exit, and the **last frame** (clean).
|
|
282
|
+
4. Check: no text collides with the subject's face, safe zones respected, contrast holds,
|
|
283
|
+
no clipped descenders (g, p, y, commas) in masked text.
|
|
284
|
+
5. Report what was verified and what still needs the user's eyes at full speed.
|
|
285
|
+
|
|
279
286
|
---
|
|
280
287
|
|
|
281
|
-
## Remotion gotchas
|
|
282
|
-
|
|
283
|
-
- Never use CSS `transition`/`animation` — drive everything from `useCurrentFrame()`.
|
|
284
|
-
- Never use `Math.random()` — use `random(seed)` from `remotion`.
|
|
285
|
-
- Never use `will-change` (render-tab flicker, see Smoothness system).
|
|
286
|
-
- **`<Presence>` owns its element's `transform` and `opacity`.** Any extra transform (centring with
|
|
287
|
-
`translateX(-50%)`) or opacity (fading content) passed in `style` is silently overwritten — put it
|
|
288
|
-
on an inner wrapper. This bug appeared twice in the first project (off-centre tag, labels that
|
|
289
|
-
never faded).
|
|
290
|
-
- em-based spacing (gaps, padding) resolves against the element's own font-size: set `fontSize`
|
|
291
|
-
on the container, or word gaps collapse.
|
|
292
|
-
- Masked text needs travel > 100% and container padding, or descenders stay visible.
|
|
293
|
-
- Load fonts through `@remotion/fonts` (brand your display typeface is local in `assets/fonts/`) so renders wait for them.
|
|
294
|
-
|
|
288
|
+
## Remotion gotchas
|
|
289
|
+
|
|
290
|
+
- Never use CSS `transition`/`animation` — drive everything from `useCurrentFrame()`.
|
|
291
|
+
- Never use `Math.random()` — use `random(seed)` from `remotion`.
|
|
292
|
+
- Never use `will-change` (render-tab flicker, see Smoothness system).
|
|
293
|
+
- **`<Presence>` owns its element's `transform` and `opacity`.** Any extra transform (centring with
|
|
294
|
+
`translateX(-50%)`) or opacity (fading content) passed in `style` is silently overwritten — put it
|
|
295
|
+
on an inner wrapper. This bug appeared twice in the first project (off-centre tag, labels that
|
|
296
|
+
never faded).
|
|
297
|
+
- em-based spacing (gaps, padding) resolves against the element's own font-size: set `fontSize`
|
|
298
|
+
on the container, or word gaps collapse.
|
|
299
|
+
- Masked text needs travel > 100% and container padding, or descenders stay visible.
|
|
300
|
+
- Load fonts through `@remotion/fonts` (brand your display typeface is local in `assets/fonts/`) so renders wait for them.
|
|
301
|
+
|
|
295
302
|
---
|
|
296
303
|
|
|
297
|
-
## Feedback format from the user
|
|
298
|
-
|
|
299
|
-
Notes arrive as timecode + specific issue ("0:14 title lands 3 frames late, bounce too strong").
|
|
300
|
-
Translate each note into token or timing changes; if a note conflicts with a law, say so and ask.
|
|
301
|
-
|
|
304
|
+
## Feedback format from the user
|
|
305
|
+
|
|
306
|
+
Notes arrive as timecode + specific issue ("0:14 title lands 3 frames late, bounce too strong").
|
|
307
|
+
Translate each note into token or timing changes; if a note conflicts with a law, say so and ask.
|
|
308
|
+
|
|
302
309
|
---
|
|
303
310
|
|
|
304
311
|
## What is already installed
|
|
@@ -1,17 +1,23 @@
|
|
|
1
|
-
# The Python half of the studio: transcription, word alignment, cut detection,
|
|
2
|
-
# green-screen cutouts and colour grading.
|
|
3
|
-
#
|
|
4
|
-
# python -m venv .venv
|
|
5
|
-
# .venv/bin/pip install -r requirements.txt (Windows: .venv\Scripts\pip)
|
|
6
|
-
#
|
|
7
|
-
# Python 3.11 or 3.12. Not 3.13+ — WhisperX's dependencies do not support it yet, and the
|
|
8
|
-
# failure is a wall of build errors rather than a clear message.
|
|
9
|
-
#
|
|
10
|
-
# certifi is not optional on macOS: torch downloads the alignment model with urllib, which
|
|
11
|
-
# on a python.org build has no root certificates at all.
|
|
12
|
-
|
|
13
|
-
whisperx==3.8.6
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
1
|
+
# The Python half of the studio: transcription, word alignment, cut detection,
|
|
2
|
+
# green-screen cutouts and colour grading.
|
|
3
|
+
#
|
|
4
|
+
# python -m venv .venv
|
|
5
|
+
# .venv/bin/pip install -r requirements.txt (Windows: .venv\Scripts\pip)
|
|
6
|
+
#
|
|
7
|
+
# Python 3.11 or 3.12. Not 3.13+ — WhisperX's dependencies do not support it yet, and the
|
|
8
|
+
# failure is a wall of build errors rather than a clear message.
|
|
9
|
+
#
|
|
10
|
+
# certifi is not optional on macOS: torch downloads the alignment model with urllib, which
|
|
11
|
+
# on a python.org build has no root certificates at all.
|
|
12
|
+
|
|
13
|
+
whisperx==3.8.6
|
|
14
|
+
# BiRefNet, for cutting the subject out of ordinary footage so text can sit behind them.
|
|
15
|
+
# MIT licensed and self-hosted: nothing uploads, and there is no per-video cost.
|
|
16
|
+
transformers>=4.44
|
|
17
|
+
timm>=1.0
|
|
18
|
+
einops>=0.8
|
|
19
|
+
kornia>=0.7
|
|
20
|
+
torch>=2.2
|
|
21
|
+
numpy>=1.26
|
|
22
|
+
Pillow>=10.0
|
|
23
|
+
certifi>=2024.2.2
|
|
@@ -1,68 +1,181 @@
|
|
|
1
|
-
"""Build a transparent cutout of the subject for a frame range
|
|
2
|
-
|
|
3
|
-
Usage (from the project root):
|
|
4
|
-
python scripts/make_cutout.py <
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
1
|
+
"""Build a transparent cutout of the subject for a frame range, from ordinary footage.
|
|
2
|
+
|
|
3
|
+
Usage (from the project root):
|
|
4
|
+
python scripts/make_cutout.py <video> <out_dir> <first_frame> <last_frame>
|
|
5
|
+
python scripts/make_cutout.py <video> <out_dir> <first> <last> --alpha-from <clip>
|
|
6
|
+
python scripts/make_cutout.py <video> <out_dir> <first> <last> --model full
|
|
7
|
+
|
|
8
|
+
Writes RGBA PNGs named f<frame>.png (absolute source frame numbers), so a composition can map
|
|
9
|
+
composition frames to cutout frames directly.
|
|
10
|
+
|
|
11
|
+
No green screen. The matte is segmented from the footage itself, because asking a customer to
|
|
12
|
+
shoot against green makes the feature unavailable to almost all of them - most have never done
|
|
13
|
+
it and do not know the term.
|
|
14
|
+
|
|
15
|
+
--alpha-from stays for the case where a background-free or green copy already exists, which is
|
|
16
|
+
faster and still sharper on hair. It is used ONLY for the alpha; colour always comes from the
|
|
17
|
+
original clip, so no green ever reaches the subject's edges.
|
|
18
|
+
|
|
19
|
+
Two things decide whether this reads as real rather than pasted on:
|
|
20
|
+
|
|
21
|
+
* The alpha is smoothed ACROSS frames. A per-frame matte shimmers along the edge, and type
|
|
22
|
+
sitting against someone's head is precisely where the eye catches it. A 3-frame median kills
|
|
23
|
+
single-frame flicker without the lag an average introduces on fast motion.
|
|
24
|
+
* The edge is choked before it is feathered, so the old background cannot leak into the seam.
|
|
25
|
+
|
|
26
|
+
It is worst where it is most visible: flyaway hair, motion blur, a hand crossing the text. If a
|
|
27
|
+
shot will not hold, move the text rather than ship a halo.
|
|
28
|
+
"""
|
|
29
|
+
|
|
30
|
+
import argparse
|
|
31
|
+
import json
|
|
32
|
+
import subprocess
|
|
33
|
+
import sys
|
|
34
|
+
from pathlib import Path
|
|
35
|
+
|
|
36
|
+
import numpy as np
|
|
37
|
+
from PIL import Image
|
|
38
|
+
import sys as _sys, pathlib as _pl
|
|
39
|
+
_sys.path.insert(0, str(_pl.Path(__file__).resolve().parent / 'lib'))
|
|
40
|
+
from media import FFMPEG, FFPROBE # resolved per platform - never hardcode a tool path
|
|
41
|
+
|
|
42
|
+
MODELS = {"lite": "ZhengPeng7/BiRefNet_lite", "full": "ZhengPeng7/BiRefNet"}
|
|
43
|
+
SEG_SIZE = 1024
|
|
44
|
+
|
|
45
|
+
|
|
46
|
+
def probe(path: str):
|
|
47
|
+
"""Real dimensions and frame rate. Never assume 1920x1080 at 30 - vertical phone footage is normal."""
|
|
48
|
+
out = subprocess.run(
|
|
49
|
+
[FFPROBE, "-v", "quiet", "-print_format", "json", "-show_streams", "-select_streams", "v:0", path],
|
|
50
|
+
capture_output=True, text=True, check=True,
|
|
51
|
+
).stdout
|
|
52
|
+
s = json.loads(out)["streams"][0]
|
|
53
|
+
num, den = (s.get("r_frame_rate") or "30/1").split("/")
|
|
54
|
+
return int(s["width"]), int(s["height"]), float(num) / float(den or 1)
|
|
55
|
+
|
|
56
|
+
|
|
57
|
+
def read_frames(path: str, first: int, count: int, w: int, h: int, fps: float):
|
|
58
|
+
proc = subprocess.Popen(
|
|
59
|
+
[FFMPEG, "-v", "error", "-ss", f"{first / fps:.6f}", "-i", path, "-frames:v", str(count),
|
|
60
|
+
"-f", "rawvideo", "-pix_fmt", "rgb24", "-"],
|
|
61
|
+
stdout=subprocess.PIPE,
|
|
62
|
+
)
|
|
63
|
+
size = w * h * 3
|
|
64
|
+
for _ in range(count):
|
|
65
|
+
buf = proc.stdout.read(size)
|
|
66
|
+
if len(buf) < size:
|
|
67
|
+
break
|
|
68
|
+
yield np.frombuffer(buf, np.uint8).reshape(h, w, 3)
|
|
69
|
+
proc.stdout.close()
|
|
70
|
+
proc.wait()
|
|
71
|
+
|
|
72
|
+
|
|
73
|
+
def green_matte(green: np.ndarray) -> np.ndarray:
|
|
74
|
+
"""Alpha 0..1 from how green each pixel is, relative to the flat background green."""
|
|
75
|
+
g = green.astype(np.float32)
|
|
76
|
+
greenness = g[..., 1] - np.maximum(g[..., 0], g[..., 2])
|
|
77
|
+
background = np.median(greenness[greenness > 60]) if (greenness > 60).any() else 150.0
|
|
78
|
+
lo, hi = 0.2 * background, 0.7 * background
|
|
79
|
+
t = np.clip((greenness - lo) / (hi - lo), 0, 1)
|
|
80
|
+
return 1 - t * t * (3 - 2 * t) # smoothstep: soft, stable edge
|
|
81
|
+
|
|
82
|
+
|
|
83
|
+
def load_segmenter(which: str):
|
|
84
|
+
"""BiRefNet, lazily. Importing torch costs seconds; do not pay it on the --alpha-from path."""
|
|
85
|
+
try:
|
|
86
|
+
import torch
|
|
87
|
+
from transformers import AutoModelForImageSegmentation
|
|
88
|
+
except ModuleNotFoundError as exc:
|
|
89
|
+
print(f"Cannot import {exc.name}. The matting model needs the studio's Python environment.", file=sys.stderr)
|
|
90
|
+
print(" see what is installed with: npm run doctor", file=sys.stderr)
|
|
91
|
+
raise SystemExit(2)
|
|
92
|
+
|
|
93
|
+
# trust_remote_code: these repos ship their own model class, which transformers executes.
|
|
94
|
+
model = AutoModelForImageSegmentation.from_pretrained(MODELS[which], trust_remote_code=True)
|
|
95
|
+
model.eval()
|
|
96
|
+
# float32 on CPU: half precision is slower there rather than faster, and can produce NaNs.
|
|
97
|
+
device = "cuda" if torch.cuda.is_available() else "cpu"
|
|
98
|
+
model.to(device)
|
|
99
|
+
return torch, model, device
|
|
100
|
+
|
|
101
|
+
|
|
102
|
+
def segment(torch, model, device, rgb: np.ndarray) -> np.ndarray:
|
|
103
|
+
"""Alpha 0..1 for one frame, returned at the frame's own resolution."""
|
|
104
|
+
h, w = rgb.shape[:2]
|
|
105
|
+
img = Image.fromarray(rgb).resize((SEG_SIZE, SEG_SIZE), Image.BILINEAR)
|
|
106
|
+
x = np.asarray(img, dtype=np.float32) / 255.0
|
|
107
|
+
x = (x - np.array([0.485, 0.456, 0.406], np.float32)) / np.array([0.229, 0.224, 0.225], np.float32)
|
|
108
|
+
t = torch.from_numpy(x).permute(2, 0, 1).unsqueeze(0).to(device)
|
|
109
|
+
with torch.no_grad():
|
|
110
|
+
pred = model(t)[-1].sigmoid()[0, 0].cpu().numpy()
|
|
111
|
+
return np.asarray(Image.fromarray((pred * 255).astype(np.uint8)).resize((w, h), Image.BILINEAR), np.float32) / 255.0
|
|
112
|
+
|
|
113
|
+
|
|
114
|
+
def smooth_in_time(alphas):
|
|
115
|
+
"""3-frame median. Removes single-frame flicker without the lag an average would add."""
|
|
116
|
+
if len(alphas) < 3:
|
|
117
|
+
return alphas
|
|
118
|
+
out = [alphas[0]]
|
|
119
|
+
for i in range(1, len(alphas) - 1):
|
|
120
|
+
out.append(np.median(np.stack([alphas[i - 1], alphas[i], alphas[i + 1]]), axis=0))
|
|
121
|
+
out.append(alphas[-1])
|
|
122
|
+
return out
|
|
123
|
+
|
|
124
|
+
|
|
125
|
+
def clean_edge(alpha: np.ndarray) -> np.ndarray:
|
|
126
|
+
"""Choke by a pixel so the old background cannot leak into the seam, then feather lightly."""
|
|
127
|
+
choked = np.minimum.reduce([alpha, np.roll(alpha, 1, 0), np.roll(alpha, -1, 0), np.roll(alpha, 1, 1), np.roll(alpha, -1, 1)])
|
|
128
|
+
return (choked * 4 + np.roll(choked, 1, 0) + np.roll(choked, -1, 0) + np.roll(choked, 1, 1) + np.roll(choked, -1, 1)) / 8
|
|
129
|
+
|
|
130
|
+
|
|
131
|
+
def main() -> None:
|
|
132
|
+
ap = argparse.ArgumentParser(description="Transparent cutout of the subject, from ordinary footage.")
|
|
133
|
+
ap.add_argument("video")
|
|
134
|
+
ap.add_argument("out_dir", type=Path)
|
|
135
|
+
ap.add_argument("first", type=int)
|
|
136
|
+
ap.add_argument("last", type=int)
|
|
137
|
+
ap.add_argument("--alpha-from", default=None,
|
|
138
|
+
help="Optional background-free or green copy of the SAME take, used only for the alpha.")
|
|
139
|
+
ap.add_argument("--model", choices=sorted(MODELS), default="lite",
|
|
140
|
+
help="lite is several times faster and good enough for most shots; full is sharper on hair.")
|
|
141
|
+
args = ap.parse_args()
|
|
142
|
+
|
|
143
|
+
w, h, fps = probe(args.video)
|
|
144
|
+
count = args.last - args.first + 1
|
|
145
|
+
args.out_dir.mkdir(parents=True, exist_ok=True)
|
|
146
|
+
|
|
147
|
+
frames = list(read_frames(args.video, args.first, count, w, h, fps))
|
|
148
|
+
if not frames:
|
|
149
|
+
print("No frames read - check the frame range against the clip's length.", file=sys.stderr)
|
|
150
|
+
raise SystemExit(1)
|
|
151
|
+
|
|
152
|
+
if args.alpha_from:
|
|
153
|
+
aw, ah, afps = probe(args.alpha_from)
|
|
154
|
+
if (aw, ah) != (w, h) or abs(afps - fps) > 0.01:
|
|
155
|
+
print(f"The alpha clip does not match: {aw}x{ah}@{afps:.2f} against {w}x{h}@{fps:.2f}.", file=sys.stderr)
|
|
156
|
+
print("A mismatched pair makes a halo or a ghost. Re-export it to match, or drop --alpha-from.", file=sys.stderr)
|
|
157
|
+
raise SystemExit(1)
|
|
158
|
+
print(f"Alpha from the supplied clip, colour from the original - {count} frames.")
|
|
159
|
+
alphas = [green_matte(g) for g in read_frames(args.alpha_from, args.first, count, w, h, fps)]
|
|
160
|
+
else:
|
|
161
|
+
print(f"Segmenting {count} frames at {w}x{h} with BiRefNet ({args.model}).")
|
|
162
|
+
print("The first run downloads the model. On a CPU expect a second or two per frame.")
|
|
163
|
+
torch, model, device = load_segmenter(args.model)
|
|
164
|
+
print(f"Running on {device}.")
|
|
165
|
+
alphas = []
|
|
166
|
+
for i, rgb in enumerate(frames):
|
|
167
|
+
alphas.append(segment(torch, model, device, rgb))
|
|
168
|
+
if (i + 1) % 10 == 0 or i + 1 == count:
|
|
169
|
+
print(f" {i + 1}/{count}", flush=True)
|
|
170
|
+
|
|
171
|
+
for i, a in enumerate(smooth_in_time(alphas)):
|
|
172
|
+
alpha = (clean_edge(a).clip(0, 1) * 255).round().astype(np.uint8)
|
|
173
|
+
Image.fromarray(np.dstack([frames[i], alpha]), "RGBA").save(
|
|
174
|
+
args.out_dir / f"f{args.first + i}.png", compress_level=1
|
|
175
|
+
)
|
|
176
|
+
|
|
177
|
+
print(f"wrote {len(alphas)} frames to {args.out_dir}")
|
|
178
|
+
|
|
179
|
+
|
|
180
|
+
if __name__ == "__main__":
|
|
181
|
+
main()
|
|
@@ -16,7 +16,12 @@ import {spawnSync} from 'node:child_process';
|
|
|
16
16
|
import {existsSync, mkdirSync, readdirSync, readFileSync, writeFileSync} from 'node:fs';
|
|
17
17
|
import {basename, extname, join} from 'node:path';
|
|
18
18
|
|
|
19
|
-
|
|
19
|
+
// Never hardcode a tool path. media.mjs finds the ffmpeg that ships with the studio
|
|
20
|
+
// (ffmpeg-static, pulled in by npm install) before it looks at PATH, which is what lets this
|
|
21
|
+
// run on a machine that has never had ffmpeg installed on it. A hardcoded path failed
|
|
22
|
+
// silently here: the sheets were simply never generated and nothing said why.
|
|
23
|
+
import {FFMPEG} from './lib/media.mjs';
|
|
24
|
+
|
|
20
25
|
const REF_DIR = 'references';
|
|
21
26
|
const OUT = join('work', 'refs');
|
|
22
27
|
const all = process.argv.includes('--all');
|
|
@@ -13,7 +13,10 @@ import {spawnSync} from 'node:child_process';
|
|
|
13
13
|
import {mkdirSync, rmSync, writeFileSync} from 'node:fs';
|
|
14
14
|
import {join, resolve} from 'node:path';
|
|
15
15
|
|
|
16
|
-
|
|
16
|
+
// Never hardcode a tool path. media.mjs finds the ffmpeg that ships with the studio
|
|
17
|
+
// (ffmpeg-static, pulled in by npm install) before it looks at PATH, which is what lets this
|
|
18
|
+
// run on a machine that has never had ffmpeg installed on it.
|
|
19
|
+
import {FFMPEG} from './lib/media.mjs';
|
|
17
20
|
|
|
18
21
|
const args = process.argv.slice(2);
|
|
19
22
|
const id = args.find((a) => !a.startsWith('--'));
|
|
@@ -1,88 +1,90 @@
|
|
|
1
|
-
# Brief: {{name}}
|
|
2
|
-
|
|
3
|
-
Fill in what you can. Plain words are fine. Leave a line blank if you don't care.
|
|
4
|
-
When done: drop your files in this folder's `footage/` (and `screenshots/`), then run
|
|
5
|
-
`npm run prep -- {{name}}` and start the session with "start {{name}}".
|
|
6
|
-
|
|
7
|
-
## Basics
|
|
8
|
-
Format: 16:9 <!-- 16:9, 9:16, or both -->
|
|
9
|
-
Language: en <!-- en, hi, ar ... the language you speak in the video -->
|
|
10
|
-
Talking head file: <!-- file name in footage/, e.g. —.mp4. Leave blank if there's only one clip. -->
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
|
66
|
-
|
|
67
|
-
|
|
|
68
|
-
|
|
|
69
|
-
|
|
|
70
|
-
|
|
|
71
|
-
|
|
|
72
|
-
|
|
|
73
|
-
|
|
|
74
|
-
|
|
|
75
|
-
|
|
|
76
|
-
|
|
|
77
|
-
|
|
|
78
|
-
|
|
|
79
|
-
|
|
|
80
|
-
|
|
|
81
|
-
|
|
|
82
|
-
|
|
|
83
|
-
|
|
|
84
|
-
|
|
|
85
|
-
|
|
|
86
|
-
|
|
|
87
|
-
|
|
|
88
|
-
|
|
|
1
|
+
# Brief: {{name}}
|
|
2
|
+
|
|
3
|
+
Fill in what you can. Plain words are fine. Leave a line blank if you don't care.
|
|
4
|
+
When done: drop your files in this folder's `footage/` (and `screenshots/`), then run
|
|
5
|
+
`npm run prep -- {{name}}` and start the session with "start {{name}}".
|
|
6
|
+
|
|
7
|
+
## Basics
|
|
8
|
+
Format: 16:9 <!-- 16:9, 9:16, or both -->
|
|
9
|
+
Language: en <!-- en, hi, ar ... the language you speak in the video -->
|
|
10
|
+
Talking head file: <!-- file name in footage/, e.g. —.mp4. Leave blank if there's only one clip. -->
|
|
11
|
+
Background-free copy: <!-- Optional. Leave blank. Only fill this in if you already export a
|
|
12
|
+
version with the background removed or replaced by flat green; the
|
|
13
|
+
cutout is built automatically from your normal footage otherwise. -->
|
|
14
|
+
Length target: <!-- e.g. 25s, 90s -->
|
|
15
|
+
|
|
16
|
+
## Script
|
|
17
|
+
<!-- Paste the script exactly as you read it. Numbers as digits: $99, 11 minutes, 3 days. -->
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
## The one thing viewers must remember
|
|
21
|
+
<!-- One sentence. Everything in the edit serves this. -->
|
|
22
|
+
|
|
23
|
+
|
|
24
|
+
## Must-have moments
|
|
25
|
+
<!-- By the words you say, not by time. Plain words are fine.
|
|
26
|
+
- When I say "look at this" → my clip shrinks and other creators' clips appear
|
|
27
|
+
- "millions of views" → a counter racing up to 1M+ -->
|
|
28
|
+
|
|
29
|
+
|
|
30
|
+
## Moments from references
|
|
31
|
+
<!-- You DON'T need the name of the effect. Clip + time + what you like about it.
|
|
32
|
+
Format: - clip-file.mp4 mm:ss-mm:ss — what you like (feel words are fine: snappy, floaty, punchy)
|
|
33
|
+
- a-reel-you-like.mp4 00:02-00:04 — the colour screens flipping with each phrase, feels punchy -->
|
|
34
|
+
|
|
35
|
+
|
|
36
|
+
## Effects from the library
|
|
37
|
+
<!-- Optional: pick any by name (see the list at the bottom), and say where.
|
|
38
|
+
- colour flip → "different person, different story" -->
|
|
39
|
+
|
|
40
|
+
|
|
41
|
+
## B-roll
|
|
42
|
+
<!-- Which clip where, by the words you say. Mark clips you DON'T want used.
|
|
43
|
+
- creators' clips → "all of these people" -->
|
|
44
|
+
|
|
45
|
+
|
|
46
|
+
## Screenshots / things to blur
|
|
47
|
+
<!-- Names, emails, prices, logos that must not be readable. -->
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
## CTA / ending
|
|
51
|
+
<!-- What the viewer should do, and what the last frame shows. -->
|
|
52
|
+
|
|
53
|
+
|
|
54
|
+
## Sound
|
|
55
|
+
SFX library: audio/ <!-- put sound effects in this folder's audio/, or say "none" -->
|
|
56
|
+
Music: added by me in Premiere / Instagram <!-- or a file name in audio/ -->
|
|
57
|
+
|
|
58
|
+
## Don't want
|
|
59
|
+
<!-- Anything you dislike: e.g. boxes over my footage, huge text, bouncy motion. -->
|
|
60
|
+
|
|
61
|
+
|
|
62
|
+
---
|
|
63
|
+
|
|
64
|
+
### Effects library (names you can use)
|
|
65
|
+
| Name | What it looks like |
|
|
66
|
+
|---|---|
|
|
67
|
+
| Text behind me | Big headline sits behind your head (works from normal footage) |
|
|
68
|
+
| Card swap | You shrink into a rounded card, turn into another clip mid-shrink, a row of cards slides in and drifts |
|
|
69
|
+
| Colour flip | Hard cuts to solid accent / ink screens on each phrase; a word rolls into the next and the colours swap |
|
|
70
|
+
| Counter race | A number races up from 0 and lands on a value (e.g. 1,000,000+) |
|
|
71
|
+
| Zoom-through cut | A jump cut hidden by a quick push-in that eases back — used on every cut by default |
|
|
72
|
+
| Gradient CTA | Dark gradient rises from the bottom, small label + big word typed on the VO, then ticked pointers |
|
|
73
|
+
| Browser mockup | Screenshots inside a clean browser window that pans and zooms to the detail |
|
|
74
|
+
| Pointer click | An arrow glides to a spot and clicks, with a ripple |
|
|
75
|
+
| Rolling number | Digits roll from one value to another (e.g. $26.50 → $48.94) |
|
|
76
|
+
| Highlight marker | Accent marker wipes behind a key phrase |
|
|
77
|
+
| Side card | You move into a card beside the graphics (16:9) or below them (9:16) |
|
|
78
|
+
| Headline to pill | A big headline shrinks into a small corner tag that stays on screen |
|
|
79
|
+
| Word-by-word captions | On by default (looks: clean, block, highlight, pop) |
|
|
80
|
+
| Whip pan | Fast blurred pan between two unrelated shots |
|
|
81
|
+
| Push | The next scene pushes the last one out |
|
|
82
|
+
| Mask wipe | A soft-edged wipe reveals the next scene |
|
|
83
|
+
| Match cut | Cut hidden inside a shared shape or movement |
|
|
84
|
+
| Dip | Quick dip through the background colour |
|
|
85
|
+
| Speed ramp | Glide into slow-mo (or speed-up) and back out |
|
|
86
|
+
| Beat cut | Cuts land on the beat of your music |
|
|
87
|
+
| Tighten it | Silence and filler words removed, every join hidden as a punch-in |
|
|
88
|
+
| Cut to the music | Whole edit built off the track: slow intro, faster build, hardest on the drop |
|
|
89
|
+
| Film look | Grain, halation, light leak, gate weave (looks: clean, film, super8, cinema) |
|
|
90
|
+
| Sound pass | Whooshes, impacts and risers placed off the cut list, ducked under the voice |
|