@slatesvideo/shared 0.5.0 → 0.5.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -54,10 +54,23 @@ Never write romantic, sexual, or suggestive content involving or directed at min
54
54
 
55
55
  If a box fails, apply the substitution table before writing the prompt.
56
56
 
57
- ## Editing real footage (Kling O3 edit) — real people in the SOURCE
57
+ ## Editing real footage (Kling O3 edit / Omni Flash edit) — real people in the SOURCE
58
58
 
59
59
  Video edit takes the user's own footage, which often contains real people. Rules:
60
60
 
61
61
  - The user must hold rights/consent for any real person's likeness in footage they edit — ask once when it's clearly someone other than the user, then proceed.
62
62
  - Kling's video-to-video filter behavior on real faces is **not yet verified** (unlike Seedance, where the consent-gated real-face route is confirmed). If an edit of real-person footage is rejected by the provider, do NOT retry-spam variations — tell the user the filter blocked it and offer a no-face crop/segment or an AI-character swap instead.
63
+ - **Omni Flash: own-footage editing of the uploader's own face PASSED live 2026-07-09** (real talking-head clip, edited on our fal route) despite Google's documented "recognizable people" restriction — treat that restriction as aimed at third-party/public figures, but expect probabilistic refusals and never promise passage.
63
64
  - Never use edit to put a real, named public figure into a scene, or to make someone appear to say/do something they didn't. Faceless b-roll (hands, products, landscapes, crowds-from-behind) edits freely.
65
+
66
+ ## Gemini / Omni Flash filter regime (video gen + edit) — receipts 2026-07-09
67
+
68
+ Google's filter is its own regime (stricter than fal-hosted Kling about harm-to-a-person, looser than BytePlus about faces). Live receipts:
69
+
70
+ | Blocked (`content_policy_violation`) | Passed |
71
+ |---|---|
72
+ | "his fingertips **ignite** with a small real flame" (fire ON a body part = harm) | "small **magical** flames appear on his fingertips … vanish when he blows on them" |
73
+
74
+ - **Harm-to-person framing is the tripwire**, not the effect itself. Reframe body-contact effects as magical / supernatural / harmless VFX: "magical flames", "a glowing aura", "sparks of light dance on". Avoid ignite / burn / on fire / catch fire applied to a person.
75
+ - **Never use a real object as a metaphor** — "candle-like flame" rendered a literal candle in the subject's hand. Describe the effect, not an object that resembles it.
76
+ - The block is a 422 refund (no credits lost) and arrives mid-generation — one reframe per the substitution mindset above, don't retry-spam.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-cost-discipline
3
- description: Mandatory pre-flight discipline before ANY generation call (image or video) — estimate cost, announce in dollars, get confirmation, aggregate batches. Read this every time before calling slates_generate_image or any future slates_generate_* op. Skipping this risks burning the user's credits on guesses.
3
+ description: Mandatory pre-flight discipline before ANY generation call (image or video) — estimate cost, announce in credits, get confirmation, aggregate batches. Read this every time before calling slates_generate_image or any future slates_generate_* op. Skipping this risks burning the user's credits on guesses.
4
4
  ---
5
5
 
6
6
  # Slates cost discipline — read before every generation
@@ -20,23 +20,25 @@ Before ANY `slates_generate_*` call, run `slates_estimate_generation_cost` first
20
20
 
21
21
  If aspect ratio or resolution isn't obvious from the user's request, **ask before estimating**. Don't guess.
22
22
 
23
- ### 2. Announce in dollars, plainly, before spending
23
+ ### 2. Announce in credits, plainly, before spending
24
24
 
25
- Format: `About to spend $X.XX on N image(s) at [resolution] [aspect ratio]. Proceed?`
25
+ Slates bills abstract **credits** (they never expire). Announce the credit total the estimate returns — never dollars.
26
+
27
+ Format: `About to spend N credits on M image(s) at [resolution] [aspect ratio]. Proceed?`
26
28
 
27
29
  Examples:
28
- - `About to spend $0.10 on 1 image at 1k 16:9. Proceed?`
29
- - `About to spend $0.40 on 4 images at 2k 9:16 (variants). Proceed?`
30
+ - `About to spend 4 credits on 1 image at 1k 16:9. Proceed?`
31
+ - `About to spend 24 credits on 4 images at 2k 9:16 (variants). Proceed?`
30
32
 
31
- Below ~$0.20 you can proceed silently after announcing once. Above $0.20 wait for explicit confirmation. Above $0.50 the server itself will gate with `requires_confirm` — pass `confirm: true` only after the user explicitly OKs.
33
+ Below ~7 credits you can proceed silently after announcing once. Above ~7 credits wait for explicit confirmation. Above ~17 credits the server itself will gate with `requires_confirm` — pass `confirm: true` only after the user explicitly OKs.
32
34
 
33
35
  ### 3. Aggregate batches into ONE upfront announcement
34
36
 
35
- If you're planning a multi-call workflow (5 storyboard frames, 3 character variants, a grid of options), **announce the total before the first call**, not five $0.10 announcements after the fact.
37
+ If you're planning a multi-call workflow (5 storyboard frames, 3 character variants, a grid of options), **announce the total before the first call**, not five small announcements after the fact.
36
38
 
37
- Format: `Plan: N generations totaling $X.XX. [Brief description of the sequence.] Proceed with the batch?`
39
+ Format: `Plan: N generations totaling C credits. [Brief description of the sequence.] Proceed with the batch?`
38
40
 
39
- Example: `Plan: 6 frame generations at 1k 16:9 totaling $0.60 — establishing wide, push-in, two-shot, reverse, OTS, insert. Proceed?`
41
+ Example: `Plan: 6 frame generations at 1k 16:9 totaling 24 credits — establishing wide, push-in, two-shot, reverse, OTS, insert. Proceed?`
40
42
 
41
43
  ### 3b. Batch authorization — one approval covers the enumerated batch
42
44
 
@@ -52,7 +54,7 @@ One approval = that plan, as enumerated, at those prices. Nothing else.
52
54
 
53
55
  ### 4. Track the running total
54
56
 
55
- After each generation completes, the response includes `cost_cents` (when available). Keep a running tally in your context. Surface it every 3 generations or whenever the user asks "how much have we spent?"
57
+ After each generation completes, the response includes `cost_credits` (when available). Keep a running tally in your context. Surface it every 3 generations or whenever the user asks "how much have we spent?"
56
58
 
57
59
  ## Resolution decision rules
58
60
 
@@ -66,6 +68,8 @@ After each generation completes, the response includes `cost_cents` (when availa
66
68
 
67
69
  Resolution is a price lever, not a free choice: on Nano Banana 2 and FLUX.2 Max, 4k costs roughly 2x 1k (Seedream 5 Lite is flat-priced regardless of resolution). Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models` for current numbers instead of assuming. Pick the cheapest resolution that serves the use case.
68
70
 
71
+ **4K VIDEO is Pro-only (2026-07-07).** The ladder above is for IMAGES (open at every tier). For VIDEO — Kling, Seedance, Veo — 4K requires a Slates Pro account; a base-tier 4K video gen is rejected server-side with `PRO_REQUIRED`. Default video to 1080p or lower and only reach for 4K when the user is on Pro and explicitly asks. 4K *images* are never gated.
72
+
69
73
  ## Aspect ratio decision rules
70
74
 
71
75
  Ask the user when ambiguous. Otherwise:
@@ -83,7 +87,7 @@ If the user prompt mixes signals (e.g. "cinematic Instagram post"), ask. Don't g
83
87
 
84
88
  ## When the gate fires
85
89
 
86
- The server returns `requires_clarification` when aspect ratio or resolution is missing. The server returns `requires_confirm` when total spend exceeds $0.50. In both cases:
90
+ The server returns `requires_clarification` when aspect ratio or resolution is missing. The server returns `requires_confirm` when total spend exceeds ~17 credits. In both cases:
87
91
 
88
92
  1. Surface the gate response to the user
89
93
  2. Get a clean answer
@@ -10,7 +10,7 @@ You are building a 30-second hyper-motion direct-response ad. The user has hande
10
10
  **Hard rules**
11
11
 
12
12
  - Always estimate cost before generating. Use `slates_estimate_generation_cost` and surface the total.
13
- - All MCP/CLI generation routes through Slates Credits, period. BYOK is desktop-UI only by design — don't suggest "use your own keys" workarounds.
13
+ - All Slates generation routes through Slates Credits, period (BYOK is retired) — don't suggest "use your own keys" workarounds.
14
14
  - Default model: `nano-banana-2-2k`. For close-up product hero frames step up to `4k` only if the user asks.
15
15
  - Hyper-motion = punchy cuts, 4 frames in 30 seconds, ~7s each. Don't over-storyboard.
16
16
 
@@ -58,7 +58,7 @@ For each frame:
58
58
 
59
59
  - **Don't** generate text overlays in the image. Slates renders captions/CTAs at the editor stage.
60
60
  - **Don't** burn credits on slot-machine prompting. If the first generation is off, refine the prompt; don't just regenerate.
61
- - **Don't** skip the cost estimate. Confirm with the user above $0.50.
61
+ - **Don't** skip the cost estimate. Confirm with the user above ~17 credits.
62
62
  - **Don't** invent visual specifics about the product (colors, textures, angles) that aren't in the reference image. Reference-anchored prompts only.
63
63
 
64
64
  ## Voice
@@ -28,7 +28,7 @@ The user's request is one of:
28
28
  | Aesthetic / compositional | `slates_generate_image` with the original in `referenceAssetIds` + a refined prompt. Don't re-roll from scratch. |
29
29
  | Wholesale | New prompt, no reference, fresh generation. Treat as a new brief. |
30
30
 
31
- **`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Default model `nano-banana-2` — the only edit model that also takes extra `referenceAssetIds`; `flux-2-max` / `seedream-5-lite` use their own edit endpoints and ignore references. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost > $0.50 gates on `confirm=true`.
31
+ **`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Default model `nano-banana-2` — the only edit model that also takes extra `referenceAssetIds`; `flux-2-max` / `seedream-5-lite` use their own edit endpoints and ignore references. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost above ~17 credits gates on `confirm=true`.
32
32
 
33
33
  ### 4. Generate, evaluate, decide
34
34
  - Estimate cost first.
@@ -14,7 +14,7 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
14
14
  | **General-purpose — the default for most shots** | **Kling 3.0 std** | Cost-effective workhorse. Strong image-to-video: preserves identity, layout, and text from the start frame. Any aspect ratio, 5–15s. |
15
15
  | Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |
16
16
  | Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
17
- | **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K. |
17
+ | **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
18
18
  | The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
19
19
  | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | The only job Veo wins. |
20
20
 
@@ -22,12 +22,16 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
22
22
 
23
23
  | Job | Tool | Why |
24
24
  |---|---|---|
25
- | **Fix a 90%-right clip** — wrong shirt, one artifact, swap the subject, change the environment | **Kling O3 Edit** (`slates_edit_video`) | The default edit tool. Element lock (frontal + angles) holds identity; original motion, camera, and AUDIO preserved. One pass, no masking, ~19¢/s. |
26
- | Style-transfer-heavy re-imagining, full relocate of the scene | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. |
27
- | AI-edit the user's OWN footage (3–15s clips) | Kling O3 Edit | Takes any MP4/MOV 3–15s, 720–3840px — not just Slates gens. |
25
+ | **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + "Keep everything else the same"; long prompts destroy it (see below). |
26
+ | **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
27
+ | **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
28
+ | Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
29
+ | AI-edit the user's OWN footage | Omni Flash Edit (3–10s) or Kling O3 Edit (3–15s, 720–3840px) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
28
30
 
29
31
  - **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
30
- - Edited clips are themselves editable ≤15s clips — chain passes; lineage links each output to its parent.
32
+ - **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the pro demos (e.g. Higgsfield's split-screen short) actually work, plus gesture-only beats with voiceover laid over in post.
33
+ - **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
34
+ - Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
31
35
 
32
36
  ## Motion Transfer & Lip Sync routing (two engines per tool)
33
37
 
@@ -35,9 +39,9 @@ Both tools have a cheap Kling utility lane and a premium Seedance lane. The capa
35
39
 
36
40
  | Job | Engine | Why |
37
41
  |---|---|---|
38
- | Quick motion retarget, budget lane, or driving clip >15s | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, $0.95–1.26 / 5s, takes up to 30s driving clips. |
42
+ | Quick motion retarget, budget lane, or driving clip >15s | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |
39
43
  | **Motion transfer where fidelity or audio matters** — dance, choreography, cinematic action | **Seedance 2.0** (`motionModel=seedance-2`) | Single-pass conditioning beats post-hoc retargeting; prompt-driven; native audio. Driving clip 2–15s; bills input+output seconds (vref keys). |
40
- | Cheap lip-sync utility (re-voice a clip, simple avatar) | Kling lip-sync / avatar (`slates_generate_lip_sync`) | $0.11–0.86 / 5s blocks. |
44
+ | Cheap lip-sync utility (re-voice a clip, simple avatar) | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |
41
45
  | **Natural speech, voice cloned from the source clip, premium delivery** | **Seedance 2.0** (`engine=seedance-2`) | The line is spoken IN the generation (no TTS layer); a video source keeps its own voice; uploaded ≤15s audio can drive it. |
42
46
 
43
47
  - Faces: Seedance tool gens default `seedanceFace=true` (sources are people). A REAL person triggers the consent cascade (`[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent`, premium realface pricing).
@@ -39,7 +39,7 @@ Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the chara
39
39
 
40
40
  **Model mixing — route per `slates-model-selection`** (details in the per-model guides):
41
41
  - **Kling V3** (`slates-prompting-kling-v3`): the DEFAULT for most shots — any aspect ratio, 5-15s, strong start-frame adherence; std is the workhorse, Omni for multi-character dialogue.
42
- - **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K.
42
+ - **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K (4K video is Pro-only).
43
43
  - **Veo 3.1** (`slates-prompting-veo-3`): niche, never the default — only when native synced audio must generate WITH the video in one gen; 16:9 only, 4/6/8s.
44
44
 
45
45
  Failed gen? Check the error via `slates_get_generation_status`, fix the prompt, resubmit that one shot (a retry beyond the plan = announce the delta cost).
@@ -9,9 +9,9 @@ Two engines. Kling is the cheap utility lane (dedicated lip-sync endpoints, 5-se
9
9
 
10
10
  | Flow | Source | Engine/Model | Cost | Use case |
11
11
  |------|--------|-------|-----------|----------|
12
- | Re-dub | video clip | kling-lip-sync-video | $0.11 / 5s | Replace dialogue on an existing talking head |
13
- | Avatar standard | still image | ai-avatar/v2/standard | $0.42 / 5s | Animate a portrait into a talking avatar |
14
- | Avatar pro | still image | ai-avatar/v2/pro | $0.86 / 5s | Higher facial fidelity for hero shots |
12
+ | Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |
13
+ | Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |
14
+ | Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |
15
15
  | **Seedance native** | image or video | `engine=seedance-2` | per second (`seedance-2-face-*`; video sources bill input+output seconds) | **Premium**: natural delivery, whole-body performance, voice cloned from a video source, audio included |
16
16
 
17
17
  Pick engine + `sourceType` deliberately — they decide the pricing tier and the underlying endpoint.
@@ -34,7 +34,7 @@ Everything below applies to the **Kling** engine.
34
34
  Use **video** (re-dub) when:
35
35
  - A talking-head clip already exists (Slates-generated, recorded, or imported)
36
36
  - The mouth/face is already moving and only the audio needs to change
37
- - $0.11 is hard to beat for short dialogue replacement
37
+ - ~4 credits is hard to beat for short dialogue replacement
38
38
 
39
39
  Use **avatar** when:
40
40
  - Only a still portrait exists
@@ -127,9 +127,9 @@ Default `"."` is fine if you have nothing useful to add.
127
127
  **Use pro** when:
128
128
  - Final ads where the avatar's face fills the screen
129
129
  - The character is named / branded — identity drift kills the take
130
- - You're already paying $1+ for the surrounding video pipeline
130
+ - You're already paying tens of credits for the surrounding video pipeline
131
131
 
132
- Don't default to pro. The $0.44 delta per take adds up across iteration.
132
+ Don't default to pro. The ~15-credit delta per take adds up across iteration.
133
133
 
134
134
  ## Common failure modes
135
135
 
@@ -145,17 +145,17 @@ Don't default to pro. The $0.44 delta per take adds up across iteration.
145
145
 
146
146
  ## Cost discipline
147
147
 
148
- - Video re-dub at $0.11 is the cheapest dialogue iteration in the entire Slates stack — use it for voice A/B testing
149
- - Avatar standard at $0.42 is fine for medium use
150
- - Avatar pro at $0.86 trips the >$0.50 confirm gate — explicit user OK required every time
148
+ - Video re-dub at ~4 credits is the cheapest dialogue iteration in the entire Slates stack — use it for voice A/B testing
149
+ - Avatar standard at ~14 credits is fine for medium use
150
+ - Avatar pro at ~29 credits trips the confirm gate — explicit user OK required every time
151
151
  - All 5s. There is no shorter option.
152
152
 
153
153
  ## Workflow patterns
154
154
 
155
155
  **Voice A/B test (cheap):**
156
- 1. Generate one base talking-head video clip with Veo or Seedance (~$1.20)
156
+ 1. Generate one base talking-head video clip with Veo or Seedance (~40 credits)
157
157
  2. Run `slates_generate_lip_sync` with `sourceType: 'video'` against 3–5 different `ttsVoice` values
158
- 3. Total cost: $1.20 + (5 × $0.11) = $1.75 to compare voices
158
+ 3. Total cost: ~40 + (5 × ~4) ≈ 60 credits to compare voices
159
159
 
160
160
  **Brand avatar from a single portrait:**
161
161
  1. Generate or upload the hero portrait (face fills frame, eyes open, neutral mouth)
@@ -171,7 +171,7 @@ Don't default to pro. The $0.44 delta per take adds up across iteration.
171
171
 
172
172
  Lip-sync is mechanical — the model re-syncs the chosen source to the chosen audio. The confirm response carries the source asset's code so you can announce it in chat.
173
173
 
174
- - ✅ "Lip-syncing **IMG-A12 — Founder Portrait** to the new line. $0.86 on avatar-pro. Confirm?"
174
+ - ✅ "Lip-syncing **IMG-A12 — Founder Portrait** to the new line. ~29 credits on avatar-pro. Confirm?"
175
175
  - ❌ "Using the founder image..." (which? Three exist.)
176
176
 
177
177
  Don't second-guess the source. If the output is wrong, iterate on source choice or audio, not on a refinement prompt (there isn't one).
@@ -9,11 +9,11 @@ Take a still **target image** (your character) and a **source video** (the motio
9
9
 
10
10
  | Engine | Cost | Use case |
11
11
  |------|-----------|----------|
12
- | Kling std (`kling-mc-std-5s`) | $0.95 / 5s | General motion transfer, budget lane |
13
- | Kling pro (`kling-mc-pro-5s`) | $1.26 / 5s | Cleaner anatomy, better identity preservation |
12
+ | Kling std (`kling-mc-std-5s`) | ~32 credits / 5s | General motion transfer, budget lane |
13
+ | Kling pro (`kling-mc-pro-5s`) | ~42 credits / 5s | Cleaner anatomy, better identity preservation |
14
14
  | **Seedance 2.0** (`motionModel=seedance-2`) | per second of input+output (`seedance-2-face-vref-*`) | **Premium lane** — single-pass generation with the driving clip as a native conditioning signal: better motion fidelity, native audio, prompt-directed |
15
15
 
16
- All tiers trip the >$0.50 confirm gate. User OK required every time.
16
+ All tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)
17
17
 
18
18
  ## Seedance engine (premium single-pass)
19
19
 
@@ -80,18 +80,18 @@ Switch to `image` when the target image's composition is the brand asset and the
80
80
 
81
81
  ## Tier choice — std vs pro
82
82
 
83
- **std ($0.95)** for:
83
+ **std (~32 credits)** for:
84
84
  - Drafts, motion exploration, blocking
85
85
  - Group scenes where the character isn't a hero shot
86
86
  - When the budget is tight and the motion is the focus
87
87
 
88
- **pro ($1.26)** for:
88
+ **pro (~42 credits)** for:
89
89
  - Final hero takes
90
90
  - Branded characters where identity drift = unacceptable
91
91
  - Anatomically complex motion (limbs crossing, fast direction changes)
92
92
  - Anime / cartoon target images — pro handles non-realistic styles better
93
93
 
94
- Don't default to pro. The $0.31 delta compounds fast across iteration.
94
+ Don't default to pro. The ~10-credit delta compounds fast across iteration.
95
95
 
96
96
  ## Prompt usage (optional)
97
97
 
@@ -126,7 +126,7 @@ Leave it empty if you don't have a specific atmospheric note.
126
126
  2. Find driving footage — a clean reference video of the dance you want
127
127
  3. Upload both as project assets
128
128
  4. Run motion transfer with `motionModel: 'kling-mc-pro'`, `characterOrientation: 'video'`
129
- 5. Total cost: $1.26 per 5s take
129
+ 5. Total cost: ~42 credits per 5s take
130
130
 
131
131
  **Subtle motion on a hero portrait:**
132
132
  1. Use the locked hero portrait as the target image
@@ -143,15 +143,15 @@ Leave it empty if you don't have a specific atmospheric note.
143
143
  ## Cost discipline
144
144
 
145
145
  - 5 seconds, no shorter option
146
- - Both tiers trip the >$0.50 confirm gate — every call needs explicit user OK
147
- - Iteration is expensive: 4 takes at pro = $5.04. Lock framing + driving video before tier-up to pro.
146
+ - Both tiers trip the confirm gate — every call needs explicit user OK
147
+ - Iteration is expensive: 4 takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.
148
148
  - Always run a single std take first to validate the motion + framing combo before committing to pro
149
149
 
150
150
  ## Confirm gate: cost + codes, no inline preview
151
151
 
152
- Motion transfer is mechanical — the model deterministically applies source motion to target image. Both tiers trip the >$0.50 confirm gate; the response includes the asset codes for source and target so you can announce them in chat.
152
+ Motion transfer is mechanical — the model deterministically applies source motion to target image. Both tiers trip the confirm gate; the response includes the asset codes for source and target so you can announce them in chat.
153
153
 
154
- - ✅ "Transferring motion from **VID-V3** onto **IMG-A12 — Detective Closeup**. $1.26, confirm?"
154
+ - ✅ "Transferring motion from **VID-V3** onto **IMG-A12 — Detective Closeup**. ~42 credits, confirm?"
155
155
  - ❌ "Using the walk video and the detective image..." (multiple of each in the project.)
156
156
 
157
157
  Don't second-guess the assets the user picked — the model executes the transfer. If the output is wrong, iterate on motion source or target choice, not on a refinement prompt.
@@ -0,0 +1,44 @@
1
+ ---
2
+ name: slates-prompting-omni-flash
3
+ description: How to prompt Gemini Omni Flash (Google, via fal). Read before calling slates_generate_video with omni-flash or slates_edit_video with omni-flash-edit. Cheap 720p tier with native synced audio included — 3-10s, 16:9/9:16 only; t2v, single-start-frame i2v, or reference-to-video with up to 7 reference images. The edit variant is the EDIT-FIDELITY WINNER for footage-synced VFX (receipt 2026-07-09) — but ONLY with short prompts: one change + "Keep everything else the same." Long descriptive prompts destroy fidelity.
4
+ ---
5
+
6
+ # Gemini Omni Flash — prompting
7
+
8
+ Google's fast video generation + editing model ("Nano Banana Pro for video" in creator slang — a nickname; it is NOT the NB Pro image model). Carried on fal (`google/gemini-omni-flash*`). 720p only, 24fps, 3–10 second clips, 16:9 or 9:16. **Audio is native and included** — dialogue, SFX, and ambient generate WITH the video at no extra cost.
9
+
10
+ ## Where it routes
11
+
12
+ - **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.
13
+ - **Cheap drafts and iteration volume** — lowest-cost audio-native video seat (~6.4 cr/s at 720p).
14
+ - **NOT hero GENERATION shots** — Kling 3.0 stays the general gen default, Seedance 2.0 the premium tier; Omni Flash's *generation* quality seat is still unproven.
15
+
16
+ ## Editing (`slates_edit_video`, model `omni-flash-edit`) — THE RULES (receipts, not theory)
17
+
18
+ 1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes."* Live receipt 2026-07-09: a long "keep every frame/word/movement identical…" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same."*
19
+ 2. **Always end with "Keep everything else the same."** — the one documented preservation lever.
20
+ 3. **Never name a real-world object as a metaphor.** "Candle-like flame" rendered a literal candle in his hand. Describe the effect itself ("small magical flames on his fingertips").
21
+ 3b. **No conditional timing cues — they HARD-FAIL, not drift.** Receipt 2026-07-09: "a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…" → deterministic `invalid_request` (2×, "could not generate with the given inputs"); collapsing to one continuous action — "A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke." — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video.
22
+ 4. **Safety filter (Google's, strict about harm-to-person):** "fingertips ignite / catch fire" → `content_policy_violation`. Frame effects as magical/harmless VFX: "small magical flames appear on his fingertips" passed. See slates-content-policy §Gemini for the substitution patterns.
23
+ 5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.
24
+ 6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.
25
+ 7. Source clip 3–10s (trim longer clips first). Output length follows the source; billing per output second, rounded up. Voice editing unsupported — never ask it to change dialogue.
26
+ 8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).
27
+ 9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.
28
+
29
+ ## Generation (`slates_generate_video`, model `omni-flash`)
30
+
31
+ - **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
32
+ - Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.
33
+ - **Name references inline** the standard Slates way ("Marcus (images 1 and 2) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
34
+ - **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language ("rain patters on the tin roof"). Negative direction as plain instructions ("Do not show text").
35
+ - Duration is an explicit 3–10s integer param; cost scales linearly per second.
36
+
37
+ ## Input conditioning (Slates handles this — know it exists)
38
+
39
+ Phone footage stores rotation as a metadata flag; models ignore it and edit the raw sideways pixels. Clips must be rotation-normalized (and oversized sources downscaled) before upload — receipt 2026-07-09: a portrait Pixel clip came back sideways until conditioned. If an edit output comes back rotated, the source wasn't normalized.
40
+
41
+ ## Content notes
42
+
43
+ - Google applies its own safety filters to input images/clips and output. Uploads containing recognizable real people are restricted by Google's policy — though own-footage editing of the uploader passed on our route 2026-07-09. See slates-content-policy.
44
+ - Output carries an invisible SynthID watermark (Google-side, programmatic detection only).
@@ -5,7 +5,7 @@ description: How to prompt Seedance 2.0 (ByteDance video model). Read before cal
5
5
 
6
6
  # Seedance 2.0 — prompting
7
7
 
8
- ByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K, default 1080p), 4–15s, first+last frame + up to 9 reference images.
8
+ ByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K — 4K video is Pro-only, default 1080p), 4–15s, first+last frame + up to 9 reference images.
9
9
 
10
10
  ## Official 6-step formula
11
11
 
@@ -5,7 +5,7 @@ description: How to prompt Veo 3.1 (Google). Read before calling slates_generate
5
5
 
6
6
  # Veo 3.1 — prompting
7
7
 
8
- Google DeepMind's video model. Two tiers: `veo-3.1-fast` (cheaper, quick) and `veo-3.1-standard` (higher quality). 4k variants exist for both.
8
+ Google DeepMind's video model. Two tiers: `veo-3.1-fast` (cheaper, quick) and `veo-3.1-standard` (higher quality). 4k variants exist for both (4K video requires Slates Pro).
9
9
 
10
10
  **Native single-shot duration: 4, 6, or 8 seconds.** Longer durations require chaining clips via Extend / last-frame reuse — quality degrades if naively requested past 8s in a single generation. Aspect ratio: **16:9 only** — `slates_generate_video` locks Veo to 16:9; anything else is ignored or fails. For 9:16 vertical, use Kling or Seedance instead.
11
11
 
@@ -33,7 +33,7 @@ Ask: **"Generate frame images now? (y/N)"**
33
33
 
34
34
  ### 3. Generate frames if requested
35
35
  For each frame:
36
- - Estimate cost (`slates_estimate_generation_cost`, `count = total frames`). Confirm with user if total > $0.50.
36
+ - Estimate cost (`slates_estimate_generation_cost`, `count = total frames`). Confirm with user if total > ~17 credits.
37
37
  - Generate sequentially, with character/environment/style references attached when present in the project (`slates_list_characters`, `slates_list_environments`).
38
38
  - Each result returns inline. Evaluate. If wrong, refine prompt + regenerate (charge once, not multiple).
39
39
  - Bind to the frame via `slates_add_frame`.
@@ -47,7 +47,7 @@ The code is the FORMAL reference. The label is human texture. Use both: `IMG-A12
47
47
 
48
48
  - Track total credits spent across the loop. Surface to the user every 3 iterations.
49
49
  - Stop after 3 failed iterations on the same prompt — escalate to the user with what you tried and what's not working. The slot machine never converges.
50
- - For high-cost generations (`> $0.50`), confirm before *every* attempt, not just the first.
50
+ - For high-cost generations (above ~17 credits), confirm before *every* attempt, not just the first.
51
51
 
52
52
  ## When to break the loop
53
53