@slatesvideo/shared 0.5.0 → 0.5.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/clients/cloud.d.ts +7 -7
- package/dist/operations/index.d.ts +10 -2
- package/dist/operations/index.js +217 -94
- package/dist/prompts/model-facts.js +18 -2
- package/dist/skills/content.js +13 -12
- package/package.json +1 -1
- package/skills/slates-content-policy.md +14 -1
- package/skills/slates-cost-discipline.md +15 -11
- package/skills/slates-direct-response-ad.md +2 -2
- package/skills/slates-edit-and-iterate.md +1 -1
- package/skills/slates-model-selection.md +11 -7
- package/skills/slates-one-prompt-film.md +1 -1
- package/skills/slates-prompting-lip-sync.md +12 -12
- package/skills/slates-prompting-motion-transfer.md +11 -11
- package/skills/slates-prompting-omni-flash.md +44 -0
- package/skills/slates-prompting-seedance.md +1 -1
- package/skills/slates-prompting-veo-3.md +1 -1
- package/skills/slates-storyboard-from-script.md +1 -1
- package/skills/slates-vision-feedback-loop.md +1 -1
|
@@ -54,10 +54,23 @@ Never write romantic, sexual, or suggestive content involving or directed at min
|
|
|
54
54
|
|
|
55
55
|
If a box fails, apply the substitution table before writing the prompt.
|
|
56
56
|
|
|
57
|
-
## Editing real footage (Kling O3 edit) — real people in the SOURCE
|
|
57
|
+
## Editing real footage (Kling O3 edit / Omni Flash edit) — real people in the SOURCE
|
|
58
58
|
|
|
59
59
|
Video edit takes the user's own footage, which often contains real people. Rules:
|
|
60
60
|
|
|
61
61
|
- The user must hold rights/consent for any real person's likeness in footage they edit — ask once when it's clearly someone other than the user, then proceed.
|
|
62
62
|
- Kling's video-to-video filter behavior on real faces is **not yet verified** (unlike Seedance, where the consent-gated real-face route is confirmed). If an edit of real-person footage is rejected by the provider, do NOT retry-spam variations — tell the user the filter blocked it and offer a no-face crop/segment or an AI-character swap instead.
|
|
63
|
+
- **Omni Flash: own-footage editing of the uploader's own face PASSED live 2026-07-09** (real talking-head clip, edited on our fal route) despite Google's documented "recognizable people" restriction — treat that restriction as aimed at third-party/public figures, but expect probabilistic refusals and never promise passage.
|
|
63
64
|
- Never use edit to put a real, named public figure into a scene, or to make someone appear to say/do something they didn't. Faceless b-roll (hands, products, landscapes, crowds-from-behind) edits freely.
|
|
65
|
+
|
|
66
|
+
## Gemini / Omni Flash filter regime (video gen + edit) — receipts 2026-07-09
|
|
67
|
+
|
|
68
|
+
Google's filter is its own regime (stricter than fal-hosted Kling about harm-to-a-person, looser than BytePlus about faces). Live receipts:
|
|
69
|
+
|
|
70
|
+
| Blocked (`content_policy_violation`) | Passed |
|
|
71
|
+
|---|---|
|
|
72
|
+
| "his fingertips **ignite** with a small real flame" (fire ON a body part = harm) | "small **magical** flames appear on his fingertips … vanish when he blows on them" |
|
|
73
|
+
|
|
74
|
+
- **Harm-to-person framing is the tripwire**, not the effect itself. Reframe body-contact effects as magical / supernatural / harmless VFX: "magical flames", "a glowing aura", "sparks of light dance on". Avoid ignite / burn / on fire / catch fire applied to a person.
|
|
75
|
+
- **Never use a real object as a metaphor** — "candle-like flame" rendered a literal candle in the subject's hand. Describe the effect, not an object that resembles it.
|
|
76
|
+
- The block is a 422 refund (no credits lost) and arrives mid-generation — one reframe per the substitution mindset above, don't retry-spam.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-cost-discipline
|
|
3
|
-
description: Mandatory pre-flight discipline before ANY generation call (image or video) — estimate cost, announce in
|
|
3
|
+
description: Mandatory pre-flight discipline before ANY generation call (image or video) — estimate cost, announce in credits, get confirmation, aggregate batches. Read this every time before calling slates_generate_image or any future slates_generate_* op. Skipping this risks burning the user's credits on guesses.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Slates cost discipline — read before every generation
|
|
@@ -20,23 +20,25 @@ Before ANY `slates_generate_*` call, run `slates_estimate_generation_cost` first
|
|
|
20
20
|
|
|
21
21
|
If aspect ratio or resolution isn't obvious from the user's request, **ask before estimating**. Don't guess.
|
|
22
22
|
|
|
23
|
-
### 2. Announce in
|
|
23
|
+
### 2. Announce in credits, plainly, before spending
|
|
24
24
|
|
|
25
|
-
|
|
25
|
+
Slates bills abstract **credits** (they never expire). Announce the credit total the estimate returns — never dollars.
|
|
26
|
+
|
|
27
|
+
Format: `About to spend N credits on M image(s) at [resolution] [aspect ratio]. Proceed?`
|
|
26
28
|
|
|
27
29
|
Examples:
|
|
28
|
-
- `About to spend
|
|
29
|
-
- `About to spend
|
|
30
|
+
- `About to spend 4 credits on 1 image at 1k 16:9. Proceed?`
|
|
31
|
+
- `About to spend 24 credits on 4 images at 2k 9:16 (variants). Proceed?`
|
|
30
32
|
|
|
31
|
-
Below
|
|
33
|
+
Below ~7 credits you can proceed silently after announcing once. Above ~7 credits wait for explicit confirmation. Above ~17 credits the server itself will gate with `requires_confirm` — pass `confirm: true` only after the user explicitly OKs.
|
|
32
34
|
|
|
33
35
|
### 3. Aggregate batches into ONE upfront announcement
|
|
34
36
|
|
|
35
|
-
If you're planning a multi-call workflow (5 storyboard frames, 3 character variants, a grid of options), **announce the total before the first call**, not five
|
|
37
|
+
If you're planning a multi-call workflow (5 storyboard frames, 3 character variants, a grid of options), **announce the total before the first call**, not five small announcements after the fact.
|
|
36
38
|
|
|
37
|
-
Format: `Plan: N generations totaling
|
|
39
|
+
Format: `Plan: N generations totaling C credits. [Brief description of the sequence.] Proceed with the batch?`
|
|
38
40
|
|
|
39
|
-
Example: `Plan: 6 frame generations at 1k 16:9 totaling
|
|
41
|
+
Example: `Plan: 6 frame generations at 1k 16:9 totaling 24 credits — establishing wide, push-in, two-shot, reverse, OTS, insert. Proceed?`
|
|
40
42
|
|
|
41
43
|
### 3b. Batch authorization — one approval covers the enumerated batch
|
|
42
44
|
|
|
@@ -52,7 +54,7 @@ One approval = that plan, as enumerated, at those prices. Nothing else.
|
|
|
52
54
|
|
|
53
55
|
### 4. Track the running total
|
|
54
56
|
|
|
55
|
-
After each generation completes, the response includes `
|
|
57
|
+
After each generation completes, the response includes `cost_credits` (when available). Keep a running tally in your context. Surface it every 3 generations or whenever the user asks "how much have we spent?"
|
|
56
58
|
|
|
57
59
|
## Resolution decision rules
|
|
58
60
|
|
|
@@ -66,6 +68,8 @@ After each generation completes, the response includes `cost_cents` (when availa
|
|
|
66
68
|
|
|
67
69
|
Resolution is a price lever, not a free choice: on Nano Banana 2 and FLUX.2 Max, 4k costs roughly 2x 1k (Seedream 5 Lite is flat-priced regardless of resolution). Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models` for current numbers instead of assuming. Pick the cheapest resolution that serves the use case.
|
|
68
70
|
|
|
71
|
+
**4K VIDEO is Pro-only (2026-07-07).** The ladder above is for IMAGES (open at every tier). For VIDEO — Kling, Seedance, Veo — 4K requires a Slates Pro account; a base-tier 4K video gen is rejected server-side with `PRO_REQUIRED`. Default video to 1080p or lower and only reach for 4K when the user is on Pro and explicitly asks. 4K *images* are never gated.
|
|
72
|
+
|
|
69
73
|
## Aspect ratio decision rules
|
|
70
74
|
|
|
71
75
|
Ask the user when ambiguous. Otherwise:
|
|
@@ -83,7 +87,7 @@ If the user prompt mixes signals (e.g. "cinematic Instagram post"), ask. Don't g
|
|
|
83
87
|
|
|
84
88
|
## When the gate fires
|
|
85
89
|
|
|
86
|
-
The server returns `requires_clarification` when aspect ratio or resolution is missing. The server returns `requires_confirm` when total spend exceeds
|
|
90
|
+
The server returns `requires_clarification` when aspect ratio or resolution is missing. The server returns `requires_confirm` when total spend exceeds ~17 credits. In both cases:
|
|
87
91
|
|
|
88
92
|
1. Surface the gate response to the user
|
|
89
93
|
2. Get a clean answer
|
|
@@ -10,7 +10,7 @@ You are building a 30-second hyper-motion direct-response ad. The user has hande
|
|
|
10
10
|
**Hard rules**
|
|
11
11
|
|
|
12
12
|
- Always estimate cost before generating. Use `slates_estimate_generation_cost` and surface the total.
|
|
13
|
-
- All
|
|
13
|
+
- All Slates generation routes through Slates Credits, period (BYOK is retired) — don't suggest "use your own keys" workarounds.
|
|
14
14
|
- Default model: `nano-banana-2-2k`. For close-up product hero frames step up to `4k` only if the user asks.
|
|
15
15
|
- Hyper-motion = punchy cuts, 4 frames in 30 seconds, ~7s each. Don't over-storyboard.
|
|
16
16
|
|
|
@@ -58,7 +58,7 @@ For each frame:
|
|
|
58
58
|
|
|
59
59
|
- **Don't** generate text overlays in the image. Slates renders captions/CTAs at the editor stage.
|
|
60
60
|
- **Don't** burn credits on slot-machine prompting. If the first generation is off, refine the prompt; don't just regenerate.
|
|
61
|
-
- **Don't** skip the cost estimate. Confirm with the user above
|
|
61
|
+
- **Don't** skip the cost estimate. Confirm with the user above ~17 credits.
|
|
62
62
|
- **Don't** invent visual specifics about the product (colors, textures, angles) that aren't in the reference image. Reference-anchored prompts only.
|
|
63
63
|
|
|
64
64
|
## Voice
|
|
@@ -28,7 +28,7 @@ The user's request is one of:
|
|
|
28
28
|
| Aesthetic / compositional | `slates_generate_image` with the original in `referenceAssetIds` + a refined prompt. Don't re-roll from scratch. |
|
|
29
29
|
| Wholesale | New prompt, no reference, fresh generation. Treat as a new brief. |
|
|
30
30
|
|
|
31
|
-
**`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Default model `nano-banana-2` — the only edit model that also takes extra `referenceAssetIds`; `flux-2-max` / `seedream-5-lite` use their own edit endpoints and ignore references. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost
|
|
31
|
+
**`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Default model `nano-banana-2` — the only edit model that also takes extra `referenceAssetIds`; `flux-2-max` / `seedream-5-lite` use their own edit endpoints and ignore references. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost above ~17 credits gates on `confirm=true`.
|
|
32
32
|
|
|
33
33
|
### 4. Generate, evaluate, decide
|
|
34
34
|
- Estimate cost first.
|
|
@@ -14,7 +14,7 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
|
|
|
14
14
|
| **General-purpose — the default for most shots** | **Kling 3.0 std** | Cost-effective workhorse. Strong image-to-video: preserves identity, layout, and text from the start frame. Any aspect ratio, 5–15s. |
|
|
15
15
|
| Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |
|
|
16
16
|
| Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
|
|
17
|
-
| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K. |
|
|
17
|
+
| **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
|
|
18
18
|
| The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
|
|
19
19
|
| Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | The only job Veo wins. |
|
|
20
20
|
|
|
@@ -22,12 +22,16 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
|
|
|
22
22
|
|
|
23
23
|
| Job | Tool | Why |
|
|
24
24
|
|---|---|---|
|
|
25
|
-
| **
|
|
26
|
-
|
|
|
27
|
-
|
|
|
25
|
+
| **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + "Keep everything else the same"; long prompts destroy it (see below). |
|
|
26
|
+
| **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
|
|
27
|
+
| **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
|
|
28
|
+
| Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
|
|
29
|
+
| AI-edit the user's OWN footage | Omni Flash Edit (3–10s) or Kling O3 Edit (3–15s, 720–3840px) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
|
|
28
30
|
|
|
29
31
|
- **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
|
|
30
|
-
-
|
|
32
|
+
- **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the pro demos (e.g. Higgsfield's split-screen short) actually work, plus gesture-only beats with voiceover laid over in post.
|
|
33
|
+
- **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
|
|
34
|
+
- Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
|
|
31
35
|
|
|
32
36
|
## Motion Transfer & Lip Sync routing (two engines per tool)
|
|
33
37
|
|
|
@@ -35,9 +39,9 @@ Both tools have a cheap Kling utility lane and a premium Seedance lane. The capa
|
|
|
35
39
|
|
|
36
40
|
| Job | Engine | Why |
|
|
37
41
|
|---|---|---|
|
|
38
|
-
| Quick motion retarget, budget lane, or driving clip >15s | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget,
|
|
42
|
+
| Quick motion retarget, budget lane, or driving clip >15s | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |
|
|
39
43
|
| **Motion transfer where fidelity or audio matters** — dance, choreography, cinematic action | **Seedance 2.0** (`motionModel=seedance-2`) | Single-pass conditioning beats post-hoc retargeting; prompt-driven; native audio. Driving clip 2–15s; bills input+output seconds (vref keys). |
|
|
40
|
-
| Cheap lip-sync utility (re-voice a clip, simple avatar) | Kling lip-sync / avatar (`slates_generate_lip_sync`) |
|
|
44
|
+
| Cheap lip-sync utility (re-voice a clip, simple avatar) | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |
|
|
41
45
|
| **Natural speech, voice cloned from the source clip, premium delivery** | **Seedance 2.0** (`engine=seedance-2`) | The line is spoken IN the generation (no TTS layer); a video source keeps its own voice; uploaded ≤15s audio can drive it. |
|
|
42
46
|
|
|
43
47
|
- Faces: Seedance tool gens default `seedanceFace=true` (sources are people). A REAL person triggers the consent cascade (`[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent`, premium realface pricing).
|
|
@@ -39,7 +39,7 @@ Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the chara
|
|
|
39
39
|
|
|
40
40
|
**Model mixing — route per `slates-model-selection`** (details in the per-model guides):
|
|
41
41
|
- **Kling V3** (`slates-prompting-kling-v3`): the DEFAULT for most shots — any aspect ratio, 5-15s, strong start-frame adherence; std is the workhorse, Omni for multi-character dialogue.
|
|
42
|
-
- **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K.
|
|
42
|
+
- **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K (4K video is Pro-only).
|
|
43
43
|
- **Veo 3.1** (`slates-prompting-veo-3`): niche, never the default — only when native synced audio must generate WITH the video in one gen; 16:9 only, 4/6/8s.
|
|
44
44
|
|
|
45
45
|
Failed gen? Check the error via `slates_get_generation_status`, fix the prompt, resubmit that one shot (a retry beyond the plan = announce the delta cost).
|
|
@@ -9,9 +9,9 @@ Two engines. Kling is the cheap utility lane (dedicated lip-sync endpoints, 5-se
|
|
|
9
9
|
|
|
10
10
|
| Flow | Source | Engine/Model | Cost | Use case |
|
|
11
11
|
|------|--------|-------|-----------|----------|
|
|
12
|
-
| Re-dub | video clip | kling-lip-sync-video |
|
|
13
|
-
| Avatar standard | still image | ai-avatar/v2/standard |
|
|
14
|
-
| Avatar pro | still image | ai-avatar/v2/pro |
|
|
12
|
+
| Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |
|
|
13
|
+
| Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |
|
|
14
|
+
| Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |
|
|
15
15
|
| **Seedance native** | image or video | `engine=seedance-2` | per second (`seedance-2-face-*`; video sources bill input+output seconds) | **Premium**: natural delivery, whole-body performance, voice cloned from a video source, audio included |
|
|
16
16
|
|
|
17
17
|
Pick engine + `sourceType` deliberately — they decide the pricing tier and the underlying endpoint.
|
|
@@ -34,7 +34,7 @@ Everything below applies to the **Kling** engine.
|
|
|
34
34
|
Use **video** (re-dub) when:
|
|
35
35
|
- A talking-head clip already exists (Slates-generated, recorded, or imported)
|
|
36
36
|
- The mouth/face is already moving and only the audio needs to change
|
|
37
|
-
-
|
|
37
|
+
- ~4 credits is hard to beat for short dialogue replacement
|
|
38
38
|
|
|
39
39
|
Use **avatar** when:
|
|
40
40
|
- Only a still portrait exists
|
|
@@ -127,9 +127,9 @@ Default `"."` is fine if you have nothing useful to add.
|
|
|
127
127
|
**Use pro** when:
|
|
128
128
|
- Final ads where the avatar's face fills the screen
|
|
129
129
|
- The character is named / branded — identity drift kills the take
|
|
130
|
-
- You're already paying
|
|
130
|
+
- You're already paying tens of credits for the surrounding video pipeline
|
|
131
131
|
|
|
132
|
-
Don't default to pro. The
|
|
132
|
+
Don't default to pro. The ~15-credit delta per take adds up across iteration.
|
|
133
133
|
|
|
134
134
|
## Common failure modes
|
|
135
135
|
|
|
@@ -145,17 +145,17 @@ Don't default to pro. The $0.44 delta per take adds up across iteration.
|
|
|
145
145
|
|
|
146
146
|
## Cost discipline
|
|
147
147
|
|
|
148
|
-
- Video re-dub at
|
|
149
|
-
- Avatar standard at
|
|
150
|
-
- Avatar pro at
|
|
148
|
+
- Video re-dub at ~4 credits is the cheapest dialogue iteration in the entire Slates stack — use it for voice A/B testing
|
|
149
|
+
- Avatar standard at ~14 credits is fine for medium use
|
|
150
|
+
- Avatar pro at ~29 credits trips the confirm gate — explicit user OK required every time
|
|
151
151
|
- All 5s. There is no shorter option.
|
|
152
152
|
|
|
153
153
|
## Workflow patterns
|
|
154
154
|
|
|
155
155
|
**Voice A/B test (cheap):**
|
|
156
|
-
1. Generate one base talking-head video clip with Veo or Seedance (
|
|
156
|
+
1. Generate one base talking-head video clip with Veo or Seedance (~40 credits)
|
|
157
157
|
2. Run `slates_generate_lip_sync` with `sourceType: 'video'` against 3–5 different `ttsVoice` values
|
|
158
|
-
3. Total cost:
|
|
158
|
+
3. Total cost: ~40 + (5 × ~4) ≈ 60 credits to compare voices
|
|
159
159
|
|
|
160
160
|
**Brand avatar from a single portrait:**
|
|
161
161
|
1. Generate or upload the hero portrait (face fills frame, eyes open, neutral mouth)
|
|
@@ -171,7 +171,7 @@ Don't default to pro. The $0.44 delta per take adds up across iteration.
|
|
|
171
171
|
|
|
172
172
|
Lip-sync is mechanical — the model re-syncs the chosen source to the chosen audio. The confirm response carries the source asset's code so you can announce it in chat.
|
|
173
173
|
|
|
174
|
-
- ✅ "Lip-syncing **IMG-A12 — Founder Portrait** to the new line.
|
|
174
|
+
- ✅ "Lip-syncing **IMG-A12 — Founder Portrait** to the new line. ~29 credits on avatar-pro. Confirm?"
|
|
175
175
|
- ❌ "Using the founder image..." (which? Three exist.)
|
|
176
176
|
|
|
177
177
|
Don't second-guess the source. If the output is wrong, iterate on source choice or audio, not on a refinement prompt (there isn't one).
|
|
@@ -9,11 +9,11 @@ Take a still **target image** (your character) and a **source video** (the motio
|
|
|
9
9
|
|
|
10
10
|
| Engine | Cost | Use case |
|
|
11
11
|
|------|-----------|----------|
|
|
12
|
-
| Kling std (`kling-mc-std-5s`) |
|
|
13
|
-
| Kling pro (`kling-mc-pro-5s`) |
|
|
12
|
+
| Kling std (`kling-mc-std-5s`) | ~32 credits / 5s | General motion transfer, budget lane |
|
|
13
|
+
| Kling pro (`kling-mc-pro-5s`) | ~42 credits / 5s | Cleaner anatomy, better identity preservation |
|
|
14
14
|
| **Seedance 2.0** (`motionModel=seedance-2`) | per second of input+output (`seedance-2-face-vref-*`) | **Premium lane** — single-pass generation with the driving clip as a native conditioning signal: better motion fidelity, native audio, prompt-directed |
|
|
15
15
|
|
|
16
|
-
All tiers trip the
|
|
16
|
+
All tiers trip the confirm gate. User OK required every time. (Prices are approximate — `slates_estimate_generation_cost` returns the exact credit total.)
|
|
17
17
|
|
|
18
18
|
## Seedance engine (premium single-pass)
|
|
19
19
|
|
|
@@ -80,18 +80,18 @@ Switch to `image` when the target image's composition is the brand asset and the
|
|
|
80
80
|
|
|
81
81
|
## Tier choice — std vs pro
|
|
82
82
|
|
|
83
|
-
**std (
|
|
83
|
+
**std (~32 credits)** for:
|
|
84
84
|
- Drafts, motion exploration, blocking
|
|
85
85
|
- Group scenes where the character isn't a hero shot
|
|
86
86
|
- When the budget is tight and the motion is the focus
|
|
87
87
|
|
|
88
|
-
**pro (
|
|
88
|
+
**pro (~42 credits)** for:
|
|
89
89
|
- Final hero takes
|
|
90
90
|
- Branded characters where identity drift = unacceptable
|
|
91
91
|
- Anatomically complex motion (limbs crossing, fast direction changes)
|
|
92
92
|
- Anime / cartoon target images — pro handles non-realistic styles better
|
|
93
93
|
|
|
94
|
-
Don't default to pro. The
|
|
94
|
+
Don't default to pro. The ~10-credit delta compounds fast across iteration.
|
|
95
95
|
|
|
96
96
|
## Prompt usage (optional)
|
|
97
97
|
|
|
@@ -126,7 +126,7 @@ Leave it empty if you don't have a specific atmospheric note.
|
|
|
126
126
|
2. Find driving footage — a clean reference video of the dance you want
|
|
127
127
|
3. Upload both as project assets
|
|
128
128
|
4. Run motion transfer with `motionModel: 'kling-mc-pro'`, `characterOrientation: 'video'`
|
|
129
|
-
5. Total cost:
|
|
129
|
+
5. Total cost: ~42 credits per 5s take
|
|
130
130
|
|
|
131
131
|
**Subtle motion on a hero portrait:**
|
|
132
132
|
1. Use the locked hero portrait as the target image
|
|
@@ -143,15 +143,15 @@ Leave it empty if you don't have a specific atmospheric note.
|
|
|
143
143
|
## Cost discipline
|
|
144
144
|
|
|
145
145
|
- 5 seconds, no shorter option
|
|
146
|
-
- Both tiers trip the
|
|
147
|
-
- Iteration is expensive: 4 takes at pro
|
|
146
|
+
- Both tiers trip the confirm gate — every call needs explicit user OK
|
|
147
|
+
- Iteration is expensive: 4 takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.
|
|
148
148
|
- Always run a single std take first to validate the motion + framing combo before committing to pro
|
|
149
149
|
|
|
150
150
|
## Confirm gate: cost + codes, no inline preview
|
|
151
151
|
|
|
152
|
-
Motion transfer is mechanical — the model deterministically applies source motion to target image. Both tiers trip the
|
|
152
|
+
Motion transfer is mechanical — the model deterministically applies source motion to target image. Both tiers trip the confirm gate; the response includes the asset codes for source and target so you can announce them in chat.
|
|
153
153
|
|
|
154
|
-
- ✅ "Transferring motion from **VID-V3** onto **IMG-A12 — Detective Closeup**.
|
|
154
|
+
- ✅ "Transferring motion from **VID-V3** onto **IMG-A12 — Detective Closeup**. ~42 credits, confirm?"
|
|
155
155
|
- ❌ "Using the walk video and the detective image..." (multiple of each in the project.)
|
|
156
156
|
|
|
157
157
|
Don't second-guess the assets the user picked — the model executes the transfer. If the output is wrong, iterate on motion source or target choice, not on a refinement prompt.
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-omni-flash
|
|
3
|
+
description: How to prompt Gemini Omni Flash (Google, via fal). Read before calling slates_generate_video with omni-flash or slates_edit_video with omni-flash-edit. Cheap 720p tier with native synced audio included — 3-10s, 16:9/9:16 only; t2v, single-start-frame i2v, or reference-to-video with up to 7 reference images. The edit variant is the EDIT-FIDELITY WINNER for footage-synced VFX (receipt 2026-07-09) — but ONLY with short prompts: one change + "Keep everything else the same." Long descriptive prompts destroy fidelity.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Gemini Omni Flash — prompting
|
|
7
|
+
|
|
8
|
+
Google's fast video generation + editing model ("Nano Banana Pro for video" in creator slang — a nickname; it is NOT the NB Pro image model). Carried on fal (`google/gemini-omni-flash*`). 720p only, 24fps, 3–10 second clips, 16:9 or 9:16. **Audio is native and included** — dialogue, SFX, and ambient generate WITH the video at no extra cost.
|
|
9
|
+
|
|
10
|
+
## Where it routes
|
|
11
|
+
|
|
12
|
+
- **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.
|
|
13
|
+
- **Cheap drafts and iteration volume** — lowest-cost audio-native video seat (~6.4 cr/s at 720p).
|
|
14
|
+
- **NOT hero GENERATION shots** — Kling 3.0 stays the general gen default, Seedance 2.0 the premium tier; Omni Flash's *generation* quality seat is still unproven.
|
|
15
|
+
|
|
16
|
+
## Editing (`slates_edit_video`, model `omni-flash-edit`) — THE RULES (receipts, not theory)
|
|
17
|
+
|
|
18
|
+
1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes."* Live receipt 2026-07-09: a long "keep every frame/word/movement identical…" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same."*
|
|
19
|
+
2. **Always end with "Keep everything else the same."** — the one documented preservation lever.
|
|
20
|
+
3. **Never name a real-world object as a metaphor.** "Candle-like flame" rendered a literal candle in his hand. Describe the effect itself ("small magical flames on his fingertips").
|
|
21
|
+
3b. **No conditional timing cues — they HARD-FAIL, not drift.** Receipt 2026-07-09: "a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…" → deterministic `invalid_request` (2×, "could not generate with the given inputs"); collapsing to one continuous action — "A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke." — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video.
|
|
22
|
+
4. **Safety filter (Google's, strict about harm-to-person):** "fingertips ignite / catch fire" → `content_policy_violation`. Frame effects as magical/harmless VFX: "small magical flames appear on his fingertips" passed. See slates-content-policy §Gemini for the substitution patterns.
|
|
23
|
+
5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.
|
|
24
|
+
6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.
|
|
25
|
+
7. Source clip 3–10s (trim longer clips first). Output length follows the source; billing per output second, rounded up. Voice editing unsupported — never ask it to change dialogue.
|
|
26
|
+
8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).
|
|
27
|
+
9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.
|
|
28
|
+
|
|
29
|
+
## Generation (`slates_generate_video`, model `omni-flash`)
|
|
30
|
+
|
|
31
|
+
- **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
|
|
32
|
+
- Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.
|
|
33
|
+
- **Name references inline** the standard Slates way ("Marcus (images 1 and 2) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
|
|
34
|
+
- **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language ("rain patters on the tin roof"). Negative direction as plain instructions ("Do not show text").
|
|
35
|
+
- Duration is an explicit 3–10s integer param; cost scales linearly per second.
|
|
36
|
+
|
|
37
|
+
## Input conditioning (Slates handles this — know it exists)
|
|
38
|
+
|
|
39
|
+
Phone footage stores rotation as a metadata flag; models ignore it and edit the raw sideways pixels. Clips must be rotation-normalized (and oversized sources downscaled) before upload — receipt 2026-07-09: a portrait Pixel clip came back sideways until conditioned. If an edit output comes back rotated, the source wasn't normalized.
|
|
40
|
+
|
|
41
|
+
## Content notes
|
|
42
|
+
|
|
43
|
+
- Google applies its own safety filters to input images/clips and output. Uploads containing recognizable real people are restricted by Google's policy — though own-footage editing of the uploader passed on our route 2026-07-09. See slates-content-policy.
|
|
44
|
+
- Output carries an invisible SynthID watermark (Google-side, programmatic detection only).
|
|
@@ -5,7 +5,7 @@ description: How to prompt Seedance 2.0 (ByteDance video model). Read before cal
|
|
|
5
5
|
|
|
6
6
|
# Seedance 2.0 — prompting
|
|
7
7
|
|
|
8
|
-
ByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K, default 1080p), 4–15s, first+last frame + up to 9 reference images.
|
|
8
|
+
ByteDance's video model — first-party via **BytePlus ModelArk** (credits only, no BYOK). Audio always generated alongside the video. Single model `seedance-2` across the full resolution ladder (480p / 720p / 1080p / native 4K — 4K video is Pro-only, default 1080p), 4–15s, first+last frame + up to 9 reference images.
|
|
9
9
|
|
|
10
10
|
## Official 6-step formula
|
|
11
11
|
|
|
@@ -5,7 +5,7 @@ description: How to prompt Veo 3.1 (Google). Read before calling slates_generate
|
|
|
5
5
|
|
|
6
6
|
# Veo 3.1 — prompting
|
|
7
7
|
|
|
8
|
-
Google DeepMind's video model. Two tiers: `veo-3.1-fast` (cheaper, quick) and `veo-3.1-standard` (higher quality). 4k variants exist for both.
|
|
8
|
+
Google DeepMind's video model. Two tiers: `veo-3.1-fast` (cheaper, quick) and `veo-3.1-standard` (higher quality). 4k variants exist for both (4K video requires Slates Pro).
|
|
9
9
|
|
|
10
10
|
**Native single-shot duration: 4, 6, or 8 seconds.** Longer durations require chaining clips via Extend / last-frame reuse — quality degrades if naively requested past 8s in a single generation. Aspect ratio: **16:9 only** — `slates_generate_video` locks Veo to 16:9; anything else is ignored or fails. For 9:16 vertical, use Kling or Seedance instead.
|
|
11
11
|
|
|
@@ -33,7 +33,7 @@ Ask: **"Generate frame images now? (y/N)"**
|
|
|
33
33
|
|
|
34
34
|
### 3. Generate frames if requested
|
|
35
35
|
For each frame:
|
|
36
|
-
- Estimate cost (`slates_estimate_generation_cost`, `count = total frames`). Confirm with user if total >
|
|
36
|
+
- Estimate cost (`slates_estimate_generation_cost`, `count = total frames`). Confirm with user if total > ~17 credits.
|
|
37
37
|
- Generate sequentially, with character/environment/style references attached when present in the project (`slates_list_characters`, `slates_list_environments`).
|
|
38
38
|
- Each result returns inline. Evaluate. If wrong, refine prompt + regenerate (charge once, not multiple).
|
|
39
39
|
- Bind to the frame via `slates_add_frame`.
|
|
@@ -47,7 +47,7 @@ The code is the FORMAL reference. The label is human texture. Use both: `IMG-A12
|
|
|
47
47
|
|
|
48
48
|
- Track total credits spent across the loop. Surface to the user every 3 iterations.
|
|
49
49
|
- Stop after 3 failed iterations on the same prompt — escalate to the user with what you tried and what's not working. The slot machine never converges.
|
|
50
|
-
- For high-cost generations (
|
|
50
|
+
- For high-cost generations (above ~17 credits), confirm before *every* attempt, not just the first.
|
|
51
51
|
|
|
52
52
|
## When to break the loop
|
|
53
53
|
|