@slatesvideo/shared 0.5.3 → 0.5.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +31 -0
- package/dist/operations/index.d.ts +4 -8
- package/dist/operations/index.js +26 -38
- package/dist/prompts/character-sheet.d.ts +19 -18
- package/dist/prompts/character-sheet.js +84 -38
- package/dist/prompts/environment-sheet.d.ts +9 -1
- package/dist/prompts/environment-sheet.js +17 -3
- package/dist/prompts/model-facts.js +4 -1
- package/dist/prompts/partials.generated.d.ts +2 -0
- package/dist/prompts/partials.generated.js +14 -0
- package/dist/prompts/prompting-tips.js +50 -19
- package/dist/prompts/reference-composer.d.ts +1 -1
- package/dist/prompts/reference-composer.js +3 -4
- package/dist/prompts/reference-rules.d.ts +43 -14
- package/dist/prompts/reference-rules.js +51 -27
- package/dist/skills/content.js +14 -14
- package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
- package/exports/slates-prompt-builder/generated/reference-character.md +78 -0
- package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
- package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
- package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +8 -3
- package/skills/_partials/decision-log.md +12 -0
- package/skills/_partials/reference-rules-core.md +12 -0
- package/skills/_partials/reference-tips-short.md +2 -0
- package/skills/_partials/references-read-literally.md +11 -0
- package/skills/_partials/still-gate.md +3 -0
- package/skills/slates-character-identity.md +100 -0
- package/skills/slates-cost-discipline.md +10 -0
- package/skills/slates-edit-and-iterate.md +17 -2
- package/skills/slates-model-selection.md +24 -1
- package/skills/slates-one-prompt-film.md +22 -3
- package/skills/slates-prompting-flux-2-max.md +36 -5
- package/skills/slates-prompting-gpt-image-2.md +1 -1
- package/skills/slates-prompting-kling-v3.md +40 -9
- package/skills/slates-prompting-nano-banana-2.md +44 -12
- package/skills/slates-prompting-omni-flash.md +1 -1
- package/skills/slates-prompting-seedance.md +295 -90
- package/skills/slates-prompting-veo-3.md +33 -4
- package/skills/slates-storyboard-from-script.md +19 -0
- package/skills/slates-vision-feedback-loop.md +49 -2
- package/skills/slates-character-turnaround.md +0 -55
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompt-builder
|
|
3
|
+
description: Turn a plain-language idea into a paste-ready AI video or image prompt using the production Slates prompting guides for Seedance 2.0, Kling 3.0, and Nano Banana 2. Use for video prompts, image prompts, shot planning, character-reference preparation, ads, brand films, product videos, talking heads, or any request that needs generation-ready visual direction.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
<!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
|
|
7
|
+
|
|
8
|
+
# Slates Prompt Builder
|
|
9
|
+
|
|
10
|
+
Turn the user's idea into the exact prompt to paste into their generation tool. The deliverable is the prompt, not a lecture about prompting.
|
|
11
|
+
|
|
12
|
+
This portable skill is deliberately thin. Its reference files are generated directly from the same production skills used by the Slates MCP server, CLI-installed Claude skills, and Studio Agent. Treat those references as authoritative; never recreate their rules from memory.
|
|
13
|
+
|
|
14
|
+
## The curated stack
|
|
15
|
+
|
|
16
|
+
<!-- @generated:model-routing -->
|
|
17
|
+
| Model | Canonical route | Guide |
|
|
18
|
+
|---|---|---|
|
|
19
|
+
| **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics. In the Motion Transfer / Lip Sync tools, Kling (MC / lip-sync / avatar) is the cheap utility lane; Seedance is the premium single-pass lane. | `reference-kling.md` |
|
|
20
|
+
| **Seedance 2.0** | PREMIUM video tier — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those). Up to 9 ingredient images. Strong I2V / own-footage restyle. Native 4K, but 4K VIDEO is a Pro-only tier gate (base maxes at 1080p; server returns PRO_REQUIRED) — default 1080p unless the user is on Pro. Also the PREMIUM engine inside the Motion Transfer and Lip Sync tools (single-pass: driving video / dialogue are native conditioning signals — better motion fidelity, natural speech, voice cloned from a video source; video references bill input+output seconds). | `reference-seedance.md` |
|
|
21
|
+
| **Nano Banana 2 (Gemini 3.1 Flash Image)** | Default image model. 14 refs hard cap (10 object + 4 character). Brief it like a creative director, not tag soup. No negativePrompt field — use positive reframing. Best image start-frame for legible text. Knowledge cutoff Jan 2025. | `reference-nano-banana.md` |
|
|
22
|
+
<!-- @end:model-routing -->
|
|
23
|
+
|
|
24
|
+
If the user names a model, use it. Otherwise route by the generated table above.
|
|
25
|
+
|
|
26
|
+
## Workflow
|
|
27
|
+
|
|
28
|
+
1. Read the brief. A sentence or a full storyboard is enough.
|
|
29
|
+
2. If intent is clear, take the fast path: choose the model and write the prompt immediately. Do not interrogate the user for optional detail.
|
|
30
|
+
3. Load the matching generated reference file before writing:
|
|
31
|
+
- `reference-seedance.md`
|
|
32
|
+
- `reference-kling.md`
|
|
33
|
+
- `reference-nano-banana.md`
|
|
34
|
+
4. For recurring characters, identity consistency, or character-sheet preparation, also load `reference-character.md`. It owns the exact sheet architecture, background plate, lighting, and evaluation gate.
|
|
35
|
+
5. For conflict, creatures, crowds, destruction, weapons, public figures, or young characters, also load `reference-content-policy.md` and construct the scene safely from the first word.
|
|
36
|
+
6. If a referenced production guide mentions Slates operations or billing and those tools are not available, use its prompting doctrine and ignore only the transport-specific instruction. Never invent a tool call.
|
|
37
|
+
7. Return one paste-ready prompt. If the concept genuinely requires multiple generations, return the smallest ordered chain (for example: Nano Banana 2 start frame, then Kling motion prompt).
|
|
38
|
+
|
|
39
|
+
## Output
|
|
40
|
+
|
|
41
|
+
```text
|
|
42
|
+
--- PROMPT (<Model>) ---
|
|
43
|
+
<paste-ready prompt>
|
|
44
|
+
--- END ---
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Then give no more than three short notes covering only decisions the user needs to understand: the model route, a non-obvious constraint, or how references should be attached. Do not expose chain-of-thought, internal scoring, density maps, or a shot table unless the user explicitly asks for one.
|
|
48
|
+
|
|
49
|
+
## Hard boundaries
|
|
50
|
+
|
|
51
|
+
- Never restate a model's syntax from memory; load its generated reference.
|
|
52
|
+
- Never hand-invent a character-sheet prompt when `reference-character.md` already defines the canonical one.
|
|
53
|
+
- Never silently add weather, props, style, or camera movement the user did not request. If you apply a sane default, name it briefly in the notes.
|
|
54
|
+
- Never carry image-model lens, aperture, film-stock, or camera-body syntax into Seedance. Follow the model reference's translation rule.
|
|
55
|
+
- Never second-stamp Seedance shots. Follow its official `Shot 1 / Shot 2 / Shot 3` structure.
|
|
56
|
+
|
|
57
|
+
## Provenance
|
|
58
|
+
|
|
59
|
+
Every `reference-*.md` file in this package is generated from `@slatesvideo/shared`. If a generated reference and this router appear to disagree, the generated reference wins and the router must be corrected at its canonical source.
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
<!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
|
|
2
|
+
|
|
3
|
+
> **This is the real thing.** Every rule below is the working doctrine Slates runs in production against this model — not a summary written for a handout. Slates automates it end to end; the doctrine works by hand too.
|
|
4
|
+
|
|
5
|
+
# Character identity sheet — Slates workflow
|
|
6
|
+
|
|
7
|
+
A character's identity sheet is attached to **every** downstream generation that mentions it, so a flaw in the sheet becomes a flaw in every shot made from it. Building it well is the highest-leverage thing you can do for a project.
|
|
8
|
+
|
|
9
|
+
<!-- @inject:references-read-literally -->
|
|
10
|
+
> **The general law: the model reads a reference literally.**
|
|
11
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
12
|
+
|
|
13
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
14
|
+
|
|
15
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
16
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
17
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
18
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
19
|
+
|
|
20
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
21
|
+
<!-- @end:references-read-literally -->
|
|
22
|
+
|
|
23
|
+
## The shape: ONE sheet, three panels
|
|
24
|
+
|
|
25
|
+
Slates generates **one identity sheet per character**, bound as the character's canonical reference:
|
|
26
|
+
|
|
27
|
+
| Panel | What it carries |
|
|
28
|
+
|---|---|
|
|
29
|
+
| **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
|
|
30
|
+
| **Full-body front, relaxed A-pose — framed from the collarbone down, head not shown** | Build, proportion, wardrobe. Headless on purpose: a front-facing body panel renders a ~40px face that can't match the portrait's, so the sheet would carry two competing identities and the model averages them. |
|
|
31
|
+
| **Full-body back, head and hair visible** | Hair fall and the back of the outfit — the only panel where either reads. Keeps its head because there's no face to compete with. |
|
|
32
|
+
|
|
33
|
+
The rule is **kill every competing rendering of the FACE, not every head** — which is why exactly one body panel is headless.
|
|
34
|
+
|
|
35
|
+
On a deep neutral-grey plate (`#3a3a3c`), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look. Quadrupeds and non-bipedal characters are carved out — natural standing stance, head shown on both body panels.
|
|
36
|
+
|
|
37
|
+
**The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
|
|
38
|
+
|
|
39
|
+
**Why one sheet.** Every `@character` mention attaches that character's canonical identity image, so each character costs one reference slot. It also reduces competing facial renderings to **one** — with the front panel headless and the back panel turned away, the portrait is the only face on the sheet, so there is nothing left to average.
|
|
40
|
+
|
|
41
|
+
## Workflow
|
|
42
|
+
|
|
43
|
+
### Get the reference
|
|
44
|
+
The user has either:
|
|
45
|
+
- Pasted/uploaded an image of the character (real person, drawing, AI render).
|
|
46
|
+
- Described the character in text only.
|
|
47
|
+
|
|
48
|
+
If image: upload it as a reference.
|
|
49
|
+
If text only: generate from prompt-only — less consistent, so warn the user.
|
|
50
|
+
|
|
51
|
+
### Generate the sheet
|
|
52
|
+
|
|
53
|
+
- Default to Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted spend.
|
|
54
|
+
- When the result returns inline, **evaluate it before binding**:
|
|
55
|
+
- Is the portrait clearly the largest panel, and is it off-frontal?
|
|
56
|
+
- **Is the front body panel cleanly headless** — an empty collar with the garment holding its shape, no partial face, no floating jaw, no smeared neck stump? A botched crop is worse than no crop.
|
|
57
|
+
- Do the body panels read as the same build, wardrobe and hair as the portrait?
|
|
58
|
+
- Catchlights present, irises readable rather than black holes?
|
|
59
|
+
- Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
|
|
60
|
+
- Plate a flat deep grey, not white and not black?
|
|
61
|
+
- If off: one focused refinement, then regenerate. The sheet is upstream of everything — it is worth a re-roll that a scene frame is not.
|
|
62
|
+
|
|
63
|
+
## How the reference gets used at scene time
|
|
64
|
+
|
|
65
|
+
Slates cites the sheet inline under the character's name — `{name} (image N)` — in the exact order it sends references. That **name** is the anti-averaging lever, and it is each model's own official mechanism (NB2: "assign a distinct name"; Seedance: `Reference <Subject_N> in <Image_N>`; Kling: reuse a fixed label verbatim).
|
|
66
|
+
|
|
67
|
+
Critically, the app injects **no** wardrobe, expression, or lighting directive. The user's scene prompt owns all of that — which is why `@{name}` dropped into a movie-still injection keeps the still's own clothing and lighting instead of dragging the sheet's.
|
|
68
|
+
|
|
69
|
+
## Anti-patterns
|
|
70
|
+
|
|
71
|
+
- **Don't** studio-light, white-background, or black-background the sheet. White bleeds into the video and washes out the location; black eats edge detail. Flat, even, shadowless light on a deep neutral grey.
|
|
72
|
+
- **Don't** hand-write the sheet prompt when the op will build it — that is how the template and the shipped prompt fork.
|
|
73
|
+
- **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
|
|
74
|
+
- **Don't** skip binding. An unbound asset doesn't help downstream.
|
|
75
|
+
- **Don't** invent character details. Stick to what's in the reference image and the user's description.
|
|
76
|
+
- **Don't** describe the headless front panel as removal or decapitation — in `userNotes` or any hand-written variant. The template asks for it as *framing* — "cropped at the collarbone, head not shown, invisible-mannequin presentation" — which is a standard e-commerce genre with deep training data. Removal phrasing is untested and invites a refusal.
|
|
77
|
+
- **Don't** use 4K — wastes credits, no quality gain at sheet scale.
|
|
78
|
+
- **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `reference-seedance.md`.
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
<!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
|
|
2
|
+
|
|
3
|
+
> **This is the real thing.** Every rule below is the working doctrine Slates runs in production against this model — not a summary written for a handout. Slates automates it end to end; the doctrine works by hand too.
|
|
4
|
+
|
|
5
|
+
# Content-policy-safe construction — read before any risk-surface prompt
|
|
6
|
+
|
|
7
|
+
Don't depict the harm — depict the energy, the aftermath, the threat, or the scale. Build the scene safe from the first word.
|
|
8
|
+
|
|
9
|
+
Write scenes that hit full cinematic impact without ever *needing* to depict prohibited content. This is a craft move, not a compromise — the substitutions below usually read as more cinematic, not less, and they keep your generation from getting silently rejected or degraded by the model's filter. A standoff is more tense than a massacre; an evacuated city is eerier than a crowd in panic; a roar lands harder than a kill. Load this whenever a prompt involves conflict, creatures, crowds, destruction, weapons, or young characters.
|
|
10
|
+
|
|
11
|
+
## Substitution table
|
|
12
|
+
|
|
13
|
+
| Avoid | Use instead |
|
|
14
|
+
|---|---|
|
|
15
|
+
| Civilians in panic, crowds fleeing under debris | An evacuated / empty city; abandoned streets; a lone figure for scale |
|
|
16
|
+
| Weapons firing into buildings or at people | Energy-discharge standoffs, searchlights sweeping, charged auras, shockwaves with no muzzle fire |
|
|
17
|
+
| Creatures tearing into each other, gore | A grapple / standoff — roars, near-misses, circling, an energy clash; combat that stays contained (in/on the water, never lifting into the air) |
|
|
18
|
+
| Destruction with people in harm's way | Destruction in uninhabited terrain — glaciers, deserts, ruins, open sea, evacuated zones |
|
|
19
|
+
| Realistic guns as the focus | Stylized / fantasy implements, weapons slung-not-fired, the weapon as silhouette or prop only |
|
|
20
|
+
| Blood, wounds, death | Impact light, dust, debris, buckling and collapse, a silhouette dropping out of frame |
|
|
21
|
+
| Real, named public figures | Original / anonymous characters |
|
|
22
|
+
| Real brand logos | Original or generalized branding — except the user's own product, which is the whole point of a brand film |
|
|
23
|
+
|
|
24
|
+
## The safe benchmark
|
|
25
|
+
|
|
26
|
+
When a scene starts drifting risky, pull it back toward this shape: **one original creature, in a generalized monument or amphitheatre, in daylight, no weapons present, performing expressive action** (rising, roaring, spreading wings). Original design, generalized location, daylight, expressive rather than violent. That's the confirmed-safe envelope — most epic ideas re-stage into it without losing the punch.
|
|
27
|
+
|
|
28
|
+
## Containment rule — it doubles as a physics win
|
|
29
|
+
|
|
30
|
+
Give any creature or combat scene a **containment rule** that grounds the physics at the same time:
|
|
31
|
+
- "the fight STAYS at the sea surface — they breach, dive, grapple, submerge, but never fly or get carried into the air"
|
|
32
|
+
- "boss scale locked ~2.5 human-heights, NOT kaiju-giant"
|
|
33
|
+
- "destruction stays in the evacuated valley"
|
|
34
|
+
|
|
35
|
+
This improves coherence (the model isn't inventing absurd escalation) AND keeps the scene inside policy — same clause buys both.
|
|
36
|
+
|
|
37
|
+
## Scale and stakes without harm
|
|
38
|
+
|
|
39
|
+
Epic stakes come from environmental danger and reaction, not depicted victims: tiny figures diving clear of *collapsing* terrain (not being crushed), a war-horn over an *empty* field, an army *scrambling* across a frozen valley as a titan tears free of a glacier. The danger is the environment; the figures are reacting, not dying. Snow plumes, glowing runes, splintering ice, shockwaves, and dust carry the chaos.
|
|
40
|
+
|
|
41
|
+
## Minors — hard rule
|
|
42
|
+
|
|
43
|
+
Never write romantic, sexual, or suggestive content involving or directed at minors, and never anything that sexualizes a young-presenting character. Any scene with children stays wholesome and age-appropriate. Non-negotiable — it overrides every stylistic goal.
|
|
44
|
+
|
|
45
|
+
## Pre-flight (run before delivering any risk-surface prompt)
|
|
46
|
+
|
|
47
|
+
- [ ] No civilians depicted in panic/harm; crowds are evacuated or absent.
|
|
48
|
+
- [ ] No weapons firing at people/buildings; threat is energy / searchlight / silhouette.
|
|
49
|
+
- [ ] No creature-on-creature or creature-on-person gore; combat is grapple / standoff / roar, contained.
|
|
50
|
+
- [ ] Destruction is in uninhabited / evacuated terrain.
|
|
51
|
+
- [ ] Creatures are original ("not based on any franchise"); no real public figures; no real brand logos except the user's own product.
|
|
52
|
+
- [ ] Anything with children is wholesome and age-appropriate.
|
|
53
|
+
|
|
54
|
+
If a box fails, apply the substitution table before writing the prompt.
|
|
55
|
+
|
|
56
|
+
## Editing real footage (Kling O3 edit / Omni Flash edit) — real people in the SOURCE
|
|
57
|
+
|
|
58
|
+
Video edit takes the user's own footage, which often contains real people. Rules:
|
|
59
|
+
|
|
60
|
+
- The user must hold rights/consent for any real person's likeness in footage they edit — ask once when it's clearly someone other than the user, then proceed.
|
|
61
|
+
- Kling's video-to-video filter behavior on real faces is **not yet verified** (unlike Seedance, where the consent-gated real-face route is confirmed). If an edit of real-person footage is rejected by the provider, do NOT retry-spam variations — tell the user the filter blocked it and offer a no-face crop/segment or an AI-character swap instead.
|
|
62
|
+
- **Omni Flash: own-footage editing of the uploader's own face PASSED live 2026-07-09** (real talking-head clip, edited on our fal route) despite Google's documented "recognizable people" restriction — treat that restriction as aimed at third-party/public figures, but expect probabilistic refusals and never promise passage.
|
|
63
|
+
- Never use edit to put a real, named public figure into a scene, or to make someone appear to say/do something they didn't. Faceless b-roll (hands, products, landscapes, crowds-from-behind) edits freely.
|
|
64
|
+
|
|
65
|
+
## Gemini / Omni Flash filter regime (video gen + edit) — receipts 2026-07-09
|
|
66
|
+
|
|
67
|
+
Google's filter is its own regime (stricter than fal-hosted Kling about harm-to-a-person, looser than BytePlus about faces). Live receipts:
|
|
68
|
+
|
|
69
|
+
| Blocked (`content_policy_violation`) | Passed |
|
|
70
|
+
|---|---|
|
|
71
|
+
| "his fingertips **ignite** with a small real flame" (fire ON a body part = harm) | "small **magical** flames appear on his fingertips … vanish when he blows on them" |
|
|
72
|
+
|
|
73
|
+
- **Harm-to-person framing is the tripwire**, not the effect itself. Reframe body-contact effects as magical / supernatural / harmless VFX: "magical flames", "a glowing aura", "sparks of light dance on". Avoid ignite / burn / on fire / catch fire applied to a person.
|
|
74
|
+
- **Never use a real object as a metaphor** — "candle-like flame" rendered a literal candle in the subject's hand. Describe the effect, not an object that resembles it.
|
|
75
|
+
- The block is a 422 refund (no credits lost) and arrives mid-generation — one reframe per the substitution mindset above, don't retry-spam.
|
|
@@ -0,0 +1,212 @@
|
|
|
1
|
+
<!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
|
|
2
|
+
|
|
3
|
+
> **This is the real thing.** Every rule below is the working doctrine Slates runs in production against this model — not a summary written for a handout. Slates automates it end to end; the doctrine works by hand too.
|
|
4
|
+
|
|
5
|
+
# Kling V3.0 — prompting
|
|
6
|
+
|
|
7
|
+
Kuaishou's video model. Three tiers: `kling-v3.0-std` (general use, no audio), `kling-v3.0-pro` (higher visual quality, no audio), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).
|
|
8
|
+
|
|
9
|
+
Up to 15s. Multi-shot supported (up to 6 cuts in 15s total). Strong on image-to-video — preserves identity, layout, and text from the input image well.
|
|
10
|
+
|
|
11
|
+
## Subject definition rule (verbatim, fal blog)
|
|
12
|
+
|
|
13
|
+
> "Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots."
|
|
14
|
+
|
|
15
|
+
## Dialogue syntax
|
|
16
|
+
|
|
17
|
+
```
|
|
18
|
+
Character says, "exact words here"
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
Use quotation marks for precise speech. Languages (Omni only): EN, ZH, JA, KO, ES.
|
|
22
|
+
|
|
23
|
+
## Voice direction formula (Omni)
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
Gender + Age Range + Voice Quality + Speech Rate + Emotional Tone + Language
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Example:
|
|
30
|
+
```
|
|
31
|
+
[Character A: Detective, mid-40s, raspy voice, slow cadence, weary]: "I've seen this before."
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
Tone phrases that fire:
|
|
35
|
+
- `speaking in a hushed, trembling whisper`
|
|
36
|
+
- `shouting with commanding authority`
|
|
37
|
+
- `clear, fearful voice`
|
|
38
|
+
- `with a trembling voice, "I'm scared"`
|
|
39
|
+
|
|
40
|
+
## The `Immediately` keyword (Omni only)
|
|
41
|
+
|
|
42
|
+
Without `Immediately`, Kling adds a natural conversational beat between speakers. With it, dialogue is back-to-back. Use when timing matters.
|
|
43
|
+
|
|
44
|
+
```
|
|
45
|
+
[Alice]: "Get down!" Immediately, [Bob]: "Where?"
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
## Speaker label discipline
|
|
49
|
+
|
|
50
|
+
Unique labels per character. **No pronouns or synonyms after first introduction** — they cause voice drift.
|
|
51
|
+
|
|
52
|
+
✅ `[Character A: Black-suited Agent]` ... `[Character A: Black-suited Agent]: "Stop."`
|
|
53
|
+
❌ `[Agent]... then he says...`
|
|
54
|
+
|
|
55
|
+
## Multi-character dialogue (Omni)
|
|
56
|
+
|
|
57
|
+
```
|
|
58
|
+
Alice says in English, "Hello!" Then Bob replies in Spanish, "¡Hola!"
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
## Sound effects, ambient noise, music
|
|
62
|
+
|
|
63
|
+
```
|
|
64
|
+
SFX: thunder cracks, footsteps approaching
|
|
65
|
+
Ambient noise: city traffic, birds chirping, ocean waves
|
|
66
|
+
Background music: tense orchestral strings, low cello
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
SFX accepts physical-cause specificity:
|
|
70
|
+
- ✅ `SFX: heavy boots on wet pavement, distant siren wailing`
|
|
71
|
+
- ❌ `SFX: footsteps`
|
|
72
|
+
|
|
73
|
+
## Image-to-video guidance
|
|
74
|
+
|
|
75
|
+
**Verbatim (fal blog):**
|
|
76
|
+
> "Treat the input image as an anchor. Kling 3.0 excels at preserving the identity, layout, and text details. Focus prompts on how the scene evolves *from* the image: subtle movements, camera motion, or environmental changes."
|
|
77
|
+
|
|
78
|
+
**Don't re-describe what's already in the image.** Focus on motion, changes, evolution.
|
|
79
|
+
|
|
80
|
+
## Multi-shot — what makes them hit
|
|
81
|
+
|
|
82
|
+
**Hard cap: total duration ≤ 15s across all shots. Max 6 cuts.**
|
|
83
|
+
|
|
84
|
+
Hit conditions:
|
|
85
|
+
- Shot labels are explicit: `Shot 1:`, `Shot 2:`
|
|
86
|
+
- One primary action per shot
|
|
87
|
+
- Subject described identically in each shot block
|
|
88
|
+
- Camera move per shot is **one verb**, not a chain
|
|
89
|
+
- Per-shot blocks: 30-60 words
|
|
90
|
+
|
|
91
|
+
Miss conditions:
|
|
92
|
+
- Compressing narrative into one paragraph
|
|
93
|
+
- Pronoun-only references after the first shot
|
|
94
|
+
- Mixing camera moves within a shot ("pan then orbit then push in")
|
|
95
|
+
- Extreme wide → extreme close in adjacent shots without reference images
|
|
96
|
+
|
|
97
|
+
## Element references (Omni)
|
|
98
|
+
|
|
99
|
+
Upload 2-4 multi-angle reference photos per character/object. Tag inline:
|
|
100
|
+
|
|
101
|
+
```
|
|
102
|
+
@element1 is the protagonist (refs: front, side, back angles).
|
|
103
|
+
@element2 is the antagonist.
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
## Reference discipline (character / environment refs)
|
|
107
|
+
|
|
108
|
+
<!-- @inject:references-read-literally -->
|
|
109
|
+
> **The general law: the model reads a reference literally.**
|
|
110
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
111
|
+
|
|
112
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
113
|
+
|
|
114
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
115
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
116
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
117
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
118
|
+
|
|
119
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
120
|
+
<!-- @end:references-read-literally -->
|
|
121
|
+
|
|
122
|
+
<!-- @inject:reference-rules-core -->
|
|
123
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
124
|
+
|
|
125
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
126
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
127
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
128
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
129
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
130
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
131
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
132
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
133
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
134
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
135
|
+
<!-- @end:reference-rules-core -->
|
|
136
|
+
|
|
137
|
+
### For Kling specifically
|
|
138
|
+
|
|
139
|
+
- **Kling's consistency lever is "lock the subject with a fixed label reused verbatim."** That is Kling's phrasing for rules 2 and 3, and it is stricter than the others: **pronoun and synonym drift breaks it**, so the exact same label must appear on every single mention — not "he", not "the detective" after you named him. Reusing the label verbatim is the whole game. Slates composes this for you from `@mentions`.
|
|
140
|
+
- **Element references are the transport for rule 1** — 2-4 multi-angle photos per character/object, tagged `@element1` / `@element2` (see Element references above). The cap is 4 combined refs on the edit path.
|
|
141
|
+
|
|
142
|
+
## Negative prompting — has a real field
|
|
143
|
+
|
|
144
|
+
Kling exposes `negative_prompt` on the fal endpoint (different from Seedance which has none). Default block to start from:
|
|
145
|
+
|
|
146
|
+
```
|
|
147
|
+
blurry, low quality, watermark, text overlay, distorted hands, extra fingers,
|
|
148
|
+
duplicate limbs, unnatural skin texture, overly saturated colors, lens flare,
|
|
149
|
+
floating objects, inconsistent shadows, jittery, flickering, morphing face
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
Layer scene-specific suppressions on top.
|
|
153
|
+
|
|
154
|
+
## Cinematic tactics
|
|
155
|
+
|
|
156
|
+
- **Motion adverb precision** modulates motion energy directly: `slowly`, `rapidly`, `gently`, `explosively`
|
|
157
|
+
- **Camera vocabulary that registers as instructions:** profile shot, tracking, following, freezing, panning, "moving in sync with the subject"
|
|
158
|
+
- **One primary camera move per shot** — never stack
|
|
159
|
+
|
|
160
|
+
## Tier choice
|
|
161
|
+
|
|
162
|
+
- **Standard**: general use, no audio
|
|
163
|
+
- **Pro**: higher visual quality, no audio
|
|
164
|
+
- **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references
|
|
165
|
+
|
|
166
|
+
Pick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — check current numbers before choosing a tier.
|
|
167
|
+
|
|
168
|
+
## Benchmark prompt structure
|
|
169
|
+
|
|
170
|
+
```
|
|
171
|
+
[Character A: <role>, <voice quality>]: "<line>." Immediately, [Character B: <role>, <voice quality>]: "<reply>."
|
|
172
|
+
Ambient noise: <soundscape>.
|
|
173
|
+
Camera <single move>.
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
Cinematic example (paraphrasing fal blog patterns):
|
|
177
|
+
> "Shot 1: Wide establishing shot of a neon-lit alleyway in heavy rain, steam rising from grates. Camera slowly tracks forward.
|
|
178
|
+
> Shot 2: Medium shot of a detective in a trench coat ducking under an awning, water dripping from his hat brim. [Detective: weary, raspy]: 'I knew she'd come back.' Ambient noise: distant traffic, rain on metal.
|
|
179
|
+
> Shot 3: Close-up on his eyes, narrowing as headlights flash across his face."
|
|
180
|
+
|
|
181
|
+
## Video-to-video EDIT — @Video1 / @ElementN / @ImageN
|
|
182
|
+
|
|
183
|
+
Kling O3 edit takes an EXISTING 3-15s clip and changes only what the prompt names — character swap, environment change, style transfer — in one pass, no masking. Original motion, camera, and audio are preserved by default. Its notation is Kling's own, different from the "image N" naming used everywhere else:
|
|
184
|
+
|
|
185
|
+
- **`@Video1`** — the source clip (always; the transport anchors the instruction to it).
|
|
186
|
+
- **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images.
|
|
187
|
+
- **`@Image1..`** — style/appearance references.
|
|
188
|
+
- Max **4 combined** element + image refs per edit.
|
|
189
|
+
|
|
190
|
+
**Prompt shape — the change, not the whole scene:**
|
|
191
|
+
|
|
192
|
+
```
|
|
193
|
+
Replace the man in @Video1 with @Element1, keeping his walk cycle, the camera move, and the rain unchanged.
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
```
|
|
197
|
+
Edit @Video1: turn the daytime street into a neon-lit Tokyo alley at night, wet asphalt reflections. Apply the visual style of @Image1. Keep the subject and camera motion exactly as they are.
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
Rules:
|
|
201
|
+
- Name what CHANGES; explicitly state what stays ("keep the motion / camera / everything else unchanged") — the model preserves better when told to.
|
|
202
|
+
- One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).
|
|
203
|
+
- Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.
|
|
204
|
+
- Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.
|
|
205
|
+
- Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings.
|
|
206
|
+
|
|
207
|
+
## Sources
|
|
208
|
+
|
|
209
|
+
- [fal.ai — Kling 3.0 Prompting Guide](https://blog.fal.ai/kling-3-0-prompting-guide/)
|
|
210
|
+
- [Vidguru — Kling 3.0 Omni Guide](https://www.vidguru.ai/blog/kling-3.0-omni-guide.html)
|
|
211
|
+
- [AcceptPrompt — Kling 3 Prompt Guide](https://www.acceptprompt.com/blog/kling-3-prompt-guide)
|
|
212
|
+
- [DataCamp — Kling 3.0 Tutorial](https://www.datacamp.com/tutorial/kling-3-0)
|
|
@@ -0,0 +1,182 @@
|
|
|
1
|
+
<!-- Generated from the Slates production prompting guides. Do not edit — this file is rebuilt from source. -->
|
|
2
|
+
|
|
3
|
+
> **This is the real thing.** Every rule below is the working doctrine Slates runs in production against this model — not a summary written for a handout. Slates automates it end to end; the doctrine works by hand too.
|
|
4
|
+
|
|
5
|
+
# Nano Banana 2 — cinematic & photorealistic prompting
|
|
6
|
+
|
|
7
|
+
Nano Banana 2 is **Gemini 3.1 Flash Image**. It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat. NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
|
|
8
|
+
|
|
9
|
+
Knowledge cutoff: January 2025. Anything after needs explicit reference images.
|
|
10
|
+
|
|
11
|
+
## Google's 4 official rules (verbatim)
|
|
12
|
+
|
|
13
|
+
1. **Be specific.** Provide concrete details on subject, lighting, and composition.
|
|
14
|
+
2. **Use positive framing.** Describe what you want, not what you don't want.
|
|
15
|
+
3. **Control the camera.** Use photographic and cinematic terms like "low angle" and "aerial view."
|
|
16
|
+
4. **Iterate.** Refine images with follow-up prompts in a conversational manner.
|
|
17
|
+
|
|
18
|
+
## Official prompt formula
|
|
19
|
+
|
|
20
|
+
```
|
|
21
|
+
[Subject] + [Action] + [Location/context] + [Composition] + [Style]
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
For the cinematic / photoreal use case, expand to:
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and action]. [3-5 specific visual details]. [LIGHTING — direction + quality]. [COLOR PALETTE]. [FILM STOCK or sensor language]. [1-2 word emotional tone].
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
## Photorealism positives — what consistently works
|
|
31
|
+
|
|
32
|
+
> ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern — do not carry it across.**
|
|
33
|
+
> Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.
|
|
34
|
+
> The leak happens in one specific way — you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `reference-seedance.md` (Part 3, "Don't cross-pollinate image-model syntax").
|
|
35
|
+
|
|
36
|
+
**Named lenses + apertures** beat generic "shallow depth of field":
|
|
37
|
+
- `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`
|
|
38
|
+
- `Panavision anamorphic` for horizontal flares + cinematic width
|
|
39
|
+
- `400mm telephoto` for compression + isolation
|
|
40
|
+
- `24mm` for environmental interiors
|
|
41
|
+
|
|
42
|
+
**Named cameras / sensors:**
|
|
43
|
+
- `ARRI Alexa 65`, `Hasselblad X2D`, `Canon EOS R5`, `Sony A7III`, `Fujifilm X-T5`
|
|
44
|
+
- "Specific gear" beats "DSLR"
|
|
45
|
+
|
|
46
|
+
**Named film stocks** (one per prompt — never mix):
|
|
47
|
+
- `Kodak Portra 400` — natural skin, warm
|
|
48
|
+
- `Fuji Velvia 50` — saturated, landscape
|
|
49
|
+
- `Ilford HP5 Plus` — black and white, gritty grain
|
|
50
|
+
- `CineStill 800T` — tungsten night, halation
|
|
51
|
+
|
|
52
|
+
**Physics-based lighting** (direction + quality):
|
|
53
|
+
- `Single key light at 45 degrees from upper left`
|
|
54
|
+
- `Late afternoon sun at 15 degrees above horizon`
|
|
55
|
+
- `Color temperature 4500K` beats `slightly warm`
|
|
56
|
+
- `Practicals only — no fill` for Deakins-style realism
|
|
57
|
+
|
|
58
|
+
**Imperfection vocabulary** (forces away from AI-clean):
|
|
59
|
+
- `visible pores`, `natural skin grain`, `peach fuzz`, `slight hyperpigmentation`
|
|
60
|
+
- `unretouched raw photography`, `ISO noise`, `sweat beading`
|
|
61
|
+
- `crisp catchlights in the eyes`, `skin micro-detail`
|
|
62
|
+
|
|
63
|
+
**Director references** (use when locking style):
|
|
64
|
+
| Director | Tone | Visual signature |
|
|
65
|
+
|---|---|---|
|
|
66
|
+
| Denis Villeneuve | Cold, vast, existential | Desaturated, overwhelming scale |
|
|
67
|
+
| Roger Deakins | Precise motivated light | Single source, deep shadows, practicals |
|
|
68
|
+
| Emmanuel Lubezki | Natural, spiritual | Available light, golden hour |
|
|
69
|
+
| Bradford Young | Warm darkness | Underexposed, rich shadows, skin tones |
|
|
70
|
+
|
|
71
|
+
**Genre cues that move the model:**
|
|
72
|
+
- `unstaged documentary photography style`
|
|
73
|
+
- `fashion magazine editorial, shot on medium-format analog film, pronounced grain`
|
|
74
|
+
- `Film still from [Director] [genre]`
|
|
75
|
+
|
|
76
|
+
## The anti-list — phrases that DEGRADE realism
|
|
77
|
+
|
|
78
|
+
These are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.
|
|
79
|
+
|
|
80
|
+
**Never use:**
|
|
81
|
+
- `8k`, `4k` (as a quality token)
|
|
82
|
+
- `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone
|
|
83
|
+
- `masterpiece`, `best quality`, `highly detailed`, `ultra-detailed`
|
|
84
|
+
- `trending on ArtStation`, `award-winning`
|
|
85
|
+
- `perfect skin`, `flawless`, `airbrushed`, `smooth skin`
|
|
86
|
+
- `cinematic` standing alone — always specify *which cinema* (director, lens, era, stock)
|
|
87
|
+
- `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue
|
|
88
|
+
|
|
89
|
+
## Negative prompting — there is no field
|
|
90
|
+
|
|
91
|
+
Nano Banana 2 has **no `negativePrompt` parameter**. Three patterns to suppress unwanted content:
|
|
92
|
+
|
|
93
|
+
1. **Positive reframing (preferred):** "empty street" not "no cars". "Unstaged documentary photography" not "not anime."
|
|
94
|
+
2. **Inline `without` / `free of`:** "without any people, vehicles, or man-made structures", "free of text overlays, logos, or watermarks."
|
|
95
|
+
3. **Constraint clauses for anatomy/quality:** "accurate anatomy with five fingers per hand, symmetrical features, natural proportions"; "sharp, well-exposed, free of blur or JPEG artifacts."
|
|
96
|
+
|
|
97
|
+
Default to #1. Reach for #2 only when positive framing can't suppress the unwanted element.
|
|
98
|
+
|
|
99
|
+
## Reference images
|
|
100
|
+
|
|
101
|
+
- **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.
|
|
102
|
+
- **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style, Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
103
|
+
|
|
104
|
+
### Reference rules (the verified ones)
|
|
105
|
+
|
|
106
|
+
<!-- @inject:references-read-literally -->
|
|
107
|
+
> **The general law: the model reads a reference literally.**
|
|
108
|
+
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
109
|
+
|
|
110
|
+
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
111
|
+
|
|
112
|
+
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
113
|
+
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
114
|
+
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
115
|
+
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
116
|
+
|
|
117
|
+
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
118
|
+
<!-- @end:references-read-literally -->
|
|
119
|
+
|
|
120
|
+
<!-- @inject:reference-rules-core -->
|
|
121
|
+
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
122
|
+
|
|
123
|
+
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
124
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
|
|
125
|
+
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
126
|
+
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
127
|
+
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
128
|
+
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
129
|
+
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
130
|
+
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
131
|
+
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
132
|
+
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
133
|
+
<!-- @end:reference-rules-core -->
|
|
134
|
+
|
|
135
|
+
### For Nano Banana 2 specifically
|
|
136
|
+
|
|
137
|
+
- **NB2's own consistency lever is "assign a distinct name to each character/object."** That is Google's phrasing for rule 3 — cite each canonical identity inline by name.
|
|
138
|
+
- **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.
|
|
139
|
+
- **Character consistency is officially "not 100% perfect"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.
|
|
140
|
+
- **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.
|
|
141
|
+
|
|
142
|
+
## Common failure modes + fixes
|
|
143
|
+
|
|
144
|
+
**Hands:** Append `accurate anatomy with five fingers per hand, symmetrical features, natural proportions, relaxed open palm`. Avoid heavy jewelry, props intersecting fingers, motion blur in references.
|
|
145
|
+
|
|
146
|
+
**Text in images:** Quote-wrap target text. Specify font (`Century Gothic, 12pt`). Long phrases work; small text degrades. Two-step works best — generate text concepts conversationally first, then ask for the image.
|
|
147
|
+
|
|
148
|
+
**Left/right confusion:** Default is **viewer's perspective**, not subject's. Append `left and right are from the character's perspective, NOT the camera's` when scene-blocking matters.
|
|
149
|
+
|
|
150
|
+
**Surreal / absurd prompts trip uncanny valley:** The model drags toward realism. If you want surrealism, lean hard into stylization keywords (`painted`, `illustrated`, `stop-motion`).
|
|
151
|
+
|
|
152
|
+
**Soft faces / dead eyes:** Add `crisp catchlights in the eyes`, `skin micro-detail`, `peach fuzz visible`. Don't stack quality enhancers — single clean prompt beats multiple re-interpretations.
|
|
153
|
+
|
|
154
|
+
**Post-cutoff content (anything after Jan 2025):** Use reference images. The model has no knowledge of recent franchises, products, events.
|
|
155
|
+
|
|
156
|
+
## Resolution tactics
|
|
157
|
+
|
|
158
|
+
- Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — check current numbers. Pick the cheapest resolution that serves the use case.
|
|
159
|
+
- **At 2K and above, the model allocates more tokens to surface detail** — explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.
|
|
160
|
+
- 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.
|
|
161
|
+
- 2K generations vary 20-60s+. Don't time-budget tightly.
|
|
162
|
+
|
|
163
|
+
## Boring vs cinema — examples
|
|
164
|
+
|
|
165
|
+
❌ **Boring:** "Wide shot of a man on a dock looking at the forest."
|
|
166
|
+
|
|
167
|
+
✅ **Cinema:** "Direct overhead drone shot on weathered dock surface. Single figure standing center frame, climbing up from frame bottom. Boot prints leading away from him toward shore. Pale winter light. Anamorphic lens flare from low sun. Desaturated blue and slate grey palette. Kodak Portra 400 grain. The path already walked by someone else. Map of threat."
|
|
168
|
+
|
|
169
|
+
❌ **Boring:** "Close up of a woman looking scared."
|
|
170
|
+
|
|
171
|
+
✅ **Cinema:** "Extreme close on subject's mouth and nose, 135mm f/2.8, shallow depth of field. Breath pluming out, catching cold light from upper-left key. Lips slightly parted, peach fuzz visible. The breath holds. CineStill 800T halation around catchlights. Waiting."
|
|
172
|
+
|
|
173
|
+
## The 3-strike rule
|
|
174
|
+
|
|
175
|
+
If three iterations on the same prompt haven't produced what the user wants, stop. Hand back to the user with what you tried and what isn't working. The slot machine doesn't converge — the prompt structure is wrong, not the seed.
|
|
176
|
+
|
|
177
|
+
## Family variants — Lite and Pro
|
|
178
|
+
|
|
179
|
+
Everything in this skill applies to the whole Nano Banana family; two variants trade speed/ceiling around NB2 full:
|
|
180
|
+
|
|
181
|
+
- **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.
|
|
182
|
+
- **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.
|