@slatesvideo/shared 0.5.4 → 0.5.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/README.md +13 -0
  2. package/dist/index.d.ts +1 -1
  3. package/dist/index.js +2 -2
  4. package/dist/operations/index.d.ts +84 -9
  5. package/dist/operations/index.js +455 -47
  6. package/dist/prompts/character-sheet.d.ts +11 -21
  7. package/dist/prompts/character-sheet.js +124 -59
  8. package/dist/prompts/model-facts.d.ts +1 -1
  9. package/dist/prompts/model-facts.js +32 -0
  10. package/dist/prompts/partials.generated.js +2 -2
  11. package/dist/prompts/prompting-tips.d.ts +1 -1
  12. package/dist/prompts/prompting-tips.js +196 -2
  13. package/dist/prompts/reference-composer.d.ts +1 -1
  14. package/dist/prompts/reference-composer.js +3 -4
  15. package/dist/prompts/reference-rules.d.ts +19 -2
  16. package/dist/prompts/reference-rules.js +21 -4
  17. package/dist/skills/content.js +14 -11
  18. package/exports/slates-prompt-builder/generated/SKILL.md +59 -0
  19. package/{skills/slates-character-turnaround.md → exports/slates-prompt-builder/generated/reference-character.md} +26 -33
  20. package/exports/slates-prompt-builder/generated/reference-content-policy.md +75 -0
  21. package/exports/slates-prompt-builder/generated/reference-kling.md +212 -0
  22. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +182 -0
  23. package/exports/slates-prompt-builder/generated/reference-seedance.md +353 -0
  24. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +79 -0
  25. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  26. package/package.json +7 -3
  27. package/skills/_partials/reference-rules-core.md +1 -1
  28. package/skills/_partials/reference-tips-short.md +1 -1
  29. package/skills/slates-character-identity.md +105 -0
  30. package/skills/slates-edit-and-iterate.md +1 -1
  31. package/skills/slates-model-selection.md +26 -0
  32. package/skills/slates-one-prompt-film.md +3 -3
  33. package/skills/slates-prompting-elevenlabs.md +131 -0
  34. package/skills/slates-prompting-flux-2-max.md +1 -1
  35. package/skills/slates-prompting-gpt-image-2.md +1 -1
  36. package/skills/slates-prompting-kling-v3.md +8 -6
  37. package/skills/slates-prompting-nano-banana-2.md +7 -5
  38. package/skills/slates-prompting-omni-flash.md +1 -1
  39. package/skills/slates-prompting-seed-audio.md +110 -0
  40. package/skills/slates-prompting-seedance.md +15 -9
  41. package/skills/slates-prompting-suno.md +110 -0
  42. package/skills/slates-prompting-veo-3.md +1 -1
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@slatesvideo/shared",
3
- "version": "0.5.4",
3
+ "version": "0.5.6",
4
4
  "description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -24,12 +24,15 @@
24
24
  "dist",
25
25
  "!dist/**/*.map",
26
26
  "skills",
27
+ "exports/slates-prompt-builder/generated",
27
28
  "README.md"
28
29
  ],
29
30
  "scripts": {
30
31
  "sync-partials": "node scripts/sync-partials.mjs",
31
- "build": "node scripts/sync-partials.mjs --check && node scripts/embed-skills.mjs && tsc",
32
- "typecheck": "node scripts/sync-partials.mjs --check && node scripts/embed-skills.mjs && tsc --noEmit",
32
+ "build-prompt-builder": "node scripts/build-prompt-builder.mjs",
33
+ "check-prompt-builder": "node scripts/build-prompt-builder.mjs --check",
34
+ "build": "node scripts/sync-partials.mjs --check && node scripts/build-prompt-builder.mjs --check && node scripts/embed-skills.mjs && tsc",
35
+ "typecheck": "node scripts/sync-partials.mjs --check && node scripts/build-prompt-builder.mjs --check && node scripts/embed-skills.mjs && tsc --noEmit",
33
36
  "prepublishOnly": "npm run build"
34
37
  },
35
38
  "repository": {
@@ -60,6 +63,7 @@
60
63
  "zod": "^3.23.0"
61
64
  },
62
65
  "devDependencies": {
66
+ "fflate": "^0.8.3",
63
67
  "typescript": "^5.7.0"
64
68
  }
65
69
  }
@@ -2,7 +2,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
2
2
 
3
3
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
4
4
  2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
5
- 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
5
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
6
6
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
7
7
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
8
8
  6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
@@ -1,2 +1,2 @@
1
1
  <!-- consumer:ts -->
2
- Name each reference inline; never write role essays. Slates does this for you: `@mention` a subject or environment and it composes `Marcus (images 1 and 2) in the cafe (image 3)`, citing them in the exact order it sends them. Citing both of a character's sheets under the SAME name is what tells the model they are one person — a "Reference Image Instructions" block does the opposite and drags the sheet's studio lighting into your scene. Start with 2-3 focused refs.
2
+ Name each reference inline; never write role essays. Slates does this for you: `@mention` a subject or environment and it composes `Marcus (image 1) in the cafe (image 2)`, citing them in the exact order it sends them. One canonical identity image avoids competing facial renderings; a "Reference Image Instructions" block drags reference lighting into your scene. Start with 2-3 focused refs.
@@ -0,0 +1,105 @@
1
+ ---
2
+ name: slates-character-identity
3
+ description: Build a Slates character from a reference image — generate one identity sheet and bind it to the character so the card updates live. Use when the user wants to create a character, build a character from an image, or starts a storyboard flow that needs consistent character references.
4
+ ---
5
+
6
+ # Character identity sheet — Slates workflow
7
+
8
+ A character's identity sheet is attached to **every** downstream generation that mentions it, so a flaw in the sheet becomes a flaw in every shot made from it. Building it well is the highest-leverage thing you can do for a project.
9
+
10
+ <!-- @inject:references-read-literally -->
11
+ > **The general law: the model reads a reference literally.**
12
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
13
+
14
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
15
+
16
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
17
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
18
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
19
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
20
+
21
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
22
+ <!-- @end:references-read-literally -->
23
+
24
+ ## The shape: ONE sheet, three panels
25
+
26
+ Slates generates **one identity sheet per character**, bound as the character's canonical reference:
27
+
28
+ | Panel | What it carries |
29
+ |---|---|
30
+ | **Chest-up portrait, three-quarter angle, largest panel (~25–30% of the sheet)** | The face. **This is the only place the model reads facial identity from** — every detail it will ever know comes from those pixels, so it gets the resolution. Off-frontal, never dead-on: an angled head reads its volume instantly. |
31
+ | **Full-body front, relaxed A-pose — cropped at the collarbone, just the face cropped out** | Build, proportion, wardrobe. The face is cropped off on purpose: a front-facing body panel renders a ~40px face that can't match the portrait's, so the sheet would carry two competing identities and the model averages them. **Only the face** — neck, arms and hands render as skin. |
32
+ | **Full-body back, head and hair visible** | Hair fall and the back of the outfit — the only panel where either reads. Keeps its head because there's no face to compete with. |
33
+
34
+ The rule is **kill every competing rendering of the FACE, not every head** — which is why exactly one body panel is headless.
35
+
36
+ On a deep neutral-grey plate (hex `3a3a3c`, emitted without the `#` — see the sigil warning in Don'ts), flat and shadowless, with catchlights in the eyes, irises never crushed to black, surface texture at the medium's own natural level of detail, broken symmetry, and no over-clean 3D-game-model look. Expression is **a slight natural smile with the teeth just visible** — a closed mouth carries no dental information, so every downstream smiling shot invents teeth, and teeth are person-specific.
37
+
38
+ **Two carve-outs, scoped differently on purpose.** Non-human characters get a natural neutral expression instead of a smile — that one is scoped by *having a human mouth*, so a bipedal robot or humanoid alien is covered. Quadrupeds and non-bipedal characters get a natural standing stance with the head shown on both body panels — that one is *anatomical*. **Both are conditionals the image model evaluates against your reference; neither is a code branch, because the op has no character-kind input.**
39
+
40
+ **The sheet inherits the source's medium** — photo, anime, illustration, painterly, 3D render — unless the user explicitly asks for a transform. None of the craft clauses above override that: they ask for *readable* eyes and *material-looking* surfaces within whatever medium the character is in, not for photorealism.
41
+
42
+ **Why one sheet.** Every `@character` mention attaches that character's canonical identity image, so each character costs one reference slot. It also reduces competing facial renderings to **one** — with the front panel headless and the back panel turned away, the portrait is the only face on the sheet, so there is nothing left to average.
43
+
44
+ ## Workflow
45
+
46
+ ### Get the reference
47
+ The user has either:
48
+ - Pasted/uploaded an image of the character (real person, drawing, AI render).
49
+ - Described the character in text only.
50
+
51
+ If image: upload it as a reference<!-- slates-only --> (`slates_upload_reference_image`)<!-- /slates-only -->.
52
+ If text only: generate from prompt-only — less consistent, so warn the user.
53
+
54
+ <!-- slates-only -->
55
+ ### Create the character record
56
+ `slates_create_character` with:
57
+ - `name` (ask if not given)
58
+ - `description` — 1-2 sentences, *visual* only ("tall, dark hair, scar over left eye"), not personality.
59
+ - `style` — leave as the source's own medium by default. Only name a transform if the user wants one (e.g. anime → realistic).
60
+ <!-- /slates-only -->
61
+
62
+ ### Generate the sheet
63
+ <!-- slates-only -->
64
+ `slates_generate_character_identity` with `characterId`, `projectId`, and `baseAssetId` (the source portrait).
65
+
66
+ **Do not hand-write the sheet prompt.** Slates builds it from the canonical template in `@slatesvideo/shared/prompts` (`buildCharacterIdentityPrompt`) — panels, plate, lighting and craft clauses included — and appends your `userNotes`. Use `userNotes` for what the template can't know: *"use the woman on the left"*, *"keep the scar on the right cheek"*. A hand-written prompt is a fork of the template and will drift from it.
67
+
68
+ - Estimate cost first with `slates_estimate_generation_cost` and announce in **credits** — never quote a price from memory.
69
+ <!-- /slates-only -->
70
+
71
+ - Default to Nano Banana 2 at 2K. **Never 4K** — no identity gain at sheet scale, wasted spend.
72
+ - When the result returns inline, **evaluate it before binding**:
73
+ - Is the portrait clearly the largest panel, and is it off-frontal?
74
+ - **Is the front body panel cleanly headless** — an empty collar above a normally rendered body, no partial face, no floating jaw, no smeared neck stump? A botched crop is worse than no crop.
75
+ - **Is the body still there?** Neck, forearms and hands rendered as skin, not an empty outfit floating on nothing. A hollow garment means the invisible-mannequin genre ran unbounded.
76
+ - Do the body panels read as the same build, wardrobe and hair as the portrait?
77
+ - Catchlights present, irises readable rather than black holes?
78
+ - Is it in the source's medium, and does it read as *that* medium done well — or has it drifted toward the over-clean game-model look?
79
+ - Plate a flat deep grey, not white and not black?
80
+ - If off: one focused refinement, then regenerate. The sheet is upstream of everything — it is worth a re-roll that a scene frame is not.
81
+ <!-- slates-only -->
82
+ - The op binds the result as the canonical identity automatically.
83
+
84
+ ### Hand back
85
+ > "Character {name} ready — identity sheet bound. Use `@{name}` in any prompt and Slates attaches it and names it inline, so the face stays consistent."
86
+ <!-- /slates-only -->
87
+
88
+ ## How the reference gets used at scene time
89
+
90
+ Slates cites the sheet inline under the character's name — `{name} (image N)` — in the exact order it sends references. That **name** is the anti-averaging lever, and it is each model's own official mechanism (NB2: "assign a distinct name"; Seedance: `Reference <Subject_N> in <Image_N>`; Kling: reuse a fixed label verbatim).
91
+
92
+ Critically, the app injects **no** wardrobe, expression, or lighting directive. The user's scene prompt owns all of that — which is why `@{name}` dropped into a movie-still injection keeps the still's own clothing and lighting instead of dragging the sheet's.
93
+
94
+ ## Anti-patterns
95
+
96
+ - **Don't** studio-light, white-background, or black-background the sheet. White bleeds into the video and washes out the location; black eats edge detail. Flat, even, shadowless light on a deep neutral grey.
97
+ - **Don't** hand-write the sheet prompt when the op will build it — that is how the template and the shipped prompt fork.
98
+ - **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
99
+ - **Don't** skip binding. An unbound asset doesn't help downstream.
100
+ - **Don't** invent character details. Stick to what's in the reference image and the user's description.
101
+ - **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on gpt-image-2** — fal returns `content_policy_violation` with `loc: ["body","prompt"]`, so the text is rejected before any image is read, because an anatomical absence reads as gore to OpenAI's classifier. It passed NB2, which is why the original receipt looked safe: **it was model-scoped.** State an exclusion as a framing choice, never as a missing body part.
102
+ - **Don't** invoke the invisible-mannequin genre without bounding it to the face. **"an invisible-mannequin presentation where the clothing holds its own shape" removed all the skin** — no neck, no hands, no forearms, a garment floating on nothing — because that *is* the e-commerce genre in full: an empty outfit. **"with just the face cropped out"** keeps the anchor and bounds it. Generalises: a genre anchor imports the whole genre, so name what STAYS, not only what goes.
103
+ - **Don't** put `#` or `@` anywhere in prompt text. Both are reference-token sigils in the desktop prompt composer and an unresolved one is **silently deleted** — no error, no log, just missing words. `#3a3a3c` reached fal as `background ()` on a real 2026-07-30 request, meaning the plate value had never been delivered to any model since the composer shipped. Write hex values bare.
104
+ - **Don't** use 4K — wastes credits, no quality gain at sheet scale.
105
+ - **Don't** feed a multi-view sheet into a Seedance shot that has **several characters in frame** without binding each character to its image and appending the anti-twin constraint — ByteDance documents multi-view assets as a cause of duplicate characters. See `slates-prompting-seedance`.
@@ -47,7 +47,7 @@ The user's request is one of:
47
47
  ### 4. Generate, evaluate, decide
48
48
  - Estimate cost first.
49
49
  - After generation, the result is inline. Compare side-by-side with the original (`slates_get_asset_image` again).
50
- - If the delta is correct: bind to the same slot (frame, character turnaround, etc.) the original was bound to.
50
+ - If the delta is correct: bind to the same role (frame, character identity, etc.) the original was bound to.
51
51
  - If the delta missed: one focused refinement, then regenerate. Cap at 3 tries.
52
52
 
53
53
  ### 5. Hand back
@@ -91,6 +91,32 @@ Both tools have a cheap Kling utility lane and a premium Seedance lane. The capa
91
91
 
92
92
  **Split rule of thumb:** readable text / panels / UI → GPT Image 2; photoreal, character-locked, widescreen, or edit-heavy → the Banana line; drafts → NB2 Lite; uncensored or odd resolutions → Seedream/FLUX.
93
93
 
94
+ ## Audio routing
95
+
96
+ **Image and video models cannot generate standalone audio, and none of the four audio models can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.
97
+
98
+ | Job | Model | Why |
99
+ |---|---|---|
100
+ | **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. 1–120s. The continuity-bed workhorse. |
101
+ | **The exact words, in a repeatable named voice** — ad reads, narration, character lines to lip-sync against | **Eleven v3** (`eleven-v3`) | Verbatim text, 20 preset voices, re-renderable after a copy tweak without the performance drifting. Billed per 100 characters. |
102
+ | **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control (0.5–22s) and a real loop mode. |
103
+ | **A song or a score** | **Suno** (`suno`) | Full music with structure. Two variations per call for one flat price, and duration is free to 360s. |
104
+
105
+ ### Named audio escalation triggers
106
+
107
+ - **"It needs to sound like a place"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.
108
+ - **"Read this line"** with copy that a client can still change → Eleven v3. Scratch dialogue while the script is moving can stay on Seed Audio.
109
+ - **"That needs a thump right there"** → Sound Effects, with the duration set to roughly the length of the event.
110
+ - **"Give it a track"** → Suno, `instrumental: true` unless a vocal is genuinely wanted (an unasked-for vocal fights dialogue).
111
+
112
+ **Rules:**
113
+
114
+ - **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.
115
+ - **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.
116
+ - **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — on Suno the extra seconds are literally free.
117
+ - **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.
118
+ - Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`, `slates-prompting-suno`.
119
+
94
120
  ## Cost is a tiebreaker, not the router
95
121
 
96
122
  Route by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.
@@ -33,7 +33,7 @@ A 4-10 shot script is where you invent the most on the user's behalf — time of
33
33
 
34
34
  ### 2. Set up the project
35
35
  - `slates_create_project` named for the piece.
36
- - Recurring character? Build it properly — `slates_create_character` + the `slates-character-turnaround` recipe — so every frame references the same turnaround.
36
+ - Recurring character? Build it properly — `slates_create_character` + the `slates-character-identity` recipe — so every frame references the same identity.
37
37
  - Recurring location? `slates_create_environment`.
38
38
  - One-off shots don't need character/environment records; skip the ceremony.
39
39
 
@@ -49,7 +49,7 @@ Price the whole batch before the first generation: frame images (count × model
49
49
  Per `slates-cost-discipline` 3b: that single OK authorizes `confirm=true` for **every enumerated call in the batch** — no per-call re-asking. Re-confirm only if a call's price overruns the plan >25% or new calls get added (extra retakes, new shots).
50
50
 
51
51
  ### 5. Generate frame images
52
- Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the character turnaround / environment / prior frames for consistency (Slates names each reference inline as "image N" — you don't hand-write role labels; reuse the same subject name across shots). Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`.
52
+ Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the character identity / environment / prior frames for consistency (Slates names each reference inline as "image N" — you don't hand-write role labels; reuse the same subject name across shots). Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`.
53
53
 
54
54
  **Multi-take where it matters:** for the hook shot and any shot the whole film hangs on, generate 2-4 variants (cheap model or 1k), pull them back with `slates_get_assets_batch`, pick the strongest on composition + identity, discard the rest. Don't multi-take filler shots.
55
55
 
@@ -83,4 +83,4 @@ Shots delivered, total spent vs. approved plan, the export path, and the single
83
83
  - **Skeleton before spend.** Project + storyboard structure are free; generation isn't.
84
84
  - **Look at everything.** Every image inline, every video via `slates_get_asset_video_frames` if a clip seems off. Never assemble a timeline from clips you haven't evaluated.
85
85
  - **3-strike rule per shot.** Three failed takes on one shot = stop, show the user what you tried, ask.
86
- - **Consistency comes from references, not luck.** Same turnaround asset on every character frame; same environment refs across a location's shots.
86
+ - **Consistency comes from references, not luck.** Same identity asset on every character frame; same environment refs across a location's shots.
@@ -0,0 +1,131 @@
1
+ ---
2
+ name: slates-prompting-elevenlabs
3
+ description: How to prompt the two ElevenLabs surfaces in Slates. Read before calling slates_generate_audio with model eleven-v3 (Eleven v3 text-to-speech - controlled, repeatable, named-voice voiceover, billed per 100 characters) or eleven-sfx (Sound Effects v2 - one short effect with an EXACT duration, 0.5-22s, billed per second). Covers the "the text field is spoken verbatim" rule, punctuation as the only timing control, stability, describing an effect by its physical cause, and when to use Seed Audio instead.
4
+ ---
5
+
6
+ # ElevenLabs — prompting (Eleven v3 TTS + Sound Effects v2)
7
+
8
+ Two separate surfaces from the same vendor, carried on fal (`fal-ai/elevenlabs/tts/eleven-v3`, `fal-ai/elevenlabs/sound-effects/v2`). They share nothing but a bill — treat them as different tools.
9
+
10
+ ## Where they route
11
+
12
+ - **`eleven-v3`** — the exact words matter and the read must be **repeatable**: ad reads, narration, character lines you will lip-sync against, anything a client will ask you to re-render after a copy tweak. Billed per 100 characters of text, rounded up.
13
+ - **`eleven-sfx`** — a single sound that has to land on a known frame, or a seamless loop. Billed per second, 0.5–22s.
14
+ - **Neither** for layered scenes. A room with dialogue *and* clatter *and* ambience is one `seed-audio` pass, not three ElevenLabs generations.
15
+ - **AUDIO-ONLY.** Neither can produce images or video.
16
+
17
+ ---
18
+
19
+ ## Eleven v3 (`eleven-v3`) — THE RULES
20
+
21
+ ### 1. 🚨 The text field is the script. Every character is spoken.
22
+
23
+ ```
24
+ ✗ (excited) Read this fast — "Grab yours today!"
25
+ ✓ Grab yours today!
26
+ ```
27
+
28
+ Stage directions, speaker names, bracketed emotion tags and markdown all get read out loud. There is no instruction channel — direction lives in `stability` and in how you punctuate.
29
+
30
+ ### 2. Punctuation is the only timing control
31
+
32
+ | You want | Write |
33
+ |---|---|
34
+ | a hard stop | `It works. Every time.` |
35
+ | a beat, not a stop | `It works — every time.` |
36
+ | a trailing hesitation | `It works… mostly.` |
37
+ | a list rhythm | `Faster, cheaper, and yours.` |
38
+
39
+ Rewrite the punctuation before you touch a setting. It moves the read more than `stability` does.
40
+
41
+ ### 3. Pick a voice and keep it
42
+
43
+ 20 presets: Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill. Default `Rachel`.
44
+
45
+ One voice per character, one per piece. A series that swaps voices between shots reads as an accident. There is **no voice cloning on this route** — a cloned ElevenLabs voice ID is not supported and must not be assumed to pass through.
46
+
47
+ ### 4. Stability
48
+
49
+ | Value | Behavior | Use for |
50
+ |---|---|---|
51
+ | ~0.3 | expressive, varies take-to-take | one dramatic line, character dialogue |
52
+ | 0.5 (default) | balanced | most reads |
53
+ | ~0.8 | flat, highly repeatable | long narration, anything you will re-render |
54
+
55
+ Raise it when re-rolls keep giving you a different performance. Lower it when the read is lifeless.
56
+
57
+ ### 5. Spell out what TTS gets wrong
58
+
59
+ Acronyms, product names, prices, years and URLs are where it embarrasses itself. Write the pronunciation:
60
+
61
+ ```
62
+ SKU → "ess kay you"
63
+ 2026 → "twenty twenty six"
64
+ $19.99 → "nineteen ninety nine"
65
+ slates.video → "slates dot video"
66
+ ```
67
+
68
+ Set `languageCode` (ISO 639-1) to force a language when the text is ambiguous or code-switched.
69
+
70
+ ### 6. Length is money, honestly
71
+
72
+ 1–5000 characters, billed in 100-character buckets rounded up. A tightened sentence costs less; a pasted stray paragraph costs more. This is the one Slates surface where editing the copy is also a cost control.
73
+
74
+ ### 7. Timestamps are free — leave them on
75
+
76
+ Word-level timestamps come back with every generation at no extra charge. They are exactly what a caption/subtitle pass consumes. There is no reason to disable them.
77
+
78
+ ---
79
+
80
+ ## Sound Effects v2 (`eleven-sfx`) — THE RULES
81
+
82
+ ### 1. Describe the physical CAUSE, not the label
83
+
84
+ ```
85
+ ✗ door sound
86
+ ✓ heavy oak door slams shut in a stone hallway
87
+
88
+ ✗ whoosh
89
+ ✓ a thick rope swung fast past a microphone, low air displacement
90
+
91
+ ✗ footsteps
92
+ ✓ boots on wet gravel, slow, one person
93
+ ```
94
+
95
+ Material + weight + surface + room. Naming all four is the difference between a usable effect and a stock-library shrug. Cap is 450 characters — you will not need them.
96
+
97
+ ### 2. One sound per generation
98
+
99
+ This surface makes a single event. A door, then footsteps, then a siren is three generations layered on the timeline — or one `seed-audio` scene, which is usually cheaper and always more coherent.
100
+
101
+ ### 3. Duration is always explicit, and it is the price
102
+
103
+ `durationSeconds` is 0.5–22 and Slates **always sends it**. (Left null the model picks, which makes the charge non-deterministic — so it is never left null.)
104
+
105
+ | Kind of sound | Ask for |
106
+ |---|---|
107
+ | impact, hit, click | 0.5–1s |
108
+ | whoosh, riser, transition | 2–4s |
109
+ | loopable bed | 8–22s + `loop: true` |
110
+
111
+ Over-asking pads the tail with room tone you then trim. Under-asking clips the decay.
112
+
113
+ ### 4. Loops
114
+
115
+ `loop: true` tiles without a seam — rain, engine hum, crowd murmur, machine noise. Combine with a longer duration so the loop point is not obvious.
116
+
117
+ ### 5. Prompt influence
118
+
119
+ `promptInfluence` 0–1, default 0.3. Higher hugs your wording with less variation between takes; lower explores. Raise it when a re-roll keeps wandering off the brief; lower it when every take sounds like the same take.
120
+
121
+ ---
122
+
123
+ ## Iterating on either surface
124
+
125
+ - TTS re-rolls that keep drifting = raise `stability`. TTS reads that sound robotic = lower it, then fix the punctuation.
126
+ - SFX re-rolls that keep missing = the prompt named a label instead of a cause. Rewrite it as a physical event.
127
+ - Three failed takes means the prompt is wrong, not the seed.
128
+
129
+ ## Content notes
130
+
131
+ ElevenLabs applies its own moderation, and voice likeness of real people is restricted by their terms. See slates-content-policy.
@@ -103,7 +103,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
103
103
 
104
104
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
105
105
  2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
106
- 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
106
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
107
107
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
108
108
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
109
109
  6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
@@ -29,7 +29,7 @@ Never rely on the provider default (it's high — the priciest tier). The Slates
29
29
 
30
30
  - State the grid explicitly and number the cells: "a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom".
31
31
  - Give each cell ONE content clause: "Panel 3: the character mid-jump, side view".
32
- - Character sheets: "character turnaround sheet: front, 3/4 left, profile, back — same character, same outfit, flat even lighting, plain background". GPT Image 2 holds the layout; the Banana line holds the *face* better — for identity-critical turnarounds prefer NB2/NB Pro and use GPT Image 2 when labels/annotations matter.
32
+ - Character identity sheets: GPT Image 2 holds structured panel layouts; the Banana line holds the *face* better. Prefer NB2/NB Pro for identity-critical sheets and GPT Image 2 when labels or annotations are the main requirement.
33
33
 
34
34
  ## References & editing
35
35
 
@@ -125,7 +125,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
125
125
 
126
126
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
127
127
  2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
128
- 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
128
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
129
129
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
130
130
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
131
131
  6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
@@ -164,7 +164,7 @@ Layer scene-specific suppressions on top.
164
164
  - **Pro**: higher visual quality, no audio
165
165
  - **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references
166
166
 
167
- Pick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models` for current numbers before choosing a tier.
167
+ Pick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.
168
168
 
169
169
  ## Benchmark prompt structure
170
170
 
@@ -179,6 +179,7 @@ Cinematic example (paraphrasing fal blog patterns):
179
179
  > Shot 2: Medium shot of a detective in a trench coat ducking under an awning, water dripping from his hat brim. [Detective: weary, raspy]: 'I knew she'd come back.' Ambient noise: distant traffic, rain on metal.
180
180
  > Shot 3: Close-up on his eyes, narrowing as headlights flash across his face."
181
181
 
182
+ <!-- slates-only -->
182
183
  ## Pre-flight: references arrive inline, refer by code
183
184
 
184
185
  When you call `slates_generate_video` with `firstFrameAssetId` or `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside cost + `requires_confirm: true`. Look at them, revise prompt if needed, then re-call with `confirm=true`. Kling Omni multi-character with several ingredient images especially benefits — confirm each character image lands cleanly before spending.
@@ -187,14 +188,15 @@ When talking to the user about the gen, refer to each reference by its short cod
187
188
 
188
189
  - ✅ "I'm anchoring on **IMG-A12** as the detective and **IMG-A18** as the alleyway environment — Omni will handle the line delivery in EN."
189
190
  - ❌ "I'm using the detective image and the alley one..." (which alley? Three exist.)
191
+ <!-- /slates-only -->
190
192
 
191
- ## Video-to-video EDIT (`slates_edit_video`) — @Video1 / @ElementN / @ImageN
193
+ ## Video-to-video EDIT<!-- slates-only --> (`slates_edit_video`)<!-- /slates-only --> — @Video1 / @ElementN / @ImageN
192
194
 
193
195
  Kling O3 edit takes an EXISTING 3-15s clip and changes only what the prompt names — character swap, environment change, style transfer — in one pass, no masking. Original motion, camera, and audio are preserved by default. Its notation is Kling's own, different from the "image N" naming used everywhere else:
194
196
 
195
197
  - **`@Video1`** — the source clip (always; the transport anchors the instruction to it).
196
- - **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically).
197
- - **`@Image1..`** — style/appearance references (pass as `styleAssetIds`).
198
+ - **`@Element1..`** — subjects to swap IN. Each element = one frontal image + up to 3 angle images<!-- slates-only --> (pass as `characterAssetIds`; @mention names in the prompt compile to @ElementN automatically)<!-- /slates-only -->.
199
+ - **`@Image1..`** — style/appearance references<!-- slates-only --> (pass as `styleAssetIds`)<!-- /slates-only -->.
198
200
  - Max **4 combined** element + image refs per edit.
199
201
 
200
202
  **Prompt shape — the change, not the whole scene:**
@@ -212,7 +214,7 @@ Rules:
212
214
  - One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).
213
215
  - Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.
214
216
  - Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.
215
- - Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings — see `slates-model-selection`.
217
+ - Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.
216
218
 
217
219
  ## Sources
218
220
 
@@ -5,7 +5,7 @@ description: How to write prompts that produce cinematic, photorealistic results
5
5
 
6
6
  # Nano Banana 2 — cinematic & photorealistic prompting
7
7
 
8
- The **default** model behind `slates_generate_image` is **Gemini 3.1 Flash Image** (Nano Banana 2) — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill. It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat. Verified against the runtime slug map in `slate/src/main/api/google.ts`. NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
8
+ Nano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the default model behind `slates_generate_image` — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
9
9
 
10
10
  Knowledge cutoff: January 2025. Anything after needs explicit reference images.
11
11
 
@@ -100,7 +100,7 @@ Default to #1. Reach for #2 only when positive framing can't suppress the unwant
100
100
  ## Reference images
101
101
 
102
102
  - **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.
103
- - **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style (or pass `referenceAssetIds`), Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (images 1 and 2) sits across from the woman (images 3 and 4) in the cafe (image 5)`, with a trailing `Render in the visual style of image 6.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**, so citing both of a subject's sheets under the SAME name ("Marcus") is what tells the model they are ONE person and stops the face averaging. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
103
+ - **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
104
104
 
105
105
  ### Reference rules (the verified ones)
106
106
 
@@ -123,7 +123,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
123
123
 
124
124
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
125
125
  2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
126
- 3. **One identity sheet per character — and whatever you do attach for a subject, NAME it as one entity.** A character's identity sheet is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is always better, because the model cannot tell which one is authoritative and averages them.** Where a character carries a second bound sheet — an explicit expression range, or a legacy turnaround+expression pair — cite BOTH under the SAME name, `Marcus (images 1 and 2)`. That shared name, not a role essay, is what tells the model they are ONE person and stops the varied expressions from averaging the face. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
126
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
127
127
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
128
128
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
129
129
  6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
@@ -135,7 +135,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
135
135
 
136
136
  ### For Nano Banana 2 specifically
137
137
 
138
- - **NB2's own consistency lever is "assign a distinct name to each character/object."** That is Google's phrasing for rule 3 — citing both of a subject's sheets under the SAME name is the officially-sanctioned mechanism, not a workaround.
138
+ - **NB2's own consistency lever is "assign a distinct name to each character/object."** That is Google's phrasing for rule 3 — cite each canonical identity inline by name.
139
139
  - **Rule 8 is a job you do, not one you delegate.** NB2 *is* the start-frame model — when a downstream video shot needs legible text, render it here and animate from this frame.
140
140
  - **Character consistency is officially "not 100% perfect"** per Google. Test before bulk generations. High-resolution, front-facing reference images help most.
141
141
  - **Injection is stochastic — budget 3-5 re-rolls per shot; re-roll, don't re-engineer.** First rolls miss faces/hands; the same prompt lands a clean one within a few tries.
@@ -156,7 +156,7 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
156
156
 
157
157
  ## Resolution tactics
158
158
 
159
- - Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — call `slates_estimate_generation_cost` for current numbers. Pick the cheapest resolution that serves the use case.
159
+ - Resolution is priced: NB2 4k costs roughly 2x 1k. Prices change — check current numbers<!-- slates-only --> by calling `slates_estimate_generation_cost`<!-- /slates-only -->. Pick the cheapest resolution that serves the use case.
160
160
  - **At 2K and above, the model allocates more tokens to surface detail** — explicit texture vocabulary (pores, fabric weave, grain) compounds at higher resolution.
161
161
  - 1k for fast iteration / drafts; 2k for hero shots; 4k only when you need print-grade detail.
162
162
  - 2K generations vary 20-60s+. Don't time-budget tightly.
@@ -182,4 +182,6 @@ Everything in this skill applies to the whole Nano Banana family; two variants t
182
182
  - **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.
183
183
  - **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.
184
184
 
185
+ <!-- slates-only -->
185
186
  Routing between them (and vs GPT Image 2 / FLUX / Seedream): `slates-model-selection`.
187
+ <!-- /slates-only -->
@@ -30,7 +30,7 @@ Google's fast video generation + editing model ("Nano Banana Pro for video" in c
30
30
 
31
31
  - **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
32
32
  - Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.
33
- - **Name references inline** the standard Slates way ("Marcus (images 1 and 2) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
33
+ - **Name references inline** the standard Slates way ("Marcus (image 1) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
34
34
  - **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language ("rain patters on the tin roof"). Negative direction as plain instructions ("Do not show text").
35
35
  - Duration is an explicit 3–10s integer param; cost scales linearly per second.
36
36