@slatesvideo/shared 0.6.10 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (85) hide show
  1. package/README.md +1 -1
  2. package/dist/auth.js +2 -2
  3. package/dist/clients/cloud.js +1 -1
  4. package/dist/index.d.ts +2 -1
  5. package/dist/index.js +4 -1
  6. package/dist/manual/content.d.ts +1 -1
  7. package/dist/manual/content.js +1 -1
  8. package/dist/operations/index.d.ts +817 -16
  9. package/dist/operations/index.js +1423 -372
  10. package/dist/operations/surface.d.ts +3 -1
  11. package/dist/operations/surface.js +37 -10
  12. package/dist/prompts/ad-presets.d.ts +77 -0
  13. package/dist/prompts/ad-presets.js +43 -0
  14. package/dist/prompts/agent-doctrine.js +5 -4
  15. package/dist/prompts/banned-tokens.d.ts +4 -29
  16. package/dist/prompts/banned-tokens.js +29 -204
  17. package/dist/prompts/craft-cards.js +2 -2
  18. package/dist/prompts/generation-policy.d.ts +41 -0
  19. package/dist/prompts/generation-policy.js +53 -0
  20. package/dist/prompts/guide-retrieval.d.ts +9 -0
  21. package/dist/prompts/guide-retrieval.js +53 -0
  22. package/dist/prompts/index.d.ts +1 -0
  23. package/dist/prompts/index.js +1 -0
  24. package/dist/prompts/model-capabilities.d.ts +18 -1
  25. package/dist/prompts/model-capabilities.js +72 -19
  26. package/dist/prompts/model-facts.d.ts +59 -0
  27. package/dist/prompts/model-facts.js +121 -15
  28. package/dist/prompts/partials.generated.js +8 -2
  29. package/dist/prompts/prompting-tips.d.ts +1 -1
  30. package/dist/prompts/prompting-tips.js +63 -18
  31. package/dist/prompts/reference-composer.d.ts +2 -0
  32. package/dist/prompts/reference-composer.js +51 -50
  33. package/dist/prompts/script-document.d.ts +165 -0
  34. package/dist/prompts/script-document.js +11 -0
  35. package/dist/prompts/shot-grammar.d.ts +4 -4
  36. package/dist/prompts/shot-grammar.js +3 -3
  37. package/dist/prompts/shot-spec.d.ts +13 -0
  38. package/dist/prompts/shot-spec.js +23 -5
  39. package/dist/skills/content.js +26 -23
  40. package/dist/update-check.d.ts +22 -0
  41. package/dist/update-check.js +109 -0
  42. package/exports/slates-chatgpt-images/generated/SKILL.md +107 -0
  43. package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
  44. package/exports/slates-prompt-builder/generated/SKILL.md +3 -3
  45. package/exports/slates-prompt-builder/generated/reference-character.md +9 -1
  46. package/exports/slates-prompt-builder/generated/reference-kling.md +3 -3
  47. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +22 -10
  48. package/exports/slates-prompt-builder/generated/reference-seedance.md +4 -4
  49. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +17 -17
  50. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  51. package/package.json +10 -4
  52. package/skills/_partials/cinematic-card.md +8 -0
  53. package/skills/_partials/cinematic-routes-short.md +2 -0
  54. package/skills/_partials/cinematic-tips-short.md +2 -0
  55. package/skills/_partials/decision-log.md +1 -13
  56. package/skills/_partials/image-defaults.md +11 -0
  57. package/skills/_partials/lens-video-split.md +1 -0
  58. package/skills/_partials/reference-rules-core.md +1 -1
  59. package/skills/_partials/sheet-tool-defaults.md +6 -0
  60. package/skills/slates-character-identity.md +9 -1
  61. package/skills/slates-chatgpt-images.md +107 -0
  62. package/skills/slates-cinematic-look.md +237 -0
  63. package/skills/slates-cost-discipline.md +18 -12
  64. package/skills/slates-direct-response-ad.md +13 -53
  65. package/skills/slates-edit-and-iterate.md +1 -1
  66. package/skills/slates-model-selection.md +139 -133
  67. package/skills/slates-one-prompt-film.md +38 -95
  68. package/skills/slates-project-organization.md +7 -3
  69. package/skills/slates-prompting-flux-2-max.md +15 -4
  70. package/skills/slates-prompting-gpt-image-2-5.md +41 -28
  71. package/skills/slates-prompting-kling-v3.md +3 -3
  72. package/skills/slates-prompting-lip-sync.md +1 -1
  73. package/skills/slates-prompting-minimax-h3.md +30 -17
  74. package/skills/slates-prompting-motion-transfer.md +1 -1
  75. package/skills/slates-prompting-nano-banana-2.md +24 -11
  76. package/skills/slates-prompting-seedance-2-5.md +12 -12
  77. package/skills/slates-prompting-seedance.md +5 -5
  78. package/skills/slates-prompting-seedream-5-lite.md +14 -3
  79. package/skills/slates-prompting-veo-3.md +1 -1
  80. package/skills/slates-script-craft.md +45 -0
  81. package/skills/slates-shot-variety.md +11 -40
  82. package/skills/slates-storyboard-from-script.md +14 -66
  83. package/skills/slates-style-prompting.md +54 -54
  84. package/skills/slates-ugc-influencer-ad.md +32 -309
  85. package/skills/slates-vision-feedback-loop.md +2 -1
@@ -24,9 +24,16 @@ description: How to prompt FLUX.2 Max (Black Forest Labs image model). Read befo
24
24
  4. **Bind every hex colour to an object.** `a #1B4D3E enamel mug` lands; an unbound colour does not.
25
25
  5. **For portraits add texture words** — `natural skin texture, realistic pores, subtle imperfections, soft diffused lighting`.
26
26
 
27
- **Examples**
28
- - `A chef plating in a steel kitchen pass. Shot on Hasselblad X2D, 80mm, f/2.8. Overhead fluorescents plus warm spill from the line. Natural skin texture, subtle imperfections. Muted steel and #7A3B2E copper.`
29
- - `An empty municipal pool at dusk, 35mm, deep focus, early digital camera with slight noise and flash falloff. Cracked #4A7C8C tiles. Candid, unstaged.`
27
+ <!-- @inject:cinematic-card -->
28
+ **For a photographic look, use only what this frame needs.** Image models default to clean, evenly lit and fully exposed. Describe what the camera sees, not just gear or mood:
29
+ - **Inspect every reference first.** Write its grade and imperfections in words: darkness, contrast, muddy or true blacks, colour, softness/noise, subject separation. Never grade cleaner or brighter than the look reference unless asked.
30
+ - **One light system** — `low sun behind her`, `her face falls into deep shadow`, `no light in front of her`.
31
+ - **Visible exposure** — `the sky burns out to white`, `dense, slightly crushed shadows`.
32
+ - **Lens name plus effect** — `200mm telephoto`, `peaks loom huge behind her and melt into soft shapes`.
33
+ - **Name every garment and close the foreground.** Omissions invite reference leakage or invented props.
34
+ Bind references inline. A scene reference owns the grade; for a look-only reference, write the new scene's light. References are optional. For owned-frame edits, describe only the change and what stays.
35
+ <!-- slates-only -->Use `slates-cinematic-look` with a technique ID or section query for more.<!-- /slates-only -->
36
+ <!-- @end:cinematic-card -->
30
37
 
31
38
  **Hard constraint:** no negative prompting. Every "no X" must be rewritten as the positive state — `no blur` becomes `sharp focus throughout`, `no people` becomes `empty scene`, `no harsh shadows` becomes `soft, diffused lighting`.
32
39
  <!-- @card:end -->
@@ -44,6 +51,10 @@ description: How to prompt FLUX.2 Max (Black Forest Labs image model). Read befo
44
51
  - `masterpiece`, `best quality`, `trending on artstation`, `8k`
45
52
  <!-- @banned:end -->
46
53
 
54
+ **Examples**
55
+ - `A chef plating in a steel kitchen pass. Shot on Hasselblad X2D, 80mm, f/2.8. Overhead fluorescents plus warm spill from the line. Natural skin texture, subtle imperfections. Muted steel and #7A3B2E copper.`
56
+ - `An empty municipal pool at dusk, 35mm, deep focus, early digital camera with slight noise and flash falloff. Cracked #4A7C8C tiles. Candid, unstaged.`
57
+
47
58
  Black Forest Labs' top image model, routed via fal.ai. In Slates: `slates_generate_image` with `model: flux-2-max` (REQUIRES projectId — no headless path), priced per resolution (1k/2k/4k — call `slates_estimate_generation_cost` for current numbers, never quote from memory). Strengths vs Nano Banana 2: photoreal texture, less censored, precise hex-color control, strong typography. Reference images route through FLUX's edit endpoint and carry a lower per-model cap than NB2's 14.
48
59
 
49
60
  ## Core structure — front-load what matters
@@ -141,7 +152,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
141
152
  Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
142
153
 
143
154
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
144
- 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
155
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
145
156
  3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
146
157
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
147
158
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-gpt-image-2-5
3
- description: Prompting GPT Image 2.5 (Flare and Sunburst) — the readable-text / character-sheet / shot-grid engine AND the photoreal front-runner. Read before calling slates_generate_image with model gpt-image-2-5-flare or gpt-image-2-5-sunburst. Covers picking the variant, the five quality tiers (high is the default and the everyday seat), resolution classes (1k/2k=1080p/3k=1440p/4k), reference-image roles, text-accuracy prompting, panel/grid layout direction, edit constraints, and when to route to the Banana line instead.
3
+ description: Prompt and edit images with GPT Image 2.5 Flare or Sunburst. Covers reference roles, realistic lighting, text, grids, quality choices and targeted edits. Use with slates_generate_image or slates_edit_image on these models.
4
4
  ---
5
5
 
6
6
  # GPT Image 2.5 — sheets, grids, and text that actually reads
@@ -15,23 +15,28 @@ description: Prompting GPT Image 2.5 (Flare and Sunburst) — the readable-text
15
15
  Keep it under 2,400 characters (the build fails above that) and keep the
16
16
  rationale, the receipts and the worked examples in the body below. -->
17
17
  <!-- /slates-only -->
18
- **Card — GPT Image 2.5.** The readable-text, ordered-panel and exact-placement engine, and the photoreal front-runner for people. Structure: subject and action, then the exact copy in quotes, then layout, then light.
19
-
20
- **Pick the seat first.** `flare` = the small/FAST seat, quality *comparable to* GPT Image 2 — drafts, exploration, volume. `sunburst` = OpenAI's *most capable*, higher quality, slower, same price — finals, hero frames, photoreal, multi-reference edits. Explore on Flare, finish on Sunburst.
21
-
22
- **The six levers**
23
- 1. **Quote every string that must render verbatim** — `the sign reads "OPEN 24 HOURS"`. Quoted strings render most reliably.
24
- 2. **Font FEEL, never a font name** — `clean geometric sans, high contrast`, `hand-painted brush lettering`.
25
- 3. **Order dense copy explicitly** — `Line 1: "..." Line 2: "..."`. It respects the ordering.
26
- 4. **Name the layout as a grid** for sheets and panels — `a 3x2 grid of panels, reading left to right, equal gutters`.
27
- 5. **Give every reference image a ROLE** — subject / style / clothing / background. New emphasis in 2.5 and the highest-leverage change for multi-reference work.
28
- 6. **Set `quality` deliberately.** Five rungs — `low`, `medium`, `high` (default), `xhigh`, `max` — spanning ~36× end to end, in UNEVEN steps: `max` is 4× `high`, but `xhigh` only ~1.8× it. `medium` is the draft seat; `high` is the everyday tier; reach past it only when tiny type, dense diagrams or many labelled elements ARE the job. 🚨 **Coming from GPT Image 2, the names moved one rung:** its `medium` is this `high`, its `high` is this `max` — same money, renamed ladder. Carrying an old value over silently buys a cheaper picture.
29
-
30
- **Examples**
31
- - `A 2x3 character turnaround sheet on a neutral grey field, equal gutters, reading left to right: front, three-quarter, profile, back, three-quarter back, top. One woman, mid-30s, cropped dark hair, olive field jacket. Flat even studio light, no cast shadows. Small caption under each panel naming the angle.`
32
- - `Photoreal portrait, natural window light from camera-left, visible skin texture and pores, 85mm compression. A man in his 50s in a charcoal knit, half-smile, looking just past lens.`
33
-
34
- **Hard constraint:** keep total on-image text under about 30 words for perfect accuracy — beyond that it degrades, gracefully but really. It has its own content filter, distinct from Gemini's.
18
+ **Card — GPT Image 2.5.** The photoreal front-runner for people, and the readable-text, ordered-panel engine. Structure: subject and action with each reference named where it is used, then any exact copy in quotes, then layout, then light.
19
+
20
+ **Pick the tier.** `flare` is Faster, quality comparable to GPT Image 2: drafts and volume. `sunburst` is Better quality, the most capable: finals, hero frames, photoreal people, multi-reference edits. Use the product default; choose Flare when speed is a stated priority.
21
+
22
+ **The levers**
23
+ 1. **Name each reference inline** — `the woman from image 1`, `lit and graded like image 2`. Never an opening paragraph about what the references are.
24
+ 2. **Quote every string that must render verbatim** — `the jacket reads "SLATES"`. Describe a font's feel, never its name; keep on-image text under about 30 words.
25
+ 3. **Name the layout as a grid** for sheets and panels — `a 3x2 grid of panels, reading left to right, equal gutters`.
26
+ 4. **Set `quality` deliberately.** `high` is the everyday tier; `max` is 4× its price, `xhigh` about 1.8×. Coming from GPT Image 2 the names moved one rung: its `medium` is this `high`.
27
+
28
+ <!-- @inject:cinematic-card -->
29
+ **For a photographic look, use only what this frame needs.** Image models default to clean, evenly lit and fully exposed. Describe what the camera sees, not just gear or mood:
30
+ - **Inspect every reference first.** Write its grade and imperfections in words: darkness, contrast, muddy or true blacks, colour, softness/noise, subject separation. Never grade cleaner or brighter than the look reference unless asked.
31
+ - **One light system** — `low sun behind her`, `her face falls into deep shadow`, `no light in front of her`.
32
+ - **Visible exposure** — `the sky burns out to white`, `dense, slightly crushed shadows`.
33
+ - **Lens name plus effect** — `200mm telephoto`, `peaks loom huge behind her and melt into soft shapes`.
34
+ - **Name every garment and close the foreground.** Omissions invite reference leakage or invented props.
35
+ Bind references inline. A scene reference owns the grade; for a look-only reference, write the new scene's light. References are optional. For owned-frame edits, describe only the change and what stays.
36
+ <!-- slates-only -->Use `slates-cinematic-look` with a technique ID or section query for more.<!-- /slates-only -->
37
+ <!-- @end:cinematic-card -->
38
+
39
+ **Hard constraint:** its own content filter, distinct from Gemini's. Never describe a reference as a photograph of a real person.
35
40
  <!-- @card:end -->
36
41
 
37
42
  <!-- @banned:start -->
@@ -42,8 +47,8 @@ description: Prompting GPT Image 2.5 (Flare and Sunburst) — the readable-text
42
47
  backticked and prose outside the backticks. -->
43
48
  <!-- /slates-only -->
44
49
  **Never use:**
45
- - a font NAME — describe the feel (`clean geometric sans, high contrast`) instead
46
- - a reference role essay (`Reference image 1 is a photograph of a woman. Use that exact woman.`) — name the subject inline instead
50
+ - a font NAME — describe the feel instead, as in: clean geometric sans, high contrast
51
+ - a reference described as a photograph of a real person (`is a photograph of a woman`), or any up-front essay about what each reference is for — name the subject inline where it is used instead, as in: the woman from image 1
47
52
  - `8k`, `masterpiece`, `best quality`, `highly detailed` — quality incantations do nothing here either
48
53
  <!-- @banned:end -->
49
54
 
@@ -55,7 +60,7 @@ GPT Image's edge is **character-level text accuracy** (~99% on English), ordered
55
60
 
56
61
  🚨 **FLARE IS NOT AN UPGRADE OVER GPT IMAGE 2 — IT IS THE FAST ONE.** OpenAI, verbatim: *"GPT Image 2.5 Flare is the small model, optimized for speed, with image quality **comparable to** GPT Image 2. GPT Image 2.5 Sunburst is the base model, optimized for quality, with **higher image quality than** GPT Image 2."* Their model pages agree: Flare is *"our fastest model for high-quality, everyday image generation"*, Sunburst *"our most capable model for image generation and editing."* **Sunburst is the seat that beats what we had; Flare is the one that holds it at half the latency.** An earlier revision of this file called Flare "better than GPT Image 2" and sent Sunburst only to multi-reference edits — both wrong, corrected 2026-09-09 against the vendor docs.
57
62
 
58
- **The production pattern: explore on Flare, finish on Sunburst.** Drafts, layout checks and volume go to Flare. Finals, hero frames, photoreal people and any edit that must preserve identity or geometry go to Sunburst.
63
+ **Choose for the task.** Use the product default for ordinary work. Flare is an option when speed matters; changing model is not a mandatory draft stage.
59
64
 
60
65
  **Sunburst's widest lead is multi-reference editing** — several references all surviving into one frame, the character-consistency-across-shots problem. Reach for it there first, but that is not the only place it belongs.
61
66
 
@@ -63,7 +68,7 @@ GPT Image's edge is **character-level text accuracy** (~99% on English), ordered
63
68
 
64
69
  🚨 **The GPT Image line is ALSO the photoreal front-runner, and this file said the opposite until 2026-08-24.** **Receipts:** Eric's direct call, plus a head-to-head on the Higgsfield rail where GPT Image 2 at `quality: high`, 2K beat both Nano Banana rails on skin realism for photoreal people — that result is why the whole AI-influencer ad lane generates its plates here. **Route photoreal to this line, not away from it.**
65
70
 
66
- ⚠️ **Which SEAT reproduces it follows from the two facts above, and it is not the obvious one.** The receipt was measured on GPT Image 2 at `high`, which is this model's **`max`** — the ladder was renamed, not repriced (see `slates-model-selection`). Flare is only *comparable* to GPT Image 2, so **Flare at `max` is the floor: it holds the measured result rather than beating it.** Sunburst is documented as higher quality than GPT Image 2, which makes **Sunburst at `max` the seat most likely to exceed it** — and a photoreal final is exactly the "quality outranks speed" case OpenAI routes to Sunburst. Nobody has re-run the head-to-head on either seat, so this is reasoning from the vendor's positioning, not a measurement. **Run Flare-max against Sunburst-max on one plate before committing the lane, and write the result here.**
71
+ **Historical receipt, not a tier recommendation:** the photoreal comparison above used GPT Image 2 at its old `high` tier. It has not been repeated on 2.5 under matched conditions. Start with the product default and test a higher tier only against an unmet requirement; the old comparison does not establish a minimum tier for this model.
67
72
 
68
73
  **What the Banana line still owns:** edit-heavy work, and holding many subjects coherently in one frame. **Not the reference ceiling any more** — that line was true until 2026-09-09, when GPT Image went to its documented 16 against Banana's 14. Route on which model keeps them all recognisable, not on the count.
69
74
 
@@ -76,18 +81,18 @@ All five rungs are exposed, and they span ~36× end to end (2k class: $0.0044
76
81
  | Tier | Use it for |
77
82
  |---|---|
78
83
  | `low` | Roughest pass — layout and composition checks, throwaway comps. |
79
- | `medium` | The draft seat. Cheaper than NB2 Lite and available up to 4K, which is why the draft lane moved here. |
80
- | `high` | **Default.** The everyday tier. Blind benchmarks on GPT Image 2 put this rung — which it called `medium` — within a hair of `max` (which it called `high`) at a quarter of the cost. Inherited from the old ladder, never re-run on 2.5, and it says nothing about `xhigh`. |
84
+ | `medium` | The draft tier. Cheaper than NB2 Lite and available up to 4K, which is why the draft lane moved here. |
85
+ | `high` | General-purpose quality tier. Blind benchmarks on GPT Image 2 put this rung — which it called `medium` — within a hair of `max` (which it called `high`) at a quarter of the cost. Inherited from the old ladder, never re-run on 2.5, and it says nothing about `xhigh`. |
81
86
  | `xhigh` | One rung short of the top at about half its price (2k: 4 cr against `max`'s 8). Worth trying before `max`. |
82
87
  | `max` | Top of the ladder. Tiny type, dense diagrams, many labelled elements. |
83
88
 
84
89
  ⚠️ **A tier label means different things on different models.** OpenAI: *"The same quality label does not imply the same image quality or response time across models."* Flare at `max` and Sunburst at `max` are not the same picture, and neither matches Nano Banana's idea of "high".
85
90
 
86
- 🚨 **The tier NAMES moved between versions and the strings did not.** GPT Image 2's `medium` is this model's `high`; its `high` is this model's `max` — same money, one rung of renaming. So a recipe, a doc or a memory that says "GPT Image at medium" means **`high` here**. Getting this backwards costs picture quality silently: nothing errors, the bill is correct for what was asked, and the image is just worse.
91
+ 🚨 **The tier NAMES moved between versions and the strings did not.** GPT Image 2's `medium` is this model's `high`; its `high` is this model's `max` — same money, one rung of renaming. For a recipe explicitly written for GPT Image 2, map the old tier before reusing it on 2.5. A current user request for `medium` still means `medium`. Getting this backwards costs picture quality silently: nothing errors, the bill is correct for what was asked, and the image is just worse.
87
92
 
88
- Never rely on the provider default. fal's default is `high`, which is correct today — but it is the third rung of five rather than the top of two, so leaning on it means a fal-side change silently reprices you. The Slates ops send `high` unless you say otherwise.
93
+ Never rely on the provider default. fal's default is `high`, which is correct today — but it is the third rung of five rather than the top of two, so leaning on it means a fal-side change silently reprices you. Slates sends its configured quality explicitly; current defaults live in `slates-model-selection`. Your explicit choice overrides them.
89
94
 
90
- **Find the tier from the top down, then walk back.** OpenAI's own procedure: *"If the output falls short, test a higher quality setting. Once it meets your requirements, test lower settings to see whether they preserve acceptable quality while reducing latency. Use `xhigh` or `max` only when they improve an unmet quality requirement within your latency budget."* A higher rung does **not** guarantee a better result on a given prompt. Compare `medium` against `high` when the job is small or dense text; that is where the rungs separate most visibly.
95
+ **Start at the default and change tiers for an unmet requirement.** OpenAI's own procedure: *"If the output falls short, test a higher quality setting. Once it meets your requirements, test lower settings to see whether they preserve acceptable quality while reducing latency. Use `xhigh` or `max` only when they improve an unmet quality requirement within your latency budget."* A higher rung does **not** guarantee a better result on a given prompt. Compare `medium` against `high` when the job is small or dense text; that is where the rungs separate most visibly.
91
96
 
92
97
  ## Resolution classes
93
98
 
@@ -101,10 +106,14 @@ Never rely on the provider default. fal's default is `high`, which is correct to
101
106
 
102
107
  🚨 **THE ASPECT RATIO CHANGES THE PRICE ON THIS MODEL, and on no other image model.** OpenAI bills image OUTPUT TOKENS and the count tracks the frame's SHAPE, so at the same resolution class **`1:1` costs about 1.8× and `4:3`/`3:4` about 1.37× what `16:9` costs**; `9:16` costs the same as `16:9`. Metered 2026-09-09 and priced into the cost key, so the quote you get before generating is the real number — but if you are choosing between shapes and the budget is tight, **16:9 or 9:16 is the cheap one.** Every other image model charges the same whatever the shape.
103
108
 
104
- ## Reference images — give every one a role
109
+ ## Reference images — give every one a role, inline, where it is used
105
110
 
106
111
  **Assign a role to every reference image: subject, style, clothing, or background.** This is new emphasis in 2.5 and the highest-leverage change for the 16-reference character lane. An unroled pile of references makes the model guess what each one is for, and it guesses differently every run — which is the drift people mistake for a consistency failure.
107
112
 
113
+ **The role rides a clause in the scene, not a paragraph in front of it.** *The woman from image 1 cooks on a rocky summit…*, *lit and graded like image 2*. Never open with sentences about what each reference is and what to take or ignore from it: that is the role essay the shared reference rules below forbid, and it drags the sheet's studio light into the scene.
114
+
115
+ **Receipt, 2026-09-15, Sunburst, IMG-A192–A198.** The up-front version returned the studio look; the inline versions were never refused and never came back as a sheet. Two costs, both fixed in words: anything the prompt does not describe is taken from the reference (name every garment), and props nobody asked for appear (say what is in the foreground and that nothing else is). One sheet-only plate kept its described location, which narrows the two-reference rule in `slates-ugc-influencer-ad`. A look reference did far less than a described light. The full ladder is the vault's `cinematic-look-research.md`; the techniques are `slates-cinematic-look`.
116
+
108
117
  Reference images route through the edit endpoint, **up to 16** — fal's documented `maxItems`, and the highest reference ceiling of any image seat in Slates (the Banana line takes 14). It was capped at 10 until 2026-09-09, which was never anybody's limit, just a number nobody had checked. The composed "image N" naming applies as everywhere else. Mask-based inpainting exists at the API level but is not surfaced: a mask is something the user has to paint, and there is no painting surface — describe the change instead.
109
118
 
110
119
  ## Editing — separate the change from the constraints
@@ -142,6 +151,8 @@ For anything with several requirements, OpenAI recommends organising the prompt
142
151
 
143
152
  Name materials, lighting, colour and medium. Mood words are cues only — "cinematic", "moody", "epic" tell the model almost nothing on their own. Give scale, atmosphere and colour instead. Camera specs (`85mm`, `f/1.4`) are appearance hints, not a physical simulation; they bias the look, they do not compute optics.
144
153
 
154
+ **Name the lens and describe its effect, every time.** A lens named alone changed nothing visible (IMG-A195, 2026-09-15); named together with what it does to the picture, it produced real compression and depth of field (IMG-A198). Wording: `slates-cinematic-look` → `compression-as-outcome`, `defocus-as-outcome`.
155
+
145
156
  **For people, state body framing and scale**: "full body visible, feet included", "hands naturally gripping the handlebars". This is also the safest way to phrase a crop — see the blocked-phrasings section below.
146
157
 
147
158
  **No special syntax is required.** Prose, JSON and tagged blocks all work equally well, so pick whatever stays maintainable in the caller.
@@ -163,6 +174,8 @@ Name materials, lighting, colour and medium. Mood words are cues only — "cinem
163
174
 
164
175
  The first reads to the filter as *recreate this real person's likeness*, which is a hard refusal regardless of what the rest of the prompt says. The second signals a fictional character and passes. **This is a wording change only — the reference image can be the same file either way.** One plate flipped from refused to accepted on this single sentence with nothing else altered.
165
176
 
177
+ **Inline naming sidesteps the question and is now the default:** never describe the reference at all, and name her where she is used (*the woman from image 1*). Six of six Sunburst plates written that way passed on 2026-09-15. Keep the sheet sentence above as the fallback if a refusal appears.
178
+
166
179
  **2. Never attach a reference sheet containing a headless body panel.** A sheet whose full-body panels are cropped above the neck is refused every time, even with the correct opener. Regenerate the sheet with the head visible in every panel. Related, and already in this file's sheet guidance: phrase a cropped panel as *framing* (`cropped at the collarbone`), never as *absence* (`the head not shown`).
167
180
 
168
181
  ⚠️ **These refusals were measured on GPT Image 2, not on 2.5.** The classifier belongs to OpenAI rather than to a model version, so the phrasing rules carry — but they are inherited, not re-measured. If Flare or Sunburst accepts one of the blocked phrasings, that is a new receipt to write down here, not a reason to delete this one.
@@ -163,7 +163,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
163
163
  Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
164
164
 
165
165
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
166
- 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
166
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
167
167
  3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
168
168
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
169
169
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
@@ -185,11 +185,11 @@ Kling exposes `negative_prompt` on the fal endpoint (different from Seedance whi
185
185
 
186
186
  ```
187
187
  blurry, low quality, watermark, text overlay, distorted hands, extra fingers,
188
- duplicate limbs, unnatural skin texture, overly saturated colors, lens flare,
188
+ duplicate limbs, unnatural skin texture, overly saturated colors,
189
189
  floating objects, inconsistent shadows, jittery, flickering, morphing face
190
190
  ```
191
191
 
192
- Layer scene-specific suppressions on top.
192
+ Layer scene-specific suppressions on top, and never suppress something the prompt asks for. This block carried `lens flare` until 2026-09-15, which silently cancelled every flare a prompt described (`slates-cinematic-look` → `source-flare`); add it back only for a shot that must have none.
193
193
 
194
194
  ## Cinematic tactics
195
195
 
@@ -60,7 +60,7 @@ Seedance can generate the performance rather than bolting a mouth onto finished
60
60
  That is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a "video 1" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.
61
61
 
62
62
  - Driving clips must be 2–15s; output duration is whatever you set (4–15s).
63
- - Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming.
63
+ - Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
64
64
  - Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.
65
65
 
66
66
  Everything below is about the Kling tool.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-minimax-h3
3
- description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09), so the seats differ on ladder and price, not on what they accept. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; pooled media tokens on Max), and audio written into the wrong section is dropped or duplicated.
3
+ description: How to prompt MiniMax H3, H3 Max and H3 Max Turbo. Read before calling slates_generate_video with model minimax-h3, minimax-h3-max or minimax-h3-max-turbo. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, runs 480p/768p plus a 1080p refinement of its 768p render, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09). minimax-h3-max-turbo is a second fal post-train with Max's ladder at half Max's rate; it takes start and end frames but has NO reference endpoint. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; pooled media tokens on Max), and audio written into the wrong section is dropped or duplicated.
4
4
  ---
5
5
 
6
6
  # MiniMax H3 — prompting
@@ -28,7 +28,7 @@ description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling sl
28
28
  - `A woman sits still at a kitchen table for a beat, then looks up. She says in English, "You said Tuesday." Scene sound: a fridge hum, a spoon set down on formica. Score: none.`
29
29
  - `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, "No es el alternador." Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`
30
30
 
31
- **Hard constraint:** the two seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 768p rather than 4K, takes the same 9+3+3 references, and costs MORE at the tier they share — it is a speed pick, never the cheap one. H3's top two resolution tiers are UPSCALES of the native render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.
31
+ **Hard constraint:** the three seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p, takes the same 9+3+3 references, and costs MORE at the tier they share — a speed pick, never the cheap one; `minimax-h3-max-turbo` has Max's ladder at half its rate and takes frames only, NO references. Every tier above 768p is built from the native 768p render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.
32
32
  <!-- @card:end -->
33
33
 
34
34
  <!-- @banned:start -->
@@ -50,21 +50,31 @@ German, Italian, Japanese, Korean, Portuguese, Russian, Spanish). That single fa
50
50
  everything below — the prompt is not a shot description with sound bolted on, it is a **timeline
51
51
  with three audio layers you author separately**.
52
52
 
53
- **Two seats, one grammar.** Everything in this file applies to both. They differ only in what the
54
- endpoint accepts:
53
+ **Three seats, one grammar.** Everything in this file applies to all three. They differ only in
54
+ what the endpoint accepts:
55
55
 
56
- | | `minimax-h3` | `minimax-h3-max` |
57
- |---|---|---|
58
- | Resolution | 480p / 768p / **2K / 4K** | 480p / 768p |
59
- | References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) |
60
- | Frames | start and/or end | start and/or end |
61
- | Price at 768p | **$0.060/s** | $0.080/s |
62
- | Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) |
56
+ | | `minimax-h3` | `minimax-h3-max` | `minimax-h3-max-turbo` |
57
+ |---|---|---|---|
58
+ | Resolution | 480p / 768p / **2K / 4K** | 480p / 768p / 1080p | 480p / 768p / 1080p |
59
+ | References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) | **none** (no reference endpoint) |
60
+ | Frames | start and/or end | start and/or end | start and/or end |
61
+ | Price at 768p | **$0.060/s** | $0.080/s | $0.040/s |
62
+ | Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) | **price** — half Max's rate at every tier |
63
63
 
64
64
  **Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it
65
65
  tops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth
66
66
  paying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.
67
67
 
68
+ **Turbo is the budget seat.** Same grammar and Max's ladder at half Max's rate, with no reference
69
+ endpoint: attach a reference and Slates refuses the call rather than dropping it. Route there for
70
+ drafts, volume and start-frame coverage, then re-run the keeper on Max or base H3 when it needs
71
+ references.
72
+
73
+ **1080p on Max and Turbo is a refinement, not a native render.** fal's schema, verbatim: *"1080P
74
+ latent refinement from a native 768P source."* It is a different stage from base H3's 2K/4K
75
+ upscaler, and it costs double the 768p second. Judge a 1080p take against the same shot at 768p
76
+ before paying for it across a batch.
77
+
68
78
  **The speed is measured, not claimed** (2026-08-27, same prompt and params on both rows): a 5-second
69
79
  768p text-to-video finished in **4.8 seconds** on Max against **57 seconds** on base H3 — roughly
70
80
  **12x**, queue to finished file. fal advertises "under 3 seconds"; the literal claim did not hold at
@@ -192,8 +202,8 @@ original wording preserved exactly: *A red neon sign reading "Open Late" glows a
192
202
 
193
203
  ## References — H3's real differentiator is the declared RELATIONSHIP
194
204
 
195
- *(BOTH rows. `minimax-h3-max` gained the reference set on 2026-09-09; its free allowance is
196
- FOUR images rather than the base row's five.)*
205
+ *(`minimax-h3` and `minimax-h3-max`. Max gained the reference set on 2026-09-09; its free allowance
206
+ is FOUR images rather than the base row's five. `minimax-h3-max-turbo` takes no references.)*
197
207
 
198
208
  <!-- @inject:references-read-literally -->
199
209
  > **The general law: the model reads a reference literally.**
@@ -213,7 +223,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
213
223
  Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
214
224
 
215
225
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
216
- 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
226
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
217
227
  3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
218
228
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
219
229
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
@@ -260,7 +270,7 @@ exactly and say so.
260
270
  **An audio reference cannot travel alone** — H3 refuses a reference set that is audio only. Pair it
261
271
  with at least one image or video reference.
262
272
 
263
- ### 💸 Reference images past the free allowance are billed — and the two rows differ
273
+ ### 💸 Reference images past the free allowance are billed — and the two reference rows differ
264
274
 
265
275
  On `minimax-h3` the first **5** are free and each additional image adds **4 credits**.
266
276
  Max pools image pixels, reference-video seconds and reference-audio seconds into one token
@@ -279,14 +289,14 @@ reference-heavy job — a quote that omits it under-reports the bill.
279
289
 
280
290
  ## Frames
281
291
 
282
- `minimax-h3` and `minimax-h3-max` both take a **start frame**, an **end frame**, or both. With an
292
+ All three rows take a **start frame**, an **end frame**, or both. With an
283
293
  end frame, land it explicitly: describe the final pose, spacing and composition as the thing the
284
294
  shot **settles into** at the end, rather than hoping the model finds it.
285
295
 
286
296
  > *"…she rotates the handle into the final angle and settles into the pose, spacing and composition
287
297
  > of image 2 at the end of the shot."*
288
298
 
289
- **Frames and references are mutually exclusive** on both rows — they are different endpoints, and
299
+ **Frames and references are mutually exclusive** on the two reference rows — they are different endpoints, and
290
300
  the reference endpoint has no frame slots at all. Slates refuses the combination rather than
291
301
  dropping one side.
292
302
 
@@ -301,6 +311,9 @@ dropping one side.
301
311
  | `minimax-h3` · 2K · 10s | 65 |
302
312
  | `minimax-h3` · 4K · 10s | 80 |
303
313
  | `minimax-h3-max` · 768p · 10s | 40 |
314
+ | `minimax-h3-max` · 1080p · 10s | 80 |
315
+ | `minimax-h3-max-turbo` · 768p · 10s | 20 |
316
+ | `minimax-h3-max-turbo` · 1080p · 10s | 40 |
304
317
  | `minimax-h3` — every reference image past the **fifth** | **+4** |
305
318
 
306
319
  **768p is the default for a reason.** It is the tier the model natively generates.
@@ -64,7 +64,7 @@ movement from video 1. Preserve the character's identity, appearance, and outfit
64
64
  That is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.
65
65
 
66
66
  - **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).
67
- - **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.
67
+ - **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
68
68
  - **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
69
69
  - `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.
70
70
 
@@ -18,20 +18,27 @@ description: How to write prompts that produce cinematic, photorealistic results
18
18
  **Card — Nano Banana 2 (Gemini 3.1 Flash Image).** Brief it like a creative director, not a tag list. Structure: `Film still from [director] [genre]. Shot on [camera] with [lens]. [Subject and action]. [3-5 specific visual details]. [Lighting — direction + quality]. [Color palette]. [Film stock]. [1-2 word tone].`
19
19
 
20
20
  **The five levers**
21
- 1. **Named lens + aperture** beats "shallow depth of field" — `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin), `Panavision anamorphic`, `400mm telephoto`.
21
+ 1. **Named lens + aperture** beats "shallow depth of field" — `85mm f/1.4`, `135mm f/2.8`, `Panavision anamorphic`, `400mm telephoto`.
22
22
  2. **Light by direction and quality**, never "good lighting" — `hard sidelight from a single window, deep falloff`, `overcast north light`, `practical tungsten spill`.
23
23
  3. **A named film stock or sensor** carries a whole palette — `Kodak Portra 400`, `Cinestill 800T`, `ARRI Alexa 65`.
24
24
  4. **Composition as a shot** — `low angle`, `aerial view`, `rule of thirds with the subject camera-left`, `foreground occlusion`.
25
25
  5. **Positive framing only.** Describe what is there. "Empty street", never "no cars"; "unstaged documentary photography", never "not anime".
26
26
 
27
- **Examples**
28
- - `Film still from a Denis Villeneuve thriller. Shot on ARRI Alexa 65, 85mm f/1.4. A woman in a charcoal wool coat stands at a rain-slick bus stop, breath visible. Hard sodium light from a single overhead lamp, deep falloff into blue night. Kodak Vision3 500T. Isolated.`
29
- - `Editorial still life on seamless bone paper. 100mm macro, f/8. A cracked ceramic bowl holding three figs. Soft north light from camera-left, one gentle shadow. Muted earth palette. Portra 400 grain. Quiet.`
27
+ <!-- @inject:cinematic-card -->
28
+ **For a photographic look, use only what this frame needs.** Image models default to clean, evenly lit and fully exposed. Describe what the camera sees, not just gear or mood:
29
+ - **Inspect every reference first.** Write its grade and imperfections in words: darkness, contrast, muddy or true blacks, colour, softness/noise, subject separation. Never grade cleaner or brighter than the look reference unless asked.
30
+ - **One light system** — `low sun behind her`, `her face falls into deep shadow`, `no light in front of her`.
31
+ - **Visible exposure** — `the sky burns out to white`, `dense, slightly crushed shadows`.
32
+ - **Lens name plus effect** — `200mm telephoto`, `peaks loom huge behind her and melt into soft shapes`.
33
+ - **Name every garment and close the foreground.** Omissions invite reference leakage or invented props.
34
+ Bind references inline. A scene reference owns the grade; for a look-only reference, write the new scene's light. References are optional. For owned-frame edits, describe only the change and what stays.
35
+ <!-- slates-only -->Use `slates-cinematic-look` with a technique ID or section query for more.<!-- /slates-only -->
36
+ <!-- @end:cinematic-card -->
30
37
 
31
38
  **Hard constraint:** there is no `negativePrompt` field. Suppress by reframing positively, or inline `without` / `free of`. Knowledge cutoff January 2025 — anything later needs reference images.
32
39
  <!-- @card:end -->
33
40
 
34
- Nano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the default model behind `slates_generate_image` — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
41
+ Nano Banana 2 is **Gemini 3.1 Flash Image**.<!-- slates-only --> It is the headless image model used when `projectId` is omitted — the op also exposes `flux-2-max` and `seedream-5-lite`, each with its own prompting skill.<!-- /slates-only --> It is **not** Gemini 3 Pro Image; that is Nano Banana **Pro** (`nano-banana-pro`), a separate model with its own seat.<!-- slates-only --> Verified against the runtime slug map in `slate/src/main/api/google.ts`.<!-- /slates-only --> NB2 is a language model that outputs pixels — brief it like a creative director, not like a Stable-Diffusion tag-soup tool. The single biggest lever for realism: **specificity that mimics how real photographers and cinematographers describe their work**.
35
42
 
36
43
  Knowledge cutoff: January 2025. Anything after needs explicit reference images.
37
44
 
@@ -56,12 +63,14 @@ Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and a
56
63
 
57
64
  ## Photorealism positives — what consistently works
58
65
 
59
- > ⚠️ **This vocabulary is an IMAGE-model lever and a video-model anti-pattern — do not carry it across.**
60
- > Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are correct and encouraged **here**. They are a **Seedance anti-pattern**: ByteDance's own guide uses shot sizes, camera moves, pacing words and its image-quality vocabulary throughout, and never once mentions fps, shutter angle, f-stop, or lens millimetres.
61
- > The leak happens in one specific way — you write an NB2 start frame, then write the video prompt to animate it and carry the look description straight across. **Translate instead of copying:** `85mm f/1.4, Portra 400` → `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. Full rule and the receipts: `slates-prompting-seedance` (Part 3, "Don't cross-pollinate image-model syntax").
66
+ ⚠️ **This vocabulary is correct here and does not carry into a video prompt.** The leak happens one way: you write an NB2 start frame, then carry its look description straight into the prompt that animates it.
67
+
68
+ <!-- @inject:lens-video-split -->
69
+ Named lenses, apertures, film stocks and camera bodies (`85mm f/1.4`, `Kodak Portra 400`, `ARRI Alexa 65`) are an image-model lever. On a video model, translate the look instead of pasting the gear list: `85mm f/1.4, Portra 400` becomes `close-up, shallow depth of field, warm natural colors, cinematic texture, film-grain texture`. ByteDance's Seedance 2.0 guide never mentions fps, shutter angle, f-stop or lens millimetres. Its Seedance 2.5 guide does, once: the visual-style line of its own storyboard example names one camera body and one 35 mm cinema lens. On 2.5 a single line like that is vendor-sanctioned; a stacked gear list still is not.
70
+ <!-- @end:lens-video-split -->
62
71
 
63
72
  **Named lenses + apertures** beat generic "shallow depth of field":
64
- - `85mm f/1.4`, `135mm f/2.8` (the cheat code for skin texture), `50mm f/1.2`, `35mm f/2`
73
+ - `85mm f/1.4`, `135mm f/2.8`, `50mm f/1.2`, `35mm f/2`
65
74
  - `Panavision anamorphic` for horizontal flares + cinematic width
66
75
  - `400mm telephoto` for compression + isolation
67
76
  - `24mm` for environmental interiors
@@ -86,6 +95,7 @@ Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and a
86
95
  - `visible pores`, `natural skin grain`, `peach fuzz`, `slight hyperpigmentation`
87
96
  - `unretouched raw photography`, `ISO noise`, `sweat beading`
88
97
  - `crisp catchlights in the eyes`, `skin micro-detail`
98
+ - Lead with the kind of photograph and the conditions on the skin (sun, wind, sweat), then add one or two of these. A bare list of flaw words read as tokens and produced plastic skin on GPT Image 2 (2026-08-24).<!-- slates-only --> Technique: `slates-cinematic-look` → `name-the-capture-context`.<!-- /slates-only -->
89
99
 
90
100
  **Director references** (use when locking style):
91
101
  | Director | Tone | Visual signature |
@@ -123,6 +133,9 @@ These are Stable-Diffusion-era tag soup. The model treats them as low-signal noi
123
133
  - `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue
124
134
  <!-- @banned:end -->
125
135
 
136
+ **Examples**
137
+ - `Film still from a Denis Villeneuve thriller. Shot on ARRI Alexa 65, 85mm f/1.4. A woman in a charcoal wool coat stands at a rain-slick bus stop, breath visible. Hard sodium light from a single overhead lamp, deep falloff into blue night. Kodak Vision3 500T. Isolated.`
138
+
126
139
  ## Negative prompting — there is no field
127
140
 
128
141
  Nano Banana 2 has **no `negativePrompt` parameter**. Three patterns to suppress unwanted content:
@@ -136,7 +149,7 @@ Default to #1. Reach for #2 only when positive framing can't suppress the unwant
136
149
  ## Reference images
137
150
 
138
151
  - **Hard limit: 14 images** (10 object-fidelity + 4 character-consistency). Categories don't trade — you can't use 14 object slots even if no characters are referenced.
139
- - **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, with a trailing `Render in the visual style of image 4.` The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
152
+ - **Name each reference inline — Slates does this for you.** When you `@mention` a subject/environment or `#mention` a style<!-- slates-only --> (or pass `referenceAssetIds`)<!-- /slates-only -->, Slates composes the prompt so each reference is named inline as "image N" — e.g. `Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3)`, or `lit and graded like image 4` where you placed the style mention. An unmentioned style attachment gets a short fallback clause The model does NOT infer a reference's role from its position; the NAME carries it. NB2's own consistency lever is literally **"assign a distinct name to each character/object"**. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render the scene's expression") — that drags the sheet's wardrobe + studio lighting into the scene. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
140
153
 
141
154
  ### Reference rules (the verified ones)
142
155
 
@@ -158,7 +171,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
158
171
  Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
159
172
 
160
173
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
161
- 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
174
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
162
175
  3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
163
176
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
164
177
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-seedance-2-5
3
- description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is a SECOND SEAT next to 2.0, not an upgrade — it buys 30-second takes, 30 image references, audio-only references and INTEGER-SECOND TIMESTAMPS, and it gives up native 4K and costs more than 2.0 at every resolution they share. Timestamps are the one grammar difference that matters: 2.0 ignores them and answers only to shot numbers, 2.5 acts on them. Otherwise it shares 2.0's grammar (read slates-prompting-seedance for subject binding, camera and constraint vocabulary); this file covers what is different, plus the two hazards unique to 2.5 — the prompt-intent task classifier and the cost trap that comes with 30-second takes.
3
+ description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calling slates_generate_video with model seedance-2.5, or slates_edit_video with model seedance-2.5-edit. 2.5 is the DEFAULT video model (Eric, 2026-09-13) — against 2.0 it buys 30-second takes, 30 image references, audio-only references and INTEGER-SECOND TIMESTAMPS, and it gives up native 4K and costs more than 2.0 at every resolution they share. Timestamps are the one grammar difference that matters: 2.0 ignores them and answers only to shot numbers, 2.5 acts on them. Otherwise it shares 2.0's grammar (read slates-prompting-seedance for subject binding, camera and constraint vocabulary); this file covers what is different, plus the two hazards unique to 2.5 — the prompt-intent task classifier and the cost trap that comes with 30-second takes.
4
4
  ---
5
5
 
6
6
  # Seedance 2.5 — prompting
@@ -28,7 +28,7 @@ description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calli
28
28
  - `[0-6] Wide shot, <Subject_1>@<Image_1> crosses an empty car park toward a idling van, slow track right. [6-12] Medium, she stops as the driver's window comes down. [12-18] Close-up, she looks off past the lens and does not answer. Rich details, natural colors. Keep it subtitle-free.`
29
29
  - `[0-10] A single continuous handheld follow behind a courier climbing a fire escape, rain. [10-20] She reaches the landing, turns, and the city opens behind her. Cinematic texture, soft lighting.`
30
30
 
31
- **Hard constraint:** it is the EXPENSIVE seat and it has NO 4K — 480p/720p/1080p only, and dearer than 2.0 at every resolution they share. It is a second seat, never an upgrade. Long takes multiply cost linearly: quote a 30-second take before you fire it.
31
+ **Hard constraint:** it is the default AND the dearer seat, and it has NO 4K — 480p/720p/1080p only, dearer than 2.0 at every resolution they share. Long takes multiply cost linearly: quote a 30-second take before you fire it.
32
32
  <!-- @card:end -->
33
33
 
34
34
  <!-- @banned:start -->
@@ -40,7 +40,7 @@ description: How to prompt Seedance 2.5 and Seedance 2.5 Edit. Read before calli
40
40
  <!-- /slates-only -->
41
41
  **Never use** (2.5 reclassifies the task and fails a fresh generation on these):
42
42
  - `edit`, `extend`, `continue the video`, `same video but` — they make the provider read a fresh generation as an edit
43
- - `85mm`, `f/1.4`, `Portra 400` and any other lens, aperture, film-stock or camera-body token — image-model vocabulary, a Seedance anti-pattern on both seats
43
+ - `f/1.4`, `Portra 400` and any other aperture or film-stock token, or a stacked list of gear — image-model vocabulary. The 2.5 guide's own example names one camera body and one 35 mm lens in a single style line, so a lone lens there is not on this list
44
44
  <!-- @banned:end -->
45
45
 
46
46
  **Read `slates-prompting-seedance` first.** The prompt GRAMMAR is the same model family: the
@@ -70,11 +70,10 @@ So 2.5 does not replace 2.0; it sits beside it, and you pay for what it buys:
70
70
  | **Timestamps in the prompt** | **✗ — ignored; shot numbers only** | **✓ — integer seconds, acted on** |
71
71
  | Multi-view image as ONE subject reference | ✗ (not recommended) | **✓ (up to 5 subjects)** |
72
72
  | Video edit as its own task type | ✗ | **✓ (`seedance-2.5-edit`)** |
73
- | Default video model | **yes** | no |
73
+ | Default video model | no | **yes** (since 2026-09-13) |
74
74
 
75
- **Route to 2.5 when the shot needs LENGTH, MANY REFERENCES, or an audio-only reference.
76
- Route to 2.0 when resolution matters at all** — which, for anything a client will see full-screen,
77
- is most of the time.
75
+ **2.5 is the default. Route to 2.0 for 4K delivery, or when the same resolution has to be
76
+ cheaper** — its 720p is $0.15/s against 2.5's $0.231/s.
78
77
 
79
78
  ---
80
79
 
@@ -128,11 +127,11 @@ Worked, at the shipped rates:
128
127
  |---|---|
129
128
  | 2.5 · 480p · 5s · faceless | 26 |
130
129
  | 2.5 · 720p · 5s · faceless | 58 |
131
- | 2.5 · 1080p · 5s · faceless | 103 |
130
+ | 2.5 · 1080p · 5s · faceless | 142 |
132
131
  | 2.5 · 720p · 30s · faceless | 347 |
133
132
  | 2.5 · 720p · 30s · AI-face route | **489** |
134
133
  | 2.5 · 720p · 30s · consented real-face route | **710** |
135
- | 2.5 · 1080p · 30s · faceless | **614** |
134
+ | 2.5 · 1080p · 30s · faceless | **853** |
136
135
  | 2.5 · 1080p · 30s · consented real-face route | **1,749** |
137
136
  | *(for scale)* 2.0 · 1080p · 15s · AI-face route | 411 |
138
137
 
@@ -365,7 +364,7 @@ not a lip-sync job. Bill it like any other edit — on the source clip's length.
365
364
 
366
365
  The three-tier face routing is identical to 2.0 — faceless → default route, an AI character's face →
367
366
  `seedanceFace: true` (the relaxed provider, a real cost premium), a real person's photo → the
368
- consent-gated premium route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`
367
+ consent-gated real-person route after a `[REAL_FACE_DETECTED]` rejection, with `realFaceConsent: true`
369
368
  set **only** after the user explicitly confirms they hold the rights to the likeness. The full rules,
370
369
  including why the real-vs-AI call is the provider's and not yours, are in
371
370
  `slates-prompting-seedance`.
@@ -373,8 +372,9 @@ including why the real-vs-AI call is the provider's and not yours, are in
373
372
  Also unchanged, and worth restating because 2.5's length makes each one more expensive to get wrong:
374
373
 
375
374
  - **One primary camera move per shot.**
376
- - **No lens / aperture / film-stock vocabulary.** That is image-model syntax and a Seedance
377
- anti-pattern.
375
+ - **No stacked lens / aperture / film-stock vocabulary.** One camera-and-lens style line is the
376
+ most the 2.5 guide itself uses; the full rule is in `slates-prompting-seedance` → "Don't
377
+ cross-pollinate image-model syntax".
378
378
  - **No `negativePrompt` field** — constraints go inline, and 2.5 acts on negative phrasing in
379
379
  exactly two dimensions: subtitles (*"no subtitles"*) and audio (*"no BGM; environmental and
380
380
  action sounds only"*, *"no audio"*). Everywhere else, describe what you want, not what you don't.