@slatesvideo/shared 0.6.11 → 0.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (83) hide show
  1. package/dist/auth.js +2 -2
  2. package/dist/clients/cloud.js +1 -1
  3. package/dist/index.d.ts +1 -1
  4. package/dist/index.js +1 -1
  5. package/dist/manual/content.d.ts +1 -1
  6. package/dist/manual/content.js +1 -1
  7. package/dist/operations/index.d.ts +817 -16
  8. package/dist/operations/index.js +1413 -360
  9. package/dist/operations/surface.d.ts +4 -1
  10. package/dist/operations/surface.js +41 -10
  11. package/dist/prompts/ad-presets.d.ts +77 -0
  12. package/dist/prompts/ad-presets.js +43 -0
  13. package/dist/prompts/agent-doctrine.js +27 -5
  14. package/dist/prompts/banned-tokens.d.ts +4 -29
  15. package/dist/prompts/banned-tokens.js +29 -204
  16. package/dist/prompts/craft-cards.js +2 -2
  17. package/dist/prompts/generation-policy.d.ts +41 -0
  18. package/dist/prompts/generation-policy.js +53 -0
  19. package/dist/prompts/guide-retrieval.d.ts +9 -0
  20. package/dist/prompts/guide-retrieval.js +53 -0
  21. package/dist/prompts/index.d.ts +1 -0
  22. package/dist/prompts/index.js +1 -0
  23. package/dist/prompts/model-capabilities.d.ts +18 -1
  24. package/dist/prompts/model-capabilities.js +72 -19
  25. package/dist/prompts/model-facts.d.ts +34 -2
  26. package/dist/prompts/model-facts.js +66 -5
  27. package/dist/prompts/partials.generated.js +8 -2
  28. package/dist/prompts/prompting-tips.d.ts +1 -1
  29. package/dist/prompts/prompting-tips.js +61 -16
  30. package/dist/prompts/reference-composer.d.ts +2 -0
  31. package/dist/prompts/reference-composer.js +51 -50
  32. package/dist/prompts/script-document.d.ts +165 -0
  33. package/dist/prompts/script-document.js +11 -0
  34. package/dist/prompts/shot-grammar.d.ts +4 -4
  35. package/dist/prompts/shot-grammar.js +3 -3
  36. package/dist/prompts/shot-spec.d.ts +13 -0
  37. package/dist/prompts/shot-spec.js +23 -5
  38. package/dist/skills/content.js +27 -24
  39. package/exports/slates-chatgpt-images/generated/SKILL.md +107 -0
  40. package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
  41. package/exports/slates-prompt-builder/generated/SKILL.md +1 -1
  42. package/exports/slates-prompt-builder/generated/reference-character.md +9 -1
  43. package/exports/slates-prompt-builder/generated/reference-kling.md +3 -3
  44. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +22 -10
  45. package/exports/slates-prompt-builder/generated/reference-seedance.md +4 -4
  46. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +17 -17
  47. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  48. package/package.json +9 -3
  49. package/skills/_partials/cinematic-card.md +8 -0
  50. package/skills/_partials/cinematic-routes-short.md +2 -0
  51. package/skills/_partials/cinematic-tips-short.md +2 -0
  52. package/skills/_partials/decision-log.md +1 -13
  53. package/skills/_partials/image-defaults.md +11 -0
  54. package/skills/_partials/lens-video-split.md +1 -0
  55. package/skills/_partials/reference-rules-core.md +1 -1
  56. package/skills/_partials/sheet-tool-defaults.md +6 -0
  57. package/skills/slates-character-identity.md +9 -1
  58. package/skills/slates-chatgpt-images.md +107 -0
  59. package/skills/slates-cinematic-look.md +237 -0
  60. package/skills/slates-cost-discipline.md +18 -12
  61. package/skills/slates-direct-response-ad.md +13 -53
  62. package/skills/slates-edit-and-iterate.md +1 -1
  63. package/skills/slates-model-selection.md +20 -14
  64. package/skills/slates-one-prompt-film.md +19 -77
  65. package/skills/slates-project-organization.md +7 -3
  66. package/skills/slates-prompting-flux-2-max.md +15 -4
  67. package/skills/slates-prompting-gpt-image-2-5.md +41 -28
  68. package/skills/slates-prompting-inworld-tts.md +174 -174
  69. package/skills/slates-prompting-kling-v3.md +3 -3
  70. package/skills/slates-prompting-lip-sync.md +1 -1
  71. package/skills/slates-prompting-minimax-h3.md +30 -17
  72. package/skills/slates-prompting-motion-transfer.md +1 -1
  73. package/skills/slates-prompting-nano-banana-2.md +24 -11
  74. package/skills/slates-prompting-seedance-2-5.md +7 -6
  75. package/skills/slates-prompting-seedance.md +5 -5
  76. package/skills/slates-prompting-seedream-5-lite.md +14 -3
  77. package/skills/slates-prompting-veo-3.md +1 -1
  78. package/skills/slates-script-craft.md +45 -0
  79. package/skills/slates-shot-variety.md +11 -40
  80. package/skills/slates-storyboard-from-script.md +14 -66
  81. package/skills/slates-style-prompting.md +4 -4
  82. package/skills/slates-ugc-influencer-ad.md +32 -309
  83. package/skills/slates-vision-feedback-loop.md +2 -1
@@ -0,0 +1,237 @@
1
+ ---
2
+ name: slates-cinematic-look
3
+ description: Use when a frame should look filmed, not generated (light, exposure, grade, lens, atmosphere, imperfection), or when an image came back too clean or studio-lit. The organized technique catalogue for every image model, and the rule for picking only what the shot needs.
4
+ ---
5
+
6
+ # Cinematic look — make a generated frame read as filmed
7
+
8
+ <!-- @card:start -->
9
+ **Image models default to clean, evenly lit and fully exposed.** Real film frames can be dark, flat, murky or burned out. Describe what the camera sees; a mood word or look reference alone does not get you there.
10
+
11
+ **Look at every reference first.** Write a look reference's grade and imperfections into the prompt: darkness, contrast, black level, colour cast/saturation, softness/noise and subject separation. Never grade cleaner, brighter or higher-contrast than that reference unless asked. Inspect the character sheet's garments. Cite references inline: `the woman from image 1`, `lit and graded like image 2`; never open with a reference-role paragraph. References are optional; these techniques work from words alone.
12
+
13
+ For a new photographic frame:
14
+ - **One physical light system:** source, position, effect on the subject; no light without a source. If using atmosphere, put it in front too.
15
+ - **Exposure as it looks:** `close to a silhouette, features only just readable`, not a stop under.
16
+ - **Lens name plus effect, every time:** `200mm telephoto`, `the peaks loom huge behind her and melt into soft shapes`.
17
+ - **Name every garment and close the foreground:** say exactly what is there; omissions invite reference leakage or invented props.
18
+ - **Only what the shot needs:** one light system, at most one exposure, atmosphere and colour choice, one or two composition moves and one moment. Add the imperfections the scene calls for.
19
+
20
+ A scene reference owns the grade; a look-only reference does not own the new scene's light. For an owned-frame edit, describe only the change and what stays. Never use a released film frame as the edit base; use it as an art-direction brief for a new scene. Use positive descriptions first; one targeted negative is enough where supported. FLUX needs positive wording.
21
+
22
+ Use query with a technique ID or section, depth "index" to browse, or depth "full" for the catalogue and examples.
23
+ <!-- @card:end -->
24
+
25
+ **Evidence scope:** most Slates receipts come from single generations of one scene on GPT Image 2.5 Sunburst. They show useful directions to test, not reliable success rates or a universal ranking. The research record owns prompts, source links and observations: `business/projects/slates/research/cinematic-look-research.md` in the vault. This skill owns active technique wording. The catalogue checker compares IDs, evidence tags and worked-example bytes.
26
+
27
+ Image models default to a clean, evenly lit, fully exposed picture. Real film frames are often graded "wrong": underexposed, backlit, silhouetted, burned out, flat or murky. A model only goes there when the prompt says what that looks like. A look reference alone does not get you there; describing the frame does.
28
+
29
+ ## Two routes to a filmed frame
30
+
31
+ 1. **Describe a new frame.** Write the scene and reference roles inline, then the light, exposure, camera and texture choices needed to realize it.
32
+ 2. **Edit a frame you own.** When a Slates plate or sheet, your own photo or footage, or a Blender render already has the composition and look, describe only the requested changes and what must remain. Example: "Take image 1 and change only the character to the character in image 2." Never use a frame from a released film as the base; a film still is an art-direction brief for route 1.
33
+
34
+ ## The rules
35
+
36
+ 1. **Describe what the camera sees, not the camera setting.** "Her face sits a stop under" was ignored; "close to a silhouette, her features only just readable" landed. Mood words carry little on their own.
37
+ 2. **Build one physical light system first.** Name the source and its position, what it does to the subject, and rule out light with no source. If the shot needs atmosphere, place it in front of the subject as well as behind. Every source and shadow must agree.
38
+ 3. **Name the lens and describe its effect, every time you describe a new photographic frame.** A lens named alone changed nothing visible in the summit test; named with its effect, it produced compression and blur.
39
+ 4. **Use only what the shot needs.** One light system (its consequence clauses count as one), at most one pick each from exposure, atmosphere and colour, one or two composition moves and one moment. Then stop adding. These are drafting limits, not model capability limits.
40
+ 5. **Name everything a reference could fill.** Every garment, and a closed list of what is in the foreground. Anything left out can come from a reference or be invented. In route 2, preserve the existing inventory and describe only the change.
41
+ 6. **One targeted negative after a positive description is fine; a list is not.** Follow the model's grammar: FLUX needs positive wording.
42
+ 7. **A scene reference owns the grade; a look-only reference does not own the light.** With a look-only reference, still write the light direction and visible exposure for the new scene. Name references inline where used, never in an opening role paragraph.
43
+ 8. **Look at every reference before writing.** First describe a look reference's own grade and imperfections: darkness, contrast, true or muddy blacks, colour cast and saturation, softness and noise, and how much the subject separates from the background. Never grade cleaner, brighter or higher-contrast than the reference unless the user asks. Inspect a character sheet's garments so the prompt names or replaces every one. Without a reference, the same techniques work from words alone.
44
+
45
+ ## Build order
46
+
47
+ Inspect references first. For a new frame: time/weather → light source and direction → effect on the subject → exposure as it looks → atmosphere in front, if needed → colour → composition and camera placement (lens plus effect) → the moment → kind of picture → inventory (every garment and the closed foreground). Carry the observed reference grade through those choices. Adapt examples to the actual scene; never inherit their props, wardrobe or light by accident.
48
+
49
+ ## Evidence tags
50
+
51
+ `receipt` measured on Slates generations · `vendor` stated in the model maker's own guide · `practitioner` a published third-party guide or test · `canon` established cinematography, never measured on a model · `untested` reasoning only: try it, then record the result in the research doc and change the tag.
52
+
53
+ ## The catalogue
54
+
55
+ Each row: the technique, its evidence, what it does to the frame, when to reach for it and when to skip it, and wording written as what the frame looks like.
56
+
57
+ <!-- @catalogue:start -->
58
+
59
+ ### 1. Exposure and tone — the "graded wrong" frame
60
+
61
+ | Technique | Evidence | What it does | Reach for · skip | Say |
62
+ |---|---|---|---|---|
63
+ | `flat-underexposure` | receipt | The whole frame sits in a narrow, dark tonal range; the subject barely separates | A dark, flat look reference or fading light · skip hard contrast, glowing practicals and crushed blacks | "Everything sits in dark, muddy navy blue; nothing is bright or truly black. Her dim face barely separates from the trees; even the focused face is slightly soft, with fine noise in the dark blues." |
64
+ | `subject-under-key` | untested | The face is clearly darker than the brightest part of the frame; features read, nothing lights them | Backlit exteriors, sunrise and sunset, window interiors · skip beauty, product, lip-sync close-ups | "Her face is in shadow, clearly darker than the sky behind her; her features are readable but nothing lights them." |
65
+ | `near-silhouette` | receipt | A dark shape with edge detail against a bright field | Wides, entrances and exits, solitude · skip shots that depend on recognising the face | "She is close to a silhouette: her face and jacket fall into deep shadow, her features only just readable." |
66
+ | `clipped-highlights` | receipt | The sky near the sun, windows and practicals go pure white | Backlit shots, windows, night practicals · skip skies that carry the story, white packaging | "The sky around the sun burns out to white." |
67
+ | `dense-shadows-with-ramp` | receipt | Shadows sink to near-black but fall off gradually, and one detail survives in the dark | Night, interiors, a backlit foreground · skip video dark work that already turns to murk | "The shadows are dense and slightly crushed, the foreground rock nearly black, and the light fades into them gradually." |
68
+ | `low-key-ratio` | vendor | One side of the face lit, the other falls away with nothing filling it | Interiors, night, close-ups that carry mood · skip bright comedy or commercial register | "Light reaches only the left side of his face; the right side falls into shadow and nothing fills it." |
69
+ | `flare-washed-contrast` | receipt | Stray light lifts the blacks and flattens contrast over part of the frame | Shooting toward the sun or a hard practical · skip crisp thriller, neon noir | "Flare washes across the upper half of the frame and lowers the contrast." |
70
+ | `lifted-matte-blacks` | canon | The darkest tones sit at charcoal, like a faded print | Daytime melancholy, period looks, overcast · never with `dense-shadows-with-ramp` | "The darkest parts of the frame are a soft charcoal rather than black." |
71
+ | `uneven-exposure-across-frame` | untested | Exposure is right for one zone only; one side runs hot, the far side goes dark | Mixed-light interiors, night streets, documentary register · skip clean product shots | "The lamp side of the room is overexposed and the far corner goes nearly black." |
72
+
73
+ ### 2. Light source and direction — one physical system
74
+
75
+ | Technique | Evidence | What it does | Reach for · skip | Say |
76
+ |---|---|---|---|---|
77
+ | `name-the-one-source` | receipt | A dominant source with shadows consistent with its distance | When the light needs control · skip when preserving existing light | "The sun is low behind her and off her right shoulder, just above the far peaks." |
78
+ | `forbid-the-phantom-key` | receipt | Removes the soft front light models invent on faces | Any backlit, side-lit or practical-lit person · skip flash, frontal sun, beauty work | "There is no light in front of her: no frontal key, no fill." |
79
+ | `hard-rim-backlight` | receipt | A bright edge on hair and shoulder while the face stays in ambient light | Golden hour, practicals behind the subject · skip overcast and blue hour, which have no hard source | "A hard orange rim of light traces her hair, the edge of her cheek and one shoulder." |
80
+ | `face-in-bounce` | receipt | The unlit side is filled only by coloured light reflected from the surroundings | Backlit exteriors, rooms with coloured walls · skip when the face must read cleanly | "Her face sits in soft, cooler bounce light off the rock." |
81
+ | `raking-side-light` | receipt | Light skims a surface so pores, weave and grain show | Close-ups, skin realism, materials · skip when the brief is flattery | "Daylight rakes across her face from one side, so texture catches along the cheekbone and the other side sits in soft shadow." |
82
+ | `practicals-only` | vendor | Lights inside the frame are the only light: pools with dark gaps between | Night interiors, cars, bars, kitchens at night · skip when the room must read | "The only light is the open fridge she is standing in; the kitchen behind her is dark." |
83
+ | `window-light-falloff` | canon | Bright near the window, then a fast drop across the room | Day interiors, quiet drama · skip big evenly lit spaces | "Grey daylight from the single window on the left; a few steps away the room drops into shadow." |
84
+ | `mixed-colour-temperatures` | receipt | Warm and cool sources side by side, uncorrected | Dusk interiors, night streets, cars at night · skip clean commercial | "Warm lamplight on his face; cold blue daylight from the window on the wall behind him." |
85
+
86
+ ### 3. Atmosphere and optics — what sits between the lens and the subject
87
+
88
+ | Technique | Evidence | What it does | Reach for · skip | Say |
89
+ |---|---|---|---|---|
90
+ | `haze-in-front` | receipt | Haze, dust or smoke between camera and subject softens the subject's edges | Exteriors, sunbeams, dusty sets · skip crisp product or text shots | "Thin haze crosses the frame in front of her, so her edges are no sharper than the rock beside her." |
91
+ | `source-flare` | vendor | Flare from a bright source in the frame | Facing the sun, headlights, stage lights · skip jargon stacks and scenes with no hard source | "The low sun at the edge of the frame throws a flare that crosses in front of her." |
92
+ | `halation-on-highlights` | canon | A soft reddish glow bleeds past bright edges | Night practicals, candles, neon, film looks · skip clean digital register | "Bright lights have a soft reddish glow bleeding past their edges." |
93
+ | `defocus-as-outcome` | receipt | Planes separate: only the subject is sharp | Close and medium shots · skip wides where the place must read | "On a 200mm telephoto lens only she and the pan are sharp; the background melts into soft shapes and the foreground rock edge falls out of focus." |
94
+ | `soft-overall-no-sharpening` | receipt | Lower microcontrast, no halos along edges | Photoreal people, film register · skip product detail and dense text | "The image is slightly soft overall; edges carry no crisp outline." |
95
+ | `grain-in-shadows` | vendor | One texture note | Film or low-light register · never stacked with other noise words | "Fine grain is visible in the shadows." |
96
+ | `edge-falloff` | canon | The corners sit darker than the centre | Film looks, night · skip flat graphic compositions | "The corners of the frame are darker than the centre." |
97
+ | `glass-or-weather-between` | receipt | Rain, smudges or condensation between the lens and the subject | Cars, cafés, storms, observed framing · skip when the face must be sharp | "Seen through a rain-streaked car window; the drops are sharp and she is soft behind them." |
98
+
99
+ ### 4. Colour and grade — preserve the reference unless a change is requested
100
+
101
+ | Technique | Evidence | What it does | Reach for · skip | Say |
102
+ |---|---|---|---|---|
103
+ | `warm-muddy` | receipt | Warm but desaturated; whites read cream, greens olive | Dusty or earthy exteriors, fatigue, nostalgia · skip crisp commercial | "The colour is warm and a little muddy rather than clean." |
104
+ | `restrained-desaturated` | canon | Muted overall; at most one colour stays strong | Drama, cold or bleak moods · skip joyful or brand-colour work | "Colours are muted and low in saturation; only the red of her jacket holds its colour." |
105
+ | `uncorrected-white-balance` | receipt | The colour cast is left in | Phone or documentary register, shade, fluorescents · skip product colour accuracy | "White balance left a little cool and uncorrected; the white mug looks bluish." |
106
+ | `tungsten-in-daylight` | canon | The daylight scene renders blue-cyan | Cold mornings, alienation · skip warm romance | "Daylight renders cold and blue, as if the camera were set for indoor lamps; skin looks pale." |
107
+ | `sodium-vapour-mono` | canon | Street light collapses colour to amber | Urban night, industrial areas, parking lots · skip scenes that need colour separation | "Orange streetlight flattens every colour to amber and brown; shadows go brown-black." |
108
+ | `bleach-bypass-look` | canon | High contrast, drained colour, silvery skin | War, grit, harsh drama · skip warm or intimate scenes | "Colour almost drained out, contrast harsh, skin grey-silver, heavy shadows." |
109
+ | `teal-orange` | vendor | Warm skin against teal shadows; the generic blockbuster look | Only when the brief asks for blockbuster register · never as a default | "Skin stays warm while the shadows and sky are pushed toward teal." |
110
+ | `era-or-device-register` | vendor | One phrase shifts colour, grain and flash together | Period or amateur looks · one per prompt | "As if shot on 1980s colour film, slightly grainy." |
111
+ | `cross-processed` | vendor | Hard colour shifts: cyan shadows, magenta highlights | Music video, fashion, 90s editorial · skip naturalism | "Colours shifted hard, cyan in the shadows and magenta in the highlights." |
112
+
113
+ ### 5. Time and weather presets — bundles of the above, and mutually exclusive
114
+
115
+ | Technique | Evidence | What it does | Reach for · skip | Say |
116
+ |---|---|---|---|---|
117
+ | `golden-hour-backlit` | receipt | Sun low behind, long shadows toward camera, rim light, face in bounce, sky near the sun white, warm haze | Warm exteriors · never with noon shadows or overcast softness | "The sun sits just above the ridge behind her; long shadows run toward the camera; a hard rim on her hair; her face in cool bounce; the sky around the sun burns white." |
118
+ | `after-sunset` | canon | No direct sun, soft shadowless light, sky fading pink to blue, practicals just on | Quiet exteriors · never with rim light or hard shadows | "The sun has just set; soft shadowless light; the sky fades from pink to deep blue; porch lights have just come on." |
119
+ | `blue-hour` | receipt | Cool even light; warm practicals run hot against it; faces dim | Streets and cafés at dusk · never with a warm key on the face | "Deep blue dusk light, almost no shadows; the café windows glow hot orange; her face is dim and blue." |
120
+ | `hard-noon` | practitioner | Bleached sky, short hard shadows, squinting | Heat, desert, exhaustion · skip anything that must flatter | "Midday sun straight overhead; tiny hard shadows under her brows and chin; she squints." |
121
+ | `overcast-flat` | practitioner | No shadows, low contrast, honest skin | Plain daylight, street realism · never with rim light or flare | "Flat grey overcast light, no shadows, low contrast, colours slightly dull." |
122
+ | `night-practicals` | canon | Pools of light, dark gaps, colour casts, blooming highlights | Night streets and interiors · never with "evenly lit" | "The street is dark between the orange streetlights; he is lit only when he passes under one." |
123
+ | `firelight` | canon | Warm flicker from below, fast falloff, black past a few metres | Campfires, candles · never with daylight fill | "The campfire is the only light: warm and flickering from below, their faces half lit, everything beyond the circle black." |
124
+ | `moonlight-day-for-night` | canon | Cool, low saturation, one faint hard shadow direction, sky darker than the ground | Night exteriors · skip when warm practicals dominate | "Cold blue moonlight from one side; faint hard shadows; colour almost gone; faces only just readable." |
125
+ | `rain-wet-night` | canon | Wet surfaces smear reflections; rain shows only where it passes a light | Urban night · rain across the whole frame, never in one corner | "Rain shows as bright streaks where it passes the streetlight; the wet asphalt reflects the lights as long smears." |
126
+ | `fog-or-dust` | canon | Depth falls off to grey; figures layer by distance | Mystery, scale · never with crisp distant detail | "Fog swallows everything past the second streetlight; the far figures are pale grey shapes." |
127
+
128
+ ### 6. Composition and camera placement
129
+
130
+ | Technique | Evidence | What it does | Reach for · skip | Say |
131
+ |---|---|---|---|---|
132
+ | `dirty-foreground-occlusion` | receipt | An out-of-focus object near the lens covers part of the frame | Anything that should feel observed rather than staged · only objects that belong in the space | "The edge of a pine branch crosses the left third of the frame, close to the lens and completely out of focus." |
133
+ | `off-centre-cropped-subject` | practitioner | The subject sits near an edge, partly cut by the frame | Candid and documentary register · skip deliberate symmetry | "She sits in the right third of the frame; her elbow is cut off by the edge." |
134
+ | `negative-space` | canon | A small subject in a large empty area | Isolation, scale · in a video plate give the empty area texture, or it moves with the foreground | "She is a small figure at the bottom right; the rest of the frame is pale, cloud-streaked sky." |
135
+ | `frame-within-frame` | canon | A doorway, window or vehicle surrounds the subject | Observed feel, confinement · skip when it hides the action | "Seen through the open barn door; the dark door frame surrounds her on three sides." |
136
+ | `over-the-shoulder-foreground` | vendor | A foreground head or shoulder, dark and soft | Dialogue, two-person scenes | "The back of his head and shoulder fill the left edge, dark and out of focus; she faces him, sharp." |
137
+ | `compression-as-outcome` | receipt | The long-lens look: the background looms huge and close | Making a background loom, crowds, heat haze · skip intimate interiors | "Shot from far away on a 200mm telephoto lens, the peaks loom huge and close behind her, stacked right up against her shoulders." |
138
+ | `camera-height` | receipt | Ground level, hip height or overhead, picked on purpose | Every shot · eye level only by choice | "The camera sits on the ground by the stove, looking up at her past the pan." |
139
+ | `grabbed-framing` | untested | Horizon slightly off, framing a beat late | Documentary or UGC register · skip composed cinema | "The horizon tilts slightly and the framing is a little late; her head is near the top edge." |
140
+ | `reflection-partial` | receipt | The subject seen in glass, steel or a puddle | Night, cities, variety across a set · never a bathroom mirror | "We see her only as a reflection in the dark shop window, overlapped by the street behind the glass." |
141
+
142
+ ### 7. Texture, wardrobe and set
143
+
144
+ | Technique | Evidence | What it does | Reach for · skip | Say |
145
+ |---|---|---|---|---|
146
+ | `name-the-capture-context` | receipt | Says what kind of picture this is, instead of listing flaws | Every photoreal shot · never a flaw inventory, which reads as tokens and turns plastic | "A frame from a film shot on location, not a studio portrait." |
147
+ | `skin-under-real-conditions` | vendor | Skin carries the environment: wind, sun, sweat | People in real conditions · never a stack of pore and blemish words | "Wind-chapped cheeks and a sunburnt nose after a day on the mountain; skin shiny with sweat at the hairline." |
148
+ | `hair-state` | receipt | Flyaways, wind, strands stuck to the forehead | Weather, action, fatigue · vary the state, never the style that carries identity | "Wind has pulled strands loose across her face." |
149
+ | `wardrobe-wear` | vendor | Creases, dust, fading | Lived-in characters · skip fashion hero shots | "Her jacket is creased at the elbows and faded at the seams; dust on the knees." |
150
+ | `full-wardrobe-spec` | receipt | Names top, legwear and footwear so a reference cannot fill the gap | Any shot built from a character reference · not optional | "Grey hiking jacket zipped up, dark hiking trousers, scuffed brown boots." |
151
+ | `closed-prop-list` | receipt | A positive, closed inventory of what is in the frame | Any frame where the model adds junk · never a "no extra props" list | "The only things on the rock are the stove and the pan." |
152
+ | `lived-in-wear-on-named-things` | practitioner | Wear goes on objects already named, never as new objects | Sets that should feel used · never "add clutter" | "The pan is blackened underneath and the rock around the stove is stained with old soot." |
153
+ | `material-specificity` | vendor | Names the physical material | Hero objects · never a generic noun | "A dented enamel mug, a scratched aluminium pot, a waxed-cotton jacket going pale at the seams." |
154
+
155
+ ### 8. Moment, performance and motion
156
+
157
+ | Technique | Evidence | What it does | Reach for · skip | Say |
158
+ |---|---|---|---|---|
159
+ | `caught-mid-action` | receipt | The action is under way and the camera goes unnoticed | People, by default · skip deliberate portraits and address-the-lens beats | "Her right hand works a wooden spatula in the pan mid-stir." |
160
+ | `eyeline-off-lens` | receipt | The eyes are on the task, another person or out of frame | Images and B-roll · never a lip-sync beat, where eyes on the lens are the point | "She looks down at the pan." |
161
+ | `unresolved-expression` | receipt | Not a stock smile: squinting, chewing, tired, mid-thought | Realism · skip brand-joy beats | "She squints against the glare, jaw set, tired." |
162
+ | `motion-blur-on-the-mover` | practitioner | Only the moving part blurs | Hands, tools, hair, passing vehicles · never on text or a face that must read | "Her hand with the spoon is a slight blur of movement; the pan and the rock are sharp." |
163
+ | `focus-slightly-missed` | untested | Focus landed just behind the subject | Documentary register · skip identity-critical shots | "Focus landed on the rock just behind her; her face is a touch soft." |
164
+ | `handheld-operator-body` | receipt | For video: write the body holding the camera, not the path | Documentary, UGC, tension · skip locked-off formal shots | "Handheld, the operator breathing; the frame sways slightly, drifts off her and corrects back." |
165
+ | `weight-and-consequence` | practitioner | For video: mass moves through the body and the world reacts | Any physical action · skip static dialogue | "She shifts her weight onto her back foot as she lifts the heavy pan; the stove wobbles." |
166
+ | `frame-zero` | receipt | In a still that feeds video, the event has not happened yet | Every image-to-video plate · never an aftermath plate | Describe the moment just before the event: the pole still straight, the glass still whole. |
167
+
168
+ <!-- @catalogue:end -->
169
+
170
+ ## Techniques that clash
171
+
172
+ - **`near-silhouette` against a recognisable face or lip-sync.** On identity-critical shots use `subject-under-key` with `face-in-bounce`; keep silhouettes for wides.
173
+ - **`dense-shadows-with-ramp` against murk.** Always say the light falls off gradually and keep one detail readable in the dark. On video dark work keep the subject readable; `slates-blocking-to-prompt`'s "no crushed blacks" is scoped to that lane.
174
+ - **`lifted-matte-blacks` against `dense-shadows-with-ramp`.** Pick one.
175
+ - **One colour-temperature story.** `warm-muddy` does not go with `tungsten-in-daylight` or a cool `uncorrected-white-balance`.
176
+ - **Presets exclude each other.** Golden hour has no noon shadows; overcast has no rim light or flare; `practicals-only` has no even room light.
177
+ - **Haze or flare against readable text or product.** Keep the atmosphere away from the text, or drop it.
178
+ - **`defocus-as-outcome` against a place that must read.** Keep it for close and medium shots.
179
+ - **Motion blur, handheld sway or missed focus against text, lip-sync or identity.** Keep them off the face and off the text.
180
+ - **`closed-prop-list` against `lived-in-wear-on-named-things`.** Wear goes on the named objects, never as new ones.
181
+
182
+ ## On video
183
+
184
+ - **The grade lives in the still.** A motion prompt describes how the light behaves as things move — the flare slides as the camera turns, haze drifts across the foreground, her face stays in shadow as she turns — and preserves the still's grade unless the user requests a change.
185
+ - **Lens and film-stock names translate; they do not paste.** The rule and ByteDance's own wording: `slates-prompting-seedance` → "Don't cross-pollinate image-model syntax".
186
+ - **A negative-prompt field can cancel a technique.** Never suppress in `negative_prompt` something the prompt asks for (Kling: `slates-prompting-kling-v3`).
187
+
188
+ ## Model notes
189
+
190
+ - **GPT Image 2.5.** Ask for a real photograph or film still outright. In the recorded summit tests, it tended to brighten faces, clean up colour and add props; visible-outcome wording gave more control than gear names alone. References route through its edit endpoint.
191
+ - **Nano Banana 2 and Pro.** Google recommends named cameras, film eras, chiaroscuro lighting and positive framing, so gear names are a sanctioned lever here: still add what they do to the picture. The recorded Nano Banana Pro edits changed more of the frame than intended; scope each requested change and name what should stay.
192
+ - **FLUX.2 Max.** No negative prompt at all, and word order is weight, so put the light system early. Its own examples use crushed shadows, blown highlights and era looks.
193
+ - **Seedream 5 Lite.** Keep it short: pick fewer techniques to fit its prompting guide's roughly 100-word ceiling. See `slates-prompting-seedream-5-lite`.
194
+
195
+ ## Worked examples
196
+
197
+ **Route 1, IMG-A197 (Sunburst, 2026-09-15).** Image 1 is the character identity sheet, image 2 a look reference. One light system, one exposure decision, the wardrobe and the foreground named. Eric: *"just so well done."*
198
+
199
+ <!-- @example:img-a197:start -->
200
+ ```text
201
+ A film still shot on location, lit and graded like image 2. The woman from image 1 cooks on a rocky summit high in the Rockies at golden hour. She stands facing camera behind a stainless steel frying pan of sliced vegetables on a camp stove on the rock in the lower foreground, framed from mid-thigh up. Her right hand works a wooden spatula in the pan mid-stir, her left rests on the pan handle. She looks down at the pan with a slight smile. Long wavy blonde hair down. Grey hiking jacket zipped up, the word "SLATES" once in small plain letters on the left chest, and dark hiking trousers. The only things on the rock are the stove and the pan. Behind her: pine tops, a deep valley and snow-capped peaks running to the horizon.
202
+
203
+ The sun sits just above the far peaks behind her right shoulder and the whole frame is exposed for that sky. She is close to a silhouette: her face and jacket fall into deep shadow, her features only just readable, lit by nothing but a faint warm bounce off the rock. A hard orange rim of light traces her hair, the edge of her cheek and one shoulder. The sky around the sun burns out to white, flare washes across the upper half of the frame and lowers the contrast, and steam off the pan glows where the sun comes through it. The shadows are dense and slightly crushed, the rock in the foreground is nearly black, and the colour is warm and a little muddy rather than clean. No other text in the image.
204
+ ```
205
+ <!-- @example:img-a197:end -->
206
+
207
+ **Route 1 with a lens, IMG-A198 (Sunburst, 2026-09-15).** The same frame, with the lens named and its effect described. Real compression and depth of field came back.
208
+
209
+ <!-- @example:img-a198:start -->
210
+ ```text
211
+ A film still shot on location from far away on a 200mm telephoto lens, lit and graded like image 2. The woman from image 1 cooks on a rocky summit high in the Rockies at golden hour. She stands facing camera behind a stainless steel frying pan of sliced vegetables on a camp stove on the rock in the lower foreground, framed from mid-thigh up. Her right hand works a wooden spatula in the pan mid-stir, her left rests on the pan handle. She looks down at the pan with a slight smile. Long wavy blonde hair down. Grey hiking jacket zipped up, the word "SLATES" once in small plain letters on the left chest, and dark hiking trousers. The only things on the rock are the stove and the pan.
212
+
213
+ The long lens compresses the distance: the snow-capped peaks loom huge and close behind her, stacked right up against her shoulders, and they melt into soft out-of-focus shapes. Only she and the pan are sharp. The pine tops between her and the peaks are a smear of dark green, and the foreground rock edge falls out of focus too.
214
+
215
+ The sun sits just above the far peaks behind her right shoulder and the whole frame is exposed for that sky. She is close to a silhouette: her face and jacket fall into deep shadow, her features only just readable, lit by nothing but a faint warm bounce off the rock. A hard orange rim of light traces her hair, the edge of her cheek and one shoulder. The sky around the sun burns out to white, flare washes across the upper half of the frame and lowers the contrast, and steam off the pan glows where the sun comes through it. The shadows are dense and slightly crushed, the rock in the foreground is nearly black, and the colour is warm and a little muddy rather than clean. No other text in the image.
216
+ ```
217
+ <!-- @example:img-a198:end -->
218
+
219
+ **Route 1, platform revision (IMG-A200 and IMG-A204).** IMG-A199 followed its written hot lamp and crushed blacks, but Eric wanted the reference's flatter, darker, softer grade. This exact revision subsequently produced A200 and A204 using neutral identity sheets and the original look reference. Eric preferred A204 to the scene-reference remixes A201–A203; A203 and A204 used the same model and quality settings. Prompt and references changed together, so their individual contributions are not isolated.
220
+
221
+ <!-- @example:platform-flat-untested:start -->
222
+ ```text
223
+ A film still shot on location from a distance on an 85mm lens, lit and graded like image 2. The young woman from image 1 waits alone at the far end of an empty country train platform at blue hour. She stands side-on to the camera in the right third of the frame, framed from the chest up, looking down the empty track toward where a train would come from, her lips slightly parted. A dark wool coat hangs open over a grey hooded sweatshirt, hood down. The wind has pulled a few strands of hair loose across her cheek. The only things near her are the edge of the concrete platform and a single old lamp post, its lamp not yet switched on.
224
+
225
+ The sun is gone and the light is almost gone with it. The whole frame is underexposed and flat: everything sits in a narrow range of dark, muddy navy blue, nothing in it is bright and nothing is truly black. The brightest thing in the picture is the dull grey-blue sky above the trees. Her face is lit only by that weak sky from behind and to her left, so it is dim and blue and barely separates from the dark trees behind her; her features are there, but you have to look for them. There is no light in front of her, no rim of light on her hair, no catchlight in her eyes, and no contrast anywhere to make her stand out.
226
+
227
+ The long lens compresses the distance: the trees sit close behind her as soft dark shapes, and only her face is in focus, and even that is slightly soft. It looks like a camera pushed to its limit in low light: low contrast, a little murky, fine noise in the dark blues, and colour drained to a cold blue-grey. No text anywhere in the image.
228
+ ```
229
+ <!-- @example:platform-flat-untested:end -->
230
+
231
+ **Route 2, IMG-A184 (Sunburst, 2026-09-09).** Image 1 was a finished blue-hour frame, image 2 a character. The whole prompt: *"take this image 1 and just change the character to the character in image 2"*. Composition, grade, light and depth of field held exactly. That base frame was not one Slates owns, so the receipt proves the technique, not a shippable workflow: use your own frame.
232
+
233
+ ## Adding or changing a technique
234
+
235
+ 1. Record the evidence first: a row in the research doc's technique table, with its tag and source.
236
+ 2. Add or change the row here, with the same id and the same tag.
237
+ 3. Run the build. The catalogue check fails until both files agree.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-cost-discipline
3
- description: Mandatory pre-flight discipline before ANY generation call (image or video) — estimate cost, announce in credits, get confirmation, aggregate batches. Read this every time before calling slates_generate_image or any future slates_generate_* op. Skipping this risks burning the user's credits on guesses.
3
+ description: Mandatory pre-flight discipline before ANY generation call (image or video) — estimate cost, announce in credits, get confirmation, aggregate batches. Use its rules when planning generation; reload only when the rules are needed. Skipping this risks burning the user's credits on guesses.
4
4
  ---
5
5
 
6
6
  # Slates cost discipline — read before every generation
@@ -14,11 +14,11 @@ Generation costs real money. Every call is on the user's credits. The user can't
14
14
  Before ANY `slates_generate_*` call, run `slates_estimate_generation_cost` first. Inputs you must lock before estimating:
15
15
 
16
16
  - **Model** — the id you are about to pass, whatever it is. `slates_estimate_generation_cost` takes the same base ids the generate ops take and resolves the billing key itself; do not build one by hand.
17
- - **Resolution** — never let the op default. Pick deliberately. Drafts → 1k. Hero → 2k. Print → 4k.
17
+ - **Resolution** — use the selected model default unless the user or delivery requires another size. Estimate and generate with the same settings.
18
18
  - **Aspect ratio** — never let the op default to 1:1. Pick from the use case (cinematic → 16:9, mobile vertical → 9:16, square feed → 1:1).
19
19
  - **Count** — explicit. Don't generate 4 when 1 will tell you if the prompt works.
20
20
 
21
- If aspect ratio or resolution isn't obvious from the user's request, **ask before estimating**. Don't guess.
21
+ If the aspect ratio cannot be inferred from the intended delivery, ask. A missing resolution uses the model default; it does not require another question.
22
22
 
23
23
  ### 2. Announce in credits, plainly, before spending
24
24
 
@@ -80,15 +80,21 @@ After each generation completes, the response includes `cost_credits` (when avai
80
80
 
81
81
  ## Resolution decision rules
82
82
 
83
- | Use case | Resolution |
84
- |---|---|
85
- | First draft of a new prompt | 1k |
86
- | Storyboard frame (will likely regenerate) | 1k |
87
- | Hero shot, locked composition | 2k |
88
- | Print, marketing asset, final delivery | 4k |
89
- | Iterating to refine | match the previous resolution |
83
+ Use the model defaults below for ordinary work. For cheap exploration, choose a supported lower setting and compare the quote. For final delivery, use the size the output needs. When comparing prompts, hold model, quality, size and references constant.
84
+
85
+ <!-- @inject:image-defaults -->
86
+ **Image default:** gpt-image-2-5-sunburst, quality `high`, 3k. User overrides take priority. Without a project, generation uses the headless Nano Banana 2 seat.
90
87
 
91
- Resolution is a price lever, not a free choice: on Nano Banana 2 and FLUX.2 Max, 4k costs roughly 2x 1k (Seedream 5 Lite is flat-priced regardless of resolution). Prices change — call `slates_estimate_generation_cost` or `slates_list_available_models` for current numbers instead of assuming. Pick the cheapest resolution that serves the use case.
88
+ | Model | Default resolution |
89
+ |---|---|
90
+ | nano-banana-2 | 2k |
91
+ | nano-banana-2-lite | 1k |
92
+ | nano-banana-pro | 2k |
93
+ | gpt-image-2-5-flare | 2k |
94
+ | gpt-image-2-5-sunburst | 3k |
95
+ | flux-2-max | 1k |
96
+ | seedream-5-lite | 2k |
97
+ <!-- @end:image-defaults -->
92
98
 
93
99
  **4K VIDEO is Pro-only (2026-07-07).** The ladder above is for IMAGES (open at every tier). For VIDEO — Kling, Seedance, Veo — 4K requires a Slates Pro account; a base-tier 4K video gen is rejected server-side with `PRO_REQUIRED`. Default video to 1080p or lower and only reach for 4K when the user is on Pro and explicitly asks. 4K *images* are never gated.
94
100
 
@@ -109,7 +115,7 @@ If the user prompt mixes signals (e.g. "cinematic Instagram post"), ask. Don't g
109
115
 
110
116
  ## When the gate fires
111
117
 
112
- The server returns `requires_clarification` when aspect ratio or resolution is missing, and `requires_confirm` when total spend crosses the gate above. In both cases:
118
+ The server returns `requires_clarification` when required composition inputs are missing, and `requires_confirm` when total spend crosses the gate above. In both cases:
113
119
 
114
120
  1. Surface the gate response to the user
115
121
  2. Get a clean answer
@@ -1,68 +1,28 @@
1
1
  ---
2
2
  name: slates-direct-response-ad
3
- description: Build a 30-second hyper-motion direct-response ad in Slates from a product image and brief. Composes upload → storyboard → frame gen → motion gen → timeline → export. Use when the user drops a product image and asks for "an ad", "a promo video", "a TikTok ad", "an Instagram ad", a launch video, or any short-form direct-response video built around a product.
3
+ description: Develop a product-led direct-response ad whose demonstration, argument and next action serve a supplied offer. Use for product-led creative direction; presenter performance belongs in slates-ugc-influencer-ad and general writing in slates-script-craft.
4
4
  ---
5
5
 
6
- # Direct-response ad — Slates workflow
6
+ # Product-led direct response
7
7
 
8
- 🚨 **Wrong skill if a PERSON talks to camera.** This file builds a product-led hyper-motion spot — product hero frames, punchy cuts, no presenter. **A creator-style ad where a synthetic person speaks to the lens is a different discipline with an inverted rulebook (ugly on purpose, one shot per generation, plate-before-video): read `slates-ugc-influencer-ad`.** Route on the presence of a talking person, not on the platform.
8
+ Use `slates-script-craft` for the words, evidence and hook/bridge variations. This guide supplies an optional product-led approach. A person can appear; a presenter, fixed duration, fixed frame count or hyper-motion treatment is not required.
9
9
 
10
- You are building a 30-second hyper-motion direct-response ad. The user has handed you a product image (or product URL) and a short brief. Slates desktop is open on the second monitor; the user watches it populate as you work.
10
+ ## Choose the product's job in the picture
11
11
 
12
- **Hard rules**
12
+ Read the brief and references. Identify the intended audience situation, the supplied claim, what can demonstrate it, and the actual next action. A close product detail, use in context, a visible problem/solution and a final product view are possible ingredients. Omit or reorder any that do not serve the piece.
13
13
 
14
- - Always estimate cost before generating. Use `slates_estimate_generation_cost` and surface the total.
15
- - All Slates generation routes through Slates Credits, period (BYOK is retired) — don't suggest "use your own keys" workarounds.
16
- - Default model: `nano-banana-2-2k`. For close-up product hero frames step up to `4k` only if the user asks.
17
- - Hyper-motion = punchy cuts, 4 frames in 30 seconds, ~7s each. Don't over-storyboard.
14
+ For example, a keys tray can interrupt a sliding key, show where it lands, then hold on the result. A quiet demonstration may work without narration. A technical product may need an explanatory exchange. Do not invent product finishes, guarantees or quantified benefits from a reference picture.
18
15
 
19
- ## Workflow
16
+ ## Keep the work editable
20
17
 
21
- ### 1. Set up the project
22
- - Create a project named for the product (`slates_create_project`).
23
- - If the user gave a product image as a file path, upload it (`slates_upload_reference_image`).
24
- - If they pasted base64 / a data URL, use the same op with `dataUrl`.
18
+ Use the current project unless another destination is requested. Save the script as document text, with non-spoken direction separate. Create shots only for the requested production units and file them into explicit scenes. Alternative openings live beside the shared body as saved versions of a section. Preserve manual edits and reference identity.
25
19
 
26
- ### 2. Generate the storyboard frames
27
- Build exactly 4 frames in this order:
20
+ Read the composed requests before any media call. Model selection, reference slots, durations and resolution come from current capabilities and `slates-model-selection`, never a copied ad recipe. Every prompt stays visible on its shot. A preset imports ordinary editable content and does not authorize generation.
28
21
 
29
- | # | Beat | Visual goal |
30
- |---|------|-------------|
31
- | 1 | Hook | Hyper-close-up of the product, dramatic light, motion blur edge |
32
- | 2 | Lifestyle | Real person using/wearing/holding the product, eye contact |
33
- | 3 | Problem→solution | The before/after moment that justifies the buy |
34
- | 4 | CTA | Clean product hero with mental room for an overlaid CTA in editing |
22
+ ## Production when requested
35
23
 
36
- For each frame:
37
- 1. Draft a tight 1-2 sentence prompt (visual only — no copy text in the image).
38
- 2. Reference the product upload's URL or asset ID for visual fidelity.
39
- 3. Call `slates_generate_image` with that prompt + reference. **You see the result inline — evaluate it.**
40
- 4. If it's wrong: refine prompt, regenerate. If it's right: bind it as a frame in the storyboard (`slates_add_frame`).
24
+ Follow `slates-cost-discipline` and the user's authorization for the exact selected requests. Estimate the set through the existing quote operation. An extra stochastic take remains an extra requested take; compatible existing media can be deliberately reused.
41
25
 
42
- ### 3. Build the storyboard
43
- - `slates_create_storyboard` named "30s ad — v1".
44
- - Default scene already exists. Add 3 more scenes ("Hook", "Lifestyle", "Problem-Solution", "CTA") via `slates_add_scene`, or just add all 4 frames to the default scene.
45
- - For each generated image, add a frame referencing the asset id (`slates_add_frame`).
26
+ Inspect results against the demonstration and supplied references. A failed job is not authorization for an unchanged reroll. Keep earlier takes accessible. Assemble into an explicitly named cut when comparing variations, inspect playback, and export the selected cuts with their results. `slates-one-prompt-film` covers this delivery task when a finished video is requested.
46
27
 
47
- ### 4. Hand back to the user
48
- - Surface estimated total credits spent.
49
- - Tell the user the storyboard is ready and they can either:
50
- - **In Slates desktop:** click each frame to generate motion (the existing UI handles motion generation).
51
- - **Continue here:** ask you to keep going.
52
-
53
- ### 5. If they say keep going — motion, assembly, export
54
- - Generate motion per frame with `slates_generate_video` (`firstFrameAssetId` = the frame's asset, `background: true`), routed per `slates-model-selection` (Kling 3.0 std 8s by default; Seedance 2 for any physics-heavy beat like the hook). Submit all four, then poll `slates_get_generation_status` until each completes (1-5 min).
55
- - Assemble: `slates_add_clip_to_timeline` for each completed clip in beat order (Hook → Lifestyle → Problem-Solution → CTA). Verify with `slates_get_timeline`; fix order with `slates_reorder_clips`.
56
- - Export: `slates_export_video` to an absolute `.mp4` path (default `<slates_get_project_directory>/exports/<product>-ad.mp4`), then `slates_reveal_file` so the user sees the file.
57
- - Full pipeline doctrine (batch cost authorization, model mixing, multi-take selection): `slates-one-prompt-film`.
58
-
59
- ## Anti-patterns
60
-
61
- - **Don't** generate text overlays in the image. Slates renders captions/CTAs at the editor stage.
62
- - **Don't** burn credits on slot-machine prompting. If the first generation is off, refine the prompt; don't just regenerate.
63
- - **Don't** skip the cost estimate. Confirm with the user above ~17 credits.
64
- - **Don't** invent visual specifics about the product (colors, textures, angles) that aren't in the reference image. Reference-anchored prompts only.
65
-
66
- ## Voice
67
-
68
- The ad lives or dies on the hook frame. Tight, sensory, no fluff. Match the user's brand. Default tone is "scroll-stopping" not "informative."
28
+ No conversion outcome is promised by this format. Creative clarity, observed distribution and measured purchases are different evidence.
@@ -42,7 +42,7 @@ The user's request is one of:
42
42
  | Aesthetic / compositional | `slates_generate_image` with the original in `referenceAssetIds` + a refined prompt. Don't re-roll from scratch. |
43
43
  | Wholesale | New prompt, no reference, fresh generation. Treat as a new brief. |
44
44
 
45
- **`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Default model `nano-banana-2` — the only edit model that also takes extra `referenceAssetIds`; `flux-2-max` / `seedream-5-lite` use their own edit endpoints and ignore references. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost above ~17 credits gates on `confirm=true`.
45
+ **`slates_edit_image` shape:** `projectId` + `sourceAssetId` + `prompt` (the edit instruction). Omit `editModel` for the app's Edit seat (the default image model). Every edit model also takes `referenceAssetIds`, up to its reference cap less one: the source is image 1. The result lands as a NEW asset (prompt prefixed `[Edit]`); the source is untouched. Cost above ~17 credits gates on `confirm=true`.
46
46
 
47
47
  ### 4. Generate, evaluate, decide
48
48
  - Estimate cost first.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-model-selection
3
- description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Seedance 2.5 is the DEFAULT video model (Eric, 2026-09-13: the best in the world — physics, effects, scale, 30s takes, 30 references, timestamps); Seedance 2.0 is the 4K seat and the cheaper one at every shared resolution; Kling 3.0 is the cost-effective seat for performances, start-frame animation and lip-sync; MiniMax H3 is the AUTHORED-AUDIO seat (three directable sound layers in one pass, declared reference relationships, 480p-4K) with MiniMax H3 Max beside it as a faster 768p-capped premium with omni-references; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 or 9:16, 4/6/8s) and never the default.
3
+ description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Seedance 2.5 is the DEFAULT video model (Eric, 2026-09-13: the best in the world — physics, effects, scale, 30s takes, 30 references, timestamps); Seedance 2.0 is the 4K seat and the cheaper one at every shared resolution; Kling 3.0 is the cost-effective seat for performances, start-frame animation and lip-sync; MiniMax H3 is the AUTHORED-AUDIO seat (three directable sound layers in one pass, declared reference relationships, 480p-4K) with MiniMax H3 Max beside it as a faster premium with omni-references and MiniMax H3 Max Turbo as its half-price, frames-only sibling; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 or 9:16, 4/6/8s) and never the default.
4
4
  ---
5
5
 
6
6
  # Model selection — the routing doctrine
@@ -28,10 +28,11 @@ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni F
28
28
  | Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |
29
29
  | Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
30
30
  | **4K delivery**, or the same resolution cheaper than 2.5 | **Seedance 2.0** | The only Seedance with native 4K (4K video is Pro-only) and cheaper than 2.5 at every shared resolution (720p $0.15/s vs $0.231/s). Same physics and effects strengths; 15s takes, 15 references, no timestamps. |
31
- | **One take longer than 15 seconds**, more than 15 references, an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | Only 2.5 does these (rules in `slates-prompting-seedance-2-5` § Timestamps); it is the default anyway. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) LENGTH is the price dial, not resolution — a 30s 720p face gen is 489 credits and a 30s 1080p faceless gen is 614, against a 1,000-credit welcome grant. Quote before any take over ~10s. |
31
+ | **One take longer than 15 seconds**, more than 15 references, an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | Only 2.5 does these (rules in `slates-prompting-seedance-2-5` § Timestamps); it is the default anyway. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) LENGTH is the price dial, not resolution — a 30s 720p face gen is 489 credits and a 30s 1080p faceless gen is 853, against a 1,000-credit welcome grant. Quote before any take over ~10s. |
32
32
  | **The SOUND has to be directed, not just present** — a specific line delivered a specific way, scene sound that has to sit under it, and score that must stay out of the characters' world | **MiniMax H3** | The only seat where audio is authored in three separate layers in ONE pass (synchronised events in the body, ambience in a soundscape section, audience-only score in its own) rather than toggled on. 5–15s, 480p / 768p / 2K / 4K, 24fps, 32kHz stereo, 11 languages. Rules in `slates-prompting-minimax-h3`. |
33
33
  | **A reference has to keep a DECLARED amount of itself** — especially moving one subject's characteristic onto a *different* subject | **MiniMax H3** | The only seat that understands a stated retention relationship (kept whole / kept in part / transferred onto another subject / loose echo). 9 images + 3 video + 3 audio, 12 files total. 🚨 The first 5 reference images are free and every one after that costs 4 credits — pass `referenceImages` to `slates_estimate_generation_cost` before a reference-heavy job. |
34
- | **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p/768p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at 768p. It still animates a start frame and an end frame — image-to-video is one of the two things it is for — and since 2026-09-09 it takes the full omni-reference set too (9 images + 3 video + 3 audio), so the seats now differ on ladder and price rather than on what they accept. Never the default; never reach for it to save money. |
34
+ | **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p to 1080p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at a 1080p refinement of its 768p render. It still animates a start frame and an end frame — image-to-video is one of the two things it is for — and since 2026-09-09 it takes the full omni-reference set too (9 images + 3 video + 3 audio), so the seats now differ on ladder and price rather than on what they accept. Never the default; never reach for it to save money. |
35
+ | **Drafts and volume** on a text-to-video or start-frame shot, where the credit budget binds and no reference is needed | **MiniMax H3 Max Turbo** | A second fal post-train of H3 with Max's ladder at **half Max's rate at every tier** ($0.040/s at 768p). It takes a start frame and an end frame but has **no reference endpoint**: a shot that needs references goes to H3 Max or base H3. Its 1080p, like Max's, is a refinement of the native 768p render. Re-run the keeper on a hero seat. |
35
36
  | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | Narrow, and now narrower: if the sound needs DIRECTING rather than merely existing, MiniMax H3 is the better seat. |
36
37
 
37
38
  ### Named Seedance escalation triggers
@@ -52,7 +53,7 @@ Concrete beats route better than an abstract category. Cost stays a tiebreaker,
52
53
  | **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + "Keep everything else the same"; long prompts destroy it (see below). |
53
54
  | **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
54
55
  | **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
55
- | Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. (2.5's relocate lane reaches 1080p too as of 2026-08-24, at $0.2457/s of combined input+output.) Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
56
+ | Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. (2.5's relocate lane reaches 1080p too as of 2026-08-24, at $0.3412/s of combined input+output.) Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
56
57
  | **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p/1080p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |
57
58
  | AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
58
59
 
@@ -72,7 +73,7 @@ Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that
72
73
 
73
74
  **Want the Seedance version of either?** It is not a switch on these tools — it is a normal `slates_generate_video` on `seedance-2` with the clip attached as a **video reference** and the motion or dialogue written into the prompt ("the character from image 1 performs the exact motion from video 1"). That routes to the same endpoint the tool would have called, with the prompt visible and editable instead of ghost-written. Single-pass conditioning genuinely beats post-hoc retargeting on fast choreography, contact, cloth and hair — and it carries native audio — so escalate there whenever fidelity matters.
74
75
 
75
- - Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane.
76
+ - Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
76
77
  - Faces on that route go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person (premium realface pricing).
77
78
 
78
79
  **Rules:**
@@ -88,18 +89,23 @@ Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that
88
89
 
89
90
  **Video models (Kling, Seedance, Veo) cannot generate standalone images — ever.** A "premium hero reference image" is still an image job: it routes to an image model below, never to Seedance.
90
91
 
91
- - **Default: Nano Banana 2** — strongest reference HANDLING (14 refs; GPT Image now takes more, at 16, but Banana is still the one that holds many subjects coherently), best legible text, the standard start-frame generator.
92
- - **NB2 Lite** — the fast/draft seat: ~half NB2's price, ~2.7× faster, 1K only. Route iteration volume and drafts here; finals go back to NB2 full (2K/4K).
93
- - **Nano Banana Pro** — the hero-frame/typography ceiling (~2× NB2). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — feed it a full subject library.
94
- - **GPT Image 2.5** — two seats, `gpt-image-2-5-flare` and `gpt-image-2-5-sunburst`, **same price**. Readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. **Also the photoreal front-runner (Eric, 2026-08-24)** — it beat both Nano Banana rails head-to-head on skin realism, which is why the AI-influencer ad lane generates every plate on this line. **The seat split is SPEED vs QUALITY, not generate vs edit** (OpenAI's own rule): Flare is the small, fast model with quality *comparable to* GPT Image 2 — drafts, exploration, volume; Sunburst is OpenAI's *most capable* image model, higher quality than GPT Image 2, deliberately slower — finals, hero frames, photoreal, and multi-reference edits, where its lead is widest. **Explore on Flare, finish on Sunburst.** Five quality tiers, cheapest first — `low` (layout checks only), `medium` (drafts), **`high` (the default)**, `xhigh`, `max` (the top). Uneven: `max` is 4× `high`, `xhigh` only ~1.8× it. **16 reference images**, the schema ceiling. **Transparent backgrounds** via `backgroundMode` — free, and the only image family that offers them.
92
+ <!-- @inject:image-defaults -->
93
+ **Image default:** gpt-image-2-5-sunburst, quality `high`, 3k. User overrides take priority. Without a project, generation uses the headless Nano Banana 2 seat.
95
94
 
96
- 🚨 **The tier names moved when 2.5 replaced GPT Image 2, and the strings did not.** GPT Image 2's `medium` is 2.5's `high`; its `high` is 2.5's `max` — same money, one rung of renaming. The 2026-08-24 photoreal result was measured at GPT Image 2 `high`, so **the tier that reproduces it is `max`**. Nobody has re-run it on 2.5; the ranking is inherited, not re-measured.
97
- - **FLUX.2 Max** — photoreal texture, hex-color binding, typography, less censored.
98
- - **Seedream 5 Lite** — uncensored + any-resolution flat price; volume exploration when the Gemini filter is in the way.
95
+ | Model | Default resolution |
96
+ |---|---|
97
+ | nano-banana-2 | 2k |
98
+ | nano-banana-2-lite | 1k |
99
+ | nano-banana-pro | 2k |
100
+ | gpt-image-2-5-flare | 2k |
101
+ | gpt-image-2-5-sunburst | 3k |
102
+ | flux-2-max | 1k |
103
+ | seedream-5-lite | 2k |
104
+ <!-- @end:image-defaults -->
99
105
 
100
- **Split rule of thumb:** readable text / panels / UI → GPT Image 2.5 (Flare to explore, Sunburst to finish); **photoreal people, finals and hero frames → Sunburst at `max`** — the 2026-08-24 result was measured at GPT Image 2's `high`, which is `max` here, and Flare only *matches* GPT Image 2 while Sunburst exceeds it; multi-reference edits where several references must all survive into one frame → Sunburst; edit-heavy work → the Banana line; drafts → GPT Image 2.5 Flare at `medium`, which now undercuts NB2 Lite on both price and resolution; uncensored or odd resolutions → Seedream/FLUX.
106
+ Use `slates_estimate_generation_cost` for the selected model's current price and craft card. Routing reasons live in the model facts returned by `slates_list_available_models`; use the model's guide for its particular strengths and limits. Choose a different seat when the brief supplies a reason, such as speed, supported output shape, or an edit that failed on the default.
101
107
 
102
- ⚠️ **This line said the opposite until 2026-08-24** — it sent photoreal *away* from GPT Image on reputation, which is the exact failure § The meta-rule above warns about. Re-run the evidence test when the roster moves. It moved again on 2026-09-09, and the ranking was carried across rather than re-measured — exactly what the meta-rule says not to trust. Treat it as a starting hypothesis for 2.5, not a receipt. **The seat choice above is likewise reasoned from OpenAI's positioning, not measured:** run Flare-`max` against Sunburst-`max` on one plate and write the answer into `slates-prompting-gpt-image-2-5`.
108
+ **Historical photoreal receipt:** the 2026-08-24 comparison favored GPT Image 2 on one skin-realism task at its old high tier. That is evidence about that comparison, not proof that 2.5 requires its most expensive tier. Raise quality only to address a specific observed shortfall and compare at the delivery crop.
103
109
 
104
110
  ## Audio routing
105
111