@slatesvideo/shared 0.6.11 → 0.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (83) hide show
  1. package/dist/auth.js +2 -2
  2. package/dist/clients/cloud.js +1 -1
  3. package/dist/index.d.ts +1 -1
  4. package/dist/index.js +1 -1
  5. package/dist/manual/content.d.ts +1 -1
  6. package/dist/manual/content.js +1 -1
  7. package/dist/operations/index.d.ts +817 -16
  8. package/dist/operations/index.js +1413 -360
  9. package/dist/operations/surface.d.ts +4 -1
  10. package/dist/operations/surface.js +41 -10
  11. package/dist/prompts/ad-presets.d.ts +77 -0
  12. package/dist/prompts/ad-presets.js +43 -0
  13. package/dist/prompts/agent-doctrine.js +27 -5
  14. package/dist/prompts/banned-tokens.d.ts +4 -29
  15. package/dist/prompts/banned-tokens.js +29 -204
  16. package/dist/prompts/craft-cards.js +2 -2
  17. package/dist/prompts/generation-policy.d.ts +41 -0
  18. package/dist/prompts/generation-policy.js +53 -0
  19. package/dist/prompts/guide-retrieval.d.ts +9 -0
  20. package/dist/prompts/guide-retrieval.js +53 -0
  21. package/dist/prompts/index.d.ts +1 -0
  22. package/dist/prompts/index.js +1 -0
  23. package/dist/prompts/model-capabilities.d.ts +18 -1
  24. package/dist/prompts/model-capabilities.js +72 -19
  25. package/dist/prompts/model-facts.d.ts +34 -2
  26. package/dist/prompts/model-facts.js +66 -5
  27. package/dist/prompts/partials.generated.js +8 -2
  28. package/dist/prompts/prompting-tips.d.ts +1 -1
  29. package/dist/prompts/prompting-tips.js +61 -16
  30. package/dist/prompts/reference-composer.d.ts +2 -0
  31. package/dist/prompts/reference-composer.js +51 -50
  32. package/dist/prompts/script-document.d.ts +165 -0
  33. package/dist/prompts/script-document.js +11 -0
  34. package/dist/prompts/shot-grammar.d.ts +4 -4
  35. package/dist/prompts/shot-grammar.js +3 -3
  36. package/dist/prompts/shot-spec.d.ts +13 -0
  37. package/dist/prompts/shot-spec.js +23 -5
  38. package/dist/skills/content.js +27 -24
  39. package/exports/slates-chatgpt-images/generated/SKILL.md +107 -0
  40. package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
  41. package/exports/slates-prompt-builder/generated/SKILL.md +1 -1
  42. package/exports/slates-prompt-builder/generated/reference-character.md +9 -1
  43. package/exports/slates-prompt-builder/generated/reference-kling.md +3 -3
  44. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +22 -10
  45. package/exports/slates-prompt-builder/generated/reference-seedance.md +4 -4
  46. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +17 -17
  47. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  48. package/package.json +9 -3
  49. package/skills/_partials/cinematic-card.md +8 -0
  50. package/skills/_partials/cinematic-routes-short.md +2 -0
  51. package/skills/_partials/cinematic-tips-short.md +2 -0
  52. package/skills/_partials/decision-log.md +1 -13
  53. package/skills/_partials/image-defaults.md +11 -0
  54. package/skills/_partials/lens-video-split.md +1 -0
  55. package/skills/_partials/reference-rules-core.md +1 -1
  56. package/skills/_partials/sheet-tool-defaults.md +6 -0
  57. package/skills/slates-character-identity.md +9 -1
  58. package/skills/slates-chatgpt-images.md +107 -0
  59. package/skills/slates-cinematic-look.md +237 -0
  60. package/skills/slates-cost-discipline.md +18 -12
  61. package/skills/slates-direct-response-ad.md +13 -53
  62. package/skills/slates-edit-and-iterate.md +1 -1
  63. package/skills/slates-model-selection.md +20 -14
  64. package/skills/slates-one-prompt-film.md +19 -77
  65. package/skills/slates-project-organization.md +7 -3
  66. package/skills/slates-prompting-flux-2-max.md +15 -4
  67. package/skills/slates-prompting-gpt-image-2-5.md +41 -28
  68. package/skills/slates-prompting-inworld-tts.md +174 -174
  69. package/skills/slates-prompting-kling-v3.md +3 -3
  70. package/skills/slates-prompting-lip-sync.md +1 -1
  71. package/skills/slates-prompting-minimax-h3.md +30 -17
  72. package/skills/slates-prompting-motion-transfer.md +1 -1
  73. package/skills/slates-prompting-nano-banana-2.md +24 -11
  74. package/skills/slates-prompting-seedance-2-5.md +7 -6
  75. package/skills/slates-prompting-seedance.md +5 -5
  76. package/skills/slates-prompting-seedream-5-lite.md +14 -3
  77. package/skills/slates-prompting-veo-3.md +1 -1
  78. package/skills/slates-script-craft.md +45 -0
  79. package/skills/slates-shot-variety.md +11 -40
  80. package/skills/slates-storyboard-from-script.md +14 -66
  81. package/skills/slates-style-prompting.md +4 -4
  82. package/skills/slates-ugc-influencer-ad.md +32 -309
  83. package/skills/slates-vision-feedback-loop.md +2 -1
@@ -1,174 +1,174 @@
1
- ---
2
- name: slates-prompting-inworld-tts
3
- description: How to use Inworld Realtime TTS-2, the VOICE seat. Read before calling slates_generate_audio with model inworld-tts-2. Speech in a SPECIFIC voice, billed per character - the prompt is the words spoken, verbatim. Covers the identity-versus-acoustics rule (what a reference clip does and does not carry), how to write a line so it is performed rather than read, when to reach for seed-audio instead, and the voice-consent rule.
4
- ---
5
-
6
- # Inworld Realtime TTS-2 — the voice seat
7
-
8
- <!-- @card:start -->
9
- <!-- slates-only -->
10
- <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
- src/prompts/craft-cards.ts and returned on every cost estimate for this
12
- model, so it is the ONE piece of positive craft guidance the agent cannot
13
- skip. Keep it under 2,400 characters (the build fails above that) and keep
14
- the rationale and the worked examples in the body below. -->
15
- <!-- /slates-only -->
16
- **Card — Inworld TTS-2.** Speech in a SPECIFIC voice. The prompt is the words spoken, verbatim — not a description of them. Text length determines the bill.
17
-
18
- **IDENTITY, NOT ACOUSTICS — the rule that decides whether cloning works**
19
- A reference carries WHO is speaking: timbre, pitch, accent, age, vowel shape. It does NOT carry WHERE they are — room tone, distance, phone EQ, reverb and mic character are *acoustics*, and this model reproduces the identity while discarding the room. So:
20
- 1. **A noisy reference does not give a noisy read — it gives a WORSE identity.** Music, a second speaker or heavy reverb corrupt what is being extracted. Use a clean single-speaker recording.
21
- 2. **You cannot get "on a payphone" by cloning a payphone recording.** Acoustics come from the MIX, or from `seed-audio` which renders a room.
22
-
23
- **DIRECTION GOES IN SQUARE BRACKETS. PARENTHESES ARE SPOKEN ALOUD.** `[whispering] I hope nobody notices` is whispered; `(quietly) I hope nobody notices` says the word "quietly" out loud. Verified by ear — the easiest way to ruin a take.
24
-
25
- - **Plain English works inside them** — it is natural-language steering, not a fixed vocabulary: `[very quiet]`, `[whisper in a hushed style]`, `[very slow]`, `[say excitedly]`. Non-verbals are their own tags: `[laugh]`, `[sigh]`, `[breathe]`, `[clear throat]`.
26
- - **A tag it does not recognise is still consumed, and still changes the read.** Never spoken, never an error — so a mistyped tag fails SILENTLY and only listening catches it.
27
- - **Tags persist across sentences** until changed; `[reset]` returns to normal.
28
- - **Punctuation is the timing.** `Wait. Stop.` differs from `Wait, stop.`
29
- - **One line, one take.** Split a paragraph so a bad clause costs one re-roll.
30
- - **Spell numbers and titles aloud:** `twenty twenty-six`, `Doctor Reyes`.
31
-
32
- **Route elsewhere when:** the scene needs dialogue mixed with effects and room tone in one pass (`seed-audio`), or it is a single non-speech sound (`eleven-sfx`). This surface makes ONE voice saying ONE thing, cleanly.
33
-
34
- **Hard constraints:** no duration parameter — length falls out of the text. Exactly one voice source: a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId` (a character's voice clip to speak AS the character, or any clean clip of one speaker), or `voiceDescription`.
35
- <!-- @card:end -->
36
-
37
- <!-- @banned:start -->
38
- <!-- slates-only -->
39
- <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
40
- extracted by src/prompts/banned-tokens.ts and returned on this model's cost
41
- estimate, and every submitted prompt is matched against it. Keep entries
42
- backticked and prose outside the backticks. -->
43
- <!-- /slates-only -->
44
- **Never use** — the prompt on this surface is SPOKEN ALOUD, so anything that describes the audio instead of being the audio gets read out as words:
45
-
46
- - `SFX`, `Ambient noise`, `Background music` as labels — this model speaks; it does not render a scene. Use `seed-audio` for those.
47
- - shot language: `wide shot`, `slow push in`, `warm tungsten` — video-prompt words, and here they would literally be said aloud
48
- - `voiceover`, `narrator says`, `he says` as stage directions wrapping the line — write only the words that should come out of the speaker
49
- <!-- @banned:end -->
50
-
51
- ## Why the prompt is not a prompt
52
-
53
- On every other surface in Slates the prompt DESCRIBES what you want and the model interprets it. Here the prompt IS the deliverable: each character is spoken aloud and each character is billed. `a gravelly man says he is tired` produces a voice saying the words "a gravelly man says he is tired".
54
-
55
- That also means the two numbers a user cares about are the same number. The text length sets the price (in 250-character buckets) and sets the length of the audio. There is nothing to choose and nothing to reconcile.
56
-
57
- ## Steering the delivery
58
-
59
- `VERIFIED BY EAR, 2026-09-05.` Every claim in this section was listened to, not
60
- inferred — an earlier draft of this skill documented tag forms that had only been
61
- probed for an HTTP 200, which proves the request was accepted and nothing about
62
- whether it was obeyed.
63
-
64
- **Square brackets are consumed. Parentheses are read aloud.** That is the whole
65
- rule, and getting it wrong is not a subtle degradation — the audience hears a
66
- narrator say the word "quietly" in the middle of your line.
67
-
68
- | Written | What comes out |
69
- |---|---|
70
- | `[whispering] I really hope nobody notices that.` | whispered, tag not spoken ✅ |
71
- | `[very quiet] I really hope nobody notices that.` | very quiet, tag not spoken ✅ |
72
- | `[whisper in a hushed style] …` | hushed, tag not spoken ✅ |
73
- | `[very slow] …` | slowed right down, tag not spoken ✅ |
74
- | `[laugh] …` | an actual laugh, then the line ✅ |
75
- | `(quietly, under his breath) …` | 🚨 **the words "quietly, under his breath" are SPOKEN** |
76
-
77
- **Plain English works — it is natural-language steering, not a fixed vocabulary.**
78
- Both the documented phrasings (`[whisper in a hushed style]`) and ordinary adverbs
79
- (`[whispering]`) were obeyed. Write the direction the way you would say it to an
80
- actor.
81
-
82
- The eight dimensions the model steers on, with a working example of each:
83
-
84
- | Dimension | Example |
85
- |---|---|
86
- | Emotion | `[say excitedly]`, `[sound sad]`, `[sound terrified]` |
87
- | Articulation | `[say with force]`, `[articulate clearly]` |
88
- | Intonation | `[say with a rising pitch]` |
89
- | Volume | `[very quiet]`, `[very loud]` |
90
- | Pitch | `[say in a low tone]` |
91
- | Range | `[say playfully]`, `[say with no pitch variation]` |
92
- | Speed | `[very fast]`, `[very slow]` |
93
- | Vocal style | `[whisper in a hushed style]`, `[give a nasal quality]` |
94
-
95
- Non-verbals sit inline where they happen: `[laugh]`, `[sigh]`, `[cough]`,
96
- `[breathe]`, `[yawn]`, `[clear throat]`.
97
-
98
- ### Four rules that are not obvious
99
-
100
- 1. 🚨 **A tag it does not recognise is still consumed, and still changes the read.**
101
- `[zzzqqq]` is not spoken and does not error — it produces a different, arbitrary
102
- delivery. So a typo in a tag is SILENT: there is no rejection, no warning, and no
103
- way to catch it except listening to the take. Treat an unexpected performance as
104
- a possible misspelled tag before you blame the voice.
105
- 2. **Tags persist across sentences.** A `[very slow]` at the top governs everything
106
- after it until something changes it. Use `[reset]` to go back to normal rather
107
- than assuming the next sentence starts clean.
108
- 3. **Do not stack opposing directions.** `[whisper in a hushed style]` together with
109
- `[very loud]` produces unpredictable results — the model is resolving a
110
- contradiction, and which side wins is not something you can rely on.
111
- 4. **Tags COUNT toward the billed characters**, even though they are never spoken.
112
- They are part of the text sent to the vendor, so the vendor charges for them and
113
- so do we — billing what was actually sent is the only honest basis. It rarely
114
- matters (a 13-character tag inside a 250-character bucket), but a line sitting
115
- just under a bucket boundary can be pushed into the next one by a long
116
- direction. Prefer `[very slow]` over `[say this one very slowly please]`.
117
-
118
- ## Identity versus acoustics, at length
119
-
120
- This is the distinction that decides whether the feature feels good, and it is worth being precise about because the failure is quiet — you get a usable clip that is subtly not the person.
121
-
122
- **What a reference clip transfers:** vocal timbre, pitch range, accent and regional vowels, apparent age, speech rate tendencies, and the particular rasp or breathiness of the source speaker.
123
-
124
- **What it does not transfer:** the room, the microphone, the codec, the distance from the mic, any processing on the source, and any other sound present in it.
125
-
126
- So the ideal reference is boring: one person, close to a microphone, no music, no second speaker, no heavy reverb, five to fifteen seconds, speaking normally rather than performing. A phone voice memo in a quiet room beats a beautifully produced clip with a music bed underneath it.
127
-
128
- **Two failure modes, both common:**
129
-
130
- - *"I cloned my podcast intro and it doesn't sound like me."* The intro had music under it. The model averaged the music into the identity. Re-clone from a clean stretch.
131
- - *"I want the line to sound like it's coming through a car radio."* Clone the clean voice, then EQ and process the returned clip on the timeline. A radio-sounding reference makes a worse voice, not a radio effect.
132
-
133
- ## Getting the voice onto the call
134
-
135
- Exactly one source per call, and none of them requires a character to exist first:
136
-
137
- - **A preset:** `slates_list_voices` lists stock voices with gender, age, accent and tags — filter by any of them, or search the descriptions ("gravelly", "narration"). Pass the chosen `voiceId`. Presets clone nothing, so they are the fastest path and avoid the clone-creation rate ceiling.
138
- - **Speak AS a character:** `voiceReferenceAssetId: <its voiceAssetId>` (the clip on the row `slates_list_characters` returns). The seat clones the clip for that take and discards the vendor voice afterwards, so there is nothing to reconcile — but cloning shares a ceiling of two new voices a minute across every Slates user, so a run of lines in one cloned voice pauses between takes rather than failing. Send each line once; do not re-send one that already came back. Any other clean clip of one speaker works the same way.
139
- - **A voice with no recording:** `voiceDescription` (7–1000 characters of words). If it will be used again, keep the returned clip on a character with `slates_update_character` (`voiceAssetId`) so later lines clone the same clip instead of designing a new voice each time — a convenience, never a requirement.
140
-
141
- ## Consent
142
-
143
- Cloning a real person's voice needs that person's explicit, documented permission, scoped to what you are making. Clone from original human recordings only — never from another model's output. This is the same gate the real-face route applies to likeness, and it applies here for the same reason.
144
-
145
- ## Worked examples
146
-
147
- **A line with a direction**
148
-
149
- ```
150
- [very quiet] I heard what you said in there. I'm not going to pretend I didn't.
151
- ```
152
-
153
- **A line that needs its numbers spoken**
154
-
155
- ```
156
- The vote was three hundred and twelve to eighty-nine. It carried at four minutes past midnight.
157
- ```
158
-
159
- **A paragraph, split into three takes** — so one bad clause costs one re-roll:
160
-
161
- ```
162
- 1. You keep asking me why I stayed.
163
- 2. It wasn't loyalty. It wasn't even fear, not by the end.
164
- 3. [very slow] It was that I couldn't picture the version of me that left.
165
- ```
166
-
167
- **What NOT to send**
168
-
169
- ```
170
- (gravelly, tired) a tired old man narrates the opening of the film, wide shot, warm tungsten
171
- ```
172
-
173
- Every word of that is spoken aloud — **including the parenthetical**, which is the
174
- trap: it looks like a stage direction and is treated as dialogue. Describe the voice when you are CHOOSING one (`voiceDescription`, or the desktop's voice picker); the prompt is only ever the words.
1
+ ---
2
+ name: slates-prompting-inworld-tts
3
+ description: How to use Inworld Realtime TTS-2, the VOICE seat. Read before calling slates_generate_audio with model inworld-tts-2. Speech in a SPECIFIC voice, billed per character - the prompt is the words spoken, verbatim. Covers the identity-versus-acoustics rule (what a reference clip does and does not carry), how to write a line so it is performed rather than read, when to reach for seed-audio instead, and the voice-consent rule.
4
+ ---
5
+
6
+ # Inworld Realtime TTS-2 — the voice seat
7
+
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Keep it under 2,400 characters (the build fails above that) and keep
14
+ the rationale and the worked examples in the body below. -->
15
+ <!-- /slates-only -->
16
+ **Card — Inworld TTS-2.** Speech in a SPECIFIC voice. The prompt is the words spoken, verbatim — not a description of them. Text length determines the bill.
17
+
18
+ **IDENTITY, NOT ACOUSTICS — the rule that decides whether cloning works**
19
+ A reference carries WHO is speaking: timbre, pitch, accent, age, vowel shape. It does NOT carry WHERE they are — room tone, distance, phone EQ, reverb and mic character are *acoustics*, and this model reproduces the identity while discarding the room. So:
20
+ 1. **A noisy reference does not give a noisy read — it gives a WORSE identity.** Music, a second speaker or heavy reverb corrupt what is being extracted. Use a clean single-speaker recording.
21
+ 2. **You cannot get "on a payphone" by cloning a payphone recording.** Acoustics come from the MIX, or from `seed-audio` which renders a room.
22
+
23
+ **DIRECTION GOES IN SQUARE BRACKETS. PARENTHESES ARE SPOKEN ALOUD.** `[whispering] I hope nobody notices` is whispered; `(quietly) I hope nobody notices` says the word "quietly" out loud. Verified by ear — the easiest way to ruin a take.
24
+
25
+ - **Plain English works inside them** — it is natural-language steering, not a fixed vocabulary: `[very quiet]`, `[whisper in a hushed style]`, `[very slow]`, `[say excitedly]`. Non-verbals are their own tags: `[laugh]`, `[sigh]`, `[breathe]`, `[clear throat]`.
26
+ - **A tag it does not recognise is still consumed, and still changes the read.** Never spoken, never an error — so a mistyped tag fails SILENTLY and only listening catches it.
27
+ - **Tags persist across sentences** until changed; `[reset]` returns to normal.
28
+ - **Punctuation is the timing.** `Wait. Stop.` differs from `Wait, stop.`
29
+ - **One line, one take.** Split a paragraph so a bad clause costs one re-roll.
30
+ - **Spell numbers and titles aloud:** `twenty twenty-six`, `Doctor Reyes`.
31
+
32
+ **Route elsewhere when:** the scene needs dialogue mixed with effects and room tone in one pass (`seed-audio`), or it is a single non-speech sound (`eleven-sfx`). This surface makes ONE voice saying ONE thing, cleanly.
33
+
34
+ **Hard constraints:** no duration parameter — length falls out of the text. Exactly one voice source: a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId` (a character's voice clip to speak AS the character, or any clean clip of one speaker), or `voiceDescription`.
35
+ <!-- @card:end -->
36
+
37
+ <!-- @banned:start -->
38
+ <!-- slates-only -->
39
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
40
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
41
+ estimate, and every submitted prompt is matched against it. Keep entries
42
+ backticked and prose outside the backticks. -->
43
+ <!-- /slates-only -->
44
+ **Never use** — the prompt on this surface is SPOKEN ALOUD, so anything that describes the audio instead of being the audio gets read out as words:
45
+
46
+ - `SFX`, `Ambient noise`, `Background music` as labels — this model speaks; it does not render a scene. Use `seed-audio` for those.
47
+ - shot language: `wide shot`, `slow push in`, `warm tungsten` — video-prompt words, and here they would literally be said aloud
48
+ - `voiceover`, `narrator says`, `he says` as stage directions wrapping the line — write only the words that should come out of the speaker
49
+ <!-- @banned:end -->
50
+
51
+ ## Why the prompt is not a prompt
52
+
53
+ On every other surface in Slates the prompt DESCRIBES what you want and the model interprets it. Here the prompt IS the deliverable: each character is spoken aloud and each character is billed. `a gravelly man says he is tired` produces a voice saying the words "a gravelly man says he is tired".
54
+
55
+ That also means the two numbers a user cares about are the same number. The text length sets the price (in 250-character buckets) and sets the length of the audio. There is nothing to choose and nothing to reconcile.
56
+
57
+ ## Steering the delivery
58
+
59
+ `VERIFIED BY EAR, 2026-09-05.` Every claim in this section was listened to, not
60
+ inferred — an earlier draft of this skill documented tag forms that had only been
61
+ probed for an HTTP 200, which proves the request was accepted and nothing about
62
+ whether it was obeyed.
63
+
64
+ **Square brackets are consumed. Parentheses are read aloud.** That is the whole
65
+ rule, and getting it wrong is not a subtle degradation — the audience hears a
66
+ narrator say the word "quietly" in the middle of your line.
67
+
68
+ | Written | What comes out |
69
+ |---|---|
70
+ | `[whispering] I really hope nobody notices that.` | whispered, tag not spoken ✅ |
71
+ | `[very quiet] I really hope nobody notices that.` | very quiet, tag not spoken ✅ |
72
+ | `[whisper in a hushed style] …` | hushed, tag not spoken ✅ |
73
+ | `[very slow] …` | slowed right down, tag not spoken ✅ |
74
+ | `[laugh] …` | an actual laugh, then the line ✅ |
75
+ | `(quietly, under his breath) …` | 🚨 **the words "quietly, under his breath" are SPOKEN** |
76
+
77
+ **Plain English works — it is natural-language steering, not a fixed vocabulary.**
78
+ Both the documented phrasings (`[whisper in a hushed style]`) and ordinary adverbs
79
+ (`[whispering]`) were obeyed. Write the direction the way you would say it to an
80
+ actor.
81
+
82
+ The eight dimensions the model steers on, with a working example of each:
83
+
84
+ | Dimension | Example |
85
+ |---|---|
86
+ | Emotion | `[say excitedly]`, `[sound sad]`, `[sound terrified]` |
87
+ | Articulation | `[say with force]`, `[articulate clearly]` |
88
+ | Intonation | `[say with a rising pitch]` |
89
+ | Volume | `[very quiet]`, `[very loud]` |
90
+ | Pitch | `[say in a low tone]` |
91
+ | Range | `[say playfully]`, `[say with no pitch variation]` |
92
+ | Speed | `[very fast]`, `[very slow]` |
93
+ | Vocal style | `[whisper in a hushed style]`, `[give a nasal quality]` |
94
+
95
+ Non-verbals sit inline where they happen: `[laugh]`, `[sigh]`, `[cough]`,
96
+ `[breathe]`, `[yawn]`, `[clear throat]`.
97
+
98
+ ### Four rules that are not obvious
99
+
100
+ 1. 🚨 **A tag it does not recognise is still consumed, and still changes the read.**
101
+ `[zzzqqq]` is not spoken and does not error — it produces a different, arbitrary
102
+ delivery. So a typo in a tag is SILENT: there is no rejection, no warning, and no
103
+ way to catch it except listening to the take. Treat an unexpected performance as
104
+ a possible misspelled tag before you blame the voice.
105
+ 2. **Tags persist across sentences.** A `[very slow]` at the top governs everything
106
+ after it until something changes it. Use `[reset]` to go back to normal rather
107
+ than assuming the next sentence starts clean.
108
+ 3. **Do not stack opposing directions.** `[whisper in a hushed style]` together with
109
+ `[very loud]` produces unpredictable results — the model is resolving a
110
+ contradiction, and which side wins is not something you can rely on.
111
+ 4. **Tags COUNT toward the billed characters**, even though they are never spoken.
112
+ They are part of the text sent to the vendor, so the vendor charges for them and
113
+ so do we — billing what was actually sent is the only honest basis. It rarely
114
+ matters (a 13-character tag inside a 250-character bucket), but a line sitting
115
+ just under a bucket boundary can be pushed into the next one by a long
116
+ direction. Prefer `[very slow]` over `[say this one very slowly please]`.
117
+
118
+ ## Identity versus acoustics, at length
119
+
120
+ This is the distinction that decides whether the feature feels good, and it is worth being precise about because the failure is quiet — you get a usable clip that is subtly not the person.
121
+
122
+ **What a reference clip transfers:** vocal timbre, pitch range, accent and regional vowels, apparent age, speech rate tendencies, and the particular rasp or breathiness of the source speaker.
123
+
124
+ **What it does not transfer:** the room, the microphone, the codec, the distance from the mic, any processing on the source, and any other sound present in it.
125
+
126
+ So the ideal reference is boring: one person, close to a microphone, no music, no second speaker, no heavy reverb, five to fifteen seconds, speaking normally rather than performing. A phone voice memo in a quiet room beats a beautifully produced clip with a music bed underneath it.
127
+
128
+ **Two failure modes, both common:**
129
+
130
+ - *"I cloned my podcast intro and it doesn't sound like me."* The intro had music under it. The model averaged the music into the identity. Re-clone from a clean stretch.
131
+ - *"I want the line to sound like it's coming through a car radio."* Clone the clean voice, then EQ and process the returned clip on the timeline. A radio-sounding reference makes a worse voice, not a radio effect.
132
+
133
+ ## Getting the voice onto the call
134
+
135
+ Exactly one source per call, and none of them requires a character to exist first:
136
+
137
+ - **A preset:** `slates_list_voices` lists stock voices with gender, age, accent and tags — filter by any of them, or search the descriptions ("gravelly", "narration"). Pass the chosen `voiceId`. Presets clone nothing, so they are the fastest path and avoid the clone-creation rate ceiling.
138
+ - **Speak AS a character:** `voiceReferenceAssetId: <its voiceAssetId>` (the clip on the row `slates_list_characters` returns). The seat clones the clip for that take and discards the vendor voice afterwards, so there is nothing to reconcile — but cloning shares a ceiling of two new voices a minute across every Slates user, so a run of lines in one cloned voice pauses between takes rather than failing. Send each line once; do not re-send one that already came back. Any other clean clip of one speaker works the same way.
139
+ - **A voice with no recording:** `voiceDescription` (7–1000 characters of words). If it will be used again, keep the returned clip on a character with `slates_update_character` (`voiceAssetId`) so later lines clone the same clip instead of designing a new voice each time — a convenience, never a requirement.
140
+
141
+ ## Consent
142
+
143
+ Cloning a real person's voice needs that person's explicit, documented permission, scoped to what you are making. Clone from original human recordings only — never from another model's output. This is the same gate the real-face route applies to likeness, and it applies here for the same reason.
144
+
145
+ ## Worked examples
146
+
147
+ **A line with a direction**
148
+
149
+ ```
150
+ [very quiet] I heard what you said in there. I'm not going to pretend I didn't.
151
+ ```
152
+
153
+ **A line that needs its numbers spoken**
154
+
155
+ ```
156
+ The vote was three hundred and twelve to eighty-nine. It carried at four minutes past midnight.
157
+ ```
158
+
159
+ **A paragraph, split into three takes** — so one bad clause costs one re-roll:
160
+
161
+ ```
162
+ 1. You keep asking me why I stayed.
163
+ 2. It wasn't loyalty. It wasn't even fear, not by the end.
164
+ 3. [very slow] It was that I couldn't picture the version of me that left.
165
+ ```
166
+
167
+ **What NOT to send**
168
+
169
+ ```
170
+ (gravelly, tired) a tired old man narrates the opening of the film, wide shot, warm tungsten
171
+ ```
172
+
173
+ Every word of that is spoken aloud — **including the parenthetical**, which is the
174
+ trap: it looks like a stage direction and is treated as dialogue. Describe the voice when you are CHOOSING one (`voiceDescription`, or the desktop's voice picker); the prompt is only ever the words.
@@ -163,7 +163,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
163
163
  Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
164
164
 
165
165
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
166
- 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
166
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
167
167
  3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
168
168
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
169
169
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
@@ -185,11 +185,11 @@ Kling exposes `negative_prompt` on the fal endpoint (different from Seedance whi
185
185
 
186
186
  ```
187
187
  blurry, low quality, watermark, text overlay, distorted hands, extra fingers,
188
- duplicate limbs, unnatural skin texture, overly saturated colors, lens flare,
188
+ duplicate limbs, unnatural skin texture, overly saturated colors,
189
189
  floating objects, inconsistent shadows, jittery, flickering, morphing face
190
190
  ```
191
191
 
192
- Layer scene-specific suppressions on top.
192
+ Layer scene-specific suppressions on top, and never suppress something the prompt asks for. This block carried `lens flare` until 2026-09-15, which silently cancelled every flare a prompt described (`slates-cinematic-look` → `source-flare`); add it back only for a shot that must have none.
193
193
 
194
194
  ## Cinematic tactics
195
195
 
@@ -60,7 +60,7 @@ Seedance can generate the performance rather than bolting a mouth onto finished
60
60
  That is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a "video 1" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.
61
61
 
62
62
  - Driving clips must be 2–15s; output duration is whatever you set (4–15s).
63
- - Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming.
63
+ - Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
64
64
  - Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.
65
65
 
66
66
  Everything below is about the Kling tool.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-minimax-h3
3
- description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09), so the seats differ on ladder and price, not on what they accept. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; pooled media tokens on Max), and audio written into the wrong section is dropped or duplicated.
3
+ description: How to prompt MiniMax H3, H3 Max and H3 Max Turbo. Read before calling slates_generate_video with model minimax-h3, minimax-h3-max or minimax-h3-max-turbo. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, runs 480p/768p plus a 1080p refinement of its 768p render, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09). minimax-h3-max-turbo is a second fal post-train with Max's ladder at half Max's rate; it takes start and end frames but has NO reference endpoint. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; pooled media tokens on Max), and audio written into the wrong section is dropped or duplicated.
4
4
  ---
5
5
 
6
6
  # MiniMax H3 — prompting
@@ -28,7 +28,7 @@ description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling sl
28
28
  - `A woman sits still at a kitchen table for a beat, then looks up. She says in English, "You said Tuesday." Scene sound: a fridge hum, a spoon set down on formica. Score: none.`
29
29
  - `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, "No es el alternador." Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`
30
30
 
31
- **Hard constraint:** the two seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 768p rather than 4K, takes the same 9+3+3 references, and costs MORE at the tier they share — it is a speed pick, never the cheap one. H3's top two resolution tiers are UPSCALES of the native render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.
31
+ **Hard constraint:** the three seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p, takes the same 9+3+3 references, and costs MORE at the tier they share — a speed pick, never the cheap one; `minimax-h3-max-turbo` has Max's ladder at half its rate and takes frames only, NO references. Every tier above 768p is built from the native 768p render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.
32
32
  <!-- @card:end -->
33
33
 
34
34
  <!-- @banned:start -->
@@ -50,21 +50,31 @@ German, Italian, Japanese, Korean, Portuguese, Russian, Spanish). That single fa
50
50
  everything below — the prompt is not a shot description with sound bolted on, it is a **timeline
51
51
  with three audio layers you author separately**.
52
52
 
53
- **Two seats, one grammar.** Everything in this file applies to both. They differ only in what the
54
- endpoint accepts:
53
+ **Three seats, one grammar.** Everything in this file applies to all three. They differ only in
54
+ what the endpoint accepts:
55
55
 
56
- | | `minimax-h3` | `minimax-h3-max` |
57
- |---|---|---|
58
- | Resolution | 480p / 768p / **2K / 4K** | 480p / 768p |
59
- | References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) |
60
- | Frames | start and/or end | start and/or end |
61
- | Price at 768p | **$0.060/s** | $0.080/s |
62
- | Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) |
56
+ | | `minimax-h3` | `minimax-h3-max` | `minimax-h3-max-turbo` |
57
+ |---|---|---|---|
58
+ | Resolution | 480p / 768p / **2K / 4K** | 480p / 768p / 1080p | 480p / 768p / 1080p |
59
+ | References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) | **none** (no reference endpoint) |
60
+ | Frames | start and/or end | start and/or end | start and/or end |
61
+ | Price at 768p | **$0.060/s** | $0.080/s | $0.040/s |
62
+ | Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) | **price** — half Max's rate at every tier |
63
63
 
64
64
  **Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it
65
65
  tops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth
66
66
  paying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.
67
67
 
68
+ **Turbo is the budget seat.** Same grammar and Max's ladder at half Max's rate, with no reference
69
+ endpoint: attach a reference and Slates refuses the call rather than dropping it. Route there for
70
+ drafts, volume and start-frame coverage, then re-run the keeper on Max or base H3 when it needs
71
+ references.
72
+
73
+ **1080p on Max and Turbo is a refinement, not a native render.** fal's schema, verbatim: *"1080P
74
+ latent refinement from a native 768P source."* It is a different stage from base H3's 2K/4K
75
+ upscaler, and it costs double the 768p second. Judge a 1080p take against the same shot at 768p
76
+ before paying for it across a batch.
77
+
68
78
  **The speed is measured, not claimed** (2026-08-27, same prompt and params on both rows): a 5-second
69
79
  768p text-to-video finished in **4.8 seconds** on Max against **57 seconds** on base H3 — roughly
70
80
  **12x**, queue to finished file. fal advertises "under 3 seconds"; the literal claim did not hold at
@@ -192,8 +202,8 @@ original wording preserved exactly: *A red neon sign reading "Open Late" glows a
192
202
 
193
203
  ## References — H3's real differentiator is the declared RELATIONSHIP
194
204
 
195
- *(BOTH rows. `minimax-h3-max` gained the reference set on 2026-09-09; its free allowance is
196
- FOUR images rather than the base row's five.)*
205
+ *(`minimax-h3` and `minimax-h3-max`. Max gained the reference set on 2026-09-09; its free allowance
206
+ is FOUR images rather than the base row's five. `minimax-h3-max-turbo` takes no references.)*
197
207
 
198
208
  <!-- @inject:references-read-literally -->
199
209
  > **The general law: the model reads a reference literally.**
@@ -213,7 +223,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
213
223
  Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
214
224
 
215
225
  1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
216
- 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
226
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
217
227
  3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
218
228
  4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
219
229
  5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
@@ -260,7 +270,7 @@ exactly and say so.
260
270
  **An audio reference cannot travel alone** — H3 refuses a reference set that is audio only. Pair it
261
271
  with at least one image or video reference.
262
272
 
263
- ### 💸 Reference images past the free allowance are billed — and the two rows differ
273
+ ### 💸 Reference images past the free allowance are billed — and the two reference rows differ
264
274
 
265
275
  On `minimax-h3` the first **5** are free and each additional image adds **4 credits**.
266
276
  Max pools image pixels, reference-video seconds and reference-audio seconds into one token
@@ -279,14 +289,14 @@ reference-heavy job — a quote that omits it under-reports the bill.
279
289
 
280
290
  ## Frames
281
291
 
282
- `minimax-h3` and `minimax-h3-max` both take a **start frame**, an **end frame**, or both. With an
292
+ All three rows take a **start frame**, an **end frame**, or both. With an
283
293
  end frame, land it explicitly: describe the final pose, spacing and composition as the thing the
284
294
  shot **settles into** at the end, rather than hoping the model finds it.
285
295
 
286
296
  > *"…she rotates the handle into the final angle and settles into the pose, spacing and composition
287
297
  > of image 2 at the end of the shot."*
288
298
 
289
- **Frames and references are mutually exclusive** on both rows — they are different endpoints, and
299
+ **Frames and references are mutually exclusive** on the two reference rows — they are different endpoints, and
290
300
  the reference endpoint has no frame slots at all. Slates refuses the combination rather than
291
301
  dropping one side.
292
302
 
@@ -301,6 +311,9 @@ dropping one side.
301
311
  | `minimax-h3` · 2K · 10s | 65 |
302
312
  | `minimax-h3` · 4K · 10s | 80 |
303
313
  | `minimax-h3-max` · 768p · 10s | 40 |
314
+ | `minimax-h3-max` · 1080p · 10s | 80 |
315
+ | `minimax-h3-max-turbo` · 768p · 10s | 20 |
316
+ | `minimax-h3-max-turbo` · 1080p · 10s | 40 |
304
317
  | `minimax-h3` — every reference image past the **fifth** | **+4** |
305
318
 
306
319
  **768p is the default for a reason.** It is the tier the model natively generates.
@@ -64,7 +64,7 @@ movement from video 1. Preserve the character's identity, appearance, and outfit
64
64
  That is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.
65
65
 
66
66
  - **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).
67
- - **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.
67
+ - **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
68
68
  - **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
69
69
  - `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.
70
70