@slatesvideo/shared 0.7.0 → 0.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,174 +1,174 @@
1
- ---
2
- name: slates-prompting-inworld-tts
3
- description: How to use Inworld Realtime TTS-2, the VOICE seat. Read before calling slates_generate_audio with model inworld-tts-2. Speech in a SPECIFIC voice, billed per character - the prompt is the words spoken, verbatim. Covers the identity-versus-acoustics rule (what a reference clip does and does not carry), how to write a line so it is performed rather than read, when to reach for seed-audio instead, and the voice-consent rule.
4
- ---
5
-
6
- # Inworld Realtime TTS-2 — the voice seat
7
-
8
- <!-- @card:start -->
9
- <!-- slates-only -->
10
- <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
- src/prompts/craft-cards.ts and returned on every cost estimate for this
12
- model, so it is the ONE piece of positive craft guidance the agent cannot
13
- skip. Keep it under 2,400 characters (the build fails above that) and keep
14
- the rationale and the worked examples in the body below. -->
15
- <!-- /slates-only -->
16
- **Card — Inworld TTS-2.** Speech in a SPECIFIC voice. The prompt is the words spoken, verbatim — not a description of them. Text length determines the bill.
17
-
18
- **IDENTITY, NOT ACOUSTICS — the rule that decides whether cloning works**
19
- A reference carries WHO is speaking: timbre, pitch, accent, age, vowel shape. It does NOT carry WHERE they are — room tone, distance, phone EQ, reverb and mic character are *acoustics*, and this model reproduces the identity while discarding the room. So:
20
- 1. **A noisy reference does not give a noisy read — it gives a WORSE identity.** Music, a second speaker or heavy reverb corrupt what is being extracted. Use a clean single-speaker recording.
21
- 2. **You cannot get "on a payphone" by cloning a payphone recording.** Acoustics come from the MIX, or from `seed-audio` which renders a room.
22
-
23
- **DIRECTION GOES IN SQUARE BRACKETS. PARENTHESES ARE SPOKEN ALOUD.** `[whispering] I hope nobody notices` is whispered; `(quietly) I hope nobody notices` says the word "quietly" out loud. Verified by ear — the easiest way to ruin a take.
24
-
25
- - **Plain English works inside them** — it is natural-language steering, not a fixed vocabulary: `[very quiet]`, `[whisper in a hushed style]`, `[very slow]`, `[say excitedly]`. Non-verbals are their own tags: `[laugh]`, `[sigh]`, `[breathe]`, `[clear throat]`.
26
- - **A tag it does not recognise is still consumed, and still changes the read.** Never spoken, never an error — so a mistyped tag fails SILENTLY and only listening catches it.
27
- - **Tags persist across sentences** until changed; `[reset]` returns to normal.
28
- - **Punctuation is the timing.** `Wait. Stop.` differs from `Wait, stop.`
29
- - **One line, one take.** Split a paragraph so a bad clause costs one re-roll.
30
- - **Spell numbers and titles aloud:** `twenty twenty-six`, `Doctor Reyes`.
31
-
32
- **Route elsewhere when:** the scene needs dialogue mixed with effects and room tone in one pass (`seed-audio`), or it is a single non-speech sound (`eleven-sfx`). This surface makes ONE voice saying ONE thing, cleanly.
33
-
34
- **Hard constraints:** no duration parameter — length falls out of the text. Exactly one voice source: a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId` (a character's voice clip to speak AS the character, or any clean clip of one speaker), or `voiceDescription`.
35
- <!-- @card:end -->
36
-
37
- <!-- @banned:start -->
38
- <!-- slates-only -->
39
- <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
40
- extracted by src/prompts/banned-tokens.ts and returned on this model's cost
41
- estimate, and every submitted prompt is matched against it. Keep entries
42
- backticked and prose outside the backticks. -->
43
- <!-- /slates-only -->
44
- **Never use** — the prompt on this surface is SPOKEN ALOUD, so anything that describes the audio instead of being the audio gets read out as words:
45
-
46
- - `SFX`, `Ambient noise`, `Background music` as labels — this model speaks; it does not render a scene. Use `seed-audio` for those.
47
- - shot language: `wide shot`, `slow push in`, `warm tungsten` — video-prompt words, and here they would literally be said aloud
48
- - `voiceover`, `narrator says`, `he says` as stage directions wrapping the line — write only the words that should come out of the speaker
49
- <!-- @banned:end -->
50
-
51
- ## Why the prompt is not a prompt
52
-
53
- On every other surface in Slates the prompt DESCRIBES what you want and the model interprets it. Here the prompt IS the deliverable: each character is spoken aloud and each character is billed. `a gravelly man says he is tired` produces a voice saying the words "a gravelly man says he is tired".
54
-
55
- That also means the two numbers a user cares about are the same number. The text length sets the price (in 250-character buckets) and sets the length of the audio. There is nothing to choose and nothing to reconcile.
56
-
57
- ## Steering the delivery
58
-
59
- `VERIFIED BY EAR, 2026-09-05.` Every claim in this section was listened to, not
60
- inferred — an earlier draft of this skill documented tag forms that had only been
61
- probed for an HTTP 200, which proves the request was accepted and nothing about
62
- whether it was obeyed.
63
-
64
- **Square brackets are consumed. Parentheses are read aloud.** That is the whole
65
- rule, and getting it wrong is not a subtle degradation — the audience hears a
66
- narrator say the word "quietly" in the middle of your line.
67
-
68
- | Written | What comes out |
69
- |---|---|
70
- | `[whispering] I really hope nobody notices that.` | whispered, tag not spoken ✅ |
71
- | `[very quiet] I really hope nobody notices that.` | very quiet, tag not spoken ✅ |
72
- | `[whisper in a hushed style] …` | hushed, tag not spoken ✅ |
73
- | `[very slow] …` | slowed right down, tag not spoken ✅ |
74
- | `[laugh] …` | an actual laugh, then the line ✅ |
75
- | `(quietly, under his breath) …` | 🚨 **the words "quietly, under his breath" are SPOKEN** |
76
-
77
- **Plain English works — it is natural-language steering, not a fixed vocabulary.**
78
- Both the documented phrasings (`[whisper in a hushed style]`) and ordinary adverbs
79
- (`[whispering]`) were obeyed. Write the direction the way you would say it to an
80
- actor.
81
-
82
- The eight dimensions the model steers on, with a working example of each:
83
-
84
- | Dimension | Example |
85
- |---|---|
86
- | Emotion | `[say excitedly]`, `[sound sad]`, `[sound terrified]` |
87
- | Articulation | `[say with force]`, `[articulate clearly]` |
88
- | Intonation | `[say with a rising pitch]` |
89
- | Volume | `[very quiet]`, `[very loud]` |
90
- | Pitch | `[say in a low tone]` |
91
- | Range | `[say playfully]`, `[say with no pitch variation]` |
92
- | Speed | `[very fast]`, `[very slow]` |
93
- | Vocal style | `[whisper in a hushed style]`, `[give a nasal quality]` |
94
-
95
- Non-verbals sit inline where they happen: `[laugh]`, `[sigh]`, `[cough]`,
96
- `[breathe]`, `[yawn]`, `[clear throat]`.
97
-
98
- ### Four rules that are not obvious
99
-
100
- 1. 🚨 **A tag it does not recognise is still consumed, and still changes the read.**
101
- `[zzzqqq]` is not spoken and does not error — it produces a different, arbitrary
102
- delivery. So a typo in a tag is SILENT: there is no rejection, no warning, and no
103
- way to catch it except listening to the take. Treat an unexpected performance as
104
- a possible misspelled tag before you blame the voice.
105
- 2. **Tags persist across sentences.** A `[very slow]` at the top governs everything
106
- after it until something changes it. Use `[reset]` to go back to normal rather
107
- than assuming the next sentence starts clean.
108
- 3. **Do not stack opposing directions.** `[whisper in a hushed style]` together with
109
- `[very loud]` produces unpredictable results — the model is resolving a
110
- contradiction, and which side wins is not something you can rely on.
111
- 4. **Tags COUNT toward the billed characters**, even though they are never spoken.
112
- They are part of the text sent to the vendor, so the vendor charges for them and
113
- so do we — billing what was actually sent is the only honest basis. It rarely
114
- matters (a 13-character tag inside a 250-character bucket), but a line sitting
115
- just under a bucket boundary can be pushed into the next one by a long
116
- direction. Prefer `[very slow]` over `[say this one very slowly please]`.
117
-
118
- ## Identity versus acoustics, at length
119
-
120
- This is the distinction that decides whether the feature feels good, and it is worth being precise about because the failure is quiet — you get a usable clip that is subtly not the person.
121
-
122
- **What a reference clip transfers:** vocal timbre, pitch range, accent and regional vowels, apparent age, speech rate tendencies, and the particular rasp or breathiness of the source speaker.
123
-
124
- **What it does not transfer:** the room, the microphone, the codec, the distance from the mic, any processing on the source, and any other sound present in it.
125
-
126
- So the ideal reference is boring: one person, close to a microphone, no music, no second speaker, no heavy reverb, five to fifteen seconds, speaking normally rather than performing. A phone voice memo in a quiet room beats a beautifully produced clip with a music bed underneath it.
127
-
128
- **Two failure modes, both common:**
129
-
130
- - *"I cloned my podcast intro and it doesn't sound like me."* The intro had music under it. The model averaged the music into the identity. Re-clone from a clean stretch.
131
- - *"I want the line to sound like it's coming through a car radio."* Clone the clean voice, then EQ and process the returned clip on the timeline. A radio-sounding reference makes a worse voice, not a radio effect.
132
-
133
- ## Getting the voice onto the call
134
-
135
- Exactly one source per call, and none of them requires a character to exist first:
136
-
137
- - **A preset:** `slates_list_voices` lists stock voices with gender, age, accent and tags — filter by any of them, or search the descriptions ("gravelly", "narration"). Pass the chosen `voiceId`. Presets clone nothing, so they are the fastest path and avoid the clone-creation rate ceiling.
138
- - **Speak AS a character:** `voiceReferenceAssetId: <its voiceAssetId>` (the clip on the row `slates_list_characters` returns). The seat clones the clip for that take and discards the vendor voice afterwards, so there is nothing to reconcile — but cloning shares a ceiling of two new voices a minute across every Slates user, so a run of lines in one cloned voice pauses between takes rather than failing. Send each line once; do not re-send one that already came back. Any other clean clip of one speaker works the same way.
139
- - **A voice with no recording:** `voiceDescription` (7–1000 characters of words). If it will be used again, keep the returned clip on a character with `slates_update_character` (`voiceAssetId`) so later lines clone the same clip instead of designing a new voice each time — a convenience, never a requirement.
140
-
141
- ## Consent
142
-
143
- Cloning a real person's voice needs that person's explicit, documented permission, scoped to what you are making. Clone from original human recordings only — never from another model's output. This is the same gate the real-face route applies to likeness, and it applies here for the same reason.
144
-
145
- ## Worked examples
146
-
147
- **A line with a direction**
148
-
149
- ```
150
- [very quiet] I heard what you said in there. I'm not going to pretend I didn't.
151
- ```
152
-
153
- **A line that needs its numbers spoken**
154
-
155
- ```
156
- The vote was three hundred and twelve to eighty-nine. It carried at four minutes past midnight.
157
- ```
158
-
159
- **A paragraph, split into three takes** — so one bad clause costs one re-roll:
160
-
161
- ```
162
- 1. You keep asking me why I stayed.
163
- 2. It wasn't loyalty. It wasn't even fear, not by the end.
164
- 3. [very slow] It was that I couldn't picture the version of me that left.
165
- ```
166
-
167
- **What NOT to send**
168
-
169
- ```
170
- (gravelly, tired) a tired old man narrates the opening of the film, wide shot, warm tungsten
171
- ```
172
-
173
- Every word of that is spoken aloud — **including the parenthetical**, which is the
174
- trap: it looks like a stage direction and is treated as dialogue. Describe the voice when you are CHOOSING one (`voiceDescription`, or the desktop's voice picker); the prompt is only ever the words.
1
+ ---
2
+ name: slates-prompting-inworld-tts
3
+ description: How to use Inworld Realtime TTS-2, the VOICE seat. Read before calling slates_generate_audio with model inworld-tts-2. Speech in a SPECIFIC voice, billed per character - the prompt is the words spoken, verbatim. Covers the identity-versus-acoustics rule (what a reference clip does and does not carry), how to write a line so it is performed rather than read, when to reach for seed-audio instead, and the voice-consent rule.
4
+ ---
5
+
6
+ # Inworld Realtime TTS-2 — the voice seat
7
+
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Keep it under 2,400 characters (the build fails above that) and keep
14
+ the rationale and the worked examples in the body below. -->
15
+ <!-- /slates-only -->
16
+ **Card — Inworld TTS-2.** Speech in a SPECIFIC voice. The prompt is the words spoken, verbatim — not a description of them. Text length determines the bill.
17
+
18
+ **IDENTITY, NOT ACOUSTICS — the rule that decides whether cloning works**
19
+ A reference carries WHO is speaking: timbre, pitch, accent, age, vowel shape. It does NOT carry WHERE they are — room tone, distance, phone EQ, reverb and mic character are *acoustics*, and this model reproduces the identity while discarding the room. So:
20
+ 1. **A noisy reference does not give a noisy read — it gives a WORSE identity.** Music, a second speaker or heavy reverb corrupt what is being extracted. Use a clean single-speaker recording.
21
+ 2. **You cannot get "on a payphone" by cloning a payphone recording.** Acoustics come from the MIX, or from `seed-audio` which renders a room.
22
+
23
+ **DIRECTION GOES IN SQUARE BRACKETS. PARENTHESES ARE SPOKEN ALOUD.** `[whispering] I hope nobody notices` is whispered; `(quietly) I hope nobody notices` says the word "quietly" out loud. Verified by ear — the easiest way to ruin a take.
24
+
25
+ - **Plain English works inside them** — it is natural-language steering, not a fixed vocabulary: `[very quiet]`, `[whisper in a hushed style]`, `[very slow]`, `[say excitedly]`. Non-verbals are their own tags: `[laugh]`, `[sigh]`, `[breathe]`, `[clear throat]`.
26
+ - **A tag it does not recognise is still consumed, and still changes the read.** Never spoken, never an error — so a mistyped tag fails SILENTLY and only listening catches it.
27
+ - **Tags persist across sentences** until changed; `[reset]` returns to normal.
28
+ - **Punctuation is the timing.** `Wait. Stop.` differs from `Wait, stop.`
29
+ - **One line, one take.** Split a paragraph so a bad clause costs one re-roll.
30
+ - **Spell numbers and titles aloud:** `twenty twenty-six`, `Doctor Reyes`.
31
+
32
+ **Route elsewhere when:** the scene needs dialogue mixed with effects and room tone in one pass (`seed-audio`), or it is a single non-speech sound (`eleven-sfx`). This surface makes ONE voice saying ONE thing, cleanly.
33
+
34
+ **Hard constraints:** no duration parameter — length falls out of the text. Exactly one voice source: a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId` (a character's voice clip to speak AS the character, or any clean clip of one speaker), or `voiceDescription`.
35
+ <!-- @card:end -->
36
+
37
+ <!-- @banned:start -->
38
+ <!-- slates-only -->
39
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
40
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
41
+ estimate, and every submitted prompt is matched against it. Keep entries
42
+ backticked and prose outside the backticks. -->
43
+ <!-- /slates-only -->
44
+ **Never use** — the prompt on this surface is SPOKEN ALOUD, so anything that describes the audio instead of being the audio gets read out as words:
45
+
46
+ - `SFX`, `Ambient noise`, `Background music` as labels — this model speaks; it does not render a scene. Use `seed-audio` for those.
47
+ - shot language: `wide shot`, `slow push in`, `warm tungsten` — video-prompt words, and here they would literally be said aloud
48
+ - `voiceover`, `narrator says`, `he says` as stage directions wrapping the line — write only the words that should come out of the speaker
49
+ <!-- @banned:end -->
50
+
51
+ ## Why the prompt is not a prompt
52
+
53
+ On every other surface in Slates the prompt DESCRIBES what you want and the model interprets it. Here the prompt IS the deliverable: each character is spoken aloud and each character is billed. `a gravelly man says he is tired` produces a voice saying the words "a gravelly man says he is tired".
54
+
55
+ That also means the two numbers a user cares about are the same number. The text length sets the price (in 250-character buckets) and sets the length of the audio. There is nothing to choose and nothing to reconcile.
56
+
57
+ ## Steering the delivery
58
+
59
+ `VERIFIED BY EAR, 2026-09-05.` Every claim in this section was listened to, not
60
+ inferred — an earlier draft of this skill documented tag forms that had only been
61
+ probed for an HTTP 200, which proves the request was accepted and nothing about
62
+ whether it was obeyed.
63
+
64
+ **Square brackets are consumed. Parentheses are read aloud.** That is the whole
65
+ rule, and getting it wrong is not a subtle degradation — the audience hears a
66
+ narrator say the word "quietly" in the middle of your line.
67
+
68
+ | Written | What comes out |
69
+ |---|---|
70
+ | `[whispering] I really hope nobody notices that.` | whispered, tag not spoken ✅ |
71
+ | `[very quiet] I really hope nobody notices that.` | very quiet, tag not spoken ✅ |
72
+ | `[whisper in a hushed style] …` | hushed, tag not spoken ✅ |
73
+ | `[very slow] …` | slowed right down, tag not spoken ✅ |
74
+ | `[laugh] …` | an actual laugh, then the line ✅ |
75
+ | `(quietly, under his breath) …` | 🚨 **the words "quietly, under his breath" are SPOKEN** |
76
+
77
+ **Plain English works — it is natural-language steering, not a fixed vocabulary.**
78
+ Both the documented phrasings (`[whisper in a hushed style]`) and ordinary adverbs
79
+ (`[whispering]`) were obeyed. Write the direction the way you would say it to an
80
+ actor.
81
+
82
+ The eight dimensions the model steers on, with a working example of each:
83
+
84
+ | Dimension | Example |
85
+ |---|---|
86
+ | Emotion | `[say excitedly]`, `[sound sad]`, `[sound terrified]` |
87
+ | Articulation | `[say with force]`, `[articulate clearly]` |
88
+ | Intonation | `[say with a rising pitch]` |
89
+ | Volume | `[very quiet]`, `[very loud]` |
90
+ | Pitch | `[say in a low tone]` |
91
+ | Range | `[say playfully]`, `[say with no pitch variation]` |
92
+ | Speed | `[very fast]`, `[very slow]` |
93
+ | Vocal style | `[whisper in a hushed style]`, `[give a nasal quality]` |
94
+
95
+ Non-verbals sit inline where they happen: `[laugh]`, `[sigh]`, `[cough]`,
96
+ `[breathe]`, `[yawn]`, `[clear throat]`.
97
+
98
+ ### Four rules that are not obvious
99
+
100
+ 1. 🚨 **A tag it does not recognise is still consumed, and still changes the read.**
101
+ `[zzzqqq]` is not spoken and does not error — it produces a different, arbitrary
102
+ delivery. So a typo in a tag is SILENT: there is no rejection, no warning, and no
103
+ way to catch it except listening to the take. Treat an unexpected performance as
104
+ a possible misspelled tag before you blame the voice.
105
+ 2. **Tags persist across sentences.** A `[very slow]` at the top governs everything
106
+ after it until something changes it. Use `[reset]` to go back to normal rather
107
+ than assuming the next sentence starts clean.
108
+ 3. **Do not stack opposing directions.** `[whisper in a hushed style]` together with
109
+ `[very loud]` produces unpredictable results — the model is resolving a
110
+ contradiction, and which side wins is not something you can rely on.
111
+ 4. **Tags COUNT toward the billed characters**, even though they are never spoken.
112
+ They are part of the text sent to the vendor, so the vendor charges for them and
113
+ so do we — billing what was actually sent is the only honest basis. It rarely
114
+ matters (a 13-character tag inside a 250-character bucket), but a line sitting
115
+ just under a bucket boundary can be pushed into the next one by a long
116
+ direction. Prefer `[very slow]` over `[say this one very slowly please]`.
117
+
118
+ ## Identity versus acoustics, at length
119
+
120
+ This is the distinction that decides whether the feature feels good, and it is worth being precise about because the failure is quiet — you get a usable clip that is subtly not the person.
121
+
122
+ **What a reference clip transfers:** vocal timbre, pitch range, accent and regional vowels, apparent age, speech rate tendencies, and the particular rasp or breathiness of the source speaker.
123
+
124
+ **What it does not transfer:** the room, the microphone, the codec, the distance from the mic, any processing on the source, and any other sound present in it.
125
+
126
+ So the ideal reference is boring: one person, close to a microphone, no music, no second speaker, no heavy reverb, five to fifteen seconds, speaking normally rather than performing. A phone voice memo in a quiet room beats a beautifully produced clip with a music bed underneath it.
127
+
128
+ **Two failure modes, both common:**
129
+
130
+ - *"I cloned my podcast intro and it doesn't sound like me."* The intro had music under it. The model averaged the music into the identity. Re-clone from a clean stretch.
131
+ - *"I want the line to sound like it's coming through a car radio."* Clone the clean voice, then EQ and process the returned clip on the timeline. A radio-sounding reference makes a worse voice, not a radio effect.
132
+
133
+ ## Getting the voice onto the call
134
+
135
+ Exactly one source per call, and none of them requires a character to exist first:
136
+
137
+ - **A preset:** `slates_list_voices` lists stock voices with gender, age, accent and tags — filter by any of them, or search the descriptions ("gravelly", "narration"). Pass the chosen `voiceId`. Presets clone nothing, so they are the fastest path and avoid the clone-creation rate ceiling.
138
+ - **Speak AS a character:** `voiceReferenceAssetId: <its voiceAssetId>` (the clip on the row `slates_list_characters` returns). The seat clones the clip for that take and discards the vendor voice afterwards, so there is nothing to reconcile — but cloning shares a ceiling of two new voices a minute across every Slates user, so a run of lines in one cloned voice pauses between takes rather than failing. Send each line once; do not re-send one that already came back. Any other clean clip of one speaker works the same way.
139
+ - **A voice with no recording:** `voiceDescription` (7–1000 characters of words). If it will be used again, keep the returned clip on a character with `slates_update_character` (`voiceAssetId`) so later lines clone the same clip instead of designing a new voice each time — a convenience, never a requirement.
140
+
141
+ ## Consent
142
+
143
+ Cloning a real person's voice needs that person's explicit, documented permission, scoped to what you are making. Clone from original human recordings only — never from another model's output. This is the same gate the real-face route applies to likeness, and it applies here for the same reason.
144
+
145
+ ## Worked examples
146
+
147
+ **A line with a direction**
148
+
149
+ ```
150
+ [very quiet] I heard what you said in there. I'm not going to pretend I didn't.
151
+ ```
152
+
153
+ **A line that needs its numbers spoken**
154
+
155
+ ```
156
+ The vote was three hundred and twelve to eighty-nine. It carried at four minutes past midnight.
157
+ ```
158
+
159
+ **A paragraph, split into three takes** — so one bad clause costs one re-roll:
160
+
161
+ ```
162
+ 1. You keep asking me why I stayed.
163
+ 2. It wasn't loyalty. It wasn't even fear, not by the end.
164
+ 3. [very slow] It was that I couldn't picture the version of me that left.
165
+ ```
166
+
167
+ **What NOT to send**
168
+
169
+ ```
170
+ (gravelly, tired) a tired old man narrates the opening of the film, wide shot, warm tungsten
171
+ ```
172
+
173
+ Every word of that is spoken aloud — **including the parenthetical**, which is the
174
+ trap: it looks like a stage direction and is treated as dialogue. Describe the voice when you are CHOOSING one (`voiceDescription`, or the desktop's voice picker); the prompt is only ever the words.
@@ -1,54 +1,54 @@
1
- ---
2
- name: slates-style-prompting
3
- description: Use when the user asks for a visual style ("make it anime", "painterly look", "like a Pixar film"), or when a style has to hold across several shots. Covers how photoreal, anime, painterly and 3d-render are prompted DIFFERENTLY per model, and the style-routing recipe (reference-first, styled start-frame → i2v).
4
- ---
5
-
6
- # Per-style prompting (photoreal · anime · painterly · 3d-render)
7
-
8
- The style library (`slates_create_style` / the app's style ids) defines what each style IS. This guide is how to PROMPT each style per model. Derived from `research/style-prompting-research.md` (second-brain) — claims marked *(hypothesis)* are untested; don't present them to users as fact.
9
-
10
- ## The four ground rules (all styles)
11
-
12
- 1. **Assign references where they contribute.** Describe the scene with inline bindings, such as "the woman from image 1, lit and graded like image 2." Preserve an existing scene reference when its look should stay. A look-only reference may need light and exposure described for the new scene; prose and references can work together.
13
- 2. **Use each model’s language, without imposing a fixed prompt template:**
14
- - **Nano Banana 2** — narrative prose; the style is the opening framing of the sentence ("A hand-drawn 2D anime cel illustration of…"), never a comma tag.
15
- - **Seedance 2.0** — the 8-part formula reserves "visual style" (slot 6) and "image quality" (slot 7). One clause each. Don't scatter style words through the action text.
16
- - **Kling V3** — prose scene direction; style rides the lighting/style tail of Scene → Subject → Action → Camera → Lighting/Style. Tag soup underperforms badly.
17
- 3. **Keep the intended look consistent across shots.** Reuse relevant references and stable descriptions, adapting the wording to each scene.
18
- 4. **A styled start frame is one video control.** Generate it with the image seat suited to the brief, then describe the motion. Preserve its look unless the user wants the light or grade to change.
19
-
20
- Never stack style buzzwords ("ARRI ALEXA, 35mm, film grain, depth-of-field mastery…"). One or two register tokens maximum — piles of specs dull the image.
21
-
22
- ## Photoreal
23
-
24
- - **NB2:** never the literal word "photorealistic". Describe *a real photograph*: natural skin texture and imperfection, motivated lighting, one lens/film register ("shot on a 50mm, soft window light"). Photographic composition terms: wide-angle / macro / low-angle.
25
- - **Seedance:** put "sharp focus, natural color, high detail" in the image-quality slot and always include a lighting clause. Keep motion slow and coherent — fast/burst action is the #1 quality killer and reads most fake in photoreal.
26
- - **Kling:** the photoreal-PEOPLE lane — convincing acting, dialogue, lip-sync. It breaks on close-up hands, fine fluids, and crowds beyond ~5 faces: route those beats to Seedance or reframe.
27
- - **Faces on Seedance:** photoreal humans trigger the face-tier routing (AI face vs consented real face — see slates-prompting-seedance §Faces). Set the face flags honestly; never skip them to save credits.
28
-
29
- ## Anime
30
-
31
- - **NB2:** open with the medium — "A hand-drawn 2D anime cel illustration of…" — then normal narrative Subject/Setting/Action. Clean line art, flat-shaded color, expressive eyes. NB2 has no negative prompt: phrase exclusions positively ("flat cel shading with uniform focus", not "no depth of field").
32
- - **Seedance:** visual-style slot = "2D anime style, clean line art, flat cel shading". The slow/coherent-motion preference still applies — burst sakuga actions are the same instability trap as in photoreal.
33
- - **Kling:** weakest anime lane (its strength is live-action-like acting); expect style drift on long prose-only shots. Prefer ground rule 4: NB2 anime start-frame → i2v with a motion-only prompt. *(hypothesis: refs hold Kling's anime better than prose — verify before promising.)*
34
- - Anime faces drift under multiple references faster than photoreal — the named-entity two-sheet doctrine applies unchanged.
35
-
36
- ## Painterly
37
-
38
- - **NB2:** medium + technique in the style framing: "digital concept-art painting, visible brushwork, painted edges". At most ONE school/era register ("classic gouache illustration") — a register, not an artist-name pile.
39
- - **Video:** the least-supported style lane. Use ground rule 4 (painterly NB2 frame → i2v, motion-only prompt) and expect some cleanup of painterliness over the clip *(hypothesis — set user expectations, don't promise a perfectly painterly clip)*.
40
- - Camera language still applies — painterly ≠ static; "slow push-in" works the same.
41
-
42
- ## 3D render
43
-
44
- - **NB2:** name the lineage register in the style framing: "stylized 3D render, soft global illumination, subsurface skin". Lighting vocabulary (GI, rim light) is unusually load-bearing for the 3D read.
45
- - **Seedance:** the physics/effects lane flatters 3D content — visual-style slot "stylized 3D animation", image-quality slot "clean render, high detail".
46
- - **Kling:** same start-frame preference as anime.
47
- - *(hypothesis)* An engine token ("Unreal Engine 5 render") may help NB2; if used, ONE token, style slot only — never on Seedance where spec-stuffing hurts.
48
-
49
- ## Routing recipe (what to actually do)
50
-
51
- 1. Style reference available → attach it, rely on inherit. Done.
52
- 2. No reference, image request → styled NB2 prose per the section above.
53
- 3. No reference, video request → NB2 styled start-frame first, then i2v with motion-only prompt. Direct styled text-to-video is the fallback when a start frame doesn't fit (e.g. dialogue-first Kling shots).
54
- 4. Multi-shot run → byte-identical style clause per shot + shared references.
1
+ ---
2
+ name: slates-style-prompting
3
+ description: Use when the user asks for a visual style ("make it anime", "painterly look", "like a Pixar film"), or when a style has to hold across several shots. Covers how photoreal, anime, painterly and 3d-render are prompted DIFFERENTLY per model, and the style-routing recipe (reference-first, styled start-frame → i2v).
4
+ ---
5
+
6
+ # Per-style prompting (photoreal · anime · painterly · 3d-render)
7
+
8
+ The style library (`slates_create_style` / the app's style ids) defines what each style IS. This guide is how to PROMPT each style per model. Derived from `research/style-prompting-research.md` (second-brain) — claims marked *(hypothesis)* are untested; don't present them to users as fact.
9
+
10
+ ## The four ground rules (all styles)
11
+
12
+ 1. **Assign references where they contribute.** Describe the scene with inline bindings, such as "the woman from image 1, lit and graded like image 2." Preserve an existing scene reference when its look should stay. A look-only reference may need light and exposure described for the new scene; prose and references can work together.
13
+ 2. **Use each model’s language, without imposing a fixed prompt template:**
14
+ - **Nano Banana 2** — narrative prose; the style is the opening framing of the sentence ("A hand-drawn 2D anime cel illustration of…"), never a comma tag.
15
+ - **Seedance 2.0** — the 8-part formula reserves "visual style" (slot 6) and "image quality" (slot 7). One clause each. Don't scatter style words through the action text.
16
+ - **Kling V3** — prose scene direction; style rides the lighting/style tail of Scene → Subject → Action → Camera → Lighting/Style. Tag soup underperforms badly.
17
+ 3. **Keep the intended look consistent across shots.** Reuse relevant references and stable descriptions, adapting the wording to each scene.
18
+ 4. **A styled start frame is one video control.** Generate it with the image seat suited to the brief, then describe the motion. Preserve its look unless the user wants the light or grade to change.
19
+
20
+ Never stack style buzzwords ("ARRI ALEXA, 35mm, film grain, depth-of-field mastery…"). One or two register tokens maximum — piles of specs dull the image.
21
+
22
+ ## Photoreal
23
+
24
+ - **NB2:** never the literal word "photorealistic". Describe *a real photograph*: natural skin texture and imperfection, motivated lighting, one lens/film register ("shot on a 50mm, soft window light"). Photographic composition terms: wide-angle / macro / low-angle.
25
+ - **Seedance:** put "sharp focus, natural color, high detail" in the image-quality slot and always include a lighting clause. Keep motion slow and coherent — fast/burst action is the #1 quality killer and reads most fake in photoreal.
26
+ - **Kling:** the photoreal-PEOPLE lane — convincing acting, dialogue, lip-sync. It breaks on close-up hands, fine fluids, and crowds beyond ~5 faces: route those beats to Seedance or reframe.
27
+ - **Faces on Seedance:** photoreal humans trigger the face-tier routing (AI face vs consented real face — see slates-prompting-seedance §Faces). Set the face flags honestly; never skip them to save credits.
28
+
29
+ ## Anime
30
+
31
+ - **NB2:** open with the medium — "A hand-drawn 2D anime cel illustration of…" — then normal narrative Subject/Setting/Action. Clean line art, flat-shaded color, expressive eyes. NB2 has no negative prompt: phrase exclusions positively ("flat cel shading with uniform focus", not "no depth of field").
32
+ - **Seedance:** visual-style slot = "2D anime style, clean line art, flat cel shading". The slow/coherent-motion preference still applies — burst sakuga actions are the same instability trap as in photoreal.
33
+ - **Kling:** weakest anime lane (its strength is live-action-like acting); expect style drift on long prose-only shots. Prefer ground rule 4: NB2 anime start-frame → i2v with a motion-only prompt. *(hypothesis: refs hold Kling's anime better than prose — verify before promising.)*
34
+ - Anime faces drift under multiple references faster than photoreal — the named-entity two-sheet doctrine applies unchanged.
35
+
36
+ ## Painterly
37
+
38
+ - **NB2:** medium + technique in the style framing: "digital concept-art painting, visible brushwork, painted edges". At most ONE school/era register ("classic gouache illustration") — a register, not an artist-name pile.
39
+ - **Video:** the least-supported style lane. Use ground rule 4 (painterly NB2 frame → i2v, motion-only prompt) and expect some cleanup of painterliness over the clip *(hypothesis — set user expectations, don't promise a perfectly painterly clip)*.
40
+ - Camera language still applies — painterly ≠ static; "slow push-in" works the same.
41
+
42
+ ## 3D render
43
+
44
+ - **NB2:** name the lineage register in the style framing: "stylized 3D render, soft global illumination, subsurface skin". Lighting vocabulary (GI, rim light) is unusually load-bearing for the 3D read.
45
+ - **Seedance:** the physics/effects lane flatters 3D content — visual-style slot "stylized 3D animation", image-quality slot "clean render, high detail".
46
+ - **Kling:** same start-frame preference as anime.
47
+ - *(hypothesis)* An engine token ("Unreal Engine 5 render") may help NB2; if used, ONE token, style slot only — never on Seedance where spec-stuffing hurts.
48
+
49
+ ## Routing recipe (what to actually do)
50
+
51
+ 1. Style reference available → attach it, rely on inherit. Done.
52
+ 2. No reference, image request → styled NB2 prose per the section above.
53
+ 3. No reference, video request → NB2 styled start-frame first, then i2v with motion-only prompt. Direct styled text-to-video is the fallback when a start frame doesn't fit (e.g. dialogue-first Kling shots).
54
+ 4. Multi-shot run → byte-identical style clause per shot + shared references.