@slatesvideo/shared 0.7.2 → 0.7.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (84) hide show
  1. package/dist/clients/cloud.d.ts +4 -0
  2. package/dist/clients/cloud.js +11 -3
  3. package/dist/index.d.ts +1 -0
  4. package/dist/index.js +1 -0
  5. package/dist/manual/content.d.ts +1 -1
  6. package/dist/manual/content.js +1 -1
  7. package/dist/operations/index.d.ts +12 -13
  8. package/dist/operations/index.js +158 -133
  9. package/dist/operations/surface.d.ts +6 -2
  10. package/dist/operations/surface.js +29 -5
  11. package/dist/prompts/agent-doctrine.d.ts +4 -4
  12. package/dist/prompts/agent-doctrine.js +17 -28
  13. package/dist/prompts/guide-discovery.d.ts +23 -0
  14. package/dist/prompts/guide-discovery.js +39 -0
  15. package/dist/prompts/guide-retrieval.js +1 -1
  16. package/dist/prompts/model-capabilities.d.ts +8 -9
  17. package/dist/prompts/model-capabilities.js +11 -51
  18. package/dist/prompts/model-facts.d.ts +2 -2
  19. package/dist/prompts/model-facts.js +15 -26
  20. package/dist/prompts/partials.generated.js +6 -3
  21. package/dist/prompts/prompting-tips.d.ts +1 -1
  22. package/dist/prompts/prompting-tips.js +21 -63
  23. package/dist/prompts/search-terms.d.ts +3 -0
  24. package/dist/prompts/search-terms.js +24 -0
  25. package/dist/skills/content.js +36 -37
  26. package/dist/skills/metadata.d.ts +7 -0
  27. package/dist/skills/metadata.js +29 -0
  28. package/exports/slates-chatgpt-images/generated/SKILL.md +7 -1
  29. package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
  30. package/exports/slates-prompt-builder/generated/SKILL.md +28 -16
  31. package/exports/slates-prompt-builder/generated/reference-character.md +12 -13
  32. package/exports/slates-prompt-builder/generated/reference-content-policy.md +2 -2
  33. package/exports/slates-prompt-builder/generated/reference-gpt-image-2-5.md +191 -0
  34. package/exports/slates-prompt-builder/generated/reference-kling.md +32 -11
  35. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +24 -6
  36. package/exports/slates-prompt-builder/generated/reference-omni-flash.md +65 -0
  37. package/exports/slates-prompt-builder/generated/reference-seedance-2-5.md +362 -0
  38. package/exports/slates-prompt-builder/generated/reference-seedance.md +34 -4
  39. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +77 -23
  40. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  41. package/package.json +2 -1
  42. package/skills/_partials/blender-action-curves.md +24 -0
  43. package/skills/_partials/iteration-diagnosis.md +5 -0
  44. package/skills/_partials/model-routing.md +35 -0
  45. package/skills/_partials/seedance-25-timestamps.md +2 -2
  46. package/skills/_partials/still-gate.md +2 -2
  47. package/skills/_partials/thresholds.md +1 -1
  48. package/skills/slates-blocking-to-prompt.md +15 -13
  49. package/skills/slates-camera-language.md +45 -7
  50. package/skills/slates-character-identity.md +8 -6
  51. package/skills/slates-chatgpt-images.md +7 -1
  52. package/skills/slates-cinematic-look.md +1 -1
  53. package/skills/slates-content-policy.md +4 -6
  54. package/skills/slates-cost-discipline.md +18 -12
  55. package/skills/slates-dialogue-blocking.md +6 -6
  56. package/skills/slates-direct-response-ad.md +1 -1
  57. package/skills/slates-edit-and-iterate.md +12 -4
  58. package/skills/slates-model-selection.md +82 -90
  59. package/skills/slates-one-prompt-film.md +1 -1
  60. package/skills/slates-previs-blocking.md +44 -13
  61. package/skills/slates-project-organization.md +2 -2
  62. package/skills/slates-prompting-elevenlabs.md +4 -4
  63. package/skills/slates-prompting-flux-2-max.md +2 -3
  64. package/skills/slates-prompting-gpt-image-2-5.md +2 -2
  65. package/skills/slates-prompting-inworld-tts.md +174 -174
  66. package/skills/slates-prompting-kling-v3.md +11 -9
  67. package/skills/slates-prompting-lip-sync.md +15 -15
  68. package/skills/slates-prompting-ltx-2-5.md +5 -6
  69. package/skills/slates-prompting-minimax-h3.md +11 -11
  70. package/skills/slates-prompting-motion-transfer.md +8 -8
  71. package/skills/slates-prompting-nano-banana-2.md +8 -4
  72. package/skills/slates-prompting-omni-flash.md +9 -9
  73. package/skills/slates-prompting-seed-audio.md +24 -4
  74. package/skills/slates-prompting-seedance-2-5.md +40 -30
  75. package/skills/slates-prompting-seedance.md +4 -4
  76. package/skills/slates-prompting-seedream-5-lite.md +6 -6
  77. package/skills/slates-restyle-from-blocking.md +2 -2
  78. package/skills/slates-script-craft.md +1 -1
  79. package/skills/slates-shot-variety.md +1 -1
  80. package/skills/slates-storyboard-from-script.md +1 -1
  81. package/skills/slates-style-prompting.md +56 -54
  82. package/skills/slates-ugc-influencer-ad.md +1 -1
  83. package/skills/slates-vision-feedback-loop.md +118 -110
  84. package/skills/slates-prompting-veo-3.md +0 -224
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-previs-blocking
3
- description: Build a 3D blocking pass in Blender, render it grey-box, and use it as a reference video so the generated shot follows a camera path you designed instead of one the model invented. Use when the user wants precise camera control, a multi-cut sequence, a one-take move, spatial consistency across shots, or says the camera keeps drifting / they keep burning credits re-rolling.
3
+ description: "Build and render a Blender blocking pass for precise camera paths, cut timing or spatial continuity, then guide video generation with the clip. Use when those controls are required or prompting has failed to hold them."
4
4
  ---
5
5
 
6
6
  # Previs blocking — design the shot, then generate it
@@ -89,17 +89,48 @@ The whole of `slates-camera-language`. Build the rig, then keyframe it. Then **r
89
89
 
90
90
  Add it after the moves are right, never before — noise on top of a wrong path just hides the wrong path.
91
91
 
92
+ <!-- @inject:blender-action-curves -->
93
+ ## Read animation curves from the active action layout
94
+
95
+ Blender 5 uses layered actions: curves belong to the channelbag for `animation_data.action_slot`, inside each layer's strips. A direct `action.fcurves` lookup failed on Blender 5.2.1 in the 2026-08-28 blocking run. Feature-detect the layout before changing interpolation or noise; an unanimated object can legitimately have no curves.
96
+
97
+ The snippets below use this small Blender-side iterator. It runs inside Blender; no add-on code is imported into the MCP package.
98
+
99
+ ```python
100
+ def action_curves(datablock):
101
+ anim = getattr(datablock, "animation_data", None)
102
+ action = getattr(anim, "action", None)
103
+ if action is None:
104
+ return
105
+ if hasattr(action, "fcurves"):
106
+ yield from action.fcurves
107
+ elif getattr(anim, "action_slot", None) is not None:
108
+ for layer in action.layers:
109
+ for strip in layer.strips:
110
+ if hasattr(strip, "channelbag"):
111
+ bag = strip.channelbag(anim.action_slot)
112
+ if bag is not None:
113
+ yield from bag.fcurves
114
+ ```
115
+
116
+ Use the datablock that owns the keyed property: the curve data for `eval_time`, the object for location and rotation, the camera data for lens. Confirm a named channel exists before assuming a keyframe operation created it.
117
+ <!-- @end:blender-action-curves -->
118
+
92
119
  ### 5. Verify the cuts
93
120
 
94
121
  The one check that catches the most damage: on a multi-cut blocking, camera position, target and focal length must all change **exactly on the cut frame, with no transition frame between**. One interpolated frame reads as a whip-pan the model will faithfully reproduce.
95
122
 
96
123
  ```python
97
- # Every camera f-curve keyframe on a cut frame must be CONSTANT out of the
98
- # previous key, or the cut smears.
99
- for fc in cam.animation_data.action.fcurves:
100
- for kp in fc.keyframe_points:
101
- if int(kp.co[0]) in CUT_FRAMES:
102
- kp.interpolation = 'CONSTANT'
124
+ # On a jump-cut camera, key the final pre-cut pose at cut_frame - 1.
125
+ # CONSTANT belongs to that preceding key: interpolation controls its OUTGOING segment.
126
+ # Check object transforms and camera data (including lens). Repeat for a keyed target.
127
+ for owner in (cam, cam.data):
128
+ for fc in action_curves(owner):
129
+ keys = list(fc.keyframe_points)
130
+ for previous, current in zip(keys, keys[1:]):
131
+ if current.co[0] in CUT_FRAMES:
132
+ assert previous.co[0] == current.co[0] - 1, "Key the final pre-cut state first"
133
+ previous.interpolation = 'CONSTANT'
103
134
  ```
104
135
 
105
136
  Also check nothing interpenetrates — proxies through floors, clones through the hero object, letters through each other. The model renders intersections as faithfully as it renders everything else.
@@ -118,27 +149,27 @@ bpy.ops.wm.save_as_mainfile(filepath=path, copy=True)
118
149
  slates_blender_render_blocking { projectId, fps: 24 }
119
150
  ```
120
151
 
121
- Renders the **scene camera** through scene settings — never the user's viewport, so the result does not depend on where they left their mouse — imports the mp4 into the project, and returns `assetId` + `durationSeconds`.
152
+ Renders the **scene camera** through scene settings and imports the mp4 into the project. The result does not depend on where the user left their viewport or mouse. It returns `asset.id` + `durationSeconds`.
122
153
 
123
154
  Then:
124
155
 
125
156
  ```
126
157
  slates_generate_video {
127
158
  model: "seedance-2.5",
128
- videoReferenceAssetIds: [<the blocking asset>],
159
+ videoReferenceAssetIds: [<asset.id>],
129
160
  videoReferenceSecondsEach: [<durationSeconds>],
130
161
  characterAssetIds: [...], environmentAssetIds: [...], styleAssetIds: [...],
131
162
  prompt: <written per slates-blocking-to-prompt>
132
163
  }
133
164
  ```
134
165
 
135
- **Four inputs, and that is the entire stack:** a character sheet each, one location/style reference, the blocking clip, and a prompt written against the blocking. Resist adding a fifth.
166
+ **A focused reference stack:** one identity sheet per character, any location or look reference the brief needs, the blocking clip, and a prompt written against it. Add a reference only for a distinct requirement; more competing references add variables rather than guaranteeing fidelity. An audio reference is valid when voice or sound continuity needs it and the selected endpoint supports it.
136
167
 
137
- Model note: seedance-2.5 is the seat for this — 10 reference videos at up to 30s each. seedance-2 and minimax-h3 take 3 at 15s. Route per `slates-model-selection`.
168
+ Choose a model that accepts video references using `slates-model-selection`, then read its current reference caps. Video duration limits apply to the combined reference clips, not to each clip independently; quote each actual input duration.
138
169
 
139
170
  ## Leaving holes on purpose
140
171
 
141
- Where the model outperforms any blockout you could build — liquid, smoke, fire, cloth — **block a black gap instead** and say so in the prompt: `CUT 7 (14.5-17.0, black gap in the reference)`. You are reserving a slot, not forgetting one.
172
+ Where the model outperforms any blockout you could build — liquid, smoke, fire, cloth — **block a black gap instead** and say so in the prompt: `CUT 7 (14.5-17.0, black gap in the reference)`. You are reserving a slot, not forgetting one. Keep these exact times in the blocking record, then translate model-facing time cues through `slates-blocking-to-prompt`; not every endpoint accepts fractional timestamps.
142
173
 
143
174
  ## What not to do
144
175
 
@@ -146,7 +177,7 @@ Where the model outperforms any blockout you could build — liquid, smoke, fire
146
177
  - **Don't animate what you don't need.** Heads especially — a proxy head turning wrong is worse than one that never turns.
147
178
  - **Don't build the camera before the geometry.** It has nothing to aim at, and every value you set gets redone.
148
179
  - **Don't skip reading the scene back.** Write timings from `slates_blender_scene`'s `cutSeconds`, never from what you intended to build.
149
- - **Don't exceed the model's reference-video ceiling.** A 40s blocking against a 30s cap silently truncates.
180
+ - **Don't exceed the model's reference-video ceiling.** A 40s blocking against a 30s cap is rejected; trim it first.
150
181
 
151
182
  ## Related
152
183
 
@@ -1,13 +1,13 @@
1
1
  ---
2
2
  name: slates-project-organization
3
- description: Use when the user names an asset by code ("use IMG-A36"), asks what a code means, or is organizing or navigating a project. Covers the asset short-code system (IMG-A12 / VID-V3 / AUD-S1 badges on every gallery card), folders for film STRUCTURE, and the typed tabs for reusable references.
3
+ description: "Navigate and organize Slates projects, asset codes such as IMG-A12 or VID-V3, folders, Library references and templates. Use when locating media, preparing a reusable production or keeping a project legible."
4
4
  ---
5
5
 
6
6
  # Organizing a Slates project
7
7
 
8
8
  Slates already gives every REUSABLE reference a home — the **Library**, in categories the user names (Characters, Locations, Products, Looks…; `slates_list_library`), each item used in a prompt as `@name`, or `#name` for a look. Do NOT recreate those as folders, and never invent a category the user did not ask for. Folders are for **structure**, never type.
9
9
 
10
- **Folders = where an asset sits in the FILM**, and they mirror to real subfolders on disk (`projects/<id>/…`), so a human can open the project in Resolve/Finder and navigate it like an edit. Use them for work product, not references.
10
+ **Folders = where an asset sits in the FILM.** They are in-app structure only. Use them for work product, not references.
11
11
 
12
12
  Create with `slates_create_folder`; file assets with `slates_move_assets_to_folder`. Generations land in the project's active folder, so set it before a batch.
13
13
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-elevenlabs
3
- description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before calling slates_generate_audio with model eleven-sfx — ONE short effect with an EXACT duration, or a seamless loop, billed per second. Covers describing an effect by its physical cause, the one-sound-per-generation rule, picking a duration, loops, prompt_influence, and when to use Seed Audio instead.
3
+ description: "Prompt ElevenLabs Sound Effects v2 (eleven-sfx) for a single effect or loop. Use with slates_generate_audio on this model; covers physical causes, duration, material, space and prompt influence."
4
4
  ---
5
5
 
6
6
  # ElevenLabs Sound Effects v2 — prompting
@@ -15,7 +15,7 @@ description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before ca
15
15
  Keep it under 2,400 characters (the build fails above that) and keep the
16
16
  rationale, the receipts and the worked examples in the body below. -->
17
17
  <!-- /slates-only -->
18
- **Card — ElevenLabs Sound Effects v2.** ONE short sound with an exact length, or a seamless loop. The only Slates audio surface with a real duration control and a real loop mode.
18
+ **Card: ElevenLabs Sound Effects v2.** ONE short sound with an exact length, or a seamless loop. A Slates audio surface with a real duration control and a real loop mode.
19
19
 
20
20
  **The five levers**
21
21
  1. **Describe the physical CAUSE, not the label** — `heavy oak door slams shut`, `boot scuffs on grit`, `a latch drops home`.
@@ -42,7 +42,7 @@ description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before ca
42
42
  - `door sound`, `whoosh`, `footsteps`, `impact`, `ambience` standing alone
43
43
  <!-- @banned:end -->
44
44
 
45
- One short sound with an exact length, carried on fal (`fal-ai/elevenlabs/sound-effects/v2`). This is the only Slates audio surface with a real duration control and a real loop mode.
45
+ One short sound with an exact length, carried on fal (`fal-ai/elevenlabs/sound-effects/v2`). This Slates audio surface has a real duration control and a real loop mode.
46
46
 
47
47
  ## Where it routes
48
48
 
@@ -87,7 +87,7 @@ Slates **always sends** `durationSeconds`. (Left null the model picks, which mak
87
87
  **The thresholds, from the code that enforces them:**
88
88
 
89
89
  - **Confirm gate:** above **17 credits** an op returns `requires_confirm` and will not
90
- proceed until you re-call with `confirm: true`. Below it, announce the cost once and go.
90
+ proceed until you re-call with `confirm: true`. This is a code gate, not permission to spend: every generation still needs the user-approved plan or quote.
91
91
  - **Deviation pause:** the desktop Studio Agent stops and re-asks when projected generation spend
92
92
  exceeds the approved plan by more than **20%**. You do not trigger this; the app does.
93
93
  - **Seed Audio duration:** **3–120 seconds.** There is no duration
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-flux-2-max
3
- description: How to prompt FLUX.2 Max (Black Forest Labs image model). Read before calling slates_generate_image with model flux-2-max, or slates_edit_image with editModel flux-2-max. FLUX.2 wants front-loaded structure, real camera vocabulary, and positive-only phrasing — no negative prompts, no tag soup.
3
+ description: "Prompt or edit images with FLUX.2 Max (flux-2-max). Use with slates_generate_image or slates_edit_image on this model; covers word order, camera vocabulary, materials, colours and positive phrasing."
4
4
  ---
5
5
 
6
6
  # FLUX.2 Max — prompting
@@ -130,7 +130,7 @@ Use natural language for exploration, JSON when the layout is locked and you're
130
130
 
131
131
  ## Reference images (edit path)
132
132
 
133
- In Slates, pass `referenceAssetIds` on `slates_generate_image` — FLUX routes them through its edit endpoint. Slates names each reference inline in the prompt ("the subject (image 1), the style (image 2)") in the order it sends them, so you don't hand-write role labels; the name carries the role and unnamed-by-position blending is avoided. For surgical changes to one existing image use `slates_edit_image` with `editModel: flux-2-max` (note: FLUX edits ignore extra referenceAssetIds — that's NB2-only).
133
+ In Slates, pass `referenceAssetIds` on `slates_generate_image` — FLUX routes them through its edit endpoint. Slates names each reference inline in the prompt ("the subject (image 1), the style (image 2)") in the order it sends them, so you don't hand-write role labels; the name carries the role and unnamed-by-position blending is avoided. For surgical changes to one existing image use `slates_edit_image` with `editModel: flux-2-max`. Extra `referenceAssetIds` are supported within the current edit-reference cap, with the source occupying one slot; read the tool schema for that cap. An older desktop without this capability refuses the request rather than silently omitting references.
134
134
 
135
135
  ### Reference rules (the verified ones)
136
136
 
@@ -166,7 +166,6 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
166
166
  ### For FLUX.2 Max specifically
167
167
 
168
168
  - **FLUX caps references well below NB2's 14, so rule 1's "2-4" is a ceiling here, not a starting point.** Be deliberate about which roles earn a slot.
169
- - **Rule 9 has a hard edge on this model:** `slates_edit_image` with `editModel: flux-2-max` ignores extra `referenceAssetIds` — that is NB2-only. A FLUX edit sees the source image and the prompt, nothing else.
170
169
  - **FLUX has no memory between generations, so rule 7 is enforced by repetition.** Define the character exhaustively once and repeat those exact descriptors verbatim in every subsequent prompt — see Character consistency across a series below.
171
170
 
172
171
  ## Character consistency across a series
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-gpt-image-2-5
3
- description: Prompt and edit images with GPT Image 2.5 Flare or Sunburst. Covers reference roles, realistic lighting, text, grids, quality choices and targeted edits. Use with slates_generate_image or slates_edit_image on these models.
3
+ description: "Prompt or edit images with GPT Image 2.5 Flare or Sunburst. Use with slates_generate_image or slates_edit_image on these models; covers reference roles, lighting, text, panels, quality choices and edits."
4
4
  ---
5
5
 
6
6
  # GPT Image 2.5 — sheets, grids, and text that actually reads
@@ -135,7 +135,7 @@ Reference images route through the edit endpoint, **up to 16** — fal's documen
135
135
 
136
136
  ## Transparent backgrounds
137
137
 
138
- If you need a cut-out rather than a scene, **ask for it explicitly and check the alpha**. OpenAI: request `background=transparent` and use PNG or WebP, then *"check the decoded image's alpha channel, including hair, glass, shadows, and object edges"* — a painted-white backdrop is the common failure and it is not transparency. Say what must NOT appear: *"no solid backdrop, no checkerboard, no scenery, no watermark"*, and do not let the product get restyled while the background is removed. **On every follow-up edit, repeat the transparency requirement** or it gets dropped. (Slates always requests PNG, so the format half is handled for you. **`background` IS surfaced now** — the Background control on the prompt bar, and `backgroundMode` on `slates_generate_image` / `slates_edit_image`. It is free: fal prices this family on size × quality alone.)
138
+ If you need a cut-out rather than a scene, **ask for it explicitly and check the alpha**. OpenAI: request `background=transparent` and use PNG or WebP, then *"check the decoded image's alpha channel, including hair, glass, shadows, and object edges"* — a painted-white backdrop is the common failure and it is not transparency. Say what must NOT appear: *"no solid backdrop, no checkerboard, no scenery, no watermark"*, and do not let the product get restyled while the background is removed. **On every follow-up edit, repeat the transparency requirement** or it gets dropped. <!-- slates-only -->(Slates always requests PNG, so the format half is handled for you. **`background` IS surfaced now** — the Background control on the prompt bar, and `backgroundMode` on `slates_generate_image` / `slates_edit_image`. It is free: fal prices this family on size × quality alone.)<!-- /slates-only -->
139
139
 
140
140
  ## When an edit must not touch a region at all
141
141
 
@@ -1,174 +1,174 @@
1
- ---
2
- name: slates-prompting-inworld-tts
3
- description: How to use Inworld Realtime TTS-2, the VOICE seat. Read before calling slates_generate_audio with model inworld-tts-2. Speech in a SPECIFIC voice, billed per character - the prompt is the words spoken, verbatim. Covers the identity-versus-acoustics rule (what a reference clip does and does not carry), how to write a line so it is performed rather than read, when to reach for seed-audio instead, and the voice-consent rule.
4
- ---
5
-
6
- # Inworld Realtime TTS-2 — the voice seat
7
-
8
- <!-- @card:start -->
9
- <!-- slates-only -->
10
- <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
- src/prompts/craft-cards.ts and returned on every cost estimate for this
12
- model, so it is the ONE piece of positive craft guidance the agent cannot
13
- skip. Keep it under 2,400 characters (the build fails above that) and keep
14
- the rationale and the worked examples in the body below. -->
15
- <!-- /slates-only -->
16
- **Card — Inworld TTS-2.** Speech in a SPECIFIC voice. The prompt is the words spoken, verbatim — not a description of them. Text length determines the bill.
17
-
18
- **IDENTITY, NOT ACOUSTICS — the rule that decides whether cloning works**
19
- A reference carries WHO is speaking: timbre, pitch, accent, age, vowel shape. It does NOT carry WHERE they are — room tone, distance, phone EQ, reverb and mic character are *acoustics*, and this model reproduces the identity while discarding the room. So:
20
- 1. **A noisy reference does not give a noisy read — it gives a WORSE identity.** Music, a second speaker or heavy reverb corrupt what is being extracted. Use a clean single-speaker recording.
21
- 2. **You cannot get "on a payphone" by cloning a payphone recording.** Acoustics come from the MIX, or from `seed-audio` which renders a room.
22
-
23
- **DIRECTION GOES IN SQUARE BRACKETS. PARENTHESES ARE SPOKEN ALOUD.** `[whispering] I hope nobody notices` is whispered; `(quietly) I hope nobody notices` says the word "quietly" out loud. Verified by ear — the easiest way to ruin a take.
24
-
25
- - **Plain English works inside them** — it is natural-language steering, not a fixed vocabulary: `[very quiet]`, `[whisper in a hushed style]`, `[very slow]`, `[say excitedly]`. Non-verbals are their own tags: `[laugh]`, `[sigh]`, `[breathe]`, `[clear throat]`.
26
- - **A tag it does not recognise is still consumed, and still changes the read.** Never spoken, never an error — so a mistyped tag fails SILENTLY and only listening catches it.
27
- - **Tags persist across sentences** until changed; `[reset]` returns to normal.
28
- - **Punctuation is the timing.** `Wait. Stop.` differs from `Wait, stop.`
29
- - **One line, one take.** Split a paragraph so a bad clause costs one re-roll.
30
- - **Spell numbers and titles aloud:** `twenty twenty-six`, `Doctor Reyes`.
31
-
32
- **Route elsewhere when:** the scene needs dialogue mixed with effects and room tone in one pass (`seed-audio`), or it is a single non-speech sound (`eleven-sfx`). This surface makes ONE voice saying ONE thing, cleanly.
33
-
34
- **Hard constraints:** no duration parameter — length falls out of the text. Exactly one voice source: a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId` (a character's voice clip to speak AS the character, or any clean clip of one speaker), or `voiceDescription`.
35
- <!-- @card:end -->
36
-
37
- <!-- @banned:start -->
38
- <!-- slates-only -->
39
- <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
40
- extracted by src/prompts/banned-tokens.ts and returned on this model's cost
41
- estimate, and every submitted prompt is matched against it. Keep entries
42
- backticked and prose outside the backticks. -->
43
- <!-- /slates-only -->
44
- **Never use** — the prompt on this surface is SPOKEN ALOUD, so anything that describes the audio instead of being the audio gets read out as words:
45
-
46
- - `SFX`, `Ambient noise`, `Background music` as labels — this model speaks; it does not render a scene. Use `seed-audio` for those.
47
- - shot language: `wide shot`, `slow push in`, `warm tungsten` — video-prompt words, and here they would literally be said aloud
48
- - `voiceover`, `narrator says`, `he says` as stage directions wrapping the line — write only the words that should come out of the speaker
49
- <!-- @banned:end -->
50
-
51
- ## Why the prompt is not a prompt
52
-
53
- On every other surface in Slates the prompt DESCRIBES what you want and the model interprets it. Here the prompt IS the deliverable: each character is spoken aloud and each character is billed. `a gravelly man says he is tired` produces a voice saying the words "a gravelly man says he is tired".
54
-
55
- That also means the two numbers a user cares about are the same number. The text length sets the price (in 250-character buckets) and sets the length of the audio. There is nothing to choose and nothing to reconcile.
56
-
57
- ## Steering the delivery
58
-
59
- `VERIFIED BY EAR, 2026-09-05.` Every claim in this section was listened to, not
60
- inferred — an earlier draft of this skill documented tag forms that had only been
61
- probed for an HTTP 200, which proves the request was accepted and nothing about
62
- whether it was obeyed.
63
-
64
- **Square brackets are consumed. Parentheses are read aloud.** That is the whole
65
- rule, and getting it wrong is not a subtle degradation — the audience hears a
66
- narrator say the word "quietly" in the middle of your line.
67
-
68
- | Written | What comes out |
69
- |---|---|
70
- | `[whispering] I really hope nobody notices that.` | whispered, tag not spoken ✅ |
71
- | `[very quiet] I really hope nobody notices that.` | very quiet, tag not spoken ✅ |
72
- | `[whisper in a hushed style] …` | hushed, tag not spoken ✅ |
73
- | `[very slow] …` | slowed right down, tag not spoken ✅ |
74
- | `[laugh] …` | an actual laugh, then the line ✅ |
75
- | `(quietly, under his breath) …` | 🚨 **the words "quietly, under his breath" are SPOKEN** |
76
-
77
- **Plain English works — it is natural-language steering, not a fixed vocabulary.**
78
- Both the documented phrasings (`[whisper in a hushed style]`) and ordinary adverbs
79
- (`[whispering]`) were obeyed. Write the direction the way you would say it to an
80
- actor.
81
-
82
- The eight dimensions the model steers on, with a working example of each:
83
-
84
- | Dimension | Example |
85
- |---|---|
86
- | Emotion | `[say excitedly]`, `[sound sad]`, `[sound terrified]` |
87
- | Articulation | `[say with force]`, `[articulate clearly]` |
88
- | Intonation | `[say with a rising pitch]` |
89
- | Volume | `[very quiet]`, `[very loud]` |
90
- | Pitch | `[say in a low tone]` |
91
- | Range | `[say playfully]`, `[say with no pitch variation]` |
92
- | Speed | `[very fast]`, `[very slow]` |
93
- | Vocal style | `[whisper in a hushed style]`, `[give a nasal quality]` |
94
-
95
- Non-verbals sit inline where they happen: `[laugh]`, `[sigh]`, `[cough]`,
96
- `[breathe]`, `[yawn]`, `[clear throat]`.
97
-
98
- ### Four rules that are not obvious
99
-
100
- 1. 🚨 **A tag it does not recognise is still consumed, and still changes the read.**
101
- `[zzzqqq]` is not spoken and does not error — it produces a different, arbitrary
102
- delivery. So a typo in a tag is SILENT: there is no rejection, no warning, and no
103
- way to catch it except listening to the take. Treat an unexpected performance as
104
- a possible misspelled tag before you blame the voice.
105
- 2. **Tags persist across sentences.** A `[very slow]` at the top governs everything
106
- after it until something changes it. Use `[reset]` to go back to normal rather
107
- than assuming the next sentence starts clean.
108
- 3. **Do not stack opposing directions.** `[whisper in a hushed style]` together with
109
- `[very loud]` produces unpredictable results — the model is resolving a
110
- contradiction, and which side wins is not something you can rely on.
111
- 4. **Tags COUNT toward the billed characters**, even though they are never spoken.
112
- They are part of the text sent to the vendor, so the vendor charges for them and
113
- so do we — billing what was actually sent is the only honest basis. It rarely
114
- matters (a 13-character tag inside a 250-character bucket), but a line sitting
115
- just under a bucket boundary can be pushed into the next one by a long
116
- direction. Prefer `[very slow]` over `[say this one very slowly please]`.
117
-
118
- ## Identity versus acoustics, at length
119
-
120
- This is the distinction that decides whether the feature feels good, and it is worth being precise about because the failure is quiet — you get a usable clip that is subtly not the person.
121
-
122
- **What a reference clip transfers:** vocal timbre, pitch range, accent and regional vowels, apparent age, speech rate tendencies, and the particular rasp or breathiness of the source speaker.
123
-
124
- **What it does not transfer:** the room, the microphone, the codec, the distance from the mic, any processing on the source, and any other sound present in it.
125
-
126
- So the ideal reference is boring: one person, close to a microphone, no music, no second speaker, no heavy reverb, five to fifteen seconds, speaking normally rather than performing. A phone voice memo in a quiet room beats a beautifully produced clip with a music bed underneath it.
127
-
128
- **Two failure modes, both common:**
129
-
130
- - *"I cloned my podcast intro and it doesn't sound like me."* The intro had music under it. The model averaged the music into the identity. Re-clone from a clean stretch.
131
- - *"I want the line to sound like it's coming through a car radio."* Clone the clean voice, then EQ and process the returned clip on the timeline. A radio-sounding reference makes a worse voice, not a radio effect.
132
-
133
- ## Getting the voice onto the call
134
-
135
- Exactly one source per call, and none of them requires a character to exist first:
136
-
137
- - **A preset:** `slates_list_voices` lists stock voices with gender, age, accent and tags — filter by any of them, or search the descriptions ("gravelly", "narration"). Pass the chosen `voiceId`. Presets clone nothing, so they are the fastest path and avoid the clone-creation rate ceiling.
138
- - **Speak AS a character:** `voiceReferenceAssetId: <its voiceAssetId>` (the clip on the row `slates_list_characters` returns). The seat clones the clip for that take and discards the vendor voice afterwards, so there is nothing to reconcile — but cloning shares a ceiling of two new voices a minute across every Slates user, so a run of lines in one cloned voice pauses between takes rather than failing. Send each line once; do not re-send one that already came back. Any other clean clip of one speaker works the same way.
139
- - **A voice with no recording:** `voiceDescription` (7–1000 characters of words). If it will be used again, keep the returned clip on a character with `slates_update_character` (`voiceAssetId`) so later lines clone the same clip instead of designing a new voice each time — a convenience, never a requirement.
140
-
141
- ## Consent
142
-
143
- Cloning a real person's voice needs that person's explicit, documented permission, scoped to what you are making. Clone from original human recordings only — never from another model's output. This is the same gate the real-face route applies to likeness, and it applies here for the same reason.
144
-
145
- ## Worked examples
146
-
147
- **A line with a direction**
148
-
149
- ```
150
- [very quiet] I heard what you said in there. I'm not going to pretend I didn't.
151
- ```
152
-
153
- **A line that needs its numbers spoken**
154
-
155
- ```
156
- The vote was three hundred and twelve to eighty-nine. It carried at four minutes past midnight.
157
- ```
158
-
159
- **A paragraph, split into three takes** — so one bad clause costs one re-roll:
160
-
161
- ```
162
- 1. You keep asking me why I stayed.
163
- 2. It wasn't loyalty. It wasn't even fear, not by the end.
164
- 3. [very slow] It was that I couldn't picture the version of me that left.
165
- ```
166
-
167
- **What NOT to send**
168
-
169
- ```
170
- (gravelly, tired) a tired old man narrates the opening of the film, wide shot, warm tungsten
171
- ```
172
-
173
- Every word of that is spoken aloud — **including the parenthetical**, which is the
174
- trap: it looks like a stage direction and is treated as dialogue. Describe the voice when you are CHOOSING one (`voiceDescription`, or the desktop's voice picker); the prompt is only ever the words.
1
+ ---
2
+ name: slates-prompting-inworld-tts
3
+ description: "Direct speech with Inworld Realtime TTS-2 (inworld-tts-2). Use with slates_generate_audio on this model; covers voice identity, reference acoustics, delivery tags, punctuation and voice consent."
4
+ ---
5
+
6
+ # Inworld Realtime TTS-2 — the voice seat
7
+
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Keep it under 2,400 characters (the build fails above that) and keep
14
+ the rationale and the worked examples in the body below. -->
15
+ <!-- /slates-only -->
16
+ **Card — Inworld TTS-2.** Speech in a SPECIFIC voice. The prompt is the words spoken, verbatim — not a description of them. Text length determines the bill.
17
+
18
+ **IDENTITY, NOT ACOUSTICS — the rule that decides whether cloning works**
19
+ A reference carries WHO is speaking: timbre, pitch, accent, age, vowel shape. It does NOT carry WHERE they are — room tone, distance, phone EQ, reverb and mic character are *acoustics*, and this model reproduces the identity while discarding the room. So:
20
+ 1. **A noisy reference does not give a noisy read — it gives a WORSE identity.** Music, a second speaker or heavy reverb corrupt what is being extracted. Use a clean single-speaker recording.
21
+ 2. **You cannot get "on a payphone" by cloning a payphone recording.** Acoustics come from the MIX, or from `seed-audio` which renders a room.
22
+
23
+ **DIRECTION GOES IN SQUARE BRACKETS. PARENTHESES ARE SPOKEN ALOUD.** `[whispering] I hope nobody notices` is whispered; `(quietly) I hope nobody notices` says the word "quietly" out loud. Verified by ear — the easiest way to ruin a take.
24
+
25
+ - **Plain English works inside them** — it is natural-language steering, not a fixed vocabulary: `[very quiet]`, `[whisper in a hushed style]`, `[very slow]`, `[say excitedly]`. Non-verbals are their own tags: `[laugh]`, `[sigh]`, `[breathe]`, `[clear throat]`.
26
+ - **A tag it does not recognise is still consumed, and still changes the read.** Never spoken, never an error — so a mistyped tag fails SILENTLY and only listening catches it.
27
+ - **Tags persist across sentences** until changed; `[reset]` returns to normal.
28
+ - **Punctuation is the timing.** `Wait. Stop.` differs from `Wait, stop.`
29
+ - **One line, one take.** Split a paragraph so a bad clause costs one re-roll.
30
+ - **Spell numbers and titles aloud:** `twenty twenty-six`, `Doctor Reyes`.
31
+
32
+ **Route elsewhere when:** the scene needs dialogue mixed with effects and room tone in one pass (`seed-audio`), or it is a single non-speech sound (`eleven-sfx`). This surface makes ONE voice saying ONE thing, cleanly.
33
+
34
+ **Hard constraints:** no duration parameter — length falls out of the text. Exactly one voice source: a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId` (a character's voice clip to speak AS the character, or any clean clip of one speaker), or `voiceDescription`.
35
+ <!-- @card:end -->
36
+
37
+ <!-- @banned:start -->
38
+ <!-- slates-only -->
39
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
40
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
41
+ estimate, and every submitted prompt is matched against it. Keep entries
42
+ backticked and prose outside the backticks. -->
43
+ <!-- /slates-only -->
44
+ **Never use** — the prompt on this surface is SPOKEN ALOUD, so anything that describes the audio instead of being the audio gets read out as words:
45
+
46
+ - `SFX`, `Ambient noise`, `Background music` as labels — this model speaks; it does not render a scene. Use `seed-audio` for those.
47
+ - shot language: `wide shot`, `slow push in`, `warm tungsten` — video-prompt words, and here they would literally be said aloud
48
+ - `voiceover`, `narrator says`, `he says` as stage directions wrapping the line — write only the words that should come out of the speaker
49
+ <!-- @banned:end -->
50
+
51
+ ## Why the prompt is not a prompt
52
+
53
+ On every other surface in Slates the prompt DESCRIBES what you want and the model interprets it. Here the prompt IS the deliverable: each character is spoken aloud and each character is billed. `a gravelly man says he is tired` produces a voice saying the words "a gravelly man says he is tired".
54
+
55
+ That also means the two numbers a user cares about are the same number. The text length sets the price (in 250-character buckets) and sets the length of the audio. There is nothing to choose and nothing to reconcile.
56
+
57
+ ## Steering the delivery
58
+
59
+ `VERIFIED BY EAR, 2026-09-05.` Every claim in this section was listened to, not
60
+ inferred — an earlier draft of this skill documented tag forms that had only been
61
+ probed for an HTTP 200, which proves the request was accepted and nothing about
62
+ whether it was obeyed.
63
+
64
+ **Square brackets are consumed. Parentheses are read aloud.** That is the whole
65
+ rule, and getting it wrong is not a subtle degradation — the audience hears a
66
+ narrator say the word "quietly" in the middle of your line.
67
+
68
+ | Written | What comes out |
69
+ |---|---|
70
+ | `[whispering] I really hope nobody notices that.` | whispered, tag not spoken ✅ |
71
+ | `[very quiet] I really hope nobody notices that.` | very quiet, tag not spoken ✅ |
72
+ | `[whisper in a hushed style] …` | hushed, tag not spoken ✅ |
73
+ | `[very slow] …` | slowed right down, tag not spoken ✅ |
74
+ | `[laugh] …` | an actual laugh, then the line ✅ |
75
+ | `(quietly, under his breath) …` | 🚨 **the words "quietly, under his breath" are SPOKEN** |
76
+
77
+ **Plain English works — it is natural-language steering, not a fixed vocabulary.**
78
+ Both the documented phrasings (`[whisper in a hushed style]`) and ordinary adverbs
79
+ (`[whispering]`) were obeyed. Write the direction the way you would say it to an
80
+ actor.
81
+
82
+ The eight dimensions the model steers on, with a working example of each:
83
+
84
+ | Dimension | Example |
85
+ |---|---|
86
+ | Emotion | `[say excitedly]`, `[sound sad]`, `[sound terrified]` |
87
+ | Articulation | `[say with force]`, `[articulate clearly]` |
88
+ | Intonation | `[say with a rising pitch]` |
89
+ | Volume | `[very quiet]`, `[very loud]` |
90
+ | Pitch | `[say in a low tone]` |
91
+ | Range | `[say playfully]`, `[say with no pitch variation]` |
92
+ | Speed | `[very fast]`, `[very slow]` |
93
+ | Vocal style | `[whisper in a hushed style]`, `[give a nasal quality]` |
94
+
95
+ Non-verbals sit inline where they happen: `[laugh]`, `[sigh]`, `[cough]`,
96
+ `[breathe]`, `[yawn]`, `[clear throat]`.
97
+
98
+ ### Four rules that are not obvious
99
+
100
+ 1. 🚨 **A tag it does not recognise is still consumed, and still changes the read.**
101
+ `[zzzqqq]` is not spoken and does not error — it produces a different, arbitrary
102
+ delivery. So a typo in a tag is SILENT: there is no rejection, no warning, and no
103
+ way to catch it except listening to the take. Treat an unexpected performance as
104
+ a possible misspelled tag before you blame the voice.
105
+ 2. **Tags persist across sentences.** A `[very slow]` at the top governs everything
106
+ after it until something changes it. Use `[reset]` to go back to normal rather
107
+ than assuming the next sentence starts clean.
108
+ 3. **Do not stack opposing directions.** `[whisper in a hushed style]` together with
109
+ `[very loud]` produces unpredictable results — the model is resolving a
110
+ contradiction, and which side wins is not something you can rely on.
111
+ 4. **Tags COUNT toward the billed characters**, even though they are never spoken.
112
+ They are part of the text sent to the vendor, so the vendor charges for them and
113
+ so do we — billing what was actually sent is the only honest basis. It rarely
114
+ matters (a 13-character tag inside a 250-character bucket), but a line sitting
115
+ just under a bucket boundary can be pushed into the next one by a long
116
+ direction. Prefer `[very slow]` over `[say this one very slowly please]`.
117
+
118
+ ## Identity versus acoustics, at length
119
+
120
+ This is the distinction that decides whether the feature feels good, and it is worth being precise about because the failure is quiet — you get a usable clip that is subtly not the person.
121
+
122
+ **What a reference clip transfers:** vocal timbre, pitch range, accent and regional vowels, apparent age, speech rate tendencies, and the particular rasp or breathiness of the source speaker.
123
+
124
+ **What it does not transfer:** the room, the microphone, the codec, the distance from the mic, any processing on the source, and any other sound present in it.
125
+
126
+ So the ideal reference is boring: one person, close to a microphone, no music, no second speaker, no heavy reverb, five to fifteen seconds, speaking normally rather than performing. A phone voice memo in a quiet room beats a beautifully produced clip with a music bed underneath it.
127
+
128
+ **Two failure modes, both common:**
129
+
130
+ - *"I cloned my podcast intro and it doesn't sound like me."* The intro had music under it. The model averaged the music into the identity. Re-clone from a clean stretch.
131
+ - *"I want the line to sound like it's coming through a car radio."* Clone the clean voice, then EQ and process the returned clip on the timeline. A radio-sounding reference makes a worse voice, not a radio effect.
132
+
133
+ ## Getting the voice onto the call
134
+
135
+ Exactly one source per call, and none of them requires a character to exist first:
136
+
137
+ - **A preset:** `slates_list_voices` lists stock voices with gender, age, accent and tags — filter by any of them, or search the descriptions ("gravelly", "narration"). Pass the chosen `voiceId`. Presets clone nothing, so they are the fastest path and avoid the clone-creation rate ceiling.
138
+ - **Speak AS a character:** `voiceReferenceAssetId: <its voiceAssetId>` (the clip on the row `slates_list_characters` returns). The seat clones the clip for that take and discards the vendor voice afterwards, so there is nothing to reconcile — but cloning shares a ceiling of two new voices a minute across every Slates user, so a run of lines in one cloned voice pauses between takes rather than failing. Send each line once; do not re-send one that already came back. Any other clean clip of one speaker works the same way.
139
+ - **A voice with no recording:** `voiceDescription` (7–1000 characters of words). If it will be used again, keep the returned clip on a character with `slates_update_character` (`voiceAssetId`) so later lines clone the same clip instead of designing a new voice each time — a convenience, never a requirement.
140
+
141
+ ## Consent
142
+
143
+ Cloning a real person's voice needs that person's explicit, documented permission, scoped to what you are making. Clone from original human recordings only — never from another model's output. This is the same gate the real-face route applies to likeness, and it applies here for the same reason.
144
+
145
+ ## Worked examples
146
+
147
+ **A line with a direction**
148
+
149
+ ```
150
+ [very quiet] I heard what you said in there. I'm not going to pretend I didn't.
151
+ ```
152
+
153
+ **A line that needs its numbers spoken**
154
+
155
+ ```
156
+ The vote was three hundred and twelve to eighty-nine. It carried at four minutes past midnight.
157
+ ```
158
+
159
+ **A paragraph, split into three takes** — so one bad clause costs one re-roll:
160
+
161
+ ```
162
+ 1. You keep asking me why I stayed.
163
+ 2. It wasn't loyalty. It wasn't even fear, not by the end.
164
+ 3. [very slow] It was that I couldn't picture the version of me that left.
165
+ ```
166
+
167
+ **What NOT to send**
168
+
169
+ ```
170
+ (gravelly, tired) a tired old man narrates the opening of the film, wide shot, warm tungsten
171
+ ```
172
+
173
+ Every word of that is spoken aloud — **including the parenthetical**, which is the
174
+ trap: it looks like a stage direction and is treated as dialogue. Describe the voice when you are CHOOSING one (`voiceDescription`, or the desktop's voice picker); the prompt is only ever the words.