@slatesvideo/shared 0.7.2 → 0.7.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (84) hide show
  1. package/dist/clients/cloud.d.ts +4 -0
  2. package/dist/clients/cloud.js +11 -3
  3. package/dist/index.d.ts +1 -0
  4. package/dist/index.js +1 -0
  5. package/dist/manual/content.d.ts +1 -1
  6. package/dist/manual/content.js +1 -1
  7. package/dist/operations/index.d.ts +12 -13
  8. package/dist/operations/index.js +158 -133
  9. package/dist/operations/surface.d.ts +6 -2
  10. package/dist/operations/surface.js +29 -5
  11. package/dist/prompts/agent-doctrine.d.ts +4 -4
  12. package/dist/prompts/agent-doctrine.js +17 -28
  13. package/dist/prompts/guide-discovery.d.ts +23 -0
  14. package/dist/prompts/guide-discovery.js +39 -0
  15. package/dist/prompts/guide-retrieval.js +1 -1
  16. package/dist/prompts/model-capabilities.d.ts +8 -9
  17. package/dist/prompts/model-capabilities.js +11 -51
  18. package/dist/prompts/model-facts.d.ts +2 -2
  19. package/dist/prompts/model-facts.js +15 -26
  20. package/dist/prompts/partials.generated.js +6 -3
  21. package/dist/prompts/prompting-tips.d.ts +1 -1
  22. package/dist/prompts/prompting-tips.js +21 -63
  23. package/dist/prompts/search-terms.d.ts +3 -0
  24. package/dist/prompts/search-terms.js +24 -0
  25. package/dist/skills/content.js +36 -37
  26. package/dist/skills/metadata.d.ts +7 -0
  27. package/dist/skills/metadata.js +29 -0
  28. package/exports/slates-chatgpt-images/generated/SKILL.md +7 -1
  29. package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
  30. package/exports/slates-prompt-builder/generated/SKILL.md +28 -16
  31. package/exports/slates-prompt-builder/generated/reference-character.md +12 -13
  32. package/exports/slates-prompt-builder/generated/reference-content-policy.md +2 -2
  33. package/exports/slates-prompt-builder/generated/reference-gpt-image-2-5.md +191 -0
  34. package/exports/slates-prompt-builder/generated/reference-kling.md +32 -11
  35. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +24 -6
  36. package/exports/slates-prompt-builder/generated/reference-omni-flash.md +65 -0
  37. package/exports/slates-prompt-builder/generated/reference-seedance-2-5.md +362 -0
  38. package/exports/slates-prompt-builder/generated/reference-seedance.md +34 -4
  39. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +77 -23
  40. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  41. package/package.json +2 -1
  42. package/skills/_partials/blender-action-curves.md +24 -0
  43. package/skills/_partials/iteration-diagnosis.md +5 -0
  44. package/skills/_partials/model-routing.md +35 -0
  45. package/skills/_partials/seedance-25-timestamps.md +2 -2
  46. package/skills/_partials/still-gate.md +2 -2
  47. package/skills/_partials/thresholds.md +1 -1
  48. package/skills/slates-blocking-to-prompt.md +15 -13
  49. package/skills/slates-camera-language.md +45 -7
  50. package/skills/slates-character-identity.md +8 -6
  51. package/skills/slates-chatgpt-images.md +7 -1
  52. package/skills/slates-cinematic-look.md +1 -1
  53. package/skills/slates-content-policy.md +4 -6
  54. package/skills/slates-cost-discipline.md +18 -12
  55. package/skills/slates-dialogue-blocking.md +6 -6
  56. package/skills/slates-direct-response-ad.md +1 -1
  57. package/skills/slates-edit-and-iterate.md +12 -4
  58. package/skills/slates-model-selection.md +82 -90
  59. package/skills/slates-one-prompt-film.md +1 -1
  60. package/skills/slates-previs-blocking.md +44 -13
  61. package/skills/slates-project-organization.md +2 -2
  62. package/skills/slates-prompting-elevenlabs.md +4 -4
  63. package/skills/slates-prompting-flux-2-max.md +2 -3
  64. package/skills/slates-prompting-gpt-image-2-5.md +2 -2
  65. package/skills/slates-prompting-inworld-tts.md +174 -174
  66. package/skills/slates-prompting-kling-v3.md +11 -9
  67. package/skills/slates-prompting-lip-sync.md +15 -15
  68. package/skills/slates-prompting-ltx-2-5.md +5 -6
  69. package/skills/slates-prompting-minimax-h3.md +11 -11
  70. package/skills/slates-prompting-motion-transfer.md +8 -8
  71. package/skills/slates-prompting-nano-banana-2.md +8 -4
  72. package/skills/slates-prompting-omni-flash.md +9 -9
  73. package/skills/slates-prompting-seed-audio.md +24 -4
  74. package/skills/slates-prompting-seedance-2-5.md +40 -30
  75. package/skills/slates-prompting-seedance.md +4 -4
  76. package/skills/slates-prompting-seedream-5-lite.md +6 -6
  77. package/skills/slates-restyle-from-blocking.md +2 -2
  78. package/skills/slates-script-craft.md +1 -1
  79. package/skills/slates-shot-variety.md +1 -1
  80. package/skills/slates-storyboard-from-script.md +1 -1
  81. package/skills/slates-style-prompting.md +56 -54
  82. package/skills/slates-ugc-influencer-ad.md +1 -1
  83. package/skills/slates-vision-feedback-loop.md +118 -110
  84. package/skills/slates-prompting-veo-3.md +0 -224
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-kling-v3
3
- description: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_generate_video with kling-v3.0-std, kling-v3.0-pro, or kling-v3.0-omni. Kling has dialogue + SFX + ambient native syntax (Omni adds multi-character dialogue and language codes). Multi-shot rules differ from Seedance/Veo — don't cross syntaxes.
3
+ description: "Prompt Kling V3.0 video generation and Kling O3 video edits. Use with Kling models on slates_generate_video or slates_edit_video; covers subjects, dialogue, sound syntax, multi-shot direction and edit fidelity."
4
4
  ---
5
5
 
6
6
  # Kling V3.0 — prompting
@@ -15,7 +15,7 @@ description: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_gen
15
15
  Keep it under 2,400 characters (the build fails above that) and keep the
16
16
  rationale, the receipts and the worked examples in the body below. -->
17
17
  <!-- /slates-only -->
18
- **Card — Kling V3.0.** The general default. Define the core subjects clearly at the START and keep those descriptions identical across shots. Up to 15s, up to 6 cuts, and the strongest image-to-video identity hold in the catalogue.
18
+ **Card — Kling V3.0.** Define the core subjects clearly at the START and keep those descriptions identical across shots. Strong image-to-video identity hold; use the current capability surface for duration and multi-shot limits, and the model catalogue for routing.
19
19
 
20
20
  **The five levers**
21
21
  1. **Dialogue in quotes** — `Character says, "exact words here"`. On Omni, direct the voice with `Gender + Age + Voice quality + Speech rate + Emotional tone + Language`: `[Character A: Detective, mid-40s, raspy, slow cadence, weary]: "I've seen this before."`
@@ -44,7 +44,7 @@ description: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_gen
44
44
  - `single continuous take` — Seedance's phrase, and it fights Kling's multi-shot
45
45
  <!-- @banned:end -->
46
46
 
47
- Kuaishou's video model. Three tiers: `kling-v3.0-std` (general use, no audio), `kling-v3.0-pro` (higher visual quality, no audio), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).
47
+ Kuaishou's video model. Three tiers: `kling-v3.0-std` (general use, sound supported), `kling-v3.0-pro` (higher visual quality, sound supported), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).
48
48
 
49
49
  Up to 15s. Multi-shot supported (up to 6 cuts in 15s total). Strong on image-to-video — preserves identity, layout, and text from the input image well.
50
50
 
@@ -134,7 +134,9 @@ Miss conditions:
134
134
  - Mixing camera moves within a shot ("pan then orbit then push in")
135
135
  - Extreme wide → extreme close in adjacent shots without reference images
136
136
 
137
- ## Element references (Omni)
137
+ ## Element references
138
+
139
+ Standard and Pro take element references with a first frame; Omni also takes references without one. 4K refuses reference images.
138
140
 
139
141
  Upload 2-4 multi-angle reference photos per character/object. Tag inline:
140
142
 
@@ -199,11 +201,11 @@ Layer scene-specific suppressions on top, and never suppress something the promp
199
201
 
200
202
  ## Tier choice
201
203
 
202
- - **Standard**: general use, no audio
203
- - **Pro**: higher visual quality, no audio
204
- - **Omni**: multi-character dialogue, audio-visual co-gen, language codes, `@elementN` references
204
+ - **Standard**: general use, sound supported
205
+ - **Pro**: higher visual quality, sound supported
206
+ - **Omni**: multi-character dialogue, audio-visual co-gen, language codes, references without a first frame
205
207
 
206
- Pick by capability: need dialogue/audio → Omni; need maximum visual quality silent → Pro; everything else → Standard. Prices change — check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.
208
+ Every tier can generate dialogue and sound. Sound is on unless `sound: false` is passed; below 4K it bills the audio key, while 4K includes audio. Pick by visual quality and reference needs. Prices change; check current numbers before choosing a tier<!-- slates-only -->; call `slates_estimate_generation_cost` or `slates_list_available_models`<!-- /slates-only -->.
207
209
 
208
210
  ## Benchmark prompt structure
209
211
 
@@ -253,7 +255,7 @@ Rules:
253
255
  - One edit intent per pass. Chain passes for compound changes (each output is itself an editable clip, linked to its parent).
254
256
  - Billing is per second of OUTPUT ≈ the clip length, rounded UP to the next second. A 7.3s clip bills as 8s.
255
257
  - Clip constraints: 3-15s, 720-3840px, MP4/MOV. Agents can pre-trim on the timeline when a clip runs long.
256
- - Routing: Kling edit is the default edit tool (element lock + audio intact); Seedance edit/relocate wins style-transfer-heavy re-imaginings<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.
258
+ - Route by the required change: this edit seat supports element/style-reference control and original-audio retention. Read the current catalogue for defaults and competing seats<!-- slates-only --> — see `slates-model-selection`<!-- /slates-only -->.
257
259
 
258
260
  ## Sources
259
261
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-lip-sync
3
- description: How to set up lip-sync — Kling-only (dedicated lip-sync and avatar endpoints, 5-second outputs). Read before calling slates_generate_lip_sync. Two flows — video→video re-dub and image→video avatar — with different inputs, pricing, and gotchas. Voice catalog, framing rules, audio file constraints, and which tier to pick. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.
3
+ description: "Prepare Kling lip-sync or avatar generation with slates_generate_lip_sync. Covers source selection, voice, framing, audio constraints and the separate Seedance video-reference alternative."
4
4
  ---
5
5
 
6
6
  # Lip-sync — setup guide
@@ -15,7 +15,7 @@ description: How to set up lip-sync — Kling-only (dedicated lip-sync and avata
15
15
  Keep it under 2,400 characters (the build fails above that) and keep the
16
16
  rationale, the receipts and the worked examples in the body below. -->
17
17
  <!-- /slates-only -->
18
- **Card — Lip-sync (Kling only).** Two different flows with different inputs and different prices; every output is 5 seconds.
18
+ **Card: Lip-sync (Kling only).** Two different flows with different inputs and different prices; output follows the source clip for video or the voice track for a still, billed per 5s block.
19
19
 
20
20
  **The five levers**
21
21
  1. **Pick `sourceType` deliberately** — `video` re-dubs an existing talking head (cheapest); `image` animates a still portrait (avatar-standard, then avatar-pro only on the final selected take).
@@ -28,7 +28,7 @@ description: How to set up lip-sync — Kling-only (dedicated lip-sync and avata
28
28
  - `Soft rim light, warm office, gentle confident smile between sentences.`
29
29
  - `Cool blue evening light through a window, focused intent expression.` (Or `.` — an empty prompt is fine when you have nothing to add.)
30
30
 
31
- **Hard constraint:** it is Kling-only and always 5 seconds. For a generated PERFORMANCE instead — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal — that is a normal Seedance video generation with the clip attached as a video reference, not a mode of this tool. A real recording, or a cloned/cast voice rendered on `inworld-tts-2`, for production; this tool's built-in TTS is for scratch.
31
+ **Hard constraint:** it is Kling-only; output follows the media and bills per 5s block. For a generated PERFORMANCE instead (head movement, gesture, delivery energy, with the dialogue as a native conditioning signal), that is a normal Seedance video generation with the clip attached as a video reference, not a mode of this tool. A real recording, or a cloned/cast voice rendered on `inworld-tts-2`, for production; this tool's built-in TTS is for scratch.
32
32
  <!-- @card:end -->
33
33
 
34
34
  <!-- @banned:start -->
@@ -43,13 +43,13 @@ description: How to set up lip-sync — Kling-only (dedicated lip-sync and avata
43
43
  - `reader_en_m-v1` — listed in fal's docs, returns "Voice id not found" in production
44
44
  <!-- @banned:end -->
45
45
 
46
- **This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; every entry is a real endpoint and every output is 5 seconds.
46
+ **This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; output follows the source clip for video or the voice track for a still, billed per 5s block.
47
47
 
48
48
  | Flow | Source | Model | Cost | Use case |
49
49
  |------|--------|-------|-----------|----------|
50
- | Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s | Replace dialogue on an existing talking head |
51
- | Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s | Animate a portrait into a talking avatar |
52
- | Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s | Higher facial fidelity for hero shots |
50
+ | Re-dub | video clip | kling-lip-sync-video | ~4 credits / 5s block | Replace dialogue on an existing talking head |
51
+ | Avatar standard | still image | ai-avatar/v2/standard | ~14 credits / 5s block with uploaded audio; typed text adds one flat voice block | Animate a portrait into a talking avatar |
52
+ | Avatar pro | still image | ai-avatar/v2/pro | ~29 credits / 5s block with uploaded audio; typed text adds one flat voice block | Higher facial fidelity for hero shots |
53
53
 
54
54
  Pick `sourceType` deliberately — it decides the pricing tier and the underlying endpoint.
55
55
 
@@ -60,7 +60,7 @@ Seedance can generate the performance rather than bolting a mouth onto finished
60
60
  That is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a "video 1" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.
61
61
 
62
62
  - Driving clips must be 2–15s; output duration is whatever you set (4–15s).
63
- - Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
63
+ - Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys); pass the clip duration and quote before confirming. On both Seedance 2.0 and 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
64
64
  - Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.
65
65
 
66
66
  Everything below is about the Kling tool.
@@ -81,7 +81,7 @@ Use **avatar** when:
81
81
 
82
82
  ### Video flow (`sourceType: 'video'`)
83
83
  - Format: mp4 or mov
84
- - Duration: 2–10s (lip-sync output is always 5s — long videos get trimmed)
84
+ - Duration: 2–10s (output follows the source clip)
85
85
  - Resolution: 720p or 1080p (480p will be rejected)
86
86
  - Max file size: 100MB
87
87
  - Face must be visible and roughly facing camera. Profile shots fail.
@@ -107,7 +107,7 @@ Two ways to drive the lips:
107
107
  - Pass `audioFilePath` — absolute path to an audio file on the user's machine
108
108
  - Format: mp3, wav, m4a, ogg, aac
109
109
  - Max 5MB
110
- - Duration: 2–60s (output is 5s — longer audio gets trimmed)
110
+ - Duration: 2–60s (avatar output follows the voice track)
111
111
  - Single clean voice. Music underneath, multiple speakers, or noisy mics produce garbage lips.
112
112
 
113
113
  Prefer upload for production-quality voice. TTS for fast iteration / placeholder dialogue.
@@ -181,15 +181,15 @@ Don't default to pro. The ~15-credit delta per take adds up across iteration.
181
181
 
182
182
  ## Cost discipline
183
183
 
184
- - Video re-dub at ~4 credits is the cheapest dialogue iteration in the entire Slates stack — use it for voice A/B testing
185
- - Avatar standard at ~14 credits is fine for medium use
186
- - Avatar pro at ~29 credits trips the confirm gate — explicit user OK required every time
187
- - All 5s. There is no shorter option.
184
+ - Video re-dub at ~4 credits per 5s block is the cheapest dialogue iteration in the entire Slates stack; use it for voice A/B testing
185
+ - Avatar standard at ~14 credits per 5s block is fine for medium use; typed text on a still adds one flat voice block
186
+ - Avatar pro at ~29 credits per 5s block trips the confirm gate; explicit user OK required every time
187
+ - Output follows the media; billing rounds up to whole 5s blocks.
188
188
 
189
189
  ## Workflow patterns
190
190
 
191
191
  **Voice A/B test (cheap):**
192
- 1. Generate one base talking-head video clip with Veo or Seedance (~40 credits)
192
+ 1. Generate one base talking-head video clip with Seedance (~40 credits)
193
193
  2. Run `slates_generate_lip_sync` with `sourceType: 'video'` against 3–5 different `ttsVoice` values
194
194
  3. Total cost: ~40 + (5 × ~4) ≈ 60 credits to compare voices
195
195
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-ltx-2-5
3
- description: How to prompt LTX-2.5 and LTX-2.5 Pro. Read before calling slates_generate_video with model ltx-2-5 or ltx-2-5-pro. LTX scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE — Lightricks ranks the prompt sound, camera, character detail, shot type and scene, then scene dressing, all in one flowing paragraph. It is also the catalogue's native MULTISHOT seat: one generation carries two to four connected shots holding character, light and voice across the cuts. Base ltx-2-5 is the distilled build — 720p/1080p/1440p/4K, clips of 6 to 20 seconds in EVEN steps, and the cheapest native 1080p second in Slates; ltx-2-5-pro is the full diffusion build and is NOT a superset, reaching only 1080p and 10 seconds for about a third more money. Three hazards live here: durations are even numbers only from six (there is no 5s or 7s clip), the model has NO reference endpoint at all so identity references are unavailable, and any sound not anchored to something in frame gets invented for you.
3
+ description: "Prompt LTX-2.5 or LTX-2.5 Pro (ltx-2-5, ltx-2-5-pro) with slates_generate_video. Covers sound-first direction, connected shots, frame inputs, variant differences and duration constraints."
4
4
  ---
5
5
 
6
6
  # LTX-2.5 — prompting
@@ -125,11 +125,10 @@ The model renders actions. It does not render adjectives.
125
125
 
126
126
  ---
127
127
 
128
- ## 4. Multishot — the thing this model is uniquely for
128
+ ## 4. Multishot
129
129
 
130
130
  **One LTX generation can carry several connected shots**, holding character, environment, lighting,
131
- voice and style across every cut. Nothing else in the catalogue does this natively; everywhere else
132
- you generate separate clips and stitch them, and identity drifts between them.
131
+ voice and style across every cut. It is one of the seats that carry several shots in one generation.
133
132
 
134
133
  **Working range is two to four shots.** Three is the comfortable stopping point.
135
134
 
@@ -182,7 +181,7 @@ Choose the length the beat needs.
182
181
 
183
182
  ### Aspect ratios: 16:9 and 9:16, and nothing else
184
183
 
185
- The narrowest set in the catalogue alongside Veo. Square, 4:5 and 21:9 are not available on this
184
+ The narrowest set in the catalogue. Square, 4:5 and 21:9 are not available on this
186
185
  model at any resolution.
187
186
 
188
187
  ### Frames, not references
@@ -210,7 +209,7 @@ Native synchronised audio is **included at every resolution on both seats**, wit
210
209
  no toggle that costs money — unlike Kling, where sound is a paid dimension. A 6-second 1080p LTX
211
210
  clip **with sound** is 39 credits.
212
211
 
213
- Combined with 1080p at $0.13/s — the cheapest native 1080p second in Slates — this makes LTX **the
212
+ Combined with 1080p at $0.13/s, the cheapest native 1080p second with sound included in Slates, this makes LTX **the
214
213
  coverage seat**: the one to reach for when the job is many takes rather than one hero shot, when a
215
214
  sequence needs its own sound, or when the credit budget is the binding constraint.
216
215
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-minimax-h3
3
- description: How to prompt MiniMax H3, H3 Max and H3 Max Turbo. Read before calling slates_generate_video with model minimax-h3, minimax-h3-max or minimax-h3-max-turbo. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, runs 480p/768p plus a 1080p refinement of its 768p render, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09). minimax-h3-max-turbo is a second fal post-train with Max's ladder at half Max's rate; it takes start and end frames but has NO reference endpoint. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; pooled media tokens on Max), and audio written into the wrong section is dropped or duplicated.
3
+ description: "Prompt MiniMax H3, H3 Max or H3 Max Turbo with slates_generate_video. Covers separately authored dialogue, scene sound and score, declared reference relationships, frame inputs and variant-specific constraints."
4
4
  ---
5
5
 
6
6
  # MiniMax H3 — prompting
@@ -15,18 +15,18 @@ description: How to prompt MiniMax H3, H3 Max and H3 Max Turbo. Read before call
15
15
  Keep it under 2,400 characters (the build fails above that) and keep the
16
16
  rationale, the receipts and the worked examples in the body below. -->
17
17
  <!-- /slates-only -->
18
- **Card — MiniMax H3.** The only seat where audio is AUTHORED rather than toggled: dialogue, scene sound and score are three separate sections of the prompt, generated in one pass, and putting a sound in the wrong section drops or doubles it.
18
+ **Card: MiniMax H3.** The only seat with three separately authored audio layers: dialogue, scene sound and score are three separate sections of the prompt, generated in one pass, and putting a sound in the wrong section drops or doubles it.
19
19
 
20
20
  **The five levers**
21
- 1. **Write the three audio layers separately** — `Scene sound:` for what is in the room, `Score:` for what only the audience hears, and the dialogue quoted inline. Section decides attribution.
21
+ 1. **Write the three audio layers separately**: `overall_soundscape:` for what is in the room, `non_diegetic_music:` for what only the audience hears, and the dialogue quoted inline. Section decides attribution.
22
22
  2. **Quote dialogue and name the language** — `says in English`, `speaks in Spanish`. Eleven languages are stably supported; the language is part of the instruction, not an afterthought.
23
23
  3. **Declare the reference RELATIONSHIP**, which no other seat has: `kept whole`, `partly kept`, `transferred`, or `a loose echo`. An undeclared reference is a guess.
24
24
  4. **Give a beat of stillness before a line** — `sits still for a beat, then looks up`. The sync needs something to lock against; a character already mid-motion when the line starts drifts.
25
25
  5. **Describe the beat structure** — `waits`, `then speaks`, `under the last three seconds`. H3 is a timeline, so write one.
26
26
 
27
27
  **Examples**
28
- - `A woman sits still at a kitchen table for a beat, then looks up. She says in English, "You said Tuesday." Scene sound: a fridge hum, a spoon set down on formica. Score: none.`
29
- - `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, "No es el alternador." Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`
28
+ - `A woman sits still at a kitchen table for a beat, then looks up. She says in English, "You said Tuesday." overall_soundscape: a fridge hum, a spoon set down on formica. non_diegetic_music: N/A.`
29
+ - `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, "No es el alternador." overall_soundscape: a socket wrench, a radio two bays over. non_diegetic_music: a low sustained cello under the last three seconds, audience only.`
30
30
 
31
31
  **Hard constraint:** the three seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p, takes the same 9+3+3 references, and costs MORE at the tier they share — a speed pick, never the cheap one; `minimax-h3-max-turbo` has Max's ladder at half its rate and takes frames only, NO references. Every tier above 768p is built from the native 768p render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.
32
32
  <!-- @card:end -->
@@ -61,7 +61,7 @@ what the endpoint accepts:
61
61
  | Price at 768p | **$0.060/s** | $0.080/s | $0.040/s |
62
62
  | Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) | **price** — half Max's rate at every tier |
63
63
 
64
- **Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it
64
+ **Max is the premium seat, not the budget one.** It is 33% dearer at 768p, equal at 480p, and it
65
65
  tops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth
66
66
  paying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.
67
67
 
@@ -91,15 +91,15 @@ reaching for 2K, which adds its own artifacting on top.
91
91
 
92
92
  ## The one thing that makes H3 different: audio is a THREE-LAYER instruction
93
93
 
94
- Every other video seat treats sound as on or off. H3 splits it, and the split is enforced by where
94
+ Kling and Seedance have their own sound syntax. H3 splits audio into three separately authored layers, enforced by where
95
95
  you write each thing. Get the section wrong and the sound is dropped, doubled, or attributed to the
96
96
  wrong source.
97
97
 
98
98
  | Layer | What belongs in it | Where it goes |
99
99
  |---|---|---|
100
100
  | **Synchronised events** | dialogue, singing, and any sound tied to a specific shot or action | the **body** of the prompt, on the beat it lands |
101
- | **Scene sound** | ambience and physical sounds that run across the whole clip — room tone, rain, traffic, a ventilation hum | the **soundscape** section |
102
- | **Score** | music the characters cannot hear; audience-only | the **music** section |
101
+ | **Scene sound** | ambience and physical sounds that run across the whole clip, room tone, rain, traffic, a ventilation hum | the **overall_soundscape** section |
102
+ | **Score** | music the characters cannot hear; audience-only | the **non_diegetic_music** section |
103
103
 
104
104
  **Three rules, all from MiniMax's own guide:**
105
105
 
@@ -125,10 +125,10 @@ the middle-aged baker with a calm, slightly raspy voice places a fresh loaf on t
125
125
  says: "First batch of the morning." [Shot 2] At 00:05.000, the camera cuts to a close-up of
126
126
  steam rising from the sliced bread while his final words carry over from the previous shot.
127
127
 
128
- Soundscape: wooden shutters scrape open over a quiet street, trays clink softly inside, a
128
+ overall_soundscape: wooden shutters scrape open over a quiet street, trays clink softly inside, a
129
129
  doorbell rings once, then light footsteps and the crisp sound of bread being sliced.
130
130
 
131
- Score: a soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes,
131
+ non_diegetic_music: a soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes,
132
132
  gentle fade at the end.
133
133
  ```
134
134
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-motion-transfer
3
- description: How to set up motion transfer — Kling Motion Control only (std and pro tiers, 5-second outputs). Read before calling slates_generate_motion_transfer. Reference image (character) + driving video (motion source) → new video of the character performing the motion. Asset selection rules, character_orientation, tiers, and prompt usage. Also covers the Seedance alternative, which is a normal video generation rather than a mode of this tool.
3
+ description: "Prepare Kling Motion Control with slates_generate_motion_transfer: a character image plus a driving clip. Covers orientation, source selection, tiers and the separate Seedance video-reference alternative."
4
4
  ---
5
5
 
6
6
  # Motion transfer — setup guide
@@ -15,7 +15,7 @@ description: How to set up motion transfer — Kling Motion Control only (std an
15
15
  Keep it under 2,400 characters (the build fails above that) and keep the
16
16
  rationale, the receipts and the worked examples in the body below. -->
17
17
  <!-- /slates-only -->
18
- **Card — Motion transfer (Kling Motion Control only).** A target IMAGE (your character) plus a source VIDEO (the motion) produces your character performing that motion. Always 5 seconds.
18
+ **Card: Motion transfer (Kling Motion Control only).** A target IMAGE (your character) plus a source VIDEO (the motion) produces your character performing that motion. Output follows the driving clip, up to 30s with video orientation or 10s with image orientation, billed per 5s block.
19
19
 
20
20
  **The five levers**
21
21
  1. **The target image must show body proportions clearly** and the character must occupy more than about 5% of the frame. A tiny figure in a wide shot has nothing to drive.
@@ -23,7 +23,7 @@ description: How to set up motion transfer — Kling Motion Control only (std an
23
23
  3. **Choose `characterOrientation` on purpose** — `video` takes the source clip's framing, `image` preserves the portrait's. It is the most-missed choice here.
24
24
 
25
25
  4. **The prompt is atmosphere only** — `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.` Motion verbs are ignored; the motion is already in the driving video.
26
- 5. **Pick the best 5 seconds of the source up front**, and write only atmosphere: `soft afternoon sunlight`, `vintage warm color grade`, `clean studio backdrop`. The output is 5s regardless, so a long driving clip just wastes the choice.
26
+ 5. **Pick the source section up front**, and write only atmosphere: `soft afternoon sunlight`, `vintage warm color grade`, `clean studio backdrop`. The output follows that section, up to the orientation's limit; longer clips cost more blocks.
27
27
 
28
28
  **Examples**
29
29
  - `Soft afternoon sunlight, dust motes in the air, vintage warm color grade.`
@@ -63,8 +63,8 @@ movement from video 1. Preserve the character's identity, appearance, and outfit
63
63
 
64
64
  That is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.
65
65
 
66
- - **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).
67
- - **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
66
+ - **Reference videos must total 2–15s on 2.0, or 2–30s on 2.5.** Longer clips: trim first. Kling MC (`characterOrientation: 'video'`) takes up to 30s.
67
+ - **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key; quote via the confirm gate before spending. On both Seedance 2.0 and 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
68
68
  - **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
69
69
  - `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.
70
70
 
@@ -180,13 +180,13 @@ Leave it empty if you don't have a specific atmospheric note.
180
180
  - Pro tier on first iteration — waste, switch to it once the motion + framing combo is locked
181
181
  - Cartoon driving videos — guaranteed failure
182
182
  - Cropped or partial target characters — identity will drift
183
- - Long driving videos when output is 5s — pick the best 5s of the source upfront
183
+ - Driving videos longer than the needed motion; pick the source section upfront, within the orientation's limit
184
184
 
185
185
  ## Cost discipline
186
186
 
187
- - 5 seconds, no shorter option
187
+ - Output follows the driving clip, up to 30s with video orientation or 10s with image orientation; billed per 5s block
188
188
  - Both tiers trip the confirm gate — every call needs explicit user OK
189
- - Iteration is expensive: 4 takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.
189
+ - Iteration is expensive: 4 five-second takes at pro ≈ 168 credits. Lock framing + driving video before tier-up to pro.
190
190
  - Always run a single std take first to validate the motion + framing combo before committing to pro
191
191
 
192
192
  ## Confirm gate: cost + codes, no inline preview
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-nano-banana-2
3
- description: How to write prompts that produce cinematic, photorealistic results from Nano Banana 2 (Google Gemini 3.1 Flash Image, accessed via fal-ai/nano-banana-2). Read this before calling slates_generate_image when the user wants film-quality, real-world, or cinematic output. Skip for stylized / illustrated / cartoon work — the rules differ.
3
+ description: "Prompt Nano Banana 2, Lite or Pro image generation and edits. Use on these models; covers photographic craft, subject and reference binding, typography, prompt structure and family differences."
4
4
  ---
5
5
 
6
6
  # Nano Banana 2 — cinematic & photorealistic prompting
@@ -220,16 +220,20 @@ Identity = a few flat-lit neutral angles; one reference per role, named inline;
220
220
 
221
221
  ✅ **Cinema:** "Extreme close on subject's mouth and nose, 135mm f/2.8, shallow depth of field. Breath pluming out, catching cold light from upper-left key. Lips slightly parted, peach fuzz visible. The breath holds. CineStill 800T halation around catchlights. Waiting."
222
222
 
223
- ## The 3-strike rule
223
+ <!-- @inject:iteration-diagnosis -->
224
+ ## Diagnose repeated failures
224
225
 
225
- If three iterations on the same prompt haven't produced what the user wants, stop. Hand back to the user with what you tried and what isn't working. The slot machine doesn't converge — the prompt structure is wrong, not the seed.
226
+ After three failed attempts at the same requirement, pause unchanged re-rolls and diagnose the source reference, prompt structure, model fit and tool result. Three is a review checkpoint, not a universal limit or proof that the seed cannot matter. Preserve the attempts and name what each test changed.
227
+
228
+ Continue autonomously when the brief is clear, a specific correction is supported and the next request is already authorized. Hand control back when taste or intent cannot be inferred, the next request needs fresh consent, or the available tool cannot meet the requirement. A failed roll never authorizes an additional charge. Follow the existing batch and per-request cost policy.
229
+ <!-- @end:iteration-diagnosis -->
226
230
 
227
231
  ## Family variants — Lite and Pro
228
232
 
229
233
  Everything in this skill applies to the whole Nano Banana family; two variants trade speed/ceiling around NB2 full:
230
234
 
231
235
  - **nano-banana-2-lite** — ~half the price, ~2.7× faster, **1K output only**, max 4 refs. The draft/iteration seat: explore compositions here, then re-run the winner on NB2 full at 2K/4K. Same Gemini filter.
232
- - **nano-banana-pro** — the hero-frame/typography ceiling (~2× NB2, 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — it takes a full subject library in one call.
236
+ - **nano-banana-pro**: the hero-frame/typography ceiling (2× NB2 at 1K, 1.33× at 2K, about 1.9× at 4K; 4K native). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs; it takes a full subject library in one call.
233
237
 
234
238
  <!-- slates-only -->
235
239
  Routing between them (and vs GPT Image 2.5 / FLUX / Seedream): `slates-model-selection`.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-omni-flash
3
- description: How to prompt Gemini Omni Flash (Google, via fal). Read before calling slates_generate_video with omni-flash or slates_edit_video with omni-flash-edit. Cheap 720p tier with native synced audio included — 3-10s, 16:9/9:16 only; t2v, single-start-frame i2v, or reference-to-video with up to 7 reference images. The edit variant is the EDIT-FIDELITY WINNER for footage-synced VFX (receipt 2026-07-09) — but ONLY with short prompts: one change + "Keep everything else the same." Long descriptive prompts destroy fidelity.
3
+ description: "Prompt Gemini Omni Flash video generation or Omni Flash edits (omni-flash, omni-flash-edit). Covers native sound, reference inputs and short change-only prompts that preserve edit fidelity."
4
4
  ---
5
5
 
6
6
  # Gemini Omni Flash — prompting
@@ -22,7 +22,7 @@ description: How to prompt Gemini Omni Flash (Google, via fal). Read before call
22
22
  2. **Editing: always end with `Keep everything else the same.`** — the one documented preservation lever.
23
23
 
24
24
  3. **Editing: describe the EFFECT, never a real object as a metaphor.** "Candle-like flame" rendered a literal candle in the subject's hand.
25
- 4. **Editing: no conditional timing cues.** "…when he calls it, as he walks…" hard-fails with `invalid_request`. Collapse to one continuous action; the model syncs to the footage's own motion.
25
+ 4. **Editing: no chained stage directions.** Several beats cued to moments ("…flies onto his shoulder when he calls it, and perches as he walks…") hard-failed with `invalid_request`. One effect tied to an action already in the footage ("…when he snaps his fingers…") passed. Collapse to one continuous action; the model syncs to the footage's own motion.
26
26
  5. **Generation: the opposite — describe fully.** Subject, action, setting, `camera tracking alongside`, `overcast flat light`, tone. Audio is prompt-driven with no parameters: dialogue in quotes, sound in plain language — `rain patters on the tin roof`, `spray from tyres`, `a horn somewhere behind`.
27
27
 
28
28
  **Examples**
@@ -42,7 +42,7 @@ description: How to prompt Gemini Omni Flash (Google, via fal). Read before call
42
42
  **Never use in an EDIT prompt** (each one has a receipt above):
43
43
  - a long preservation preamble — it produces WORSE drift than `Keep everything else the same.`
44
44
  - a real object as a metaphor: `candle-like`, `flame-like`, `laser-like`
45
- - a conditional timing cue: `when he`, `as she`, `once they` — these hard-fail, they do not merely drift
45
+ - several staged beats cued to moments (`when he calls it, and perches as he walks`): these hard-fail, they do not merely drift. One effect tied to an action already in the clip (`when he snaps his fingers`) passed
46
46
  - harm-to-person framing: `ignite`, `catch fire`, `on fire` applied to a person trips the safety filter
47
47
  <!-- @banned:end -->
48
48
 
@@ -51,15 +51,15 @@ Google's fast video generation + editing model ("Nano Banana Pro for video" in c
51
51
  ## Where it routes
52
52
 
53
53
  - **Video editing (`omni-flash-edit`) — its headline strength and the edit-lane default** for footage-synced VFX: verified 2026-07-09 head-to-head vs Kling O3 Edit on real phone footage (fire-on-fingertips on a talking take) — Omni Flash held lip movement perfectly, audio near-identical, and executed both action beats; Kling kept audio verbatim but drifted lips and missed the second beat. Full routing: slates-model-selection.
54
- - **Cheap drafts and iteration volume** — lowest-cost audio-native video seat (~6.4 cr/s at 720p).
55
- - **NOT hero GENERATION shots** — Kling 3.0 stays the general gen default, Seedance 2.0 the premium tier; Omni Flash's *generation* quality seat is still unproven.
54
+ - **Drafts and iteration volume with sound**: an audio-native video seat (~6.4 cr/s at 720p).
55
+ - **Hero-generation quality is still unproven.** Use the current model-routing guide for the final generation seat; the edit-fidelity receipt does not establish generation quality.
56
56
 
57
- ## Editing (`slates_edit_video`, model `omni-flash-edit`) — THE RULES (receipts, not theory)
57
+ ## Editing (<!-- slates-only -->`slates_edit_video`, <!-- /slates-only -->model `omni-flash-edit`) — THE RULES (receipts, not theory)
58
58
 
59
59
  1. **SHORT PROMPT. One change. Nothing else.** Google's own doc: *"Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes."* Live receipt 2026-07-09: a long "keep every frame/word/movement identical…" preamble produced WORSE drift (re-synthesized performance, wrong timing); the winning prompt was two sentences: *"Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same."*
60
60
  2. **Always end with "Keep everything else the same."** — the one documented preservation lever.
61
61
  3. **Never name a real-world object as a metaphor.** "Candle-like flame" rendered a literal candle in his hand. Describe the effect itself ("small magical flames on his fingertips").
62
- 3b. **No conditional timing cues — they HARD-FAIL, not drift.** Receipt 2026-07-09: "a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…" → deterministic `invalid_request` (2×, "could not generate with the given inputs"); collapsing to one continuous action — "A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke." — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video.
62
+ 3b. **No chained stage directions: they HARD-FAIL, not drift.** Receipt 2026-07-09: "a dragon appears behind him, flies onto his shoulder WHEN HE CALLS IT, and perches AS HE WALKS…" → deterministic `invalid_request` (2×, "could not generate with the given inputs"); collapsing to one continuous action — "A small photorealistic dragon flies in and perches on his shoulder, puffing a small breath of flame and smoke." — succeeded first try. The model syncs the change to the footage's own motion; it cannot take beat-by-beat stage directions cued to moments in the video. The winning flames prompt above ("when he snaps his fingers") shows one effect tied to an action the footage already contains is fine; which part of the dragon prompt triggered the refusal is untested beyond that.
63
63
  4. **Safety filter (Google's, strict about harm-to-person):** "fingertips ignite / catch fire" → `content_policy_violation`. Frame effects as magical/harmless VFX: "small magical flames appear on his fingertips" passed. See slates-content-policy §Gemini for the substitution patterns.
64
64
  5. **Expect a possible tail artifact** — jitter or a doubled final speech beat in the last ~0.5s. Plan to trim the tail on the timeline; don't burn a re-roll on it.
65
65
  6. **Prompt + source clip ONLY.** No element/style reference images — identity swaps that need refs go to `kling-v3.0-omni-edit`.
@@ -67,9 +67,9 @@ Google's fast video generation + editing model ("Nano Banana Pro for video" in c
67
67
  8. **Ship via segment-splice** (the workflow, not the model): edit only the seconds where the change happens, splice back over the original on the timeline with the original audio underneath. Most of the deliverable stays untouched original footage — this is how the pro demos are actually assembled (gesture-only edited beats + voiceover in post).
68
68
  9. Chain edits one change at a time — each edit saves as a new asset linked to its parent.
69
69
 
70
- ## Generation (`slates_generate_video`, model `omni-flash`)
70
+ ## Generation (<!-- slates-only -->`slates_generate_video`, <!-- /slates-only -->model `omni-flash`)
71
71
 
72
- - **Inputs:** prompt only (t2v), prompt + ONE start frame (`firstFrameAssetId`, i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
72
+ - **Inputs:** prompt only (t2v), prompt + ONE start frame (<!-- slates-only -->`firstFrameAssetId`, <!-- /slates-only -->i2v), or prompt + up to **7 reference images** (ingredient/character/environment/style asset params — they merge into one reference list). No last frame, no video/audio references — the op rejects them.
73
73
  - Descriptive prompts are fine for GENERATION (the short-prompt law above is edit-specific). Structure like a shot brief: subject + action + setting + camera + lighting + tone.
74
74
  - **Name references inline** the standard Slates way ("Marcus (image 1) walks…"). The endpoint also accepts explicit `<IMAGE_REF_0>`-style binding tags (zero-indexed) — useful when a specific image must bind to a specific role.
75
75
  - **Audio is prompt-driven** — no audio parameters. Dialogue in quotes; direct sound in plain language ("rain patters on the tin roof"). Negative direction as plain instructions ("Do not show text").
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-seed-audio
3
- description: How to prompt Seed Audio 1.0 (ByteDance, via fal). Read before calling slates_generate_audio with model seed-audio. The one-pass audio SCENE model — dialogue, SFX and ambience together from ONE plain sentence. CRITICAL - it has NO duration parameter, so length must be named IN THE PROMPT TEXT and Slates bills the duration you request. Covers the one-sentence doctrine, the crowd-size rule, why Kling "SFX:" syntax hurts here, and the audio-refs-XOR-image input rule.
3
+ description: "Prompt Seed Audio 1.0 (seed-audio) with slates_generate_audio. Covers scene sentences, dialogue, crowd size, spatial sound, duration in prompt text and mutually exclusive audio or image references."
4
4
  ---
5
5
 
6
6
  # Seed Audio 1.0 — prompting
@@ -43,7 +43,29 @@ description: How to prompt Seed Audio 1.0 (ByteDance, via fal). Read before call
43
43
  - shot language: `wide shot`, `slow push in`, `warm tungsten` — camera and lighting words are video-prompt words the model has to ignore
44
44
  <!-- @banned:end -->
45
45
 
46
- ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence. 1–120 seconds. It is the default audio model in Slates and the workhorse for continuity beds.
46
+ ByteDance's one-pass audio scene model, carried on fal (`bytedance/seed-audio-1.0`). It generates dialogue, sound effects and ambience **together**, from a single plain sentence. Use the generated duration bounds below and the current catalogue for routing; it is useful for continuity beds.
47
+
48
+ <!-- @inject:thresholds -->
49
+ <!-- GENERATED from @slatesvideo/shared — do not edit between the markers.
50
+ Source: CONFIRM_CREDITS, DEVIATION_FACTOR and the audio bounds in
51
+ packages/shared/src/operations/index.ts. Every number here is REFUSED by an
52
+ op when a prompt gets it wrong, which is why none of them is typed by hand
53
+ any more: this block replaced four claims that contradicted the code. -->
54
+
55
+ **The thresholds, from the code that enforces them:**
56
+
57
+ - **Confirm gate:** above **17 credits** an op returns `requires_confirm` and will not
58
+ proceed until you re-call with `confirm: true`. This is a code gate, not permission to spend: every generation still needs the user-approved plan or quote.
59
+ - **Deviation pause:** the desktop Studio Agent stops and re-asks when projected generation spend
60
+ exceeds the approved plan by more than **20%**. You do not trigger this; the app does.
61
+ - **Seed Audio duration:** **3–120 seconds.** There is no duration
62
+ parameter on the model — the number you pass is written into the prompt AND is what the user is
63
+ billed. Outside that range the op refuses rather than clamping.
64
+ - **Sound Effects duration:** **1–22 seconds**, billed per second, never left for the
65
+ model to pick.
66
+
67
+ Never quote a credit figure from memory: `slates_estimate_generation_cost` returns the real one.
68
+ <!-- @end:thresholds -->
47
69
 
48
70
  ## Where it routes
49
71
 
@@ -134,8 +156,6 @@ Set `multilingual: true` for non-English or mixed-language lines.
134
156
  | `volume` | 0.5–2.0 | Rarely — normalize on the timeline instead. |
135
157
  | `pitch` | −12…+12 semitones | Ageing or shifting a voice. Small moves only; ±3 is already a lot. |
136
158
  | `multilingual` | bool | Non-English or code-switched lines. |
137
- | `sampleRate` | 8k–48k | Leave at 24000 unless you are matching an existing stem. |
138
- | `outputFormat` | mp3 / wav / pcm / ogg_opus | wav when this is going into a mix; mp3 otherwise. |
139
159
 
140
160
  ## Iterating
141
161