@slatesvideo/shared 0.6.3 → 0.6.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. package/dist/api-url.d.ts +11 -0
  2. package/dist/api-url.js +11 -0
  3. package/dist/auth.d.ts +13 -1
  4. package/dist/auth.js +9 -5
  5. package/dist/clients/cloud.d.ts +3 -0
  6. package/dist/clients/cloud.js +34 -3
  7. package/dist/clients/desktop.js +3 -0
  8. package/dist/index.d.ts +7 -2
  9. package/dist/index.js +20 -2
  10. package/dist/manual/content.d.ts +2 -0
  11. package/dist/manual/content.js +3 -0
  12. package/dist/manual/index.d.ts +5 -0
  13. package/dist/manual/index.js +20 -0
  14. package/dist/operations/index.d.ts +253 -30
  15. package/dist/operations/index.js +1387 -141
  16. package/dist/operations/surface.d.ts +69 -0
  17. package/dist/operations/surface.js +227 -0
  18. package/dist/prompts/agent-doctrine.js +8 -0
  19. package/dist/prompts/asset-label.d.ts +23 -0
  20. package/dist/prompts/asset-label.js +70 -0
  21. package/dist/prompts/banned-tokens.d.ts +15 -3
  22. package/dist/prompts/banned-tokens.js +76 -9
  23. package/dist/prompts/character-sheet.d.ts +0 -2
  24. package/dist/prompts/character-sheet.js +0 -2
  25. package/dist/prompts/craft-cards.d.ts +20 -0
  26. package/dist/prompts/craft-cards.js +82 -0
  27. package/dist/prompts/environment-sheet.js +16 -0
  28. package/dist/prompts/index.d.ts +1 -0
  29. package/dist/prompts/index.js +4 -0
  30. package/dist/prompts/model-capabilities.d.ts +52 -0
  31. package/dist/prompts/model-capabilities.js +42 -0
  32. package/dist/prompts/model-facts.d.ts +0 -4
  33. package/dist/prompts/model-facts.js +8 -4
  34. package/dist/prompts/partials.generated.js +2 -1
  35. package/dist/prompts/prompting-tips.d.ts +1 -1
  36. package/dist/prompts/prompting-tips.js +58 -0
  37. package/dist/prompts/reference-composer.d.ts +36 -7
  38. package/dist/prompts/reference-composer.js +75 -20
  39. package/dist/prompts/reference-rules.d.ts +15 -26
  40. package/dist/prompts/reference-rules.js +15 -93
  41. package/dist/prompts/shot-grammar.d.ts +154 -0
  42. package/dist/prompts/shot-grammar.js +184 -0
  43. package/dist/prompts/shot-spec.d.ts +278 -0
  44. package/dist/prompts/shot-spec.js +319 -0
  45. package/dist/skills/content.js +25 -23
  46. package/exports/slates-prompt-builder/generated/reference-content-policy.md +6 -0
  47. package/exports/slates-prompt-builder/generated/reference-kling.md +22 -0
  48. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +17 -0
  49. package/exports/slates-prompt-builder/generated/reference-seedance.md +17 -0
  50. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +15 -15
  51. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  52. package/package.json +83 -73
  53. package/skills/_partials/decision-log.md +5 -4
  54. package/skills/_partials/thresholds.md +19 -0
  55. package/skills/slates-content-policy.md +15 -1
  56. package/skills/slates-cost-discipline.md +26 -4
  57. package/skills/slates-model-selection.md +4 -3
  58. package/skills/slates-one-prompt-film.md +20 -12
  59. package/skills/slates-project-organization.md +1 -1
  60. package/skills/slates-prompting-elevenlabs.md +61 -2
  61. package/skills/slates-prompting-flux-2-max.md +39 -0
  62. package/skills/slates-prompting-gpt-image-2.md +109 -70
  63. package/skills/slates-prompting-inworld-tts.md +174 -0
  64. package/skills/slates-prompting-kling-v3.md +39 -0
  65. package/skills/slates-prompting-lip-sync.md +38 -0
  66. package/skills/slates-prompting-ltx-2-5.md +38 -0
  67. package/skills/slates-prompting-minimax-h3.md +39 -0
  68. package/skills/slates-prompting-motion-transfer.md +38 -0
  69. package/skills/slates-prompting-nano-banana-2.md +26 -0
  70. package/skills/slates-prompting-omni-flash.md +41 -0
  71. package/skills/slates-prompting-seed-audio.md +39 -1
  72. package/skills/slates-prompting-seedance-2-5.md +38 -0
  73. package/skills/slates-prompting-seedance.md +26 -0
  74. package/skills/slates-prompting-seedream-5-lite.md +38 -0
  75. package/skills/slates-prompting-veo-3.md +39 -0
  76. package/skills/slates-shot-variety.md +53 -0
  77. package/skills/slates-storyboard-from-script.md +31 -15
  78. package/skills/slates-style-prompting.md +1 -1
  79. package/skills/slates-vision-feedback-loop.md +1 -1
@@ -1,10 +1,47 @@
1
1
  ---
2
2
  name: slates-prompting-elevenlabs
3
- description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before calling slates_generate_audio with model eleven-sfx — ONE short effect with an EXACT duration (0.5-22s), or a seamless loop, billed per second. Covers describing an effect by its physical cause, the one-sound-per-generation rule, picking a duration, loops, prompt_influence, and when to use Seed Audio instead.
3
+ description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before calling slates_generate_audio with model eleven-sfx — ONE short effect with an EXACT duration, or a seamless loop, billed per second. Covers describing an effect by its physical cause, the one-sound-per-generation rule, picking a duration, loops, prompt_influence, and when to use Seed Audio instead.
4
4
  ---
5
5
 
6
6
  # ElevenLabs Sound Effects v2 — prompting
7
7
 
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved
14
+ compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.
15
+ Keep it under 2,400 characters (the build fails above that) and keep the
16
+ rationale, the receipts and the worked examples in the body below. -->
17
+ <!-- /slates-only -->
18
+ **Card — ElevenLabs Sound Effects v2.** ONE short sound with an exact length, or a seamless loop. The only Slates audio surface with a real duration control and a real loop mode.
19
+
20
+ **The five levers**
21
+ 1. **Describe the physical CAUSE, not the label** — `heavy oak door slams shut`, `boot scuffs on grit`, `a latch drops home`.
22
+ 2. **Name the material and the space.** The material decides the timbre and the space decides the tail: `on wet concrete`, `in a tiled stairwell`, `across an empty warehouse`.
23
+ 3. **One sound per generation.** A room with dialogue AND clatter AND ambience is one Seed Audio pass, not three effects.
24
+ 4. **Pick the duration from the cut**, not from a feeling: roughly 0.5-1s for an `impact`, 2-4s for a `whoosh`, 8-22s for a `loopable bed`.
25
+ 5. **Ask for a loop explicitly** — `seamless loop` — when the sound has to lie under a whole scene, and keep it featureless enough to survive the seam.
26
+
27
+ **Examples**
28
+ - `A heavy oak door slams shut in a stone hallway, brief reverberant tail.` (1.5s)
29
+ - `Steady rain on a tin awning, no thunder, no wind gusts, seamless loop.` (18s)
30
+
31
+ **Hard constraint:** it is billed per second and the duration is never left for the model to pick — that would make the charge non-deterministic. It is NOT a speech surface: a line in a specific voice is `inworld-tts-2`, and dialogue inside a scene is Seed Audio, which casts and performs the line in the room.
32
+ <!-- @card:end -->
33
+
34
+ <!-- @banned:start -->
35
+ <!-- slates-only -->
36
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
37
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
38
+ estimate, and every submitted prompt is matched against it. Keep entries
39
+ backticked and prose outside the backticks. -->
40
+ <!-- /slates-only -->
41
+ **Never use** — a label is not a sound; describe the physical cause:
42
+ - `door sound`, `whoosh`, `footsteps`, `impact`, `ambience` standing alone
43
+ <!-- @banned:end -->
44
+
8
45
  One short sound with an exact length, carried on fal (`fal-ai/elevenlabs/sound-effects/v2`). This is the only Slates audio surface with a real duration control and a real loop mode.
9
46
 
10
47
  ## Where it routes
@@ -38,7 +75,29 @@ This surface makes a single event. A door, then footsteps, then a siren is three
38
75
 
39
76
  ### 3. Duration is always explicit, and it is the price
40
77
 
41
- `durationSeconds` is 0.5–22 and Slates **always sends it**. (Left null the model picks, which makes the charge non-deterministic — so it is never left null.)
78
+ Slates **always sends** `durationSeconds`. (Left null the model picks, which makes the charge non-deterministic — so it is never left null.) The window it must fall in:
79
+
80
+ <!-- @inject:thresholds -->
81
+ <!-- GENERATED from @slatesvideo/shared — do not edit between the markers.
82
+ Source: CONFIRM_CREDITS, DEVIATION_FACTOR and the audio bounds in
83
+ packages/shared/src/operations/index.ts. Every number here is REFUSED by an
84
+ op when a prompt gets it wrong, which is why none of them is typed by hand
85
+ any more: this block replaced four claims that contradicted the code. -->
86
+
87
+ **The thresholds, from the code that enforces them:**
88
+
89
+ - **Confirm gate:** above **17 credits** an op returns `requires_confirm` and will not
90
+ proceed until you re-call with `confirm: true`. Below it, announce the cost once and go.
91
+ - **Deviation pause:** the desktop Studio Agent stops and re-asks when projected generation spend
92
+ exceeds the approved plan by more than **20%**. You do not trigger this; the app does.
93
+ - **Seed Audio duration:** **3–120 seconds.** There is no duration
94
+ parameter on the model — the number you pass is written into the prompt AND is what the user is
95
+ billed. Outside that range the op refuses rather than clamping.
96
+ - **Sound Effects duration:** **1–22 seconds**, billed per second, never left for the
97
+ model to pick.
98
+
99
+ Never quote a credit figure from memory: `slates_estimate_generation_cost` returns the real one.
100
+ <!-- @end:thresholds -->
42
101
 
43
102
  | Kind of sound | Ask for |
44
103
  |---|---|
@@ -5,6 +5,45 @@ description: How to prompt FLUX.2 Max (Black Forest Labs image model). Read befo
5
5
 
6
6
  # FLUX.2 Max — prompting
7
7
 
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved
14
+ compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.
15
+ Keep it under 2,400 characters (the build fails above that) and keep the
16
+ rationale, the receipts and the worked examples in the body below. -->
17
+ <!-- /slates-only -->
18
+ **Card — FLUX.2 Max.** Word order is weight: it attends hardest to the start. Structure: `Subject + Action + Style + Context`, then secondary detail. Length 10-30 words for a concept test, 30-80 for most work, 80+ only for a genuinely complex scene.
19
+
20
+ **The five levers**
21
+ 1. **Front-load the subject and the one action.** Anything after the first clause is a modifier, and it is read as one.
22
+ 2. **Name real gear** — `Shot on Hasselblad X2D, 80mm, f/2.8, natural light`, `Kodak Portra 400, natural grain`. This is the single biggest realism lever.
23
+ 3. **Era cues as a package** — `early digital camera, slight noise, flash photography, candid` reads 2000s; `film grain, warm cast, soft focus` reads 80s.
24
+ 4. **Bind every hex colour to an object.** `a #1B4D3E enamel mug` lands; an unbound colour does not.
25
+ 5. **For portraits add texture words** — `natural skin texture, realistic pores, subtle imperfections, soft diffused lighting`.
26
+
27
+ **Examples**
28
+ - `A chef plating in a steel kitchen pass. Shot on Hasselblad X2D, 80mm, f/2.8. Overhead fluorescents plus warm spill from the line. Natural skin texture, subtle imperfections. Muted steel and #7A3B2E copper.`
29
+ - `An empty municipal pool at dusk, 35mm, deep focus, early digital camera with slight noise and flash falloff. Cracked #4A7C8C tiles. Candid, unstaged.`
30
+
31
+ **Hard constraint:** no negative prompting. Every "no X" must be rewritten as the positive state — `no blur` becomes `sharp focus throughout`, `no people` becomes `empty scene`, `no harsh shadows` becomes `soft, diffused lighting`.
32
+ <!-- @card:end -->
33
+
34
+ <!-- @banned:start -->
35
+ <!-- slates-only -->
36
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
37
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
38
+ estimate, and every submitted prompt is matched against it. Keep entries
39
+ backticked and prose outside the backticks. -->
40
+ <!-- /slates-only -->
41
+ **Never use** — there is no negative prompting, so each of these has a positive form:
42
+ - `no blur` (say `sharp focus throughout`), `no people` (say `empty scene`), `no harsh shadows` (say `soft, diffused lighting`)
43
+ - an unbound hex colour — bind it to an object or it lands inconsistently
44
+ - `masterpiece`, `best quality`, `trending on artstation`, `8k`
45
+ <!-- @banned:end -->
46
+
8
47
  Black Forest Labs' top image model, routed via fal.ai. In Slates: `slates_generate_image` with `model: flux-2-max` (REQUIRES projectId — no headless path), priced per resolution (1k/2k/4k — call `slates_estimate_generation_cost` for current numbers, never quote from memory). Strengths vs Nano Banana 2: photoreal texture, less censored, precise hex-color control, strong typography. Reference images route through FLUX's edit endpoint and carry a lower per-model cap than NB2's 14.
9
48
 
10
49
  ## Core structure — front-load what matters
@@ -1,70 +1,109 @@
1
- ---
2
- name: slates-prompting-gpt-image-2
3
- description: Prompting GPT Image 2 — the readable-text / character-sheet / shot-grid engine AND the current photoreal front-runner. Read before calling slates_generate_image with model gpt-image-2. Covers the quality tiers (medium default, high for max text precision), resolution classes (1k/2k=1080p/3k=1440p/4k), text-accuracy prompting, panel/grid layout direction, and when to route to the Banana line instead.
4
- ---
5
-
6
- # GPT Image 2 — sheets, grids, and text that actually reads
7
-
8
- GPT Image 2's edge is **character-level text accuracy** (~99% on English), ordered panels, and exact element placement — the jobs where every other model garbles a word or shuffles a layout.
9
-
10
- 🚨 **It is ALSO the photoreal front-runner, and this file said the opposite until 2026-08-24.** **Receipts:** Eric's direct call, plus a head-to-head on the Higgsfield rail where GPT Image 2 at `quality: high`, 2K beat both Nano Banana rails on skin realism for photoreal people — that result is why the whole AI-influencer ad lane generates its plates here. **Route photoreal to this model, not away from it.**
11
-
12
- **What the Banana line still owns:** edit-heavy work and the 14-reference ceiling.
13
-
14
- **What would kill this:** a head-to-head at the intended crop going the other way. Per `slates-model-selection` § The meta-rule, re-run the evidence test when the roster changes — never carry a ranking forward on reputation. That rule is exactly what this correction failed.
15
-
16
- ## Quality tiers — always set explicitly
17
-
18
- - **medium** (default) — sharp text, fast, the value seat: half NB2's price at the 1080p class. Blind benchmarks put it within a hair of high at a quarter of the cost. Start here.
19
- - **high** — ~4× the price; max text precision + reasoning. A deliberate premium pick when tiny type, dense diagrams, or many labeled elements ARE the job.
20
-
21
- Never rely on the provider default (it's high — the priciest tier). The Slates ops send medium unless you say otherwise.
22
-
23
- ## Resolution classes
24
-
25
- `1k` = 1024²-class · `2k` = 1920×1080-class · `3k` = 2560×1440-class · `4k` = 3840×2160-class. Pick 2k for most sheets/panels; 4k for print-density grids. 4K exists at BOTH tiers and is API-only — even paid ChatGPT can't render it.
26
-
27
- ## Prompting for text accuracy
28
-
29
- - **Quote every string that must render verbatim**: `the sign reads "OPEN 24 HOURS"` — quoted strings render most reliably.
30
- - Specify font *feel*, not font names: "clean geometric sans, high contrast", "hand-painted brush lettering".
31
- - For dense text (posters, UI mocks), list the copy as ordered lines: `Line 1: "..." Line 2: "..."` — GPT Image 2 respects ordering.
32
- - Keep total on-image text under ~30 words for perfect accuracy; beyond that, accuracy degrades gracefully but degrades.
33
-
34
- ## Panels, sheets, and grids
35
-
36
- - State the grid explicitly and number the cells: "a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom".
37
- - Give each cell ONE content clause: "Panel 3: the character mid-jump, side view".
38
- - Character identity sheets: GPT Image 2 holds both the structured panel layout AND photoreal skin, which is why the influencer-ad lane builds its sheets here at `quality: high`, 2K. Reach for NB2/NB Pro when the sheet needs many reference images folded in (14-ref ceiling) or when it is an edit of an existing sheet.
39
-
40
- ## References & editing
41
-
42
- Reference images route through the edit endpoint (up to ~10). The composed "image N" naming applies as everywhere else. Mask-based inpainting exists at the API level but isn't surfaced — describe the change instead.
43
-
44
- ## 🚨 WHAT GETS YOU BLOCKED — read before writing a prompt with a person in it
45
-
46
- **Receipt: 24 consecutive attempts on one character, 2026-08-24, same project and same rail.** Eleven were refused with `content_policy_violation` on the fal edit endpoint. The refusals were never about the scene — one of the blocked prompts was a woman standing at a kitchen counter with her hand on it. **Two phrasings were hard blocks, 5 for 5 each, and neither ever passed:**
47
-
48
- **1. Never describe the reference as a photograph of a real person.**
49
-
50
- > ❌ `Reference image 1 is a photograph of a woman. Use that exact woman.`
51
- > ✅ `Reference image 1 is a character identity sheet showing one woman across several panels — the face in the large portrait panel is the authority for her identity. Use that exact woman.`
52
-
53
- The first reads to the filter as *recreate this real person's likeness*, which is a hard refusal regardless of what the rest of the prompt says. The second signals a fictional character and passes. **This is a wording change only — the reference image can be the same file either way.** One plate flipped from refused to accepted on this single sentence with nothing else altered.
54
-
55
- **2. Never attach a reference sheet containing a headless body panel.** A sheet whose full-body panels are cropped above the neck is refused every time, even with the correct opener. Regenerate the sheet with the head visible in every panel. Related, and already in this file's sheet guidance: phrase a cropped panel as *framing* (`cropped at the collarbone`), never as *absence* (`the head not shown`).
56
-
57
- **On top of those, ordinary content triggers still apply** and they stack independently — a correct opener does not rescue them:
58
-
59
- | Refused | Why, and the fix |
60
- |---|---|
61
- | A woman sitting on a bed in a bedroom | Domestic + bed reads as intimate. Move her to a chair, a rug, another room. |
62
- | A knife, even lying flat on a chopping board next to a lemon | The object is the trigger, not the framing. Swap it — a cast-iron pan cleared instantly. |
63
-
64
- **🚨 Refusals are PROBABILISTIC. Retry once before rewriting a word.** In the same session an identical prompt, identical reference, identical params was refused and then accepted on a straight re-fire. A rejected job returns no file and costs nothing, so a retry is free and a rewrite is not — rewriting first is how you end up changing four variables and learning nothing. **Only redesign after two or three refusals.**
65
-
66
- **And change ONE thing at a time.** The eleven refusals above took far longer to diagnose than they should have because a reference swap and an opener rewrite shipped in the same call. Isolate on the prompt you actually want, so a pass leaves you with a usable asset instead of a data point.
67
-
68
- ## Filter regime
69
-
70
- OpenAI moderate — a third regime distinct from Gemini (NB family) and ByteDance (Seedream). Real-face references pass more readily than Gemini; violence/brand rules are similar. `slates-content-policy` applies unchanged.
1
+ ---
2
+ name: slates-prompting-gpt-image-2
3
+ description: Prompting GPT Image 2 — the readable-text / character-sheet / shot-grid engine AND the current photoreal front-runner. Read before calling slates_generate_image with model gpt-image-2. Covers the quality tiers (medium default, high for max text precision), resolution classes (1k/2k=1080p/3k=1440p/4k), text-accuracy prompting, panel/grid layout direction, and when to route to the Banana line instead.
4
+ ---
5
+
6
+ # GPT Image 2 — sheets, grids, and text that actually reads
7
+
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved
14
+ compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.
15
+ Keep it under 2,400 characters (the build fails above that) and keep the
16
+ rationale, the receipts and the worked examples in the body below. -->
17
+ <!-- /slates-only -->
18
+ **Card — GPT Image 2.** The readable-text, ordered-panel and exact-placement engine, and (measured 2026-08-24 at `quality: high`) the photoreal front-runner for people. Structure: subject and action, then the exact copy in quotes, then layout, then light.
19
+
20
+ **The five levers**
21
+ 1. **Quote every string that must render verbatim** — `the sign reads "OPEN 24 HOURS"`. Quoted strings render most reliably.
22
+ 2. **Font FEEL, never a font name** — `clean geometric sans, high contrast`, `hand-painted brush lettering`.
23
+ 3. **Order dense copy explicitly** — `Line 1: "..." Line 2: "..."`. It respects the ordering.
24
+ 4. **Name the layout as a grid** for sheets and panels — `a 3x2 grid of panels, reading left to right, equal gutters`.
25
+ 5. **Set `quality` deliberately.** `medium` is sharp text at a quarter of the price and is within a hair of `high` in blind tests; reach for `high` only when tiny type, dense diagrams or many labelled elements ARE the job.
26
+
27
+ **Examples**
28
+ - `A 2x3 character turnaround sheet on a neutral grey field, equal gutters, reading left to right: front, three-quarter, profile, back, three-quarter back, top. One woman, mid-30s, cropped dark hair, olive field jacket. Flat even studio light, no cast shadows. Small caption under each panel naming the angle.`
29
+ - `Photoreal portrait, natural window light from camera-left, visible skin texture and pores, 85mm compression. A man in his 50s in a charcoal knit, half-smile, looking just past lens.`
30
+
31
+ **Hard constraint:** keep total on-image text under about 30 words for perfect accuracy — beyond that it degrades, gracefully but really. It has its own content filter, distinct from Gemini's.
32
+ <!-- @card:end -->
33
+
34
+ <!-- @banned:start -->
35
+ <!-- slates-only -->
36
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
37
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
38
+ estimate, and every submitted prompt is matched against it. Keep entries
39
+ backticked and prose outside the backticks. -->
40
+ <!-- /slates-only -->
41
+ **Never use:**
42
+ - a font NAME — describe the feel (`clean geometric sans, high contrast`) instead
43
+ - a reference role essay (`Reference image 1 is a photograph of a woman. Use that exact woman.`) — name the subject inline instead
44
+ - `8k`, `masterpiece`, `best quality`, `highly detailed` — quality incantations do nothing here either
45
+ <!-- @banned:end -->
46
+
47
+ GPT Image 2's edge is **character-level text accuracy** (~99% on English), ordered panels, and exact element placement — the jobs where every other model garbles a word or shuffles a layout.
48
+
49
+ 🚨 **It is ALSO the photoreal front-runner, and this file said the opposite until 2026-08-24.** **Receipts:** Eric's direct call, plus a head-to-head on the Higgsfield rail where GPT Image 2 at `quality: high`, 2K beat both Nano Banana rails on skin realism for photoreal people — that result is why the whole AI-influencer ad lane generates its plates here. **Route photoreal to this model, not away from it.**
50
+
51
+ **What the Banana line still owns:** edit-heavy work and the 14-reference ceiling.
52
+
53
+ **What would kill this:** a head-to-head at the intended crop going the other way. Per `slates-model-selection` § The meta-rule, re-run the evidence test when the roster changes — never carry a ranking forward on reputation. That rule is exactly what this correction failed.
54
+
55
+ ## Quality tiers — always set explicitly
56
+
57
+ - **medium** (default) — sharp text, fast, the value seat: half NB2's price at the 1080p class. Blind benchmarks put it within a hair of high at a quarter of the cost. Start here.
58
+ - **high** — ~4× the price; max text precision + reasoning. A deliberate premium pick when tiny type, dense diagrams, or many labeled elements ARE the job.
59
+
60
+ Never rely on the provider default (it's high — the priciest tier). The Slates ops send medium unless you say otherwise.
61
+
62
+ ## Resolution classes
63
+
64
+ `1k` = 1024²-class · `2k` = 1920×1080-class · `3k` = 2560×1440-class · `4k` = 3840×2160-class. Pick 2k for most sheets/panels; 4k for print-density grids. 4K exists at BOTH tiers and is API-only — even paid ChatGPT can't render it.
65
+
66
+ ## Prompting for text accuracy
67
+
68
+ - **Quote every string that must render verbatim**: `the sign reads "OPEN 24 HOURS"` — quoted strings render most reliably.
69
+ - Specify font *feel*, not font names: "clean geometric sans, high contrast", "hand-painted brush lettering".
70
+ - For dense text (posters, UI mocks), list the copy as ordered lines: `Line 1: "..." Line 2: "..."` — GPT Image 2 respects ordering.
71
+ - Keep total on-image text under ~30 words for perfect accuracy; beyond that, accuracy degrades gracefully but degrades.
72
+
73
+ ## Panels, sheets, and grids
74
+
75
+ - State the grid explicitly and number the cells: "a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom".
76
+ - Give each cell ONE content clause: "Panel 3: the character mid-jump, side view".
77
+ - Character identity sheets: GPT Image 2 holds both the structured panel layout AND photoreal skin, which is why the influencer-ad lane builds its sheets here at `quality: high`, 2K. Reach for NB2/NB Pro when the sheet needs many reference images folded in (14-ref ceiling) or when it is an edit of an existing sheet.
78
+
79
+ ## References & editing
80
+
81
+ Reference images route through the edit endpoint (up to ~10). The composed "image N" naming applies as everywhere else. Mask-based inpainting exists at the API level but isn't surfaced — describe the change instead.
82
+
83
+ ## 🚨 WHAT GETS YOU BLOCKED — read before writing a prompt with a person in it
84
+
85
+ **Receipt: 24 consecutive attempts on one character, 2026-08-24, same project and same rail.** Eleven were refused with `content_policy_violation` on the fal edit endpoint. The refusals were never about the scene — one of the blocked prompts was a woman standing at a kitchen counter with her hand on it. **Two phrasings were hard blocks, 5 for 5 each, and neither ever passed:**
86
+
87
+ **1. Never describe the reference as a photograph of a real person.**
88
+
89
+ > ❌ `Reference image 1 is a photograph of a woman. Use that exact woman.`
90
+ > ✅ `Reference image 1 is a character identity sheet showing one woman across several panels — the face in the large portrait panel is the authority for her identity. Use that exact woman.`
91
+
92
+ The first reads to the filter as *recreate this real person's likeness*, which is a hard refusal regardless of what the rest of the prompt says. The second signals a fictional character and passes. **This is a wording change only — the reference image can be the same file either way.** One plate flipped from refused to accepted on this single sentence with nothing else altered.
93
+
94
+ **2. Never attach a reference sheet containing a headless body panel.** A sheet whose full-body panels are cropped above the neck is refused every time, even with the correct opener. Regenerate the sheet with the head visible in every panel. Related, and already in this file's sheet guidance: phrase a cropped panel as *framing* (`cropped at the collarbone`), never as *absence* (`the head not shown`).
95
+
96
+ **On top of those, ordinary content triggers still apply** and they stack independently — a correct opener does not rescue them:
97
+
98
+ | Refused | Why, and the fix |
99
+ |---|---|
100
+ | A woman sitting on a bed in a bedroom | Domestic + bed reads as intimate. Move her to a chair, a rug, another room. |
101
+ | A knife, even lying flat on a chopping board next to a lemon | The object is the trigger, not the framing. Swap it — a cast-iron pan cleared instantly. |
102
+
103
+ **🚨 Refusals are PROBABILISTIC. Retry once before rewriting a word.** In the same session an identical prompt, identical reference, identical params was refused and then accepted on a straight re-fire. A rejected job returns no file and costs nothing, so a retry is free and a rewrite is not — rewriting first is how you end up changing four variables and learning nothing. **Only redesign after two or three refusals.**
104
+
105
+ **And change ONE thing at a time.** The eleven refusals above took far longer to diagnose than they should have because a reference swap and an opener rewrite shipped in the same call. Isolate on the prompt you actually want, so a pass leaves you with a usable asset instead of a data point.
106
+
107
+ ## Filter regime
108
+
109
+ OpenAI moderate — a third regime distinct from Gemini (NB family) and ByteDance (Seedream). Real-face references pass more readily than Gemini; violence/brand rules are similar. `slates-content-policy` applies unchanged.
@@ -0,0 +1,174 @@
1
+ ---
2
+ name: slates-prompting-inworld-tts
3
+ description: How to use Inworld Realtime TTS-2, the VOICE seat. Read before calling slates_generate_audio with model inworld-tts-2. Speech in a SPECIFIC voice, billed per character - the prompt is the words spoken, verbatim. Covers the identity-versus-acoustics rule (what a reference clip does and does not carry), how to write a line so it is performed rather than read, when to reach for seed-audio instead, and the voice-consent rule.
4
+ ---
5
+
6
+ # Inworld Realtime TTS-2 — the voice seat
7
+
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Keep it under 2,400 characters (the build fails above that) and keep
14
+ the rationale and the worked examples in the body below. -->
15
+ <!-- /slates-only -->
16
+ **Card — Inworld TTS-2.** Speech in a SPECIFIC voice. The prompt is the words spoken, verbatim — not a description of them. Text length determines the bill.
17
+
18
+ **IDENTITY, NOT ACOUSTICS — the rule that decides whether cloning works**
19
+ A reference carries WHO is speaking: timbre, pitch, accent, age, vowel shape. It does NOT carry WHERE they are — room tone, distance, phone EQ, reverb and mic character are *acoustics*, and this model reproduces the identity while discarding the room. So:
20
+ 1. **A noisy reference does not give a noisy read — it gives a WORSE identity.** Music, a second speaker or heavy reverb corrupt what is being extracted. Use a clean single-speaker recording.
21
+ 2. **You cannot get "on a payphone" by cloning a payphone recording.** Acoustics come from the MIX, or from `seed-audio` which renders a room.
22
+
23
+ **DIRECTION GOES IN SQUARE BRACKETS. PARENTHESES ARE SPOKEN ALOUD.** `[whispering] I hope nobody notices` is whispered; `(quietly) I hope nobody notices` says the word "quietly" out loud. Verified by ear — the easiest way to ruin a take.
24
+
25
+ - **Plain English works inside them** — it is natural-language steering, not a fixed vocabulary: `[very quiet]`, `[whisper in a hushed style]`, `[very slow]`, `[say excitedly]`. Non-verbals are their own tags: `[laugh]`, `[sigh]`, `[breathe]`, `[clear throat]`.
26
+ - **A tag it does not recognise is still consumed, and still changes the read.** Never spoken, never an error — so a mistyped tag fails SILENTLY and only listening catches it.
27
+ - **Tags persist across sentences** until changed; `[reset]` returns to normal.
28
+ - **Punctuation is the timing.** `Wait. Stop.` differs from `Wait, stop.`
29
+ - **One line, one take.** Split a paragraph so a bad clause costs one re-roll.
30
+ - **Spell numbers and titles aloud:** `twenty twenty-six`, `Doctor Reyes`.
31
+
32
+ **Route elsewhere when:** the scene needs dialogue mixed with effects and room tone in one pass (`seed-audio`), or it is a single non-speech sound (`eleven-sfx`). This surface makes ONE voice saying ONE thing, cleanly.
33
+
34
+ **Hard constraints:** no duration parameter — length falls out of the text. Exactly one voice source: a preset `voiceId` from `slates_list_voices`, a clip as `voiceReferenceAssetId` (a character's voice clip to speak AS the character, or any clean clip of one speaker), or `voiceDescription`.
35
+ <!-- @card:end -->
36
+
37
+ <!-- @banned:start -->
38
+ <!-- slates-only -->
39
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
40
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
41
+ estimate, and every submitted prompt is matched against it. Keep entries
42
+ backticked and prose outside the backticks. -->
43
+ <!-- /slates-only -->
44
+ **Never use** — the prompt on this surface is SPOKEN ALOUD, so anything that describes the audio instead of being the audio gets read out as words:
45
+
46
+ - `SFX`, `Ambient noise`, `Background music` as labels — this model speaks; it does not render a scene. Use `seed-audio` for those.
47
+ - shot language: `wide shot`, `slow push in`, `warm tungsten` — video-prompt words, and here they would literally be said aloud
48
+ - `voiceover`, `narrator says`, `he says` as stage directions wrapping the line — write only the words that should come out of the speaker
49
+ <!-- @banned:end -->
50
+
51
+ ## Why the prompt is not a prompt
52
+
53
+ On every other surface in Slates the prompt DESCRIBES what you want and the model interprets it. Here the prompt IS the deliverable: each character is spoken aloud and each character is billed. `a gravelly man says he is tired` produces a voice saying the words "a gravelly man says he is tired".
54
+
55
+ That also means the two numbers a user cares about are the same number. The text length sets the price (in 250-character buckets) and sets the length of the audio. There is nothing to choose and nothing to reconcile.
56
+
57
+ ## Steering the delivery
58
+
59
+ `VERIFIED BY EAR, 2026-09-05.` Every claim in this section was listened to, not
60
+ inferred — an earlier draft of this skill documented tag forms that had only been
61
+ probed for an HTTP 200, which proves the request was accepted and nothing about
62
+ whether it was obeyed.
63
+
64
+ **Square brackets are consumed. Parentheses are read aloud.** That is the whole
65
+ rule, and getting it wrong is not a subtle degradation — the audience hears a
66
+ narrator say the word "quietly" in the middle of your line.
67
+
68
+ | Written | What comes out |
69
+ |---|---|
70
+ | `[whispering] I really hope nobody notices that.` | whispered, tag not spoken ✅ |
71
+ | `[very quiet] I really hope nobody notices that.` | very quiet, tag not spoken ✅ |
72
+ | `[whisper in a hushed style] …` | hushed, tag not spoken ✅ |
73
+ | `[very slow] …` | slowed right down, tag not spoken ✅ |
74
+ | `[laugh] …` | an actual laugh, then the line ✅ |
75
+ | `(quietly, under his breath) …` | 🚨 **the words "quietly, under his breath" are SPOKEN** |
76
+
77
+ **Plain English works — it is natural-language steering, not a fixed vocabulary.**
78
+ Both the documented phrasings (`[whisper in a hushed style]`) and ordinary adverbs
79
+ (`[whispering]`) were obeyed. Write the direction the way you would say it to an
80
+ actor.
81
+
82
+ The eight dimensions the model steers on, with a working example of each:
83
+
84
+ | Dimension | Example |
85
+ |---|---|
86
+ | Emotion | `[say excitedly]`, `[sound sad]`, `[sound terrified]` |
87
+ | Articulation | `[say with force]`, `[articulate clearly]` |
88
+ | Intonation | `[say with a rising pitch]` |
89
+ | Volume | `[very quiet]`, `[very loud]` |
90
+ | Pitch | `[say in a low tone]` |
91
+ | Range | `[say playfully]`, `[say with no pitch variation]` |
92
+ | Speed | `[very fast]`, `[very slow]` |
93
+ | Vocal style | `[whisper in a hushed style]`, `[give a nasal quality]` |
94
+
95
+ Non-verbals sit inline where they happen: `[laugh]`, `[sigh]`, `[cough]`,
96
+ `[breathe]`, `[yawn]`, `[clear throat]`.
97
+
98
+ ### Four rules that are not obvious
99
+
100
+ 1. 🚨 **A tag it does not recognise is still consumed, and still changes the read.**
101
+ `[zzzqqq]` is not spoken and does not error — it produces a different, arbitrary
102
+ delivery. So a typo in a tag is SILENT: there is no rejection, no warning, and no
103
+ way to catch it except listening to the take. Treat an unexpected performance as
104
+ a possible misspelled tag before you blame the voice.
105
+ 2. **Tags persist across sentences.** A `[very slow]` at the top governs everything
106
+ after it until something changes it. Use `[reset]` to go back to normal rather
107
+ than assuming the next sentence starts clean.
108
+ 3. **Do not stack opposing directions.** `[whisper in a hushed style]` together with
109
+ `[very loud]` produces unpredictable results — the model is resolving a
110
+ contradiction, and which side wins is not something you can rely on.
111
+ 4. **Tags COUNT toward the billed characters**, even though they are never spoken.
112
+ They are part of the text sent to the vendor, so the vendor charges for them and
113
+ so do we — billing what was actually sent is the only honest basis. It rarely
114
+ matters (a 13-character tag inside a 250-character bucket), but a line sitting
115
+ just under a bucket boundary can be pushed into the next one by a long
116
+ direction. Prefer `[very slow]` over `[say this one very slowly please]`.
117
+
118
+ ## Identity versus acoustics, at length
119
+
120
+ This is the distinction that decides whether the feature feels good, and it is worth being precise about because the failure is quiet — you get a usable clip that is subtly not the person.
121
+
122
+ **What a reference clip transfers:** vocal timbre, pitch range, accent and regional vowels, apparent age, speech rate tendencies, and the particular rasp or breathiness of the source speaker.
123
+
124
+ **What it does not transfer:** the room, the microphone, the codec, the distance from the mic, any processing on the source, and any other sound present in it.
125
+
126
+ So the ideal reference is boring: one person, close to a microphone, no music, no second speaker, no heavy reverb, five to fifteen seconds, speaking normally rather than performing. A phone voice memo in a quiet room beats a beautifully produced clip with a music bed underneath it.
127
+
128
+ **Two failure modes, both common:**
129
+
130
+ - *"I cloned my podcast intro and it doesn't sound like me."* The intro had music under it. The model averaged the music into the identity. Re-clone from a clean stretch.
131
+ - *"I want the line to sound like it's coming through a car radio."* Clone the clean voice, then EQ and process the returned clip on the timeline. A radio-sounding reference makes a worse voice, not a radio effect.
132
+
133
+ ## Getting the voice onto the call
134
+
135
+ Exactly one source per call, and none of them requires a character to exist first:
136
+
137
+ - **A preset:** `slates_list_voices` lists stock voices with gender, age, accent and tags — filter by any of them, or search the descriptions ("gravelly", "narration"). Pass the chosen `voiceId`. Presets clone nothing, so they are the fastest path and avoid the clone-creation rate ceiling.
138
+ - **Speak AS a character:** `voiceReferenceAssetId: <its voiceAssetId>` (the clip on the row `slates_list_characters` returns). The seat clones the clip for that take and discards the vendor voice afterwards, so there is nothing to reconcile — but cloning shares a ceiling of two new voices a minute across every Slates user, so a run of lines in one cloned voice pauses between takes rather than failing. Send each line once; do not re-send one that already came back. Any other clean clip of one speaker works the same way.
139
+ - **A voice with no recording:** `voiceDescription` (7–1000 characters of words). If it will be used again, keep the returned clip on a character with `slates_update_character` (`voiceAssetId`) so later lines clone the same clip instead of designing a new voice each time — a convenience, never a requirement.
140
+
141
+ ## Consent
142
+
143
+ Cloning a real person's voice needs that person's explicit, documented permission, scoped to what you are making. Clone from original human recordings only — never from another model's output. This is the same gate the real-face route applies to likeness, and it applies here for the same reason.
144
+
145
+ ## Worked examples
146
+
147
+ **A line with a direction**
148
+
149
+ ```
150
+ [very quiet] I heard what you said in there. I'm not going to pretend I didn't.
151
+ ```
152
+
153
+ **A line that needs its numbers spoken**
154
+
155
+ ```
156
+ The vote was three hundred and twelve to eighty-nine. It carried at four minutes past midnight.
157
+ ```
158
+
159
+ **A paragraph, split into three takes** — so one bad clause costs one re-roll:
160
+
161
+ ```
162
+ 1. You keep asking me why I stayed.
163
+ 2. It wasn't loyalty. It wasn't even fear, not by the end.
164
+ 3. [very slow] It was that I couldn't picture the version of me that left.
165
+ ```
166
+
167
+ **What NOT to send**
168
+
169
+ ```
170
+ (gravelly, tired) a tired old man narrates the opening of the film, wide shot, warm tungsten
171
+ ```
172
+
173
+ Every word of that is spoken aloud — **including the parenthetical**, which is the
174
+ trap: it looks like a stage direction and is treated as dialogue. Describe the voice when you are CHOOSING one (`voiceDescription`, or the desktop's voice picker); the prompt is only ever the words.
@@ -5,6 +5,45 @@ description: How to prompt Kling V3.0 (Kuaishou). Read before calling slates_gen
5
5
 
6
6
  # Kling V3.0 — prompting
7
7
 
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved
14
+ compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.
15
+ Keep it under 2,400 characters (the build fails above that) and keep the
16
+ rationale, the receipts and the worked examples in the body below. -->
17
+ <!-- /slates-only -->
18
+ **Card — Kling V3.0.** The general default. Define the core subjects clearly at the START and keep those descriptions identical across shots. Up to 15s, up to 6 cuts, and the strongest image-to-video identity hold in the catalogue.
19
+
20
+ **The five levers**
21
+ 1. **Dialogue in quotes** — `Character says, "exact words here"`. On Omni, direct the voice with `Gender + Age + Voice quality + Speech rate + Emotional tone + Language`: `[Character A: Detective, mid-40s, raspy, slow cadence, weary]: "I've seen this before."`
22
+ 2. **Unique speaker labels, no pronouns after the introduction.** `he`, `the agent`, any synonym causes voice drift.
23
+ 3. **Sound has real syntax** — `SFX: heavy boots on wet pavement, distant siren wailing`, `Ambient noise: city traffic`, `Background music: low cello`. Always physical-cause specific; `SFX: footsteps` is not enough.
24
+ 4. **Motion adverbs modulate energy directly** — `slowly`, `rapidly`, `gently`, `explosively`. One primary camera move per shot, never stacked.
25
+ 5. **On image-to-video, do NOT re-describe the image.** It is an anchor; prompt how the scene EVOLVES from it — movement, camera, environmental change.
26
+
27
+ **Examples**
28
+ - `A detective in a wet grey overcoat stands under a stairwell light. He steps forward slowly as the light flickers. [Character A: Detective, mid-40s, raspy voice, slow cadence, weary]: "I've seen this before." SFX: heavy boots on wet concrete, distant siren wailing. Ambient noise: rain on metal.`
29
+ - `Camera tracks right alongside a cyclist crossing a bridge at dusk. She rises out of the saddle rapidly as the grade steepens. Ambient noise: wind, tyres on wet asphalt, distant traffic.`
30
+
31
+ **Hard constraint:** `Immediately` (Omni only) removes the natural conversational beat between speakers — use it when timing matters and leave it out when it does not. Kling has a real `negativePrompt` field, unlike Seedance; start from the standard block and layer scene-specific suppressions.
32
+ <!-- @card:end -->
33
+
34
+ <!-- @banned:start -->
35
+ <!-- slates-only -->
36
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
37
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
38
+ estimate, and every submitted prompt is matched against it. Keep entries
39
+ backticked and prose outside the backticks. -->
40
+ <!-- /slates-only -->
41
+ **Never use:**
42
+ - `SFX: footsteps` and any label-only effect — physical-cause specificity or nothing
43
+ - a pronoun or synonym for a speaker after the first introduction (`he`, `the agent`) — it causes voice drift; repeat the full label
44
+ - `single continuous take` — Seedance's phrase, and it fights Kling's multi-shot
45
+ <!-- @banned:end -->
46
+
8
47
  Kuaishou's video model. Three tiers: `kling-v3.0-std` (general use, no audio), `kling-v3.0-pro` (higher visual quality, no audio), `kling-v3.0-omni` (multi-character dialogue + audio-visual co-generation).
9
48
 
10
49
  Up to 15s. Multi-shot supported (up to 6 cuts in 15s total). Strong on image-to-video — preserves identity, layout, and text from the input image well.
@@ -5,6 +5,44 @@ description: How to set up lip-sync — Kling-only (dedicated lip-sync and avata
5
5
 
6
6
  # Lip-sync — setup guide
7
7
 
8
+ <!-- @card:start -->
9
+ <!-- slates-only -->
10
+ <!-- MACHINE-READ. Everything between the @card markers is extracted by
11
+ src/prompts/craft-cards.ts and returned on every cost estimate for this
12
+ model, so it is the ONE piece of positive craft guidance the agent cannot
13
+ skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved
14
+ compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.
15
+ Keep it under 2,400 characters (the build fails above that) and keep the
16
+ rationale, the receipts and the worked examples in the body below. -->
17
+ <!-- /slates-only -->
18
+ **Card — Lip-sync (Kling only).** Two different flows with different inputs and different prices; every output is 5 seconds.
19
+
20
+ **The five levers**
21
+ 1. **Pick `sourceType` deliberately** — `video` re-dubs an existing talking head (cheapest); `image` animates a still portrait (avatar-standard, then avatar-pro only on the final selected take).
22
+ 2. **The `prompt` on the avatar flows is SCENE CONTEXT, not motion direction.** Ambience, lighting, micro-expression: `Soft rim light`, `warm office`, `cool blue evening light through a window`, `gentle confident smile between sentences`, `focused intent expression`.
23
+ 3. **Clean the audio before uploading** — `noise-reduced`, `levelled`. Lip detection is sensitive, and a raw recording is the most common cause of a bad take.
24
+ 4. **Iterate on the SOURCE or the AUDIO, never on a refinement prompt** — there is not one. If the output is wrong, change the input.
25
+ 5. **Use avatar-standard for first-pass dialogue takes**, and switch to pro only once the line is locked. Facial fidelity is not visible until then.
26
+
27
+ **Examples**
28
+ - `Soft rim light, warm office, gentle confident smile between sentences.`
29
+ - `Cool blue evening light through a window, focused intent expression.` (Or `.` — an empty prompt is fine when you have nothing to add.)
30
+
31
+ **Hard constraint:** it is Kling-only and always 5 seconds. For a generated PERFORMANCE instead — head movement, gesture, delivery energy, with the dialogue as a native conditioning signal — that is a normal Seedance video generation with the clip attached as a video reference, not a mode of this tool. A real recording, or a cloned/cast voice rendered on `inworld-tts-2`, for production; this tool's built-in TTS is for scratch.
32
+ <!-- @card:end -->
33
+
34
+ <!-- @banned:start -->
35
+ <!-- slates-only -->
36
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is
37
+ extracted by src/prompts/banned-tokens.ts and returned on this model's cost
38
+ estimate, and every submitted prompt is matched against it. Keep entries
39
+ backticked and prose outside the backticks. -->
40
+ <!-- /slates-only -->
41
+ **Never use** — the avatar prompt is scene context and motion verbs are ignored:
42
+ - `turns her head`, `raises an eyebrow`, `hand gestures`, `nods`, `walks`
43
+ - `reader_en_m-v1` — listed in fal's docs, returns "Voice id not found" in production
44
+ <!-- @banned:end -->
45
+
8
46
  **This tool is Kling-only.** It wraps Kling's dedicated lip-sync and avatar endpoints; every entry is a real endpoint and every output is 5 seconds.
9
47
 
10
48
  | Flow | Source | Model | Cost | Use case |