@slatesvideo/shared 0.6.2 → 0.6.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -16,9 +16,9 @@ This portable skill is deliberately thin. Its reference files are generated dire
16
16
  <!-- @generated:model-routing -->
17
17
  | Model | Canonical route | Guide |
18
18
  |---|---|---|
19
- | **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics. Kling is also the ONLY engine behind the Motion Transfer and Lip Sync tools (MC std/pro, lip-sync, avatar) — those two tools are Kling-only. | `reference-kling.md` |
20
- | **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those). 4-15s, up to 9 ingredient images. Strong I2V / own-footage restyle. Native 4K, but 4K VIDEO is a Pro-only tier gate (base maxes at 1080p; server returns PRO_REQUIRED) — default 1080p unless the user is on Pro. Attaching a clip as a video reference (own-footage restyle, motion or dialogue conditioning) bills combined input+output seconds — at a DISCOUNTED per-second rate on every provider, roughly 0.6x the plain rate. 2.0 STAYS THE DEFAULT over 2.5 for two reasons, and neither is 1080p any more (2.5 gained 1080p on 2026-08-24): it is the only Seedance with native 4K, and it is cheaper at every shared tier (720p $0.15/s vs $0.231/s). | `reference-seedance.md` |
21
- | **Nano Banana 2 (Gemini 3.1 Flash Image)** | Default image model. 14 refs hard cap (10 object + 4 character). Brief it like a creative director, not tag soup. No negativePrompt field — use positive reframing. Best image start-frame for legible text. Knowledge cutoff Jan 2025. | `reference-nano-banana.md` |
19
+ | **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity, layout, text), acting, dialogue, lip-sync, and the widest aspect-ratio set. Escalate to Seedance for physics. Kling is also the ONLY engine behind the Motion Transfer and Lip Sync tools. | `reference-kling.md` |
20
+ | **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction or scale matter, and for hero shots. VIDEO-ONLY. Strong image-to-video and own-footage restyle. 4K is Pro-gated (base accounts get PRO_REQUIRED). Stays the default over 2.5: it is the only Seedance with native 4K and it is cheaper at every tier the two share. | `reference-seedance.md` |
21
+ | **Nano Banana 2 (Gemini 3.1 Flash Image)** | DEFAULT image model and the all-rounder — route here unless another seat's speciality is the point. Best start-frame for legible in-scene text. Knowledge cutoff Jan 2025: anything later needs reference images. | `reference-nano-banana.md` |
22
22
  <!-- @end:model-routing -->
23
23
 
24
24
  If the user names a model, use it. Otherwise route by the generated table above.
@@ -304,7 +304,8 @@ Speed ramps and slow-motion are supported in natural language, and `fast` is wid
304
304
  the lid opens in slow-motion · the blade whips through the air
305
305
  ```
306
306
 
307
- **Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`. These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
307
+ **Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.
308
+ These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
308
309
 
309
310
  ## Style block at the end
310
311
 
@@ -12,7 +12,7 @@
12
12
  },
13
13
  {
14
14
  "path": "skills/slates-prompting-seedance.md",
15
- "sha256": "ad9f3d5a20ef3bf52ebe916c81f67b9570fe1a43b54d64819da8b73972eb1a8e"
15
+ "sha256": "c82ecbd6e169f74146e2553479bdc3115d550a332ee6ebb23a1068d6fe977d02"
16
16
  },
17
17
  {
18
18
  "path": "skills/slates-prompting-kling-v3.md",
@@ -20,7 +20,7 @@
20
20
  },
21
21
  {
22
22
  "path": "skills/slates-prompting-nano-banana-2.md",
23
- "sha256": "fd47d89ffa5ac16c2d06e2d3c0e891813996dc15636a8988f7da6e8ee1816565"
23
+ "sha256": "302c316ab7ce562a56db28ed480c4ed6bbffea1dd44ec468d122873b98491661"
24
24
  },
25
25
  {
26
26
  "path": "skills/slates-content-policy.md",
@@ -28,14 +28,14 @@
28
28
  },
29
29
  {
30
30
  "path": "src/prompts/model-facts.ts",
31
- "sha256": "89a899444c021975df0a03a87b446217cdfc3b9086ef501e9e847936d7cfec2d"
31
+ "sha256": "1d6a8144643ab4ecd91fce7b16160b2152602c7de7079eab93eac700db433afc"
32
32
  }
33
33
  ],
34
34
  "outputs": [
35
35
  {
36
36
  "path": "SKILL.md",
37
- "bytes": 5174,
38
- "sha256": "6fecadc73933aa9e5001b3561aec31b3ca9edcc5146d79302720f1225bad1215"
37
+ "bytes": 4598,
38
+ "sha256": "47cd166742efe1b15fb25603922695c5397ab9d24a839fb524d69cde426a6f89"
39
39
  },
40
40
  {
41
41
  "path": "reference-character.md",
@@ -45,7 +45,7 @@
45
45
  {
46
46
  "path": "reference-seedance.md",
47
47
  "bytes": 32558,
48
- "sha256": "29b9408feec4ba91e5cd6f048ff6e6771fd2d96120378d4b6d52dec2a35eacd3"
48
+ "sha256": "c206db40577838baaaf7bc27ce0adfb1e0feb275ed9a8aea5fed377a74f0dacf"
49
49
  },
50
50
  {
51
51
  "path": "reference-kling.md",
@@ -65,8 +65,8 @@
65
65
  ],
66
66
  "archive": {
67
67
  "path": "slates-prompt-builder.skill",
68
- "bytes": 38420,
69
- "sha256": "869398265aaf8f8fba2987d1a0da24be8932f7d79bf7b67d9e866bd7ca6f134f",
68
+ "bytes": 38089,
69
+ "sha256": "4abe20453ab78bf4193db69d14a21f0a328238fec0b7e7dcb75bcffec160daee",
70
70
  "entries": [
71
71
  "SKILL.md",
72
72
  "reference-character.md",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@slatesvideo/shared",
3
- "version": "0.6.2",
3
+ "version": "0.6.3",
4
4
  "description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -0,0 +1,180 @@
1
+ ---
2
+ name: slates-prompting-ltx-2-5
3
+ description: How to prompt LTX-2.5 and LTX-2.5 Pro. Read before calling slates_generate_video with model ltx-2-5 or ltx-2-5-pro. LTX scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE — Lightricks ranks the prompt sound, camera, character detail, shot type and scene, then scene dressing, all in one flowing paragraph. It is also the catalogue's native MULTISHOT seat: one generation carries two to four connected shots holding character, light and voice across the cuts. Base ltx-2-5 is the distilled build — 720p/1080p/1440p/4K, clips of 6 to 20 seconds in EVEN steps, and the cheapest native 1080p second in Slates; ltx-2-5-pro is the full diffusion build and is NOT a superset, reaching only 1080p and 10 seconds for about a third more money. Three hazards live here: durations are even numbers only from six (there is no 5s or 7s clip), the model has NO reference endpoint at all so identity references are unavailable, and any sound not anchored to something in frame gets invented for you.
4
+ ---
5
+
6
+ # LTX-2.5 — prompting
7
+
8
+ LTX-2.5 generates picture and sound **in a single pass**, with a Gemma-4 12B text encoder reading
9
+ one flowing paragraph. That single fact drives everything below: the prompt is not a shot
10
+ description with audio bolted on, it is **a scene where the sound is load-bearing** — and
11
+ Lightricks' own priority order puts sound first, ahead of the camera.
12
+
13
+ Two seats, and the naming is a trap:
14
+
15
+ | | `ltx-2-5` (base) | `ltx-2-5-pro` |
16
+ |---|---|---|
17
+ | Build | Distilled, 8-step | Full diffusion ("Diffusion Fidelity Rendering") |
18
+ | Resolutions | 720p / 1080p / **1440p** / 4K | 720p / 1080p |
19
+ | Durations | 6–20s, even steps | 6 / 8 / 10s |
20
+ | Price | $0.09–$0.30 per second | $0.12–$0.17 per second |
21
+ | Reach for it when | iterating, long takes, 4K delivery, batch volume | one dense final render inside 1080p and 10s |
22
+
23
+ **Pro is not "base plus more."** It buys picture quality on a *narrower* envelope — it cannot make
24
+ a 1440p frame and it cannot make a 12-second clip. Reaching for it out of habit costs a third more
25
+ *and* takes away the reach.
26
+
27
+ ---
28
+
29
+ ## 1. The six parts, in priority order, in one paragraph
30
+
31
+ Lightricks ranks the elements of an LTX prompt like this. When a prompt sprawls, **cut from the
32
+ bottom.**
33
+
34
+ 1. **Sound** — highest priority; the model scores the picture as it draws it.
35
+ 2. **Camera** — framing decides visual weight and the feel of the shot.
36
+ 3. **Character detail** — expressed as physical action.
37
+ 4. **Shot type and scene** — the action itself.
38
+ 5. **Scene dressing** — the first thing to trim.
39
+
40
+ Write it as **one flowing paragraph**, not a list of labelled sections. LTX is not Seedance (eight
41
+ engineering slots) and not H3 (three separate audio layers) — it wants continuous prose.
42
+
43
+ ---
44
+
45
+ ## 2. Sound: anchor it or it gets invented
46
+
47
+ **Write the audio line last, then go back and check every cue has a source you could point at.**
48
+ Anything unanchored, the model invents for you.
49
+
50
+ The test is **"visible, or at least locatable."** A distant whistle is fine *if* you have named the
51
+ marshal's post it comes from. A "distant whistle" with nothing to attach to is a coin flip.
52
+
53
+ > the rope creaks against the cleat as she leans back, gulls calling somewhere off the port bow,
54
+ > the hull knocking hollow against the fenders
55
+
56
+ **Never write mood adjectives as sound.** "Tense atmosphere", "a sense of dread" and "ominous
57
+ ambience" produce nothing usable. If a scene feels thin, the fix is **one more moving object in
58
+ frame with a sound attached to it** — never another adjective.
59
+
60
+ ### Dialogue
61
+
62
+ Quote it, and name the language and accent:
63
+
64
+ > "We should not have come back," in English with a slight German accent.
65
+
66
+ Two rules that decide whether the lip sync lands:
67
+
68
+ - **Give the character a beat of stillness before they speak.** The sync needs something to lock
69
+ against; a character already mid-motion when the line starts drifts.
70
+ - **Describe the beat structure** — when they look, how long they wait, when they speak, where they
71
+ look afterwards.
72
+
73
+ Slates pins the frame rate at 25fps, which is also what Lightricks recommends for dialogue: at 50fps
74
+ the performance "pulls toward a video look."
75
+
76
+ ---
77
+
78
+ ## 3. Character emotion is physical
79
+
80
+ The model renders actions. It does not render adjectives.
81
+
82
+ | Instead of | Write |
83
+ |---|---|
84
+ | she looks anxious | her jaw sets, she turns the ring on her finger twice |
85
+ | he seems exhausted | he blinks slowly and lets his shoulder take the doorframe |
86
+ | a tense standoff | neither moves; his thumb finds the strap and stays there |
87
+
88
+ ---
89
+
90
+ ## 4. Multishot — the thing this model is uniquely for
91
+
92
+ **One LTX generation can carry several connected shots**, holding character, environment, lighting,
93
+ voice and style across every cut. Nothing else in the catalogue does this natively; everywhere else
94
+ you generate separate clips and stitch them, and identity drifts between them.
95
+
96
+ **Working range is two to four shots.** Three is the comfortable stopping point.
97
+
98
+ At **every** transition you must supply four things:
99
+
100
+ 1. **Name the edit in the prose** — "hard cut", "dissolve", "match cut".
101
+ 2. **Re-establish the shot completely** — scale, angle, lens and light all reset at a cut. A cut is
102
+ not a continuation.
103
+ 3. **Re-identify recurring characters by their original descriptor.** "The woman in the bronze
104
+ gown", never "she". Pronouns lose the character across a cut — this is the single most common
105
+ multishot failure.
106
+ 4. **State what the sound does at the cut.** Silence is not assumed; if the room tone should drop
107
+ out, say so.
108
+
109
+ A shape that works:
110
+
111
+ > Wide establishing shot of the workshop, dust in the window light, a lathe turning somewhere off
112
+ > frame — hard cut — macro close-up of the brass fitting as it seats, the turning noise gone,
113
+ > replaced by a single dry click — match cut — medium shot of the woman in the bronze gown stepping
114
+ > back, the room tone returning underneath her.
115
+
116
+ ---
117
+
118
+ ## 5. Camera: write it, don't enumerate it
119
+
120
+ fal exposes a `camera_motion` enum (dolly in/out/left/right, jib up/down, static, focus shift).
121
+ **Slates does not surface it, deliberately** — and prose is the better instrument anyway:
122
+
123
+ - **A written move can be tied to a specific moment.** "A slow push-in that settles as she reaches
124
+ the door, then holds" is not expressible as an enum value.
125
+ - **For multishot it would be actively wrong** — one enum value would impose a single camera
126
+ behaviour on three shots that each want their own.
127
+
128
+ So name the lens, the framing, the move, and **the moment the move resolves**.
129
+
130
+ ---
131
+
132
+ ## 6. The hard constraints
133
+
134
+ ### Durations are even numbers only, starting at six
135
+
136
+ **6, 8, 10, 12, 14, 16, 18, 20.** There is no 5-second LTX clip and no odd duration of any length.
137
+ Asking for 7s is not a rounding matter — that generation does not exist.
138
+
139
+ **And the long end is 1080p-and-below only.** At 1440p and 4K the ceiling drops to **6, 8 or 10**.
140
+
141
+ fal's own default is `auto`, which lets the model pick the length from the described action.
142
+ **Slates always sends an explicit length instead**, so what you choose is what you are billed for.
143
+ Choose the length the beat needs.
144
+
145
+ ### Aspect ratios: 16:9 and 9:16, and nothing else
146
+
147
+ The narrowest set in the catalogue alongside Veo. Square, 4:5 and 21:9 are not available on this
148
+ model at any resolution.
149
+
150
+ ### Frames, not references
151
+
152
+ LTX takes a **start frame** and an **optional end frame** (which generates a transition between the
153
+ two). It has **no reference-to-video endpoint at all** — no identity references, no style
154
+ references, no environment references, no reference video, no reference audio.
155
+
156
+ **For character consistency across separate shots, use MiniMax H3 or Kling.** Within a single LTX
157
+ generation, use multishot instead — that is precisely the gap it fills.
158
+
159
+ In image-to-video, **do not cut away from the opening frame too early.** You have paid for that
160
+ frame; let it play before the first move.
161
+
162
+ ### Do not ask for text on screen
163
+
164
+ Neither the spelling nor its stability from frame to frame can be relied on. Signage, labels,
165
+ captions and lower-thirds belong in post.
166
+
167
+ ---
168
+
169
+ ## 7. Audio is free here, and that changes the routing
170
+
171
+ Native synchronised audio is **included at every resolution on both seats**, with no surcharge and
172
+ no toggle that costs money — unlike Kling, where sound is a paid dimension. A 6-second 1080p LTX
173
+ clip **with sound** is 39 credits.
174
+
175
+ Combined with 1080p at $0.13/s — the cheapest native 1080p second in Slates — this makes LTX **the
176
+ coverage seat**: the one to reach for when the job is many takes rather than one hero shot, when a
177
+ sequence needs its own sound, or when the credit budget is the binding constraint.
178
+
179
+ Route away from it when you need identity references (H3, Kling), a ratio other than 16:9 or 9:16
180
+ (Seedance, Kling), or authored multi-layer audio direction (H3).
@@ -78,6 +78,15 @@ Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and a
78
78
 
79
79
  These are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.
80
80
 
81
+ <!-- @banned:start -->
82
+ <!-- slates-only -->
83
+ <!-- MACHINE-READ. Every `backticked` token between the @banned markers is extracted
84
+ by src/prompts/banned-tokens.ts, inlined verbatim into the slates_generate_image
85
+ op description (always in context on both surfaces), and matched against every
86
+ submitted prompt. Editing this list changes what the agent is told AND what it
87
+ is warned about — keep every entry backticked, and keep prose outside the
88
+ backticks. -->
89
+ <!-- /slates-only -->
81
90
  **Never use:**
82
91
  - `8k`, `4k` (as a quality token)
83
92
  - `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone
@@ -86,6 +95,7 @@ These are Stable-Diffusion-era tag soup. The model treats them as low-signal noi
86
95
  - `perfect skin`, `flawless`, `airbrushed`, `smooth skin`
87
96
  - `cinematic` standing alone — always specify *which cinema* (director, lens, era, stock)
88
97
  - `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue
98
+ <!-- @banned:end -->
89
99
 
90
100
  ## Negative prompting — there is no field
91
101
 
@@ -339,7 +339,16 @@ Speed ramps and slow-motion are supported in natural language, and `fast` is wid
339
339
  the lid opens in slow-motion · the blade whips through the air
340
340
  ```
341
341
 
342
- **Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`. These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
342
+ <!-- @banned:start -->
343
+ <!-- slates-only -->
344
+ <!-- MACHINE-READ — same contract as the anti-list in slates-prompting-nano-banana-2.
345
+ Extracted by src/prompts/banned-tokens.ts into the slates_generate_video op
346
+ description and matched against submitted prompts. The RECOMMENDED vocabulary
347
+ below sits OUTSIDE the markers on purpose — it is backticked too. -->
348
+ <!-- /slates-only -->
349
+ **Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.
350
+ <!-- @banned:end -->
351
+ These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
343
352
 
344
353
  ## Style block at the end
345
354