@slatesvideo/shared 0.6.2 → 0.6.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.d.ts +3 -1
- package/dist/index.js +26 -0
- package/dist/operations/index.d.ts +1 -1
- package/dist/operations/index.js +211 -24
- package/dist/prompts/agent-doctrine.d.ts +36 -0
- package/dist/prompts/agent-doctrine.js +194 -0
- package/dist/prompts/banned-tokens.d.ts +28 -0
- package/dist/prompts/banned-tokens.js +152 -0
- package/dist/prompts/model-capabilities.d.ts +13 -1
- package/dist/prompts/model-capabilities.js +97 -2
- package/dist/prompts/model-facts.d.ts +20 -0
- package/dist/prompts/model-facts.js +87 -23
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +65 -0
- package/dist/skills/content.js +3 -2
- package/exports/slates-prompt-builder/generated/SKILL.md +3 -3
- package/exports/slates-prompt-builder/generated/reference-seedance.md +2 -1
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +8 -8
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +1 -1
- package/skills/slates-prompting-ltx-2-5.md +180 -0
- package/skills/slates-prompting-nano-banana-2.md +10 -0
- package/skills/slates-prompting-seedance.md +10 -1
|
@@ -16,9 +16,9 @@ This portable skill is deliberately thin. Its reference files are generated dire
|
|
|
16
16
|
<!-- @generated:model-routing -->
|
|
17
17
|
| Model | Canonical route | Guide |
|
|
18
18
|
|---|---|---|
|
|
19
|
-
| **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity
|
|
20
|
-
| **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction
|
|
21
|
-
| **Nano Banana 2 (Gemini 3.1 Flash Image)** |
|
|
19
|
+
| **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity, layout, text), acting, dialogue, lip-sync, and the widest aspect-ratio set. Escalate to Seedance for physics. Kling is also the ONLY engine behind the Motion Transfer and Lip Sync tools. | `reference-kling.md` |
|
|
20
|
+
| **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction or scale matter, and for hero shots. VIDEO-ONLY. Strong image-to-video and own-footage restyle. 4K is Pro-gated (base accounts get PRO_REQUIRED). Stays the default over 2.5: it is the only Seedance with native 4K and it is cheaper at every tier the two share. | `reference-seedance.md` |
|
|
21
|
+
| **Nano Banana 2 (Gemini 3.1 Flash Image)** | DEFAULT image model and the all-rounder — route here unless another seat's speciality is the point. Best start-frame for legible in-scene text. Knowledge cutoff Jan 2025: anything later needs reference images. | `reference-nano-banana.md` |
|
|
22
22
|
<!-- @end:model-routing -->
|
|
23
23
|
|
|
24
24
|
If the user names a model, use it. Otherwise route by the generated table above.
|
|
@@ -304,7 +304,8 @@ Speed ramps and slow-motion are supported in natural language, and `fast` is wid
|
|
|
304
304
|
the lid opens in slow-motion · the blade whips through the air
|
|
305
305
|
```
|
|
306
306
|
|
|
307
|
-
**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.
|
|
307
|
+
**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.
|
|
308
|
+
These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
|
|
308
309
|
|
|
309
310
|
## Style block at the end
|
|
310
311
|
|
|
@@ -12,7 +12,7 @@
|
|
|
12
12
|
},
|
|
13
13
|
{
|
|
14
14
|
"path": "skills/slates-prompting-seedance.md",
|
|
15
|
-
"sha256": "
|
|
15
|
+
"sha256": "c82ecbd6e169f74146e2553479bdc3115d550a332ee6ebb23a1068d6fe977d02"
|
|
16
16
|
},
|
|
17
17
|
{
|
|
18
18
|
"path": "skills/slates-prompting-kling-v3.md",
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
},
|
|
21
21
|
{
|
|
22
22
|
"path": "skills/slates-prompting-nano-banana-2.md",
|
|
23
|
-
"sha256": "
|
|
23
|
+
"sha256": "302c316ab7ce562a56db28ed480c4ed6bbffea1dd44ec468d122873b98491661"
|
|
24
24
|
},
|
|
25
25
|
{
|
|
26
26
|
"path": "skills/slates-content-policy.md",
|
|
@@ -28,14 +28,14 @@
|
|
|
28
28
|
},
|
|
29
29
|
{
|
|
30
30
|
"path": "src/prompts/model-facts.ts",
|
|
31
|
-
"sha256": "
|
|
31
|
+
"sha256": "1d6a8144643ab4ecd91fce7b16160b2152602c7de7079eab93eac700db433afc"
|
|
32
32
|
}
|
|
33
33
|
],
|
|
34
34
|
"outputs": [
|
|
35
35
|
{
|
|
36
36
|
"path": "SKILL.md",
|
|
37
|
-
"bytes":
|
|
38
|
-
"sha256": "
|
|
37
|
+
"bytes": 4598,
|
|
38
|
+
"sha256": "47cd166742efe1b15fb25603922695c5397ab9d24a839fb524d69cde426a6f89"
|
|
39
39
|
},
|
|
40
40
|
{
|
|
41
41
|
"path": "reference-character.md",
|
|
@@ -45,7 +45,7 @@
|
|
|
45
45
|
{
|
|
46
46
|
"path": "reference-seedance.md",
|
|
47
47
|
"bytes": 32558,
|
|
48
|
-
"sha256": "
|
|
48
|
+
"sha256": "c206db40577838baaaf7bc27ce0adfb1e0feb275ed9a8aea5fed377a74f0dacf"
|
|
49
49
|
},
|
|
50
50
|
{
|
|
51
51
|
"path": "reference-kling.md",
|
|
@@ -65,8 +65,8 @@
|
|
|
65
65
|
],
|
|
66
66
|
"archive": {
|
|
67
67
|
"path": "slates-prompt-builder.skill",
|
|
68
|
-
"bytes":
|
|
69
|
-
"sha256": "
|
|
68
|
+
"bytes": 38089,
|
|
69
|
+
"sha256": "4abe20453ab78bf4193db69d14a21f0a328238fec0b7e7dcb75bcffec160daee",
|
|
70
70
|
"entries": [
|
|
71
71
|
"SKILL.md",
|
|
72
72
|
"reference-character.md",
|
|
Binary file
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@slatesvideo/shared",
|
|
3
|
-
"version": "0.6.
|
|
3
|
+
"version": "0.6.3",
|
|
4
4
|
"description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -0,0 +1,180 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-ltx-2-5
|
|
3
|
+
description: How to prompt LTX-2.5 and LTX-2.5 Pro. Read before calling slates_generate_video with model ltx-2-5 or ltx-2-5-pro. LTX scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE — Lightricks ranks the prompt sound, camera, character detail, shot type and scene, then scene dressing, all in one flowing paragraph. It is also the catalogue's native MULTISHOT seat: one generation carries two to four connected shots holding character, light and voice across the cuts. Base ltx-2-5 is the distilled build — 720p/1080p/1440p/4K, clips of 6 to 20 seconds in EVEN steps, and the cheapest native 1080p second in Slates; ltx-2-5-pro is the full diffusion build and is NOT a superset, reaching only 1080p and 10 seconds for about a third more money. Three hazards live here: durations are even numbers only from six (there is no 5s or 7s clip), the model has NO reference endpoint at all so identity references are unavailable, and any sound not anchored to something in frame gets invented for you.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# LTX-2.5 — prompting
|
|
7
|
+
|
|
8
|
+
LTX-2.5 generates picture and sound **in a single pass**, with a Gemma-4 12B text encoder reading
|
|
9
|
+
one flowing paragraph. That single fact drives everything below: the prompt is not a shot
|
|
10
|
+
description with audio bolted on, it is **a scene where the sound is load-bearing** — and
|
|
11
|
+
Lightricks' own priority order puts sound first, ahead of the camera.
|
|
12
|
+
|
|
13
|
+
Two seats, and the naming is a trap:
|
|
14
|
+
|
|
15
|
+
| | `ltx-2-5` (base) | `ltx-2-5-pro` |
|
|
16
|
+
|---|---|---|
|
|
17
|
+
| Build | Distilled, 8-step | Full diffusion ("Diffusion Fidelity Rendering") |
|
|
18
|
+
| Resolutions | 720p / 1080p / **1440p** / 4K | 720p / 1080p |
|
|
19
|
+
| Durations | 6–20s, even steps | 6 / 8 / 10s |
|
|
20
|
+
| Price | $0.09–$0.30 per second | $0.12–$0.17 per second |
|
|
21
|
+
| Reach for it when | iterating, long takes, 4K delivery, batch volume | one dense final render inside 1080p and 10s |
|
|
22
|
+
|
|
23
|
+
**Pro is not "base plus more."** It buys picture quality on a *narrower* envelope — it cannot make
|
|
24
|
+
a 1440p frame and it cannot make a 12-second clip. Reaching for it out of habit costs a third more
|
|
25
|
+
*and* takes away the reach.
|
|
26
|
+
|
|
27
|
+
---
|
|
28
|
+
|
|
29
|
+
## 1. The six parts, in priority order, in one paragraph
|
|
30
|
+
|
|
31
|
+
Lightricks ranks the elements of an LTX prompt like this. When a prompt sprawls, **cut from the
|
|
32
|
+
bottom.**
|
|
33
|
+
|
|
34
|
+
1. **Sound** — highest priority; the model scores the picture as it draws it.
|
|
35
|
+
2. **Camera** — framing decides visual weight and the feel of the shot.
|
|
36
|
+
3. **Character detail** — expressed as physical action.
|
|
37
|
+
4. **Shot type and scene** — the action itself.
|
|
38
|
+
5. **Scene dressing** — the first thing to trim.
|
|
39
|
+
|
|
40
|
+
Write it as **one flowing paragraph**, not a list of labelled sections. LTX is not Seedance (eight
|
|
41
|
+
engineering slots) and not H3 (three separate audio layers) — it wants continuous prose.
|
|
42
|
+
|
|
43
|
+
---
|
|
44
|
+
|
|
45
|
+
## 2. Sound: anchor it or it gets invented
|
|
46
|
+
|
|
47
|
+
**Write the audio line last, then go back and check every cue has a source you could point at.**
|
|
48
|
+
Anything unanchored, the model invents for you.
|
|
49
|
+
|
|
50
|
+
The test is **"visible, or at least locatable."** A distant whistle is fine *if* you have named the
|
|
51
|
+
marshal's post it comes from. A "distant whistle" with nothing to attach to is a coin flip.
|
|
52
|
+
|
|
53
|
+
> the rope creaks against the cleat as she leans back, gulls calling somewhere off the port bow,
|
|
54
|
+
> the hull knocking hollow against the fenders
|
|
55
|
+
|
|
56
|
+
**Never write mood adjectives as sound.** "Tense atmosphere", "a sense of dread" and "ominous
|
|
57
|
+
ambience" produce nothing usable. If a scene feels thin, the fix is **one more moving object in
|
|
58
|
+
frame with a sound attached to it** — never another adjective.
|
|
59
|
+
|
|
60
|
+
### Dialogue
|
|
61
|
+
|
|
62
|
+
Quote it, and name the language and accent:
|
|
63
|
+
|
|
64
|
+
> "We should not have come back," in English with a slight German accent.
|
|
65
|
+
|
|
66
|
+
Two rules that decide whether the lip sync lands:
|
|
67
|
+
|
|
68
|
+
- **Give the character a beat of stillness before they speak.** The sync needs something to lock
|
|
69
|
+
against; a character already mid-motion when the line starts drifts.
|
|
70
|
+
- **Describe the beat structure** — when they look, how long they wait, when they speak, where they
|
|
71
|
+
look afterwards.
|
|
72
|
+
|
|
73
|
+
Slates pins the frame rate at 25fps, which is also what Lightricks recommends for dialogue: at 50fps
|
|
74
|
+
the performance "pulls toward a video look."
|
|
75
|
+
|
|
76
|
+
---
|
|
77
|
+
|
|
78
|
+
## 3. Character emotion is physical
|
|
79
|
+
|
|
80
|
+
The model renders actions. It does not render adjectives.
|
|
81
|
+
|
|
82
|
+
| Instead of | Write |
|
|
83
|
+
|---|---|
|
|
84
|
+
| she looks anxious | her jaw sets, she turns the ring on her finger twice |
|
|
85
|
+
| he seems exhausted | he blinks slowly and lets his shoulder take the doorframe |
|
|
86
|
+
| a tense standoff | neither moves; his thumb finds the strap and stays there |
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## 4. Multishot — the thing this model is uniquely for
|
|
91
|
+
|
|
92
|
+
**One LTX generation can carry several connected shots**, holding character, environment, lighting,
|
|
93
|
+
voice and style across every cut. Nothing else in the catalogue does this natively; everywhere else
|
|
94
|
+
you generate separate clips and stitch them, and identity drifts between them.
|
|
95
|
+
|
|
96
|
+
**Working range is two to four shots.** Three is the comfortable stopping point.
|
|
97
|
+
|
|
98
|
+
At **every** transition you must supply four things:
|
|
99
|
+
|
|
100
|
+
1. **Name the edit in the prose** — "hard cut", "dissolve", "match cut".
|
|
101
|
+
2. **Re-establish the shot completely** — scale, angle, lens and light all reset at a cut. A cut is
|
|
102
|
+
not a continuation.
|
|
103
|
+
3. **Re-identify recurring characters by their original descriptor.** "The woman in the bronze
|
|
104
|
+
gown", never "she". Pronouns lose the character across a cut — this is the single most common
|
|
105
|
+
multishot failure.
|
|
106
|
+
4. **State what the sound does at the cut.** Silence is not assumed; if the room tone should drop
|
|
107
|
+
out, say so.
|
|
108
|
+
|
|
109
|
+
A shape that works:
|
|
110
|
+
|
|
111
|
+
> Wide establishing shot of the workshop, dust in the window light, a lathe turning somewhere off
|
|
112
|
+
> frame — hard cut — macro close-up of the brass fitting as it seats, the turning noise gone,
|
|
113
|
+
> replaced by a single dry click — match cut — medium shot of the woman in the bronze gown stepping
|
|
114
|
+
> back, the room tone returning underneath her.
|
|
115
|
+
|
|
116
|
+
---
|
|
117
|
+
|
|
118
|
+
## 5. Camera: write it, don't enumerate it
|
|
119
|
+
|
|
120
|
+
fal exposes a `camera_motion` enum (dolly in/out/left/right, jib up/down, static, focus shift).
|
|
121
|
+
**Slates does not surface it, deliberately** — and prose is the better instrument anyway:
|
|
122
|
+
|
|
123
|
+
- **A written move can be tied to a specific moment.** "A slow push-in that settles as she reaches
|
|
124
|
+
the door, then holds" is not expressible as an enum value.
|
|
125
|
+
- **For multishot it would be actively wrong** — one enum value would impose a single camera
|
|
126
|
+
behaviour on three shots that each want their own.
|
|
127
|
+
|
|
128
|
+
So name the lens, the framing, the move, and **the moment the move resolves**.
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
132
|
+
## 6. The hard constraints
|
|
133
|
+
|
|
134
|
+
### Durations are even numbers only, starting at six
|
|
135
|
+
|
|
136
|
+
**6, 8, 10, 12, 14, 16, 18, 20.** There is no 5-second LTX clip and no odd duration of any length.
|
|
137
|
+
Asking for 7s is not a rounding matter — that generation does not exist.
|
|
138
|
+
|
|
139
|
+
**And the long end is 1080p-and-below only.** At 1440p and 4K the ceiling drops to **6, 8 or 10**.
|
|
140
|
+
|
|
141
|
+
fal's own default is `auto`, which lets the model pick the length from the described action.
|
|
142
|
+
**Slates always sends an explicit length instead**, so what you choose is what you are billed for.
|
|
143
|
+
Choose the length the beat needs.
|
|
144
|
+
|
|
145
|
+
### Aspect ratios: 16:9 and 9:16, and nothing else
|
|
146
|
+
|
|
147
|
+
The narrowest set in the catalogue alongside Veo. Square, 4:5 and 21:9 are not available on this
|
|
148
|
+
model at any resolution.
|
|
149
|
+
|
|
150
|
+
### Frames, not references
|
|
151
|
+
|
|
152
|
+
LTX takes a **start frame** and an **optional end frame** (which generates a transition between the
|
|
153
|
+
two). It has **no reference-to-video endpoint at all** — no identity references, no style
|
|
154
|
+
references, no environment references, no reference video, no reference audio.
|
|
155
|
+
|
|
156
|
+
**For character consistency across separate shots, use MiniMax H3 or Kling.** Within a single LTX
|
|
157
|
+
generation, use multishot instead — that is precisely the gap it fills.
|
|
158
|
+
|
|
159
|
+
In image-to-video, **do not cut away from the opening frame too early.** You have paid for that
|
|
160
|
+
frame; let it play before the first move.
|
|
161
|
+
|
|
162
|
+
### Do not ask for text on screen
|
|
163
|
+
|
|
164
|
+
Neither the spelling nor its stability from frame to frame can be relied on. Signage, labels,
|
|
165
|
+
captions and lower-thirds belong in post.
|
|
166
|
+
|
|
167
|
+
---
|
|
168
|
+
|
|
169
|
+
## 7. Audio is free here, and that changes the routing
|
|
170
|
+
|
|
171
|
+
Native synchronised audio is **included at every resolution on both seats**, with no surcharge and
|
|
172
|
+
no toggle that costs money — unlike Kling, where sound is a paid dimension. A 6-second 1080p LTX
|
|
173
|
+
clip **with sound** is 39 credits.
|
|
174
|
+
|
|
175
|
+
Combined with 1080p at $0.13/s — the cheapest native 1080p second in Slates — this makes LTX **the
|
|
176
|
+
coverage seat**: the one to reach for when the job is many takes rather than one hero shot, when a
|
|
177
|
+
sequence needs its own sound, or when the credit budget is the binding constraint.
|
|
178
|
+
|
|
179
|
+
Route away from it when you need identity references (H3, Kling), a ratio other than 16:9 or 9:16
|
|
180
|
+
(Seedance, Kling), or authored multi-layer audio direction (H3).
|
|
@@ -78,6 +78,15 @@ Film still from [DIRECTOR] [GENRE]. Shot on [CAMERA] with [LENS]. [SUBJECT and a
|
|
|
78
78
|
|
|
79
79
|
These are Stable-Diffusion-era tag soup. The model treats them as low-signal noise. Measured success rate: ~60-70% with these vs ~95%+ with positive description.
|
|
80
80
|
|
|
81
|
+
<!-- @banned:start -->
|
|
82
|
+
<!-- slates-only -->
|
|
83
|
+
<!-- MACHINE-READ. Every `backticked` token between the @banned markers is extracted
|
|
84
|
+
by src/prompts/banned-tokens.ts, inlined verbatim into the slates_generate_image
|
|
85
|
+
op description (always in context on both surfaces), and matched against every
|
|
86
|
+
submitted prompt. Editing this list changes what the agent is told AND what it
|
|
87
|
+
is warned about — keep every entry backticked, and keep prose outside the
|
|
88
|
+
backticks. -->
|
|
89
|
+
<!-- /slates-only -->
|
|
81
90
|
**Never use:**
|
|
82
91
|
- `8k`, `4k` (as a quality token)
|
|
83
92
|
- `hyperrealistic`, `ultra-realistic`, `photorealistic` standing alone
|
|
@@ -86,6 +95,7 @@ These are Stable-Diffusion-era tag soup. The model treats them as low-signal noi
|
|
|
86
95
|
- `perfect skin`, `flawless`, `airbrushed`, `smooth skin`
|
|
87
96
|
- `cinematic` standing alone — always specify *which cinema* (director, lens, era, stock)
|
|
88
97
|
- `not anime, not cartoon, not 3D` — negation tag soup, replace with a positive style cue
|
|
98
|
+
<!-- @banned:end -->
|
|
89
99
|
|
|
90
100
|
## Negative prompting — there is no field
|
|
91
101
|
|
|
@@ -339,7 +339,16 @@ Speed ramps and slow-motion are supported in natural language, and `fast` is wid
|
|
|
339
339
|
the lid opens in slow-motion · the blade whips through the air
|
|
340
340
|
```
|
|
341
341
|
|
|
342
|
-
|
|
342
|
+
<!-- @banned:start -->
|
|
343
|
+
<!-- slates-only -->
|
|
344
|
+
<!-- MACHINE-READ — same contract as the anti-list in slates-prompting-nano-banana-2.
|
|
345
|
+
Extracted by src/prompts/banned-tokens.ts into the slates_generate_video op
|
|
346
|
+
description and matched against submitted prompts. The RECOMMENDED vocabulary
|
|
347
|
+
below sits OUTSIDE the markers on purpose — it is backticked too. -->
|
|
348
|
+
<!-- /slates-only -->
|
|
349
|
+
**Slop tokens to avoid:** `epic`, `amazing`, `beautiful`, `lots of movement`, `8K`, `masterpiece`, `trending on artstation`.
|
|
350
|
+
<!-- @banned:end -->
|
|
351
|
+
These are quality *incantations* — the officially sanctioned way to ask for quality is the image-quality slot vocabulary in Part 1 (`HD`, `rich details`, `cinematic texture`, `natural colors`, `soft lighting`).
|
|
343
352
|
|
|
344
353
|
## Style block at the end
|
|
345
354
|
|