@slatesvideo/shared 0.6.8 → 0.6.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-prompting-minimax-h3
3
- description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09), so the seats differ on ladder and price, not on what they accept. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; 4 free then +1 on Max), and audio written into the wrong section is dropped or duplicated.
3
+ description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09), so the seats differ on ladder and price, not on what they accept. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; pooled media tokens on Max), and audio written into the wrong section is dropped or duplicated.
4
4
  ---
5
5
 
6
6
  # MiniMax H3 — prompting
@@ -28,7 +28,7 @@ description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling sl
28
28
  - `A woman sits still at a kitchen table for a beat, then looks up. She says in English, "You said Tuesday." Scene sound: a fridge hum, a spoon set down on formica. Score: none.`
29
29
  - `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, "No es el alternador." Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`
30
30
 
31
- **Hard constraint:** the two seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p rather than 4K, takes the same 9+3+3 references, and costs MORE at the tier they share — it is a speed pick, never the cheap one. H3's top two resolution tiers are UPSCALES of the native render: judge at native. Reference images past the free allowance are a paid dimension of the cost key (5 free on base, 4 on Max) — declare the count when quoting.
31
+ **Hard constraint:** the two seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 768p rather than 4K, takes the same 9+3+3 references, and costs MORE at the tier they share — it is a speed pick, never the cheap one. H3's top two resolution tiers are UPSCALES of the native render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.
32
32
  <!-- @card:end -->
33
33
 
34
34
  <!-- @banned:start -->
@@ -55,7 +55,7 @@ endpoint accepts:
55
55
 
56
56
  | | `minimax-h3` | `minimax-h3-max` |
57
57
  |---|---|---|
58
- | Resolution | 480p / 768p / **2K / 4K** | 480p / 768p / **1080p** |
58
+ | Resolution | 480p / 768p / **2K / 4K** | 480p / 768p |
59
59
  | References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) |
60
60
  | Frames | start and/or end | start and/or end |
61
61
  | Price at 768p | **$0.060/s** | $0.080/s |
@@ -71,6 +71,12 @@ paying for; route to base H3 for anything needing resolution, references, or the
71
71
  4.8s wall-clock, but the order of magnitude did. For iteration loops and client-present work that gap
72
72
  is the entire reason the seat exists.
73
73
 
74
+ 🚨 **Max's known weakness: colour banding in low light (Eric, 2026-09-09).** Certain shots —
75
+ especially dark or low-key ones — come back with low-bitrate-looking banding across gradients (skies,
76
+ walls, shadow falloff). It is the one place the seat visibly gives something up. If a shot is dark
77
+ and gradient-heavy, either light it up in the prompt or route to base H3 at 768p; do not fix it by
78
+ reaching for 2K, which adds its own artifacting on top.
79
+
74
80
  ---
75
81
 
76
82
  ## The one thing that makes H3 different: audio is a THREE-LAYER instruction
@@ -256,11 +262,12 @@ with at least one image or video reference.
256
262
 
257
263
  ### 💸 Reference images past the free allowance are billed — and the two rows differ
258
264
 
259
- On `minimax-h3` the first **5** are free and each additional image adds **4 credits**. On
260
- `minimax-h3-max` the first **4** are free and each additional image adds **1 credit** — fal prices
261
- Max's references by token rather than per image, and Slates normalises every Max reference to
262
- 1024x1024 so that per-image number is exact. Both rows take **9** images, at every resolution and
263
- every length. Four extra images on a 10s
265
+ On `minimax-h3` the first **5** are free and each additional image adds **4 credits**.
266
+ Max pools image pixels, reference-video seconds and reference-audio seconds into one token
267
+ allowance. Include `referenceImages`, `videoRefSeconds` and `audioRefSeconds` when estimating;
268
+ character voices count as audio. The generation preflight resolves the actual attached media.
269
+
270
+ Four extra images on a 10s
264
271
  768p clip add 16 credits to a 30-credit generation: **more than half again**, for references that
265
272
  often make the output worse rather than better (see the 2–4 rule above).
266
273
 
@@ -294,9 +301,7 @@ dropping one side.
294
301
  | `minimax-h3` · 2K · 10s | 65 |
295
302
  | `minimax-h3` · 4K · 10s | 80 |
296
303
  | `minimax-h3-max` · 768p · 10s | 40 |
297
- | `minimax-h3-max` · 1080p · 10s | 80 |
298
304
  | `minimax-h3` — every reference image past the **fifth** | **+4** |
299
- | `minimax-h3-max` — every reference image past the **fourth** | **+1** |
300
305
 
301
306
  **768p is the default for a reason.** It is the tier the model natively generates.
302
307
 
@@ -307,8 +312,13 @@ cannot add information.
307
312
 
308
313
  **In our own test (2026-08-27, same prompt, same seed) the 2K pass came back with MORE artifacting
309
314
  than the 768p original it was built from**, while costing 33 credits for a 5-second take against 15,
310
- and taking nearly twice as long to return. One shot, so treat it as a warning rather than a law —
311
- but the mechanism explains it, and the burden of proof is on 2K.
315
+ and taking nearly twice as long to return.
316
+
317
+ 🚨 **Confirmed independently (Eric, 2026-09-09): 2K and 4K carry visible AI noise artifacting and
318
+ "just look bad".** That is now TWO separate observations, months apart, pointing the same way — it is
319
+ no longer a single-shot warning. The tiers stay available because a delivery spec sometimes demands
320
+ the pixels, but **do not route to 2K/4K for quality**: you are paying more, waiting longer, and
321
+ adding artifacts to a 768p render. Upscale in post from a clean 768p master instead.
312
322
 
313
323
  **So: generate at 768p and judge it at 768p.** Reach for 2K or 4K only when a delivery spec demands
314
324
  the pixels, and expect to be paying for size rather than quality — a post-production upscale from a