@slatesvideo/shared 0.5.10 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -17,7 +17,7 @@ This portable skill is deliberately thin. Its reference files are generated dire
17
17
  | Model | Canonical route | Guide |
18
18
  |---|---|---|
19
19
  | **Kling 3.0** | DEFAULT general-purpose video model — cost-effective, strong start-frame adherence (identity/layout/text), acting, dialogue, lip-sync, any aspect ratio. Escalate to Seedance for physics. Kling is also the ONLY engine behind the Motion Transfer and Lip Sync tools (MC std/pro, lip-sync, avatar) — those two tools are Kling-only. | `reference-kling.md` |
20
- | **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those). 4-15s, up to 9 ingredient images. Strong I2V / own-footage restyle. Native 4K, but 4K VIDEO is a Pro-only tier gate (base maxes at 1080p; server returns PRO_REQUIRED) — default 1080p unless the user is on Pro. Attaching a clip as a video reference (own-footage restyle, motion or dialogue conditioning) bills combined input+output seconds. 2.0 STAYS THE DEFAULT over 2.5 because it is the only Seedance with 1080p and 4K. | `reference-seedance.md` |
20
+ | **Seedance 2.0** | PREMIUM video tier and the DEFAULT video model — route here the moment physics, effects, destruction, or scale matter, and for hero shots. VIDEO-ONLY: cannot generate standalone images (use NB2/FLUX.2/Seedream for those). 4-15s, up to 9 ingredient images. Strong I2V / own-footage restyle. Native 4K, but 4K VIDEO is a Pro-only tier gate (base maxes at 1080p; server returns PRO_REQUIRED) — default 1080p unless the user is on Pro. Attaching a clip as a video reference (own-footage restyle, motion or dialogue conditioning) bills combined input+output seconds — at a DISCOUNTED per-second rate on every provider, roughly 0.6x the plain rate. 2.0 STAYS THE DEFAULT over 2.5 for two reasons, and neither is 1080p any more (2.5 gained 1080p on 2026-08-24): it is the only Seedance with native 4K, and it is cheaper at every shared tier (720p $0.15/s vs $0.231/s). | `reference-seedance.md` |
21
21
  | **Nano Banana 2 (Gemini 3.1 Flash Image)** | Default image model. 14 refs hard cap (10 object + 4 character). Brief it like a creative director, not tag soup. No negativePrompt field — use positive reframing. Best image start-frame for legible text. Knowledge cutoff Jan 2025. | `reference-nano-banana.md` |
22
22
  <!-- @end:model-routing -->
23
23
 
@@ -28,14 +28,14 @@
28
28
  },
29
29
  {
30
30
  "path": "src/prompts/model-facts.ts",
31
- "sha256": "998772c92eb0b443007c49a23f11c1ec80d9a0b4a28baacbecfe27cda8acbd21"
31
+ "sha256": "98f77646a75ae2078c5015638c416d1eb84de4fe057079f705439f1a16e30631"
32
32
  }
33
33
  ],
34
34
  "outputs": [
35
35
  {
36
36
  "path": "SKILL.md",
37
- "bytes": 4954,
38
- "sha256": "8d161adfb51a6da1800d9cacd93181e13966807a3490e75bd89a0093d0311cde"
37
+ "bytes": 5174,
38
+ "sha256": "6fecadc73933aa9e5001b3561aec31b3ca9edcc5146d79302720f1225bad1215"
39
39
  },
40
40
  {
41
41
  "path": "reference-character.md",
@@ -65,8 +65,8 @@
65
65
  ],
66
66
  "archive": {
67
67
  "path": "slates-prompt-builder.skill",
68
- "bytes": 38295,
69
- "sha256": "2cfe2ee4f80b0f64b53704a377c35d95ad15bedd5af3314f6b5242c8f6fb9bf1",
68
+ "bytes": 38420,
69
+ "sha256": "869398265aaf8f8fba2987d1a0da24be8932f7d79bf7b67d9e866bd7ca6f134f",
70
70
  "entries": [
71
71
  "SKILL.md",
72
72
  "reference-character.md",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@slatesvideo/shared",
3
- "version": "0.5.10",
3
+ "version": "0.6.0",
4
4
  "description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -5,6 +5,8 @@ description: Build a 30-second hyper-motion direct-response ad in Slates from a
5
5
 
6
6
  # Direct-response ad — Slates workflow
7
7
 
8
+ 🚨 **Wrong skill if a PERSON talks to camera.** This file builds a product-led hyper-motion spot — product hero frames, punchy cuts, no presenter. **A creator-style ad where a synthetic person speaks to the lens is a different discipline with an inverted rulebook (ugly on purpose, one shot per generation, plate-before-video): read `slates-ugc-influencer-ad`.** Route on the presence of a talking person, not on the platform.
9
+
8
10
  You are building a 30-second hyper-motion direct-response ad. The user has handed you a product image (or product URL) and a short brief. Slates desktop is open on the second monitor; the user watches it populate as you work.
9
11
 
10
12
  **Hard rules**
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: slates-model-selection
3
- description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Seedance 2.5 is a SECOND SEAT beside 2.0 (30s takes, 30 references and timestamp control, but 480p/720p only — never an upgrade); Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 or 9:16, 4/6/8s) and never the default.
3
+ description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Seedance 2.5 is a SECOND SEAT beside 2.0 (30s takes, 30 references and timestamp control, but no 4K and dearer at every shared resolution — never an upgrade); MiniMax H3 is the AUTHORED-AUDIO seat (three directable sound layers in one pass, declared reference relationships, 480p-4K) with MiniMax H3 Max beside it as a faster 768p-capped, reference-free premium; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 or 9:16, 4/6/8s) and never the default.
4
4
  ---
5
5
 
6
6
  # Model selection — the routing doctrine
@@ -28,7 +28,11 @@ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni F
28
28
  | Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
29
29
  | **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
30
30
  | The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
31
- | **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs, and the only Seedance seat that **acts on timestamps** (rules in `slates-prompting-seedance-2-5` § Timestamps) — and **480p or 720p on every route Slates offers, no 1080p and no 4K**. If resolution matters at all, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) 720p is NOT the cheap seat here — a 30s 720p face gen is 484 credits, more than a 15s 1080p Seedance 2.0 face gen (411), against a 1,000-credit welcome grant. Quote before any take over ~10s. |
31
+ | **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs, and the only Seedance seat that **acts on timestamps** (rules in `slates-prompting-seedance-2-5` § Timestamps) — 480p / 720p / 1080p, **no 4K**, and **dearer than 2.0 at every shared resolution** (720p $0.231/s vs $0.15/s, +54%). If you want 4K, or the same resolution cheaper, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) LENGTH is the price dial, not resolution — a 30s 720p face gen is 489 credits and a 30s 1080p faceless gen is 614, against a 1,000-credit welcome grant. Quote before any take over ~10s. |
32
+ | **The SOUND has to be directed, not just present** — a specific line delivered a specific way, scene sound that has to sit under it, and score that must stay out of the characters' world | **MiniMax H3** | The only seat where audio is authored in three separate layers in ONE pass (synchronised events in the body, ambience in a soundscape section, audience-only score in its own) rather than toggled on. 5–15s, 480p / 768p / 2K / 4K, 24fps, 32kHz stereo, 11 languages. Rules in `slates-prompting-minimax-h3`. |
33
+ | **A reference has to keep a DECLARED amount of itself** — especially moving one subject's characteristic onto a *different* subject | **MiniMax H3** | The only seat that understands a stated retention relationship (kept whole / kept in part / transferred onto another subject / loose echo). 9 images + 3 video + 3 audio, 12 files total. 🚨 The first 5 reference images are free and every one after that costs 4 credits — pass `referenceImages` to `slates_estimate_generation_cost` before a reference-heavy job. |
34
+ | **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p/768p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at 768p with NO references of any kind. Never the default; never reach for it to save money. |
35
+ | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | Narrow, and now narrower: if the sound needs DIRECTING rather than merely existing, MiniMax H3 is the better seat. |
32
36
 
33
37
  ### Named Seedance escalation triggers
34
38
 
@@ -40,7 +44,6 @@ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni F
40
44
  - **One continuous unbroken take.**
41
45
 
42
46
  Concrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).
43
- | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | The only job Veo wins. |
44
47
 
45
48
  ## Video EDIT routing (changing an existing clip)
46
49
 
@@ -49,8 +52,8 @@ Concrete beats route better than an abstract category. Cost stays a tiebreaker,
49
52
  | **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + "Keep everything else the same"; long prompts destroy it (see below). |
50
53
  | **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
51
54
  | **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
52
- | Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
53
- | **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |
55
+ | Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. (2.5's relocate lane reaches 1080p too as of 2026-08-24, at $0.2457/s of combined input+output.) Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
56
+ | **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p/1080p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |
54
57
  | AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
55
58
 
56
59
  - **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
@@ -79,7 +82,7 @@ Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that
79
82
  - **9:16 vertical → Kling or Seedance by preference**, not by necessity: Veo does take 9:16 on the route Slates uses. Route away from it because it is the niche seat, not because it can't.
80
83
  - **Ratios and durations are enforced before submit.** `slates_generate_video` validates the aspect ratio, resolution and duration against the model you picked and refuses out-of-set values with the legal list — it will not silently ignore or downgrade them. The authoritative per-model sets are in the op's own param descriptions, which are generated from the capability SSOT; prefer those over any list written in prose here.
81
84
  - **Image-to-video from an NB2 start frame** (the standard pipeline) → Kling by default, Seedance when the motion is physics-heavy. Not Veo.
82
- - **User names a model explicitly → use it.** But if it's a mismatch for the job (crazy physics on Kling std, a 30s take on anything but Seedance 2.5, 1080p on Seedance 2.5 which has none), say so in one line and offer the right route before generating.
85
+ - **User names a model explicitly → use it.** But if it's a mismatch for the job (crazy physics on Kling std, a 30s take on anything but Seedance 2.5, 4K on Seedance 2.5 which has none), say so in one line and offer the right route before generating.
83
86
 
84
87
  ## Image routing
85
88
 
@@ -88,11 +91,13 @@ Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that
88
91
  - **Default: Nano Banana 2** — best reference handling (14 refs), best legible text, the standard start-frame generator.
89
92
  - **NB2 Lite** — the fast/draft seat: ~half NB2's price, ~2.7× faster, 1K only. Route iteration volume and drafts here; finals go back to NB2 full (2K/4K).
90
93
  - **Nano Banana Pro** — the hero-frame/typography ceiling (~2× NB2). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — feed it a full subject library.
91
- - **GPT Image 2** — readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. Medium quality is the default (half NB2's price at 1080p); high (~4×) only when text precision is the whole job. 4K at both tiers is API-only — even paid ChatGPT can't render it.
94
+ - **GPT Image 2** — readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. **Also the photoreal front-runner (Eric, 2026-08-24)** — at `quality: high` it beat both Nano Banana rails head-to-head on skin realism, which is why the AI-influencer ad lane generates every plate here. Medium is the value seat (half NB2's price at 1080p); **high is the seat for photoreal skin and for text precision**. 4K at both tiers is API-only — even paid ChatGPT can't render it.
92
95
  - **FLUX.2 Max** — photoreal texture, hex-color binding, typography, less censored.
93
96
  - **Seedream 5 Lite** — uncensored + any-resolution flat price; volume exploration when the Gemini filter is in the way.
94
97
 
95
- **Split rule of thumb:** readable text / panels / UI → GPT Image 2; photoreal, character-locked, widescreen, or edit-heavy → the Banana line; drafts → NB2 Lite; uncensored or odd resolutions → Seedream/FLUX.
98
+ **Split rule of thumb:** readable text / panels / UI **and photoreal people** → GPT Image 2 (`high` for photoreal); edit-heavy work, or anything needing the 14-reference ceiling → the Banana line; drafts → NB2 Lite; uncensored or odd resolutions → Seedream/FLUX.
99
+
100
+ ⚠️ **This line said the opposite until 2026-08-24** — it sent photoreal *away* from GPT Image 2 on reputation, which is the exact failure § The meta-rule above warns about. Re-run the evidence test when the roster moves.
96
101
 
97
102
  ## Audio routing
98
103
 
@@ -59,6 +59,7 @@ Per shot: `slates_generate_image` with `referenceAssetIds` pointing at the chara
59
59
  **Model mixing — route per `slates-model-selection`** (details in the per-model guides):
60
60
  - **Kling V3** (`slates-prompting-kling-v3`): the DEFAULT for most shots — 16:9 / 9:16 / 1:1, 3-15s, strong start-frame adherence; std is the workhorse, Omni for multi-character dialogue.
61
61
  - **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K (4K video is Pro-only).
62
+ - **MiniMax H3** (`slates-prompting-minimax-h3`): route here when a shot's SOUND is part of the writing — a line delivered a particular way, scene sound under it, score that must stay outside the characters' world. It authors all three in one pass, which **collapses a shot's audio pass into its video pass** and removes the separate `slates_generate_audio` step for that shot. 5-15s, 480p/768p/2K/4K. Its sibling `minimax-h3-max` is faster but capped at 768p, takes no references, and costs MORE at 768p — a deliberate speed pick, never a saving.
62
63
  - **Veo 3.1** (`slates-prompting-veo-3`): niche, never the default — only when native synced audio must generate WITH the video in one gen; 16:9 or 9:16, 4/6/8s (8s only at 1080p/4K or with reference images).
63
64
 
64
65
  Failed gen? Check the error via `slates_get_generation_status`, fix the prompt, resubmit that one shot (a retry beyond the plan = announce the delta cost).
@@ -1,40 +1,70 @@
1
- ---
2
- name: slates-prompting-gpt-image-2
3
- description: Prompting GPT Image 2 — the readable-text / character-sheet / shot-grid engine. Read before calling slates_generate_image with model gpt-image-2. Covers the quality tiers (medium default, high for max text precision), resolution classes (1k/2k=1080p/3k=1440p/4k), text-accuracy prompting, panel/grid layout direction, and when to route to the Banana line instead.
4
- ---
5
-
6
- # GPT Image 2 — sheets, grids, and text that actually reads
7
-
8
- GPT Image 2's edge is **character-level text accuracy** (~99% on English), ordered panels, and exact element placement — the jobs where every other model garbles a word or shuffles a layout. It is NOT the photoreal or character-locked pick: route those to the Banana line (`slates-model-selection` has the split).
9
-
10
- ## Quality tiers — always set explicitly
11
-
12
- - **medium** (default) — sharp text, fast, the value seat: half NB2's price at the 1080p class. Blind benchmarks put it within a hair of high at a quarter of the cost. Start here.
13
- - **high** — ~4× the price; max text precision + reasoning. A deliberate premium pick when tiny type, dense diagrams, or many labeled elements ARE the job.
14
-
15
- Never rely on the provider default (it's high — the priciest tier). The Slates ops send medium unless you say otherwise.
16
-
17
- ## Resolution classes
18
-
19
- `1k` = 1024²-class · `2k` = 1920×1080-class · `3k` = 2560×1440-class · `4k` = 3840×2160-class. Pick 2k for most sheets/panels; 4k for print-density grids. 4K exists at BOTH tiers and is API-only — even paid ChatGPT can't render it.
20
-
21
- ## Prompting for text accuracy
22
-
23
- - **Quote every string that must render verbatim**: `the sign reads "OPEN 24 HOURS"` — quoted strings render most reliably.
24
- - Specify font *feel*, not font names: "clean geometric sans, high contrast", "hand-painted brush lettering".
25
- - For dense text (posters, UI mocks), list the copy as ordered lines: `Line 1: "..." Line 2: "..."` — GPT Image 2 respects ordering.
26
- - Keep total on-image text under ~30 words for perfect accuracy; beyond that, accuracy degrades gracefully but degrades.
27
-
28
- ## Panels, sheets, and grids
29
-
30
- - State the grid explicitly and number the cells: "a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom".
31
- - Give each cell ONE content clause: "Panel 3: the character mid-jump, side view".
32
- - Character identity sheets: GPT Image 2 holds structured panel layouts; the Banana line holds the *face* better. Prefer NB2/NB Pro for identity-critical sheets and GPT Image 2 when labels or annotations are the main requirement.
33
-
34
- ## References & editing
35
-
36
- Reference images route through the edit endpoint (up to ~10). The composed "image N" naming applies as everywhere else. Mask-based inpainting exists at the API level but isn't surfaced — describe the change instead.
37
-
38
- ## Filter regime
39
-
40
- OpenAI moderate — a third regime distinct from Gemini (NB family) and ByteDance (Seedream). Real-face references pass more readily than Gemini; violence/brand rules are similar. `slates-content-policy` applies unchanged.
1
+ ---
2
+ name: slates-prompting-gpt-image-2
3
+ description: Prompting GPT Image 2 — the readable-text / character-sheet / shot-grid engine AND the current photoreal front-runner. Read before calling slates_generate_image with model gpt-image-2. Covers the quality tiers (medium default, high for max text precision), resolution classes (1k/2k=1080p/3k=1440p/4k), text-accuracy prompting, panel/grid layout direction, and when to route to the Banana line instead.
4
+ ---
5
+
6
+ # GPT Image 2 — sheets, grids, and text that actually reads
7
+
8
+ GPT Image 2's edge is **character-level text accuracy** (~99% on English), ordered panels, and exact element placement — the jobs where every other model garbles a word or shuffles a layout.
9
+
10
+ 🚨 **It is ALSO the photoreal front-runner, and this file said the opposite until 2026-08-24.** **Receipts:** Eric's direct call, plus a head-to-head on the Higgsfield rail where GPT Image 2 at `quality: high`, 2K beat both Nano Banana rails on skin realism for photoreal people — that result is why the whole AI-influencer ad lane generates its plates here. **Route photoreal to this model, not away from it.**
11
+
12
+ **What the Banana line still owns:** edit-heavy work and the 14-reference ceiling.
13
+
14
+ **What would kill this:** a head-to-head at the intended crop going the other way. Per `slates-model-selection` § The meta-rule, re-run the evidence test when the roster changes — never carry a ranking forward on reputation. That rule is exactly what this correction failed.
15
+
16
+ ## Quality tiers — always set explicitly
17
+
18
+ - **medium** (default) — sharp text, fast, the value seat: half NB2's price at the 1080p class. Blind benchmarks put it within a hair of high at a quarter of the cost. Start here.
19
+ - **high** — ~4× the price; max text precision + reasoning. A deliberate premium pick when tiny type, dense diagrams, or many labeled elements ARE the job.
20
+
21
+ Never rely on the provider default (it's high — the priciest tier). The Slates ops send medium unless you say otherwise.
22
+
23
+ ## Resolution classes
24
+
25
+ `1k` = 1024²-class · `2k` = 1920×1080-class · `3k` = 2560×1440-class · `4k` = 3840×2160-class. Pick 2k for most sheets/panels; 4k for print-density grids. 4K exists at BOTH tiers and is API-only — even paid ChatGPT can't render it.
26
+
27
+ ## Prompting for text accuracy
28
+
29
+ - **Quote every string that must render verbatim**: `the sign reads "OPEN 24 HOURS"` — quoted strings render most reliably.
30
+ - Specify font *feel*, not font names: "clean geometric sans, high contrast", "hand-painted brush lettering".
31
+ - For dense text (posters, UI mocks), list the copy as ordered lines: `Line 1: "..." Line 2: "..."` — GPT Image 2 respects ordering.
32
+ - Keep total on-image text under ~30 words for perfect accuracy; beyond that, accuracy degrades gracefully but degrades.
33
+
34
+ ## Panels, sheets, and grids
35
+
36
+ - State the grid explicitly and number the cells: "a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom".
37
+ - Give each cell ONE content clause: "Panel 3: the character mid-jump, side view".
38
+ - Character identity sheets: GPT Image 2 holds both the structured panel layout AND photoreal skin, which is why the influencer-ad lane builds its sheets here at `quality: high`, 2K. Reach for NB2/NB Pro when the sheet needs many reference images folded in (14-ref ceiling) or when it is an edit of an existing sheet.
39
+
40
+ ## References & editing
41
+
42
+ Reference images route through the edit endpoint (up to ~10). The composed "image N" naming applies as everywhere else. Mask-based inpainting exists at the API level but isn't surfaced — describe the change instead.
43
+
44
+ ## 🚨 WHAT GETS YOU BLOCKED — read before writing a prompt with a person in it
45
+
46
+ **Receipt: 24 consecutive attempts on one character, 2026-08-24, same project and same rail.** Eleven were refused with `content_policy_violation` on the fal edit endpoint. The refusals were never about the scene — one of the blocked prompts was a woman standing at a kitchen counter with her hand on it. **Two phrasings were hard blocks, 5 for 5 each, and neither ever passed:**
47
+
48
+ **1. Never describe the reference as a photograph of a real person.**
49
+
50
+ > ❌ `Reference image 1 is a photograph of a woman. Use that exact woman.`
51
+ > ✅ `Reference image 1 is a character identity sheet showing one woman across several panels — the face in the large portrait panel is the authority for her identity. Use that exact woman.`
52
+
53
+ The first reads to the filter as *recreate this real person's likeness*, which is a hard refusal regardless of what the rest of the prompt says. The second signals a fictional character and passes. **This is a wording change only — the reference image can be the same file either way.** One plate flipped from refused to accepted on this single sentence with nothing else altered.
54
+
55
+ **2. Never attach a reference sheet containing a headless body panel.** A sheet whose full-body panels are cropped above the neck is refused every time, even with the correct opener. Regenerate the sheet with the head visible in every panel. Related, and already in this file's sheet guidance: phrase a cropped panel as *framing* (`cropped at the collarbone`), never as *absence* (`the head not shown`).
56
+
57
+ **On top of those, ordinary content triggers still apply** and they stack independently — a correct opener does not rescue them:
58
+
59
+ | Refused | Why, and the fix |
60
+ |---|---|
61
+ | A woman sitting on a bed in a bedroom | Domestic + bed reads as intimate. Move her to a chair, a rug, another room. |
62
+ | A knife, even lying flat on a chopping board next to a lemon | The object is the trigger, not the framing. Swap it — a cast-iron pan cleared instantly. |
63
+
64
+ **🚨 Refusals are PROBABILISTIC. Retry once before rewriting a word.** In the same session an identical prompt, identical reference, identical params was refused and then accepted on a straight re-fire. A rejected job returns no file and costs nothing, so a retry is free and a rewrite is not — rewriting first is how you end up changing four variables and learning nothing. **Only redesign after two or three refusals.**
65
+
66
+ **And change ONE thing at a time.** The eleven refusals above took far longer to diagnose than they should have because a reference swap and an opener rewrite shipped in the same call. Isolate on the prompt you actually want, so a pass leaves you with a usable asset instead of a data point.
67
+
68
+ ## Filter regime
69
+
70
+ OpenAI moderate — a third regime distinct from Gemini (NB family) and ByteDance (Seedream). Real-face references pass more readily than Gemini; violence/brand rules are similar. `slates-content-policy` applies unchanged.
@@ -0,0 +1,287 @@
1
+ ---
2
+ name: slates-prompting-minimax-h3
3
+ description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling slates_generate_video with model minimax-h3 or minimax-h3-max. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, capped at 768p, takes NO references of any kind, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one. Two hazards live here: reference images past the fifth cost 4 credits each on the base row, and audio written into the wrong section is dropped or duplicated.
4
+ ---
5
+
6
+ # MiniMax H3 — prompting
7
+
8
+ H3 is an **omni transformer**: it generates picture and sound in the same pass, at 24fps with
9
+ 32kHz stereo, 5–15 seconds, in 11 stably-supported languages (Arabic, Chinese, English, French,
10
+ German, Italian, Japanese, Korean, Portuguese, Russian, Spanish). That single fact drives
11
+ everything below — the prompt is not a shot description with sound bolted on, it is a **timeline
12
+ with three audio layers you author separately**.
13
+
14
+ **Two seats, one grammar.** Everything in this file applies to both. They differ only in what the
15
+ endpoint accepts:
16
+
17
+ | | `minimax-h3` | `minimax-h3-max` |
18
+ |---|---|---|
19
+ | Resolution | 480p / 768p / **2K / 4K** | 480p / 768p |
20
+ | References | 9 images + 3 video + 3 audio (12 files) | **none — no reference endpoint exists** |
21
+ | Frames | start and/or end | start and/or end |
22
+ | Price at 768p | **$0.060/s** | $0.080/s |
23
+ | Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) |
24
+
25
+ **Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it
26
+ tops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth
27
+ paying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.
28
+
29
+ **The speed is measured, not claimed** (2026-08-27, same prompt and params on both rows): a 5-second
30
+ 768p text-to-video finished in **4.8 seconds** on Max against **57 seconds** on base H3 — roughly
31
+ **12x**, queue to finished file. fal advertises "under 3 seconds"; the literal claim did not hold at
32
+ 4.8s wall-clock, but the order of magnitude did. For iteration loops and client-present work that gap
33
+ is the entire reason the seat exists.
34
+
35
+ ---
36
+
37
+ ## The one thing that makes H3 different: audio is a THREE-LAYER instruction
38
+
39
+ Every other video seat treats sound as on or off. H3 splits it, and the split is enforced by where
40
+ you write each thing. Get the section wrong and the sound is dropped, doubled, or attributed to the
41
+ wrong source.
42
+
43
+ | Layer | What belongs in it | Where it goes |
44
+ |---|---|---|
45
+ | **Synchronised events** | dialogue, singing, and any sound tied to a specific shot or action | the **body** of the prompt, on the beat it lands |
46
+ | **Scene sound** | ambience and physical sounds that run across the whole clip — room tone, rain, traffic, a ventilation hum | the **soundscape** section |
47
+ | **Score** | music the characters cannot hear; audience-only | the **music** section |
48
+
49
+ **Three rules, all from MiniMax's own guide:**
50
+
51
+ 1. **Dialogue and singing NEVER go in the soundscape section.** They are synchronised events; they
52
+ belong in the body, at the moment they happen.
53
+ 2. **Diegetic music — music the characters can hear** (a radio in the scene, a busker) — also
54
+ belongs in the **body**, not in the score section. The score section is audience-only.
55
+ 3. **Write the score in instrumental terms, not mood words.** Name the instruments, the tempo, and
56
+ how it develops. *"A restrained solo-piano score at a slow tempo, sustained low cello underneath,
57
+ no swell"* — not *"emotional music"*.
58
+
59
+ Use **N/A** for a section only when silence or absence is genuinely what the shot wants. An empty
60
+ score section is a real choice; a vague one is a wasted layer.
61
+
62
+ ### The shape, in the one prompt field
63
+
64
+ Slates sends one prompt string, so write the three layers as labelled paragraphs in this order:
65
+
66
+ ```
67
+ [Shot 1] Live-action, cinematic. A medium-wide shot frames a baker opening the shutters of a
68
+ small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as
69
+ the middle-aged baker with a calm, slightly raspy voice places a fresh loaf on the counter and
70
+ says: "First batch of the morning." [Shot 2] At 00:05.000, the camera cuts to a close-up of
71
+ steam rising from the sliced bread while his final words carry over from the previous shot.
72
+
73
+ Soundscape: wooden shutters scrape open over a quiet street, trays clink softly inside, a
74
+ doorbell rings once, then light footsteps and the crisp sound of bread being sliced.
75
+
76
+ Score: a soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes,
77
+ gentle fade at the end.
78
+ ```
79
+
80
+ **Body target: 350–500 words** for a reference-carrying shot. Dialogue-heavy content prioritises
81
+ fitting the complete spoken timeline over hitting a word count.
82
+
83
+ 🚨 **Slates disables the provider's prompt expander.** H3's API can rewrite your prompt before
84
+ generation; Slates turns that off, because a model rewriting the user's words invisibly is banned
85
+ outright (prompt transparency: what the composer shows is what the model gets). The practical
86
+ consequence is on you: **nothing will pad a thin prompt.** Write the whole body.
87
+
88
+ ---
89
+
90
+ ## Shots and timing
91
+
92
+ The first shot carries **no timestamp**. Every later shot opens with the bracket and a cut time
93
+ that increases and stays inside the clip length:
94
+
95
+ ```
96
+ [Shot 2] At 00:03.500, the camera cuts to ...
97
+ ```
98
+
99
+ Transition verbs the model knows: **cuts to · transitions to · changes to · switches to**.
100
+
101
+ **Dialogue that continues across a cut** needs the continuity said out loud — *"his final words
102
+ carry over from the previous shot"* — or the line restarts. **Speech that ends abruptly** should be
103
+ described as cut off rather than trailed off.
104
+
105
+ ---
106
+
107
+ ## Camera — write the move into the sentence
108
+
109
+ The model has a named motion vocabulary:
110
+
111
+ > Zoom In / Zoom Out · Push In / Pull Out · Pan Left / Pan Right · Truck Left / Truck Right ·
112
+ > Tilt Up / Tilt Down · Pedestal Up / Pedestal Down · Arc Shot · Tracking Shot · Static Shot ·
113
+ > Shake Slightly / Shake Strongly · POV · Roll Clockwise / Roll Counterclockwise
114
+
115
+ Modify with **amplitude** (`with small amplitude` / `with large amplitude`) and **speed**
116
+ (`at slow speed` / `at fast speed`).
117
+
118
+ 🚨 **Integrate the motion into the sentence — never stack labels.** MiniMax's own example:
119
+ *"The camera pushes in with small amplitude at slow speed toward the folded letter in her hands."*
120
+ Not *"Push In. Small amplitude. Slow."*
121
+
122
+ ---
123
+
124
+ ## Speakers and dialogue
125
+
126
+ Give each speaking character a stable identity in the prose and keep it: describe the voice once
127
+ (*"a young woman with a quiet, breathy voice"*), then refer back to the same description at every
128
+ line. Identification, delivery and action sit **outside** the quoted line; the line itself is only
129
+ the words.
130
+
131
+ ```
132
+ The young woman with a quiet, breathy voice says: "I get off at the next station."
133
+ ```
134
+
135
+ **Voiceover** needs two things — the phrase *"says in an off-screen voiceover"* **and** an explicit
136
+ statement that the lips stay closed. Without the second half the model animates a mouth.
137
+
138
+ ```
139
+ The man says in an off-screen voiceover: "I still remember that road." — his lips remain
140
+ completely closed.
141
+ ```
142
+
143
+ **On-screen text** — signs, banners, labels, subtitles, neon — goes in double quotes with the
144
+ original wording preserved exactly: *A red neon sign reading "Open Late" glows above the doorway.*
145
+
146
+ ---
147
+
148
+ ## References — H3's real differentiator is the declared RELATIONSHIP
149
+
150
+ *(Base `minimax-h3` only. `minimax-h3-max` has no reference endpoint — Slates refuses references
151
+ on that row rather than dropping them silently.)*
152
+
153
+ <!-- @inject:references-read-literally -->
154
+ > **The general law: the model reads a reference literally.**
155
+ > A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
156
+
157
+ Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
158
+
159
+ - **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
160
+ - **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
161
+ - **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
162
+ - **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
163
+
164
+ **What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
165
+ <!-- @end:references-read-literally -->
166
+
167
+ <!-- @inject:reference-rules-core -->
168
+ Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
169
+
170
+ 1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
171
+ 2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates composes the naming for you from your `@mentions` / `#tags` — you never hand-write role labels.
172
+ 3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
173
+ 4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
174
+ 5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
175
+ 6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
176
+ 7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
177
+ 8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
178
+ 9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
179
+ 10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
180
+ <!-- @end:reference-rules-core -->
181
+
182
+ ### Cite references by number — Slates already does it for you
183
+
184
+ H3 on fal takes references as **typed slots** and expects the prompt to name them by modality and
185
+ order: **`image 1`, `image 2`, `video 1`, `audio 1`**. That is exactly what the Slates composer
186
+ emits from your `@mentions` and `#tags` (`Marcus (image 1) in the workshop (image 2)`), in the
187
+ exact order it sends them.
188
+
189
+ 🚨 **Do NOT hand-write angle-bracket reference tags.** MiniMax's own model-card grammar uses
190
+ `<Subject N>` / `<Picture N>` / `<Video N>` / `<Audio N>` labels; the fal endpoints Slates calls do
191
+ not — they build the binding from the typed slots and ask for plain numbered prose. Typing the tags
192
+ yourself puts literal angle brackets in the prompt the model reads.
193
+
194
+ ### State how much of each reference survives
195
+
196
+ This is the lever no other model in the catalogue gives you. Say, in plain words, what each
197
+ reference is FOR and how much of it should carry through:
198
+
199
+ | Intent | Say something like |
200
+ |---|---|
201
+ | Keep it whole | *"Keep the woman in image 1 exactly as she appears — hair, cardigan, necklace."* |
202
+ | Keep part of it | *"Use the café in image 2 for the brick wall and the sofa; the lighting is late evening, not daylight."* |
203
+ | **Move a trait onto someone else** | *"Give the man in image 3 the weathered leather texture of the jacket in image 4."* |
204
+ | Loose echo | *"Match the general palette and grain of image 5; nothing else from it."* |
205
+
206
+ The third row is the one with no equivalent anywhere else in Slates: **transferring a characteristic
207
+ onto a different subject** is a first-class thing H3 understands. Reach for H3 when that is the job.
208
+
209
+ **Audio references** bind a voice or a texture without copying the words. Say which speaker an
210
+ audio reference is for (*"the woman in image 1 speaks in the voice timbre of audio 1"*), and when
211
+ you are referencing only the timbre, **do not carry the reference clip's original dialogue into your
212
+ prompt** — write the new line. When you genuinely want the same words re-performed, quote them
213
+ exactly and say so.
214
+
215
+ **An audio reference cannot travel alone** — H3 refuses a reference set that is audio only. Pair it
216
+ with at least one image or video reference.
217
+
218
+ ### 💸 Reference images past the fifth cost 4 credits each
219
+
220
+ The first **5** reference images are free. Each additional image — the model takes **9** — adds
221
+ **4 credits** to the generation, at every resolution and every length. Four extra images on a 10s
222
+ 768p clip add 16 credits to a 30-credit generation: **more than half again**, for references that
223
+ often make the output worse rather than better (see the 2–4 rule above).
224
+
225
+ Attach the references the shot needs, not the ceiling. Call
226
+ `slates_estimate_generation_cost` with `referenceImages` set to the real count before a
227
+ reference-heavy job — a quote that omits it under-reports the bill.
228
+
229
+ ---
230
+
231
+ ## Frames
232
+
233
+ `minimax-h3` and `minimax-h3-max` both take a **start frame**, an **end frame**, or both. With an
234
+ end frame, land it explicitly: describe the final pose, spacing and composition as the thing the
235
+ shot **settles into** at the end, rather than hoping the model finds it.
236
+
237
+ > *"…she rotates the handle into the final angle and settles into the pose, spacing and composition
238
+ > of image 2 at the end of the shot."*
239
+
240
+ **Frames and references are mutually exclusive** on both rows — they are different endpoints, and
241
+ the reference endpoint has no frame slots at all. Slates refuses the combination rather than
242
+ dropping one side.
243
+
244
+ ---
245
+
246
+ ## Cost discipline
247
+
248
+ | Combination | Credits |
249
+ |---|---:|
250
+ | `minimax-h3` · 768p · 5s | 15 |
251
+ | `minimax-h3` · 768p · 10s | 30 |
252
+ | `minimax-h3` · 2K · 10s | 65 |
253
+ | `minimax-h3` · 4K · 10s | 80 |
254
+ | `minimax-h3-max` · 768p · 10s | 40 |
255
+ | every reference image past the fifth | **+4** |
256
+
257
+ **768p is the default for a reason.** It is the tier the model natively generates.
258
+
259
+ 🚨 **2K and 4K are UPSCALES of a 768p render, not larger generations.** fal's own schema says so:
260
+ *"480P and 768P are native generation modes; 2K and 4K upscale a 768P base result."* The upscaler
261
+ (H3-Regenerate-2K) is a separate stage bolted onto a finished take — it can enlarge detail but it
262
+ cannot add information.
263
+
264
+ **In our own test (2026-08-27, same prompt, same seed) the 2K pass came back with MORE artifacting
265
+ than the 768p original it was built from**, while costing 33 credits for a 5-second take against 15,
266
+ and taking nearly twice as long to return. One shot, so treat it as a warning rather than a law —
267
+ but the mechanism explains it, and the burden of proof is on 2K.
268
+
269
+ **So: generate at 768p and judge it at 768p.** Reach for 2K or 4K only when a delivery spec demands
270
+ the pixels, and expect to be paying for size rather than quality — a post-production upscale from a
271
+ clean 768p master is very often the better result. **4K video is Pro-only** (the server returns
272
+ `PRO_REQUIRED` for a base account); 2K is open to every tier.
273
+
274
+ ---
275
+
276
+ ## Quick checklist
277
+
278
+ - Body written as a timeline, first shot untimestamped, later shots on `[Shot N] At MM:SS.mmm`.
279
+ - Camera motion written **into** a sentence with amplitude and speed.
280
+ - Dialogue and diegetic music in the body; ambience in the soundscape section; audience-only score
281
+ in the score section, described by instrument and tempo.
282
+ - Voiceover carries both the off-screen phrase and the closed-lips statement.
283
+ - References cited as `image 1` / `video 1` / `audio 1`, each with a stated job and a stated degree
284
+ of retention. No angle-bracket tags.
285
+ - Reference count is deliberate — you are paying 4 credits for each one past the fifth.
286
+ - Frames **or** references, never both.
287
+ - The prompt is the prompt: no expander will fill it out for you.