@slatesvideo/shared 0.6.11 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/auth.js +2 -2
- package/dist/clients/cloud.js +1 -1
- package/dist/index.d.ts +1 -1
- package/dist/index.js +1 -1
- package/dist/manual/content.d.ts +1 -1
- package/dist/manual/content.js +1 -1
- package/dist/operations/index.d.ts +817 -16
- package/dist/operations/index.js +1410 -360
- package/dist/operations/surface.d.ts +3 -1
- package/dist/operations/surface.js +37 -10
- package/dist/prompts/ad-presets.d.ts +77 -0
- package/dist/prompts/ad-presets.js +43 -0
- package/dist/prompts/agent-doctrine.js +5 -4
- package/dist/prompts/banned-tokens.d.ts +4 -29
- package/dist/prompts/banned-tokens.js +29 -204
- package/dist/prompts/craft-cards.js +2 -2
- package/dist/prompts/generation-policy.d.ts +41 -0
- package/dist/prompts/generation-policy.js +53 -0
- package/dist/prompts/guide-retrieval.d.ts +9 -0
- package/dist/prompts/guide-retrieval.js +53 -0
- package/dist/prompts/index.d.ts +1 -0
- package/dist/prompts/index.js +1 -0
- package/dist/prompts/model-capabilities.d.ts +18 -1
- package/dist/prompts/model-capabilities.js +72 -19
- package/dist/prompts/model-facts.d.ts +34 -2
- package/dist/prompts/model-facts.js +66 -5
- package/dist/prompts/partials.generated.js +8 -2
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +61 -16
- package/dist/prompts/reference-composer.d.ts +2 -0
- package/dist/prompts/reference-composer.js +51 -50
- package/dist/prompts/script-document.d.ts +165 -0
- package/dist/prompts/script-document.js +11 -0
- package/dist/prompts/shot-grammar.d.ts +4 -4
- package/dist/prompts/shot-grammar.js +3 -3
- package/dist/prompts/shot-spec.d.ts +13 -0
- package/dist/prompts/shot-spec.js +23 -5
- package/dist/skills/content.js +26 -23
- package/exports/slates-chatgpt-images/generated/SKILL.md +107 -0
- package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
- package/exports/slates-prompt-builder/generated/SKILL.md +1 -1
- package/exports/slates-prompt-builder/generated/reference-character.md +9 -1
- package/exports/slates-prompt-builder/generated/reference-kling.md +3 -3
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +22 -10
- package/exports/slates-prompt-builder/generated/reference-seedance.md +4 -4
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +17 -17
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +10 -4
- package/skills/_partials/cinematic-card.md +8 -0
- package/skills/_partials/cinematic-routes-short.md +2 -0
- package/skills/_partials/cinematic-tips-short.md +2 -0
- package/skills/_partials/decision-log.md +1 -13
- package/skills/_partials/image-defaults.md +11 -0
- package/skills/_partials/lens-video-split.md +1 -0
- package/skills/_partials/reference-rules-core.md +1 -1
- package/skills/_partials/sheet-tool-defaults.md +6 -0
- package/skills/slates-character-identity.md +9 -1
- package/skills/slates-chatgpt-images.md +107 -0
- package/skills/slates-cinematic-look.md +237 -0
- package/skills/slates-cost-discipline.md +18 -12
- package/skills/slates-direct-response-ad.md +13 -53
- package/skills/slates-edit-and-iterate.md +1 -1
- package/skills/slates-model-selection.md +20 -14
- package/skills/slates-one-prompt-film.md +19 -77
- package/skills/slates-project-organization.md +7 -3
- package/skills/slates-prompting-flux-2-max.md +15 -4
- package/skills/slates-prompting-gpt-image-2-5.md +41 -28
- package/skills/slates-prompting-kling-v3.md +3 -3
- package/skills/slates-prompting-lip-sync.md +1 -1
- package/skills/slates-prompting-minimax-h3.md +30 -17
- package/skills/slates-prompting-motion-transfer.md +1 -1
- package/skills/slates-prompting-nano-banana-2.md +24 -11
- package/skills/slates-prompting-seedance-2-5.md +7 -6
- package/skills/slates-prompting-seedance.md +5 -5
- package/skills/slates-prompting-seedream-5-lite.md +14 -3
- package/skills/slates-prompting-veo-3.md +1 -1
- package/skills/slates-script-craft.md +45 -0
- package/skills/slates-shot-variety.md +11 -40
- package/skills/slates-storyboard-from-script.md +14 -66
- package/skills/slates-style-prompting.md +54 -54
- package/skills/slates-ugc-influencer-ad.md +32 -309
- package/skills/slates-vision-feedback-loop.md +2 -1
|
@@ -1,96 +1,38 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-one-prompt-film
|
|
3
|
-
description:
|
|
3
|
+
description: Deliver a finished video when the user explicitly asks for one, coordinating editable writing, selected media production, a named cut and a verified export. Preserve existing work and follow generation authorization.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
#
|
|
6
|
+
# Idea to finished video
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Carry the requested piece through to an exported file. Use existing work whenever it serves the brief. The creator can enter at any point: writing, importing footage, comparing takes or changing an edit. No mandatory stage sequence, shot count or number of approval checkpoints follows from this guide.
|
|
9
9
|
|
|
10
|
-
##
|
|
10
|
+
## Make the intended piece visible
|
|
11
11
|
|
|
12
|
-
|
|
13
|
-
Turn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.
|
|
12
|
+
Use the current project and document unless another destination is requested. Save words through the revision-checked document tools; `slates-script-craft` covers writing and versions. Add production bindings only where needed. A shot needs no image, and a cut can use imported footage with no script.
|
|
14
13
|
|
|
15
|
-
|
|
14
|
+
Preserve fixed passages, explicit creative choices and custom prompt bytes. Record production choices in editable shots. Explain only consequential judgments not already visible there. Recurring cast, repeated framing, silence and dependent scenes are valid when they serve the piece.
|
|
16
15
|
|
|
17
|
-
|
|
16
|
+
## Inspect the actual requests and estimate
|
|
18
17
|
|
|
19
|
-
|
|
20
|
-
When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify **and that no row already records**:
|
|
21
|
-
|
|
22
|
-
```
|
|
23
|
-
source phrase or declared default → what you wrote → what it resolves
|
|
24
|
-
"in a diner" → warm, and the light is the reason → why the anchor was chosen, not what it is
|
|
25
|
-
(no time of day) → late afternoon, low warm key → default; say the word and it changes
|
|
26
|
-
```
|
|
27
|
-
|
|
28
|
-
🚨 **Keep it to what is NOT already data — and almost everything now IS.** A Shot holds the references and their roles, the model, every param, the shot size, the camera, the prop, the action and the spoken line, and `slates_list_shots` reads the whole board back in order with its variety counts. Narrating any of those is retelling a row the user can open. **Write the Shot, and let the log carry only the judgement no field holds** — why this world, why this light, why this register.
|
|
29
|
-
|
|
30
|
-
**Hard rule: never silently add weather, props, style, or camera movement.** Four of those are now FIELDS: put the value on the Shot (`prop`, `camera`, `shotSize`, `action`) so the user can read and change it, and put the *reason* in the log only when you invented it rather than being told it. The rule has not softened — it moved from narration into data, which is stronger, because a field can be corrected and a sentence in chat cannot.
|
|
31
|
-
|
|
32
|
-
> ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
|
|
33
|
-
<!-- @end:decision-log -->
|
|
34
|
-
|
|
35
|
-
A 4-10 shot script is where you invent the most on the user's behalf — time of day, wardrobe, weather, lens feel, camera moves the brief never mentioned. The log is what makes those visible while they are still free to change.
|
|
36
|
-
|
|
37
|
-
### 2. Set up the project
|
|
38
|
-
- `slates_create_project` named for the piece.
|
|
39
|
-
- Recurring character? Build it properly — `slates_create_character` + the `slates-character-identity` recipe — so every frame references the same identity.
|
|
40
|
-
- Recurring location? `slates_create_environment`.
|
|
41
|
-
- One-off shots don't need character/environment records; skip the ceremony.
|
|
18
|
+
Use `slates_get_shot` to inspect the composed prompt, settings and references. Model choices and supported settings come from `slates-model-selection`, the current capability surface and the selected model's guide. Do not carry limits or prices from an old example.
|
|
42
19
|
|
|
43
|
-
|
|
44
|
-
- `slates_create_storyboard`, `slates_add_scene` per script scene.
|
|
45
|
-
- `slates_create_shot` per beat — the prompt, the model, the params and the references, with the roles they carry. **A Shot needs no image**, so the entire film exists as rows before anything is paid for.
|
|
46
|
-
- `slates_get_shot` reads one back COMPOSED: the prompt the model will actually receive, its numbered references, and its exact quote. Audit your own work there — you cannot approve something the request will not contain.
|
|
47
|
-
- Structure first, spend second — the user catches script problems on the free skeleton, not on burned credits.
|
|
20
|
+
Follow `slates-cost-discipline` and the user's generation policy. Quote the exact requested set with `slates_generate_from_shots` before confirming it. Existing authorization covers its enumerated requests, not extra takes or changed inputs. Editing, choosing versions, importing and building a cut do not spend generation credits.
|
|
48
21
|
|
|
49
|
-
|
|
50
|
-
The Shots ARE the quote. `slates_generate_from_shots` without `confirm` returns one itemised total for the set plus the largest single item — no hand arithmetic, no `slates_estimate_generation_cost` per call:
|
|
22
|
+
Keep reusable historical media separate from new requests. Matching words alone do not prove matching voice, references or settings. An explicitly requested extra take is never deduplicated away.
|
|
51
23
|
|
|
52
|
-
|
|
24
|
+
## Generate only the authorized material
|
|
53
25
|
|
|
54
|
-
|
|
26
|
+
Submit the chosen requests and inspect each returned state. On an uncertain timeout, read the shot's generation IDs and job status before any retry. Diagnose a failure and follow the existing consent policy for added requests. Never discard other takes merely because a new one was selected.
|
|
55
27
|
|
|
56
|
-
|
|
57
|
-
Fire the image Shots with `slates_generate_from_shots` (`confirm: true` — step 4 authorized it). Slates names each reference inline as "image N"; you never hand-write a role label or a number. Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`, then `slates_update_shot` with `attachFrameId` so the recipe travels with the picture.
|
|
28
|
+
Inspect image composition and reference fidelity. Inspect video performance, motion and sound across playback; frame samples alone cannot establish speech or motion quality. If an edit can resolve dead air or order, use the existing take rather than assuming another generation is needed.
|
|
58
29
|
|
|
59
|
-
|
|
30
|
+
## Arrange and deliver
|
|
60
31
|
|
|
61
|
-
|
|
62
|
-
Fork each bound frame's image Shot with `slates_duplicate_shot` (`model:` the video model — that is the A/B lever the op takes inline), then `slates_update_shot` the copy with `firstFrameAssetId` = the bound frame. Two calls, because `slates_duplicate_shot` forks the prompt, the model and the params; **attachments are changed with `slates_update_shot`.** Then fire the set with `slates_generate_from_shots`.
|
|
32
|
+
Read the available timelines. Name the destination cut explicitly for a variation; independent comparisons use independent cuts. Add selected media in the intended order, preserve trim/level/transform choices and inspect the actual timeline.
|
|
63
33
|
|
|
64
|
-
|
|
34
|
+
Use supported video/XML exports and their stated fidelity limits. For selected named cuts, `slates_export_cuts` records distinct outputs and a manifest; retry unfinished outputs with the same manifest identity. Verify the returned files and playback before reporting success. Describe the completed piece, actual spend where available and output paths. Do not label a render complete merely because its submission succeeded.
|
|
65
35
|
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
- **Kling V3** (`slates-prompting-kling-v3`): the cost-effective seat — 16:9 / 9:16 / 1:1, 3-15s, strong start-frame adherence; std is the workhorse, Omni for multi-character dialogue.
|
|
70
|
-
- **MiniMax H3** (`slates-prompting-minimax-h3`): route here when a shot's SOUND is part of the writing — a line delivered a particular way, scene sound under it, score that must stay outside the characters' world. It authors all three in one pass, which **collapses a shot's audio pass into its video pass** and removes the separate `slates_generate_audio` step for that shot. 5-15s, 480p/768p/2K/4K. Its sibling `minimax-h3-max` is faster, tops out at 768p, takes the same references, and costs MORE at 768p — a deliberate speed pick, never a saving.
|
|
71
|
-
- **Veo 3.1** (`slates-prompting-veo-3`): niche, never the default — only when native synced audio must generate WITH the video in one gen; 16:9 or 9:16, 4/6/8s (8s only at 1080p/4K or with reference images).
|
|
72
|
-
|
|
73
|
-
Failed gen? The run continues past it and **nothing is retried automatically**. Read the per-Shot error in the result, fix that Shot with `slates_update_shot`, and re-fire only it (a retry beyond the plan = announce the delta cost).
|
|
74
|
-
|
|
75
|
-
### 7. Assemble the timeline
|
|
76
|
-
- `slates_get_timeline` once to get the lay of the land.
|
|
77
|
-
- `slates_add_clip_to_timeline` for each completed video asset **in story order** — defaults append back-to-back on the first video track, which is exactly an assembly cut.
|
|
78
|
-
- Order wrong? `slates_reorder_clips` with the full clip-id list. Dropped a shot? `slates_remove_clip`, then reorder to close the gap.
|
|
79
|
-
|
|
80
|
-
### 8. Export + deliver
|
|
81
|
-
- Output path: ask the user, or default to `<slates_get_project_directory>/exports/<name>.mp4`.
|
|
82
|
-
- `slates_export_video` (absolute path, `.mp4`; blocks while ffmpeg renders — minutes for long timelines).
|
|
83
|
-
- `slates_reveal_file` so the file is literally in front of them.
|
|
84
|
-
- Offer the finishing path: `slates_export_timeline_xml` → DaVinci Resolve (File → Import → Timeline) for grading, sound, and titles.
|
|
85
|
-
|
|
86
|
-
### 9. Report
|
|
87
|
-
Shots delivered, total spent vs. approved plan, the export path, and the single best next lever ("re-take shot 3 with a tighter prompt" / "add a CTA end-card").
|
|
88
|
-
|
|
89
|
-
## Hard rules
|
|
90
|
-
|
|
91
|
-
- **Two checkpoints only.** Creative plan (step 1) and total cost (step 4). Everything else runs without asking — that's the product promise.
|
|
92
|
-
- **Skeleton before spend.** Project + storyboard structure are free; generation isn't.
|
|
93
|
-
- **Look at everything.** Every image inline, every video via `slates_get_asset_video_frames` if a clip seems off. Never assemble a timeline from clips you haven't evaluated.
|
|
94
|
-
- **3-strike rule per shot.** Three failed takes on one shot = stop, show the user what you tried, ask.
|
|
95
|
-
- **Consistency comes from references, not luck.** Same identity asset on every character frame; same environment refs across a location's shots.
|
|
96
|
-
- **Plan in Shots, not in chat.** Every decision that ends up in a sentence you have to remember is a decision the user cannot see, price, fork or re-fire. A Shot is a row: it survives the conversation, and the user can open it in the app and fix one reference without you.
|
|
36
|
+
<!-- @inject:decision-log -->
|
|
37
|
+
Record production choices in the editable shot fields. Explain only consequential judgments the user did not specify and no field already records: for example, why a particular light or performance register supports the brief. Do not repeat the shot list in prose or turn this explanation into an approval gate. Follow the separate generation authorization policy before spending.
|
|
38
|
+
<!-- @end:decision-log -->
|
|
@@ -5,21 +5,25 @@ description: Use when the user names an asset by code ("use IMG-A36"), asks what
|
|
|
5
5
|
|
|
6
6
|
# Organizing a Slates project
|
|
7
7
|
|
|
8
|
-
Slates already gives
|
|
8
|
+
Slates already gives every REUSABLE reference a home — the **Library**, in categories the user names (Characters, Locations, Products, Looks…; `slates_list_library`), each item used in a prompt as `@name`, or `#name` for a look. Do NOT recreate those as folders, and never invent a category the user did not ask for. Folders are for **structure**, never type.
|
|
9
9
|
|
|
10
10
|
**Folders = where an asset sits in the FILM**, and they mirror to real subfolders on disk (`projects/<id>/…`), so a human can open the project in Resolve/Finder and navigate it like an edit. Use them for work product, not references.
|
|
11
11
|
|
|
12
12
|
Create with `slates_create_folder`; file assets with `slates_move_assets_to_folder`. Generations land in the project's active folder, so set it before a batch.
|
|
13
13
|
|
|
14
|
+
A favorite is a keeper, not a folder: `slates_set_asset_favorite` marks one asset (the heart on its card, `isFavorite` in `slates_list_assets`) without moving it. To hand files out of Slates, `slates_export_assets` copies the ORIGINALS of the assets you name into a directory (named by code, never overwriting) — the way to deliver an image-only job that never needed a shot or a timeline. Use it to flag the takes worth a second look; file with folders once the pick is made.
|
|
15
|
+
|
|
16
|
+
To reuse a whole piece rather than one reference, use a TEMPLATE: `slates_export_template` saves a board, a scene or one Shot (recipes, script words, references and the Library items they mention; never takes) as a file, `slates_get_template` reads what a file holds and its swap slots, and `slates_import_template` adds it to a project, optionally swapping a slot for one of that project's own assets. An import generates nothing; quote and fire the returned Shots as usual.
|
|
17
|
+
|
|
14
18
|
Conventions by project type:
|
|
15
19
|
- **Short film / narrative:** `Shots` (scene stills) · `Clips` (generated video) · `Final` (the export). Use one folder per scene (`Scene 1`, `Scene 2`, …) instead when the piece has distinct locations/beats.
|
|
16
20
|
- **Ad / UGC:** `Hooks` · `B-roll` · `Talking-head` · `Final`.
|
|
17
21
|
|
|
18
22
|
Rules of thumb:
|
|
19
|
-
- Reusable cast / sets / look → leave in the
|
|
23
|
+
- Reusable cast / sets / products / look → leave in the Library. Don't fold them.
|
|
20
24
|
- Scene stills, clips, and the final cut → file into the structural folder they belong to, as you make them.
|
|
21
25
|
- One folder per asset (folders are structure). Cross-cutting status (hero take, reject, variant) is a tag concern, not a folder.
|
|
22
|
-
- Keep the gallery legible: work product lives in folders; the reference scaffolding (sheets, plates, style images) stays in
|
|
26
|
+
- Keep the gallery legible: work product lives in folders; the reference scaffolding (sheets, plates, style images) stays in the Library.
|
|
23
27
|
|
|
24
28
|
## Asset codes — the shared vocabulary (IMG-A12 / VID-V3 / AUD-S1)
|
|
25
29
|
|
|
@@ -24,9 +24,16 @@ description: How to prompt FLUX.2 Max (Black Forest Labs image model). Read befo
|
|
|
24
24
|
4. **Bind every hex colour to an object.** `a #1B4D3E enamel mug` lands; an unbound colour does not.
|
|
25
25
|
5. **For portraits add texture words** — `natural skin texture, realistic pores, subtle imperfections, soft diffused lighting`.
|
|
26
26
|
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
-
|
|
27
|
+
<!-- @inject:cinematic-card -->
|
|
28
|
+
**For a photographic look, use only what this frame needs.** Image models default to clean, evenly lit and fully exposed. Describe what the camera sees, not just gear or mood:
|
|
29
|
+
- **Inspect every reference first.** Write its grade and imperfections in words: darkness, contrast, muddy or true blacks, colour, softness/noise, subject separation. Never grade cleaner or brighter than the look reference unless asked.
|
|
30
|
+
- **One light system** — `low sun behind her`, `her face falls into deep shadow`, `no light in front of her`.
|
|
31
|
+
- **Visible exposure** — `the sky burns out to white`, `dense, slightly crushed shadows`.
|
|
32
|
+
- **Lens name plus effect** — `200mm telephoto`, `peaks loom huge behind her and melt into soft shapes`.
|
|
33
|
+
- **Name every garment and close the foreground.** Omissions invite reference leakage or invented props.
|
|
34
|
+
Bind references inline. A scene reference owns the grade; for a look-only reference, write the new scene's light. References are optional. For owned-frame edits, describe only the change and what stays.
|
|
35
|
+
<!-- slates-only -->Use `slates-cinematic-look` with a technique ID or section query for more.<!-- /slates-only -->
|
|
36
|
+
<!-- @end:cinematic-card -->
|
|
30
37
|
|
|
31
38
|
**Hard constraint:** no negative prompting. Every "no X" must be rewritten as the positive state — `no blur` becomes `sharp focus throughout`, `no people` becomes `empty scene`, `no harsh shadows` becomes `soft, diffused lighting`.
|
|
32
39
|
<!-- @card:end -->
|
|
@@ -44,6 +51,10 @@ description: How to prompt FLUX.2 Max (Black Forest Labs image model). Read befo
|
|
|
44
51
|
- `masterpiece`, `best quality`, `trending on artstation`, `8k`
|
|
45
52
|
<!-- @banned:end -->
|
|
46
53
|
|
|
54
|
+
**Examples**
|
|
55
|
+
- `A chef plating in a steel kitchen pass. Shot on Hasselblad X2D, 80mm, f/2.8. Overhead fluorescents plus warm spill from the line. Natural skin texture, subtle imperfections. Muted steel and #7A3B2E copper.`
|
|
56
|
+
- `An empty municipal pool at dusk, 35mm, deep focus, early digital camera with slight noise and flash falloff. Cracked #4A7C8C tiles. Candid, unstaged.`
|
|
57
|
+
|
|
47
58
|
Black Forest Labs' top image model, routed via fal.ai. In Slates: `slates_generate_image` with `model: flux-2-max` (REQUIRES projectId — no headless path), priced per resolution (1k/2k/4k — call `slates_estimate_generation_cost` for current numbers, never quote from memory). Strengths vs Nano Banana 2: photoreal texture, less censored, precise hex-color control, strong typography. Reference images route through FLUX's edit endpoint and carry a lower per-model cap than NB2's 14.
|
|
48
59
|
|
|
49
60
|
## Core structure — front-load what matters
|
|
@@ -141,7 +152,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
|
|
|
141
152
|
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
142
153
|
|
|
143
154
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
144
|
-
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates
|
|
155
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
|
|
145
156
|
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
146
157
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
147
158
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-gpt-image-2-5
|
|
3
|
-
description:
|
|
3
|
+
description: Prompt and edit images with GPT Image 2.5 Flare or Sunburst. Covers reference roles, realistic lighting, text, grids, quality choices and targeted edits. Use with slates_generate_image or slates_edit_image on these models.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# GPT Image 2.5 — sheets, grids, and text that actually reads
|
|
@@ -15,23 +15,28 @@ description: Prompting GPT Image 2.5 (Flare and Sunburst) — the readable-text
|
|
|
15
15
|
Keep it under 2,400 characters (the build fails above that) and keep the
|
|
16
16
|
rationale, the receipts and the worked examples in the body below. -->
|
|
17
17
|
<!-- /slates-only -->
|
|
18
|
-
**Card — GPT Image 2.5.** The
|
|
19
|
-
|
|
20
|
-
**Pick the
|
|
21
|
-
|
|
22
|
-
**The
|
|
23
|
-
1. **
|
|
24
|
-
2. **
|
|
25
|
-
3. **
|
|
26
|
-
4. **
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
**
|
|
31
|
-
-
|
|
32
|
-
-
|
|
33
|
-
|
|
34
|
-
|
|
18
|
+
**Card — GPT Image 2.5.** The photoreal front-runner for people, and the readable-text, ordered-panel engine. Structure: subject and action with each reference named where it is used, then any exact copy in quotes, then layout, then light.
|
|
19
|
+
|
|
20
|
+
**Pick the tier.** `flare` is Faster, quality comparable to GPT Image 2: drafts and volume. `sunburst` is Better quality, the most capable: finals, hero frames, photoreal people, multi-reference edits. Use the product default; choose Flare when speed is a stated priority.
|
|
21
|
+
|
|
22
|
+
**The levers**
|
|
23
|
+
1. **Name each reference inline** — `the woman from image 1`, `lit and graded like image 2`. Never an opening paragraph about what the references are.
|
|
24
|
+
2. **Quote every string that must render verbatim** — `the jacket reads "SLATES"`. Describe a font's feel, never its name; keep on-image text under about 30 words.
|
|
25
|
+
3. **Name the layout as a grid** for sheets and panels — `a 3x2 grid of panels, reading left to right, equal gutters`.
|
|
26
|
+
4. **Set `quality` deliberately.** `high` is the everyday tier; `max` is 4× its price, `xhigh` about 1.8×. Coming from GPT Image 2 the names moved one rung: its `medium` is this `high`.
|
|
27
|
+
|
|
28
|
+
<!-- @inject:cinematic-card -->
|
|
29
|
+
**For a photographic look, use only what this frame needs.** Image models default to clean, evenly lit and fully exposed. Describe what the camera sees, not just gear or mood:
|
|
30
|
+
- **Inspect every reference first.** Write its grade and imperfections in words: darkness, contrast, muddy or true blacks, colour, softness/noise, subject separation. Never grade cleaner or brighter than the look reference unless asked.
|
|
31
|
+
- **One light system** — `low sun behind her`, `her face falls into deep shadow`, `no light in front of her`.
|
|
32
|
+
- **Visible exposure** — `the sky burns out to white`, `dense, slightly crushed shadows`.
|
|
33
|
+
- **Lens name plus effect** — `200mm telephoto`, `peaks loom huge behind her and melt into soft shapes`.
|
|
34
|
+
- **Name every garment and close the foreground.** Omissions invite reference leakage or invented props.
|
|
35
|
+
Bind references inline. A scene reference owns the grade; for a look-only reference, write the new scene's light. References are optional. For owned-frame edits, describe only the change and what stays.
|
|
36
|
+
<!-- slates-only -->Use `slates-cinematic-look` with a technique ID or section query for more.<!-- /slates-only -->
|
|
37
|
+
<!-- @end:cinematic-card -->
|
|
38
|
+
|
|
39
|
+
**Hard constraint:** its own content filter, distinct from Gemini's. Never describe a reference as a photograph of a real person.
|
|
35
40
|
<!-- @card:end -->
|
|
36
41
|
|
|
37
42
|
<!-- @banned:start -->
|
|
@@ -42,8 +47,8 @@ description: Prompting GPT Image 2.5 (Flare and Sunburst) — the readable-text
|
|
|
42
47
|
backticked and prose outside the backticks. -->
|
|
43
48
|
<!-- /slates-only -->
|
|
44
49
|
**Never use:**
|
|
45
|
-
- a font NAME — describe the feel
|
|
46
|
-
- a reference
|
|
50
|
+
- a font NAME — describe the feel instead, as in: clean geometric sans, high contrast
|
|
51
|
+
- a reference described as a photograph of a real person (`is a photograph of a woman`), or any up-front essay about what each reference is for — name the subject inline where it is used instead, as in: the woman from image 1
|
|
47
52
|
- `8k`, `masterpiece`, `best quality`, `highly detailed` — quality incantations do nothing here either
|
|
48
53
|
<!-- @banned:end -->
|
|
49
54
|
|
|
@@ -55,7 +60,7 @@ GPT Image's edge is **character-level text accuracy** (~99% on English), ordered
|
|
|
55
60
|
|
|
56
61
|
🚨 **FLARE IS NOT AN UPGRADE OVER GPT IMAGE 2 — IT IS THE FAST ONE.** OpenAI, verbatim: *"GPT Image 2.5 Flare is the small model, optimized for speed, with image quality **comparable to** GPT Image 2. GPT Image 2.5 Sunburst is the base model, optimized for quality, with **higher image quality than** GPT Image 2."* Their model pages agree: Flare is *"our fastest model for high-quality, everyday image generation"*, Sunburst *"our most capable model for image generation and editing."* **Sunburst is the seat that beats what we had; Flare is the one that holds it at half the latency.** An earlier revision of this file called Flare "better than GPT Image 2" and sent Sunburst only to multi-reference edits — both wrong, corrected 2026-09-09 against the vendor docs.
|
|
57
62
|
|
|
58
|
-
**
|
|
63
|
+
**Choose for the task.** Use the product default for ordinary work. Flare is an option when speed matters; changing model is not a mandatory draft stage.
|
|
59
64
|
|
|
60
65
|
**Sunburst's widest lead is multi-reference editing** — several references all surviving into one frame, the character-consistency-across-shots problem. Reach for it there first, but that is not the only place it belongs.
|
|
61
66
|
|
|
@@ -63,7 +68,7 @@ GPT Image's edge is **character-level text accuracy** (~99% on English), ordered
|
|
|
63
68
|
|
|
64
69
|
🚨 **The GPT Image line is ALSO the photoreal front-runner, and this file said the opposite until 2026-08-24.** **Receipts:** Eric's direct call, plus a head-to-head on the Higgsfield rail where GPT Image 2 at `quality: high`, 2K beat both Nano Banana rails on skin realism for photoreal people — that result is why the whole AI-influencer ad lane generates its plates here. **Route photoreal to this line, not away from it.**
|
|
65
70
|
|
|
66
|
-
|
|
71
|
+
**Historical receipt, not a tier recommendation:** the photoreal comparison above used GPT Image 2 at its old `high` tier. It has not been repeated on 2.5 under matched conditions. Start with the product default and test a higher tier only against an unmet requirement; the old comparison does not establish a minimum tier for this model.
|
|
67
72
|
|
|
68
73
|
**What the Banana line still owns:** edit-heavy work, and holding many subjects coherently in one frame. **Not the reference ceiling any more** — that line was true until 2026-09-09, when GPT Image went to its documented 16 against Banana's 14. Route on which model keeps them all recognisable, not on the count.
|
|
69
74
|
|
|
@@ -76,18 +81,18 @@ All five rungs are exposed, and they span ~36× end to end (2k class: $0.0044
|
|
|
76
81
|
| Tier | Use it for |
|
|
77
82
|
|---|---|
|
|
78
83
|
| `low` | Roughest pass — layout and composition checks, throwaway comps. |
|
|
79
|
-
| `medium` | The draft
|
|
80
|
-
| `high` |
|
|
84
|
+
| `medium` | The draft tier. Cheaper than NB2 Lite and available up to 4K, which is why the draft lane moved here. |
|
|
85
|
+
| `high` | General-purpose quality tier. Blind benchmarks on GPT Image 2 put this rung — which it called `medium` — within a hair of `max` (which it called `high`) at a quarter of the cost. Inherited from the old ladder, never re-run on 2.5, and it says nothing about `xhigh`. |
|
|
81
86
|
| `xhigh` | One rung short of the top at about half its price (2k: 4 cr against `max`'s 8). Worth trying before `max`. |
|
|
82
87
|
| `max` | Top of the ladder. Tiny type, dense diagrams, many labelled elements. |
|
|
83
88
|
|
|
84
89
|
⚠️ **A tier label means different things on different models.** OpenAI: *"The same quality label does not imply the same image quality or response time across models."* Flare at `max` and Sunburst at `max` are not the same picture, and neither matches Nano Banana's idea of "high".
|
|
85
90
|
|
|
86
|
-
🚨 **The tier NAMES moved between versions and the strings did not.** GPT Image 2's `medium` is this model's `high`; its `high` is this model's `max` — same money, one rung of renaming.
|
|
91
|
+
🚨 **The tier NAMES moved between versions and the strings did not.** GPT Image 2's `medium` is this model's `high`; its `high` is this model's `max` — same money, one rung of renaming. For a recipe explicitly written for GPT Image 2, map the old tier before reusing it on 2.5. A current user request for `medium` still means `medium`. Getting this backwards costs picture quality silently: nothing errors, the bill is correct for what was asked, and the image is just worse.
|
|
87
92
|
|
|
88
|
-
Never rely on the provider default. fal's default is `high`, which is correct today — but it is the third rung of five rather than the top of two, so leaning on it means a fal-side change silently reprices you.
|
|
93
|
+
Never rely on the provider default. fal's default is `high`, which is correct today — but it is the third rung of five rather than the top of two, so leaning on it means a fal-side change silently reprices you. Slates sends its configured quality explicitly; current defaults live in `slates-model-selection`. Your explicit choice overrides them.
|
|
89
94
|
|
|
90
|
-
**
|
|
95
|
+
**Start at the default and change tiers for an unmet requirement.** OpenAI's own procedure: *"If the output falls short, test a higher quality setting. Once it meets your requirements, test lower settings to see whether they preserve acceptable quality while reducing latency. Use `xhigh` or `max` only when they improve an unmet quality requirement within your latency budget."* A higher rung does **not** guarantee a better result on a given prompt. Compare `medium` against `high` when the job is small or dense text; that is where the rungs separate most visibly.
|
|
91
96
|
|
|
92
97
|
## Resolution classes
|
|
93
98
|
|
|
@@ -101,10 +106,14 @@ Never rely on the provider default. fal's default is `high`, which is correct to
|
|
|
101
106
|
|
|
102
107
|
🚨 **THE ASPECT RATIO CHANGES THE PRICE ON THIS MODEL, and on no other image model.** OpenAI bills image OUTPUT TOKENS and the count tracks the frame's SHAPE, so at the same resolution class **`1:1` costs about 1.8× and `4:3`/`3:4` about 1.37× what `16:9` costs**; `9:16` costs the same as `16:9`. Metered 2026-09-09 and priced into the cost key, so the quote you get before generating is the real number — but if you are choosing between shapes and the budget is tight, **16:9 or 9:16 is the cheap one.** Every other image model charges the same whatever the shape.
|
|
103
108
|
|
|
104
|
-
## Reference images — give every one a role
|
|
109
|
+
## Reference images — give every one a role, inline, where it is used
|
|
105
110
|
|
|
106
111
|
**Assign a role to every reference image: subject, style, clothing, or background.** This is new emphasis in 2.5 and the highest-leverage change for the 16-reference character lane. An unroled pile of references makes the model guess what each one is for, and it guesses differently every run — which is the drift people mistake for a consistency failure.
|
|
107
112
|
|
|
113
|
+
**The role rides a clause in the scene, not a paragraph in front of it.** *The woman from image 1 cooks on a rocky summit…*, *lit and graded like image 2*. Never open with sentences about what each reference is and what to take or ignore from it: that is the role essay the shared reference rules below forbid, and it drags the sheet's studio light into the scene.
|
|
114
|
+
|
|
115
|
+
**Receipt, 2026-09-15, Sunburst, IMG-A192–A198.** The up-front version returned the studio look; the inline versions were never refused and never came back as a sheet. Two costs, both fixed in words: anything the prompt does not describe is taken from the reference (name every garment), and props nobody asked for appear (say what is in the foreground and that nothing else is). One sheet-only plate kept its described location, which narrows the two-reference rule in `slates-ugc-influencer-ad`. A look reference did far less than a described light. The full ladder is the vault's `cinematic-look-research.md`; the techniques are `slates-cinematic-look`.
|
|
116
|
+
|
|
108
117
|
Reference images route through the edit endpoint, **up to 16** — fal's documented `maxItems`, and the highest reference ceiling of any image seat in Slates (the Banana line takes 14). It was capped at 10 until 2026-09-09, which was never anybody's limit, just a number nobody had checked. The composed "image N" naming applies as everywhere else. Mask-based inpainting exists at the API level but is not surfaced: a mask is something the user has to paint, and there is no painting surface — describe the change instead.
|
|
109
118
|
|
|
110
119
|
## Editing — separate the change from the constraints
|
|
@@ -142,6 +151,8 @@ For anything with several requirements, OpenAI recommends organising the prompt
|
|
|
142
151
|
|
|
143
152
|
Name materials, lighting, colour and medium. Mood words are cues only — "cinematic", "moody", "epic" tell the model almost nothing on their own. Give scale, atmosphere and colour instead. Camera specs (`85mm`, `f/1.4`) are appearance hints, not a physical simulation; they bias the look, they do not compute optics.
|
|
144
153
|
|
|
154
|
+
**Name the lens and describe its effect, every time.** A lens named alone changed nothing visible (IMG-A195, 2026-09-15); named together with what it does to the picture, it produced real compression and depth of field (IMG-A198). Wording: `slates-cinematic-look` → `compression-as-outcome`, `defocus-as-outcome`.
|
|
155
|
+
|
|
145
156
|
**For people, state body framing and scale**: "full body visible, feet included", "hands naturally gripping the handlebars". This is also the safest way to phrase a crop — see the blocked-phrasings section below.
|
|
146
157
|
|
|
147
158
|
**No special syntax is required.** Prose, JSON and tagged blocks all work equally well, so pick whatever stays maintainable in the caller.
|
|
@@ -163,6 +174,8 @@ Name materials, lighting, colour and medium. Mood words are cues only — "cinem
|
|
|
163
174
|
|
|
164
175
|
The first reads to the filter as *recreate this real person's likeness*, which is a hard refusal regardless of what the rest of the prompt says. The second signals a fictional character and passes. **This is a wording change only — the reference image can be the same file either way.** One plate flipped from refused to accepted on this single sentence with nothing else altered.
|
|
165
176
|
|
|
177
|
+
**Inline naming sidesteps the question and is now the default:** never describe the reference at all, and name her where she is used (*the woman from image 1*). Six of six Sunburst plates written that way passed on 2026-09-15. Keep the sheet sentence above as the fallback if a refusal appears.
|
|
178
|
+
|
|
166
179
|
**2. Never attach a reference sheet containing a headless body panel.** A sheet whose full-body panels are cropped above the neck is refused every time, even with the correct opener. Regenerate the sheet with the head visible in every panel. Related, and already in this file's sheet guidance: phrase a cropped panel as *framing* (`cropped at the collarbone`), never as *absence* (`the head not shown`).
|
|
167
180
|
|
|
168
181
|
⚠️ **These refusals were measured on GPT Image 2, not on 2.5.** The classifier belongs to OpenAI rather than to a model version, so the phrasing rules carry — but they are inherited, not re-measured. If Flare or Sunburst accepts one of the blocked phrasings, that is a new receipt to write down here, not a reason to delete this one.
|
|
@@ -163,7 +163,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
|
|
|
163
163
|
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
164
164
|
|
|
165
165
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
166
|
-
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates
|
|
166
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
|
|
167
167
|
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
168
168
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
169
169
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
@@ -185,11 +185,11 @@ Kling exposes `negative_prompt` on the fal endpoint (different from Seedance whi
|
|
|
185
185
|
|
|
186
186
|
```
|
|
187
187
|
blurry, low quality, watermark, text overlay, distorted hands, extra fingers,
|
|
188
|
-
duplicate limbs, unnatural skin texture, overly saturated colors,
|
|
188
|
+
duplicate limbs, unnatural skin texture, overly saturated colors,
|
|
189
189
|
floating objects, inconsistent shadows, jittery, flickering, morphing face
|
|
190
190
|
```
|
|
191
191
|
|
|
192
|
-
Layer scene-specific suppressions on top.
|
|
192
|
+
Layer scene-specific suppressions on top, and never suppress something the prompt asks for. This block carried `lens flare` until 2026-09-15, which silently cancelled every flare a prompt described (`slates-cinematic-look` → `source-flare`); add it back only for a shot that must have none.
|
|
193
193
|
|
|
194
194
|
## Cinematic tactics
|
|
195
195
|
|
|
@@ -60,7 +60,7 @@ Seedance can generate the performance rather than bolting a mouth onto finished
|
|
|
60
60
|
That is the same endpoint the old `engine=seedance-2` branch called — it just built the sentence for you, invisibly, and it presupposed a "video 1" that might not exist. Writing the prompt is the whole difference, and it is the part you want control of.
|
|
61
61
|
|
|
62
62
|
- Driving clips must be 2–15s; output duration is whatever you set (4–15s).
|
|
63
|
-
- Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming.
|
|
63
|
+
- Video references bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
|
|
64
64
|
- Faces go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person.
|
|
65
65
|
|
|
66
66
|
Everything below is about the Kling tool.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: slates-prompting-minimax-h3
|
|
3
|
-
description: How to prompt MiniMax H3 and
|
|
3
|
+
description: How to prompt MiniMax H3, H3 Max and H3 Max Turbo. Read before calling slates_generate_video with model minimax-h3, minimax-h3-max or minimax-h3-max-turbo. H3 is the only Slates video seat where AUDIO IS AUTHORED rather than toggled — synchronised dialogue, scene sound and an audience-only score are three separate sections of the prompt, generated in one pass — and the only one where a reference carries a DECLARED RELATIONSHIP (kept whole, partly kept, transferred onto a different subject, or a loose echo). Base minimax-h3 runs 480p/768p/2K/4K and reads 9 images + 3 video + 3 audio references; minimax-h3-max is fal's faster post-train, runs 480p/768p plus a 1080p refinement of its 768p render, and costs MORE than base H3 at 768p — a deliberate speed pick, never the default and never the cheap one; it animates start and end frames AND takes the same 9+3+3 omni-reference set (corrected 2026-09-09). minimax-h3-max-turbo is a second fal post-train with Max's ladder at half Max's rate; it takes start and end frames but has NO reference endpoint. Two hazards live here: reference images past the free allowance are billed (5 free then +4 credits on base H3; pooled media tokens on Max), and audio written into the wrong section is dropped or duplicated.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# MiniMax H3 — prompting
|
|
@@ -28,7 +28,7 @@ description: How to prompt MiniMax H3 and MiniMax H3 Max. Read before calling sl
|
|
|
28
28
|
- `A woman sits still at a kitchen table for a beat, then looks up. She says in English, "You said Tuesday." Scene sound: a fridge hum, a spoon set down on formica. Score: none.`
|
|
29
29
|
- `Two mechanics either side of an open bonnet. The younger one wipes his hands, waits, then speaks in Spanish, "No es el alternador." Scene sound: a socket wrench, a radio two bays over. Score: a low sustained cello under the last three seconds, audience only.`
|
|
30
30
|
|
|
31
|
-
**Hard constraint:** the
|
|
31
|
+
**Hard constraint:** the three seats differ in what the ENDPOINT accepts, not in grammar. Base H3 reaches 2K/4K and takes references; `minimax-h3-max` tops out at 1080p, takes the same 9+3+3 references, and costs MORE at the tier they share — a speed pick, never the cheap one; `minimax-h3-max-turbo` has Max's ladder at half its rate and takes frames only, NO references. Every tier above 768p is built from the native 768p render: judge at native. Reference inputs affect the quote; include every attached modality when estimating.
|
|
32
32
|
<!-- @card:end -->
|
|
33
33
|
|
|
34
34
|
<!-- @banned:start -->
|
|
@@ -50,21 +50,31 @@ German, Italian, Japanese, Korean, Portuguese, Russian, Spanish). That single fa
|
|
|
50
50
|
everything below — the prompt is not a shot description with sound bolted on, it is a **timeline
|
|
51
51
|
with three audio layers you author separately**.
|
|
52
52
|
|
|
53
|
-
**
|
|
54
|
-
endpoint accepts:
|
|
53
|
+
**Three seats, one grammar.** Everything in this file applies to all three. They differ only in
|
|
54
|
+
what the endpoint accepts:
|
|
55
55
|
|
|
56
|
-
| | `minimax-h3` | `minimax-h3-max` |
|
|
57
|
-
|
|
58
|
-
| Resolution | 480p / 768p / **2K / 4K** | 480p / 768p |
|
|
59
|
-
| References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) |
|
|
60
|
-
| Frames | start and/or end | start and/or end |
|
|
61
|
-
| Price at 768p | **$0.060/s** | $0.080/s |
|
|
62
|
-
| Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) |
|
|
56
|
+
| | `minimax-h3` | `minimax-h3-max` | `minimax-h3-max-turbo` |
|
|
57
|
+
|---|---|---|---|
|
|
58
|
+
| Resolution | 480p / 768p / **2K / 4K** | 480p / 768p / 1080p | 480p / 768p / 1080p |
|
|
59
|
+
| References | 9 images + 3 video + 3 audio (12 files) | 9 images + 3 video + 3 audio (12 files) | **none** (no reference endpoint) |
|
|
60
|
+
| Frames | start and/or end | start and/or end | start and/or end |
|
|
61
|
+
| Price at 768p | **$0.060/s** | $0.080/s | $0.040/s |
|
|
62
|
+
| Why pick it | resolution, references, and the cheaper second | **speed** — a 5s 768p clip in **4.8s** vs **57s** (measured) | **price** — half Max's rate at every tier |
|
|
63
63
|
|
|
64
64
|
**Max is the premium seat, not the budget one.** It is 33% dearer at the one tier they share and it
|
|
65
65
|
tops out lower. Route there when a fast turnaround on a text-to-video or start-frame shot is worth
|
|
66
66
|
paying for; route to base H3 for anything needing resolution, references, or the same tier cheaper.
|
|
67
67
|
|
|
68
|
+
**Turbo is the budget seat.** Same grammar and Max's ladder at half Max's rate, with no reference
|
|
69
|
+
endpoint: attach a reference and Slates refuses the call rather than dropping it. Route there for
|
|
70
|
+
drafts, volume and start-frame coverage, then re-run the keeper on Max or base H3 when it needs
|
|
71
|
+
references.
|
|
72
|
+
|
|
73
|
+
**1080p on Max and Turbo is a refinement, not a native render.** fal's schema, verbatim: *"1080P
|
|
74
|
+
latent refinement from a native 768P source."* It is a different stage from base H3's 2K/4K
|
|
75
|
+
upscaler, and it costs double the 768p second. Judge a 1080p take against the same shot at 768p
|
|
76
|
+
before paying for it across a batch.
|
|
77
|
+
|
|
68
78
|
**The speed is measured, not claimed** (2026-08-27, same prompt and params on both rows): a 5-second
|
|
69
79
|
768p text-to-video finished in **4.8 seconds** on Max against **57 seconds** on base H3 — roughly
|
|
70
80
|
**12x**, queue to finished file. fal advertises "under 3 seconds"; the literal claim did not hold at
|
|
@@ -192,8 +202,8 @@ original wording preserved exactly: *A red neon sign reading "Open Late" glows a
|
|
|
192
202
|
|
|
193
203
|
## References — H3's real differentiator is the declared RELATIONSHIP
|
|
194
204
|
|
|
195
|
-
*(
|
|
196
|
-
FOUR images rather than the base row's five.)*
|
|
205
|
+
*(`minimax-h3` and `minimax-h3-max`. Max gained the reference set on 2026-09-09; its free allowance
|
|
206
|
+
is FOUR images rather than the base row's five. `minimax-h3-max-turbo` takes no references.)*
|
|
197
207
|
|
|
198
208
|
<!-- @inject:references-read-literally -->
|
|
199
209
|
> **The general law: the model reads a reference literally.**
|
|
@@ -213,7 +223,7 @@ Every reference rule below is a corollary of that one sentence, which is why "pr
|
|
|
213
223
|
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
214
224
|
|
|
215
225
|
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
216
|
-
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates
|
|
226
|
+
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
|
|
217
227
|
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
218
228
|
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
219
229
|
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
@@ -260,7 +270,7 @@ exactly and say so.
|
|
|
260
270
|
**An audio reference cannot travel alone** — H3 refuses a reference set that is audio only. Pair it
|
|
261
271
|
with at least one image or video reference.
|
|
262
272
|
|
|
263
|
-
### 💸 Reference images past the free allowance are billed — and the two rows differ
|
|
273
|
+
### 💸 Reference images past the free allowance are billed — and the two reference rows differ
|
|
264
274
|
|
|
265
275
|
On `minimax-h3` the first **5** are free and each additional image adds **4 credits**.
|
|
266
276
|
Max pools image pixels, reference-video seconds and reference-audio seconds into one token
|
|
@@ -279,14 +289,14 @@ reference-heavy job — a quote that omits it under-reports the bill.
|
|
|
279
289
|
|
|
280
290
|
## Frames
|
|
281
291
|
|
|
282
|
-
|
|
292
|
+
All three rows take a **start frame**, an **end frame**, or both. With an
|
|
283
293
|
end frame, land it explicitly: describe the final pose, spacing and composition as the thing the
|
|
284
294
|
shot **settles into** at the end, rather than hoping the model finds it.
|
|
285
295
|
|
|
286
296
|
> *"…she rotates the handle into the final angle and settles into the pose, spacing and composition
|
|
287
297
|
> of image 2 at the end of the shot."*
|
|
288
298
|
|
|
289
|
-
**Frames and references are mutually exclusive** on
|
|
299
|
+
**Frames and references are mutually exclusive** on the two reference rows — they are different endpoints, and
|
|
290
300
|
the reference endpoint has no frame slots at all. Slates refuses the combination rather than
|
|
291
301
|
dropping one side.
|
|
292
302
|
|
|
@@ -301,6 +311,9 @@ dropping one side.
|
|
|
301
311
|
| `minimax-h3` · 2K · 10s | 65 |
|
|
302
312
|
| `minimax-h3` · 4K · 10s | 80 |
|
|
303
313
|
| `minimax-h3-max` · 768p · 10s | 40 |
|
|
314
|
+
| `minimax-h3-max` · 1080p · 10s | 80 |
|
|
315
|
+
| `minimax-h3-max-turbo` · 768p · 10s | 20 |
|
|
316
|
+
| `minimax-h3-max-turbo` · 1080p · 10s | 40 |
|
|
304
317
|
| `minimax-h3` — every reference image past the **fifth** | **+4** |
|
|
305
318
|
|
|
306
319
|
**768p is the default for a reason.** It is the tier the model natively generates.
|
|
@@ -64,7 +64,7 @@ movement from video 1. Preserve the character's identity, appearance, and outfit
|
|
|
64
64
|
That is the same endpoint the old `motionModel=seedance-2` branch called — it just wrote that sentence for you, invisibly. Add style/setting/camera direction freely; Seedance re-generates the whole shot.
|
|
65
65
|
|
|
66
66
|
- **Driving clip must be 2–15s** (all providers cap reference video at 15s). Longer clips: trim first, or use Kling MC (`characterOrientation: 'video'` takes up to 30s).
|
|
67
|
-
- **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending.
|
|
67
|
+
- **Billing = combined input+output seconds** (the vref keys). The server probes the clip and corrects an understated key — quote via the confirm gate before spending. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
|
|
68
68
|
- **Faces route through the face cascade**: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → confirm consent → `seedanceRealFace=true, realFaceConsent=true` (premium realface vref pricing).
|
|
69
69
|
- `characterOrientation` has no Seedance equivalent; framing follows the prompt + `aspectRatio`.
|
|
70
70
|
|