@slatesvideo/shared 0.6.4 → 0.6.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/api-url.d.ts +5 -3
- package/dist/api-url.js +5 -3
- package/dist/index.d.ts +1 -0
- package/dist/index.js +1 -0
- package/dist/manual/content.d.ts +2 -0
- package/dist/manual/content.js +3 -0
- package/dist/manual/index.d.ts +5 -0
- package/dist/manual/index.js +20 -0
- package/dist/operations/index.d.ts +67 -17
- package/dist/operations/index.js +310 -87
- package/dist/prompts/agent-doctrine.js +1 -0
- package/dist/prompts/character-sheet.js +10 -0
- package/dist/prompts/model-capabilities.js +52 -3
- package/dist/prompts/model-facts.js +12 -4
- package/dist/prompts/prompting-tips.js +9 -3
- package/dist/prompts/reference-composer.js +14 -0
- package/dist/prompts/shot-spec.d.ts +30 -10
- package/dist/prompts/shot-spec.js +41 -9
- package/dist/skills/content.js +12 -12
- package/exports/slates-prompt-builder/generated/reference-character.md +1 -1
- package/exports/slates-prompt-builder/generated/reference-seedance.md +3 -1
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +10 -10
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +1 -1
- package/skills/slates-character-identity.md +1 -1
- package/skills/slates-model-selection.md +10 -7
- package/skills/slates-prompting-elevenlabs.md +1 -1
- package/skills/slates-prompting-gpt-image-2-5.md +183 -0
- package/skills/slates-prompting-inworld-tts.md +174 -166
- package/skills/slates-prompting-lip-sync.md +1 -1
- package/skills/slates-prompting-nano-banana-2.md +1 -1
- package/skills/slates-prompting-seed-audio.md +1 -1
- package/skills/slates-prompting-seedance-2-5.md +3 -1
- package/skills/slates-prompting-seedance.md +3 -1
- package/skills/slates-ugc-influencer-ad.md +5 -3
- package/skills/slates-vision-feedback-loop.md +2 -2
- package/skills/slates-prompting-gpt-image-2.md +0 -109
|
@@ -76,7 +76,7 @@ Critically, the app injects **no** wardrobe, expression, or lighting directive.
|
|
|
76
76
|
- **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
|
|
77
77
|
- **Don't** skip binding. An unbound asset doesn't help downstream.
|
|
78
78
|
- **Don't** invent character details. Stick to what's in the reference image and the user's description.
|
|
79
|
-
- **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on gpt-image-2
|
|
79
|
+
- **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on GPT Image** (measured on `gpt-image-2`, the model 2.5 replaced; the classifier is OpenAI's, not the version's, so the rule carries — but nobody has re-run it on Flare or Sunburst) — fal returns `content_policy_violation` with `loc: ["body","prompt"]`, so the text is rejected before any image is read, because an anatomical absence reads as gore to OpenAI's classifier. It passed NB2, which is why the original receipt looked safe: **it was model-scoped.** State an exclusion as a framing choice, never as a missing body part.
|
|
80
80
|
- **Don't** invoke the invisible-mannequin genre without bounding it to the face. **"an invisible-mannequin presentation where the clothing holds its own shape" removed all the skin** — no neck, no hands, no forearms, a garment floating on nothing — because that *is* the e-commerce genre in full: an empty outfit. **"with just the face cropped out"** keeps the anchor and bounds it. Generalises: a genre anchor imports the whole genre, so name what STAYS, not only what goes.
|
|
81
81
|
- **Don't** put `#` or `@` anywhere in prompt text. Both are reference-token sigils in the desktop prompt composer and an unresolved one is **silently deleted** — no error, no log, just missing words. `#3a3a3c` reached fal as `background ()` on a real 2026-07-30 request, meaning the plate value had never been delivered to any model since the composer shipped. Write hex values bare.
|
|
82
82
|
- **Don't** use 4K — wastes credits, no quality gain at sheet scale.
|
|
@@ -228,9 +228,11 @@ Cite each by type and index, in the order they were attached — `image 1`, `vid
|
|
|
228
228
|
|
|
229
229
|
```
|
|
230
230
|
Marcus (image 1) performs the motion from video 1, in the workshop from image 2,
|
|
231
|
-
|
|
231
|
+
using the voice timbre from audio 1. Preserve his identity, appearance and outfit.
|
|
232
232
|
```
|
|
233
233
|
|
|
234
|
+
🚨 **SAY WHAT AN AUDIO REFERENCE IS FOR.** It can mean music, dialogue, voice, tone or timbre — five roles on one attachment — so an unroled clip falls back to **dialogue**: the model re-transcribes it and speaks ITS words. A real take came back as *"a map called Slates"* for *"an app called Slates"*. Name it as the voice timbre and the clip carries the voice while the prompt carries the words. ByteDance's own sentence: *"Image 1 depicts the protagonist John and uses the voice timbre from Audio 1."* Bind each speaker in a sentence, never by attachment order — position carries nothing.
|
|
235
|
+
|
|
234
236
|
### Motion transfer & lip-sync recipes (reference video / audio)
|
|
235
237
|
|
|
236
238
|
These aren't separate Seedance features — they're prompting strategies over reference media.
|
|
@@ -8,11 +8,11 @@
|
|
|
8
8
|
},
|
|
9
9
|
{
|
|
10
10
|
"path": "skills/slates-character-identity.md",
|
|
11
|
-
"sha256": "
|
|
11
|
+
"sha256": "87ede794637553db0074dd64ea9a9cb27bfc3afda52fdc6489702d81e7f94b53"
|
|
12
12
|
},
|
|
13
13
|
{
|
|
14
14
|
"path": "skills/slates-prompting-seedance.md",
|
|
15
|
-
"sha256": "
|
|
15
|
+
"sha256": "e3210015ed5079fdd1cd21e3e5fea186ad12126015af5f0fd53d65e833004382"
|
|
16
16
|
},
|
|
17
17
|
{
|
|
18
18
|
"path": "skills/slates-prompting-kling-v3.md",
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
},
|
|
21
21
|
{
|
|
22
22
|
"path": "skills/slates-prompting-nano-banana-2.md",
|
|
23
|
-
"sha256": "
|
|
23
|
+
"sha256": "9b746abb7bb3726e39d054a704e9a71f699e365e27e40ced14bca06f922f8a24"
|
|
24
24
|
},
|
|
25
25
|
{
|
|
26
26
|
"path": "skills/slates-content-policy.md",
|
|
@@ -28,7 +28,7 @@
|
|
|
28
28
|
},
|
|
29
29
|
{
|
|
30
30
|
"path": "src/prompts/model-facts.ts",
|
|
31
|
-
"sha256": "
|
|
31
|
+
"sha256": "63d6709c85d2f11d4244e46bd1a595094f1cf7066e2ac69c8a2b46783e4dc4b6"
|
|
32
32
|
}
|
|
33
33
|
],
|
|
34
34
|
"outputs": [
|
|
@@ -39,13 +39,13 @@
|
|
|
39
39
|
},
|
|
40
40
|
{
|
|
41
41
|
"path": "reference-character.md",
|
|
42
|
-
"bytes":
|
|
43
|
-
"sha256": "
|
|
42
|
+
"bytes": 10249,
|
|
43
|
+
"sha256": "75853dcf6b793df82924bf96bee1795b0a33014bb8cc8a2d18543455c0232761"
|
|
44
44
|
},
|
|
45
45
|
{
|
|
46
46
|
"path": "reference-seedance.md",
|
|
47
|
-
"bytes":
|
|
48
|
-
"sha256": "
|
|
47
|
+
"bytes": 35031,
|
|
48
|
+
"sha256": "19f26396f7fc9f688d0084cbae8570594a2eab8035113ecf853dd16e6d02d520"
|
|
49
49
|
},
|
|
50
50
|
{
|
|
51
51
|
"path": "reference-kling.md",
|
|
@@ -65,8 +65,8 @@
|
|
|
65
65
|
],
|
|
66
66
|
"archive": {
|
|
67
67
|
"path": "slates-prompt-builder.skill",
|
|
68
|
-
"bytes":
|
|
69
|
-
"sha256": "
|
|
68
|
+
"bytes": 40494,
|
|
69
|
+
"sha256": "4e412da5d81d93ec4f93f88f8478855021df4c8f64292ec84018855078d9a184",
|
|
70
70
|
"entries": [
|
|
71
71
|
"SKILL.md",
|
|
72
72
|
"reference-character.md",
|
|
Binary file
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@slatesvideo/shared",
|
|
3
|
-
"version": "0.6.
|
|
3
|
+
"version": "0.6.6",
|
|
4
4
|
"description": "Shared operations layer for the Slates MCP server and CLI: auth, cloud/desktop clients, and the single tool surface both consume. Most users want @slatesvideo/mcp-server or @slatesvideo/cli instead.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -98,7 +98,7 @@ Critically, the app injects **no** wardrobe, expression, or lighting directive.
|
|
|
98
98
|
- **Don't** create a second character image. One canonical identity is what the storyboard pipeline reads.
|
|
99
99
|
- **Don't** skip binding. An unbound asset doesn't help downstream.
|
|
100
100
|
- **Don't** invent character details. Stick to what's in the reference image and the user's description.
|
|
101
|
-
- **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on gpt-image-2
|
|
101
|
+
- **Don't** describe the front panel's crop as an absent head — in `userNotes` or any hand-written variant. The template asks for it as *framing*: **"cropped at the collarbone, an invisible-mannequin presentation with just the face cropped out"**, a standard e-commerce genre with deep training data. **"the head not shown" is a hard 422 on GPT Image** (measured on `gpt-image-2`, the model 2.5 replaced; the classifier is OpenAI's, not the version's, so the rule carries — but nobody has re-run it on Flare or Sunburst) — fal returns `content_policy_violation` with `loc: ["body","prompt"]`, so the text is rejected before any image is read, because an anatomical absence reads as gore to OpenAI's classifier. It passed NB2, which is why the original receipt looked safe: **it was model-scoped.** State an exclusion as a framing choice, never as a missing body part.
|
|
102
102
|
- **Don't** invoke the invisible-mannequin genre without bounding it to the face. **"an invisible-mannequin presentation where the clothing holds its own shape" removed all the skin** — no neck, no hands, no forearms, a garment floating on nothing — because that *is* the e-commerce genre in full: an empty outfit. **"with just the face cropped out"** keeps the anchor and bounds it. Generalises: a genre anchor imports the whole genre, so name what STAYS, not only what goes.
|
|
103
103
|
- **Don't** put `#` or `@` anywhere in prompt text. Both are reference-token sigils in the desktop prompt composer and an unresolved one is **silently deleted** — no error, no log, just missing words. `#3a3a3c` reached fal as `background ()` on a real 2026-07-30 request, meaning the plate value had never been delivered to any model since the composer shipped. Write hex values bare.
|
|
104
104
|
- **Don't** use 4K — wastes credits, no quality gain at sheet scale.
|
|
@@ -9,7 +9,7 @@ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. M
|
|
|
9
9
|
|
|
10
10
|
## 🔑 The meta-rule — above the table
|
|
11
11
|
|
|
12
|
-
The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:
|
|
12
|
+
The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2.5 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:
|
|
13
13
|
|
|
14
14
|
> **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.
|
|
15
15
|
>
|
|
@@ -88,16 +88,18 @@ Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that
|
|
|
88
88
|
|
|
89
89
|
**Video models (Kling, Seedance, Veo) cannot generate standalone images — ever.** A "premium hero reference image" is still an image job: it routes to an image model below, never to Seedance.
|
|
90
90
|
|
|
91
|
-
- **Default: Nano Banana 2** —
|
|
91
|
+
- **Default: Nano Banana 2** — strongest reference HANDLING (14 refs; GPT Image now takes more, at 16, but Banana is still the one that holds many subjects coherently), best legible text, the standard start-frame generator.
|
|
92
92
|
- **NB2 Lite** — the fast/draft seat: ~half NB2's price, ~2.7× faster, 1K only. Route iteration volume and drafts here; finals go back to NB2 full (2K/4K).
|
|
93
93
|
- **Nano Banana Pro** — the hero-frame/typography ceiling (~2× NB2). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — feed it a full subject library.
|
|
94
|
-
- **GPT Image 2** —
|
|
94
|
+
- **GPT Image 2.5** — two seats, `gpt-image-2-5-flare` and `gpt-image-2-5-sunburst`, **same price**. Readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. **Also the photoreal front-runner (Eric, 2026-08-24)** — it beat both Nano Banana rails head-to-head on skin realism, which is why the AI-influencer ad lane generates every plate on this line. **The seat split is SPEED vs QUALITY, not generate vs edit** (OpenAI's own rule): Flare is the small, fast model with quality *comparable to* GPT Image 2 — drafts, exploration, volume; Sunburst is OpenAI's *most capable* image model, higher quality than GPT Image 2, deliberately slower — finals, hero frames, photoreal, and multi-reference edits, where its lead is widest. **Explore on Flare, finish on Sunburst.** Five quality tiers, cheapest first — `low` (layout checks only), `medium` (drafts), **`high` (the default)**, `xhigh`, `max` (the top). Uneven: `max` is 4× `high`, `xhigh` only ~1.8× it. **16 reference images**, the schema ceiling. **Transparent backgrounds** via `backgroundMode` — free, and the only image family that offers them.
|
|
95
|
+
|
|
96
|
+
🚨 **The tier names moved when 2.5 replaced GPT Image 2, and the strings did not.** GPT Image 2's `medium` is 2.5's `high`; its `high` is 2.5's `max` — same money, one rung of renaming. The 2026-08-24 photoreal result was measured at GPT Image 2 `high`, so **the tier that reproduces it is `max`**. Nobody has re-run it on 2.5; the ranking is inherited, not re-measured.
|
|
95
97
|
- **FLUX.2 Max** — photoreal texture, hex-color binding, typography, less censored.
|
|
96
98
|
- **Seedream 5 Lite** — uncensored + any-resolution flat price; volume exploration when the Gemini filter is in the way.
|
|
97
99
|
|
|
98
|
-
**Split rule of thumb:** readable text / panels / UI **
|
|
100
|
+
**Split rule of thumb:** readable text / panels / UI → GPT Image 2.5 (Flare to explore, Sunburst to finish); **photoreal people, finals and hero frames → Sunburst at `max`** — the 2026-08-24 result was measured at GPT Image 2's `high`, which is `max` here, and Flare only *matches* GPT Image 2 while Sunburst exceeds it; multi-reference edits where several references must all survive into one frame → Sunburst; edit-heavy work → the Banana line; drafts → GPT Image 2.5 Flare at `medium`, which now undercuts NB2 Lite on both price and resolution; uncensored or odd resolutions → Seedream/FLUX.
|
|
99
101
|
|
|
100
|
-
⚠️ **This line said the opposite until 2026-08-24** — it sent photoreal *away* from GPT Image
|
|
102
|
+
⚠️ **This line said the opposite until 2026-08-24** — it sent photoreal *away* from GPT Image on reputation, which is the exact failure § The meta-rule above warns about. Re-run the evidence test when the roster moves. It moved again on 2026-09-09, and the ranking was carried across rather than re-measured — exactly what the meta-rule says not to trust. Treat it as a starting hypothesis for 2.5, not a receipt. **The seat choice above is likewise reasoned from OpenAI's positioning, not measured:** run Flare-`max` against Sunburst-`max` on one plate and write the answer into `slates-prompting-gpt-image-2-5`.
|
|
101
103
|
|
|
102
104
|
## Audio routing
|
|
103
105
|
|
|
@@ -105,10 +107,11 @@ Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that
|
|
|
105
107
|
|
|
106
108
|
| Job | Model | Why |
|
|
107
109
|
|---|---|---|
|
|
108
|
-
| **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. The continuity-bed workhorse
|
|
110
|
+
| **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines inside a scene | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. The continuity-bed workhorse; dialogue is performed inside the room, not cast. |
|
|
111
|
+
| **One named voice saying one line** — a character's own voice, a narrator, a clean VO to lip-sync against | **Inworld TTS-2** (`inworld-tts-2`) | The prompt IS the words, spoken verbatim and billed per character. Voice = the character's clip (cloned for the take), a description, or a preset. No room tone — mix it on the timeline. |
|
|
109
112
|
| **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control and a real loop mode. |
|
|
110
113
|
|
|
111
|
-
**There is no music model
|
|
114
|
+
**There is no music model.** A song is imported (Slates reads audio files and puts them on the timeline), not generated. A line that has to be spoken in a SPECIFIC voice is generated on Inworld TTS-2 and lip-synced against; a line that belongs to a scene is performed by Seed Audio inside it.
|
|
112
115
|
|
|
113
116
|
### Named audio escalation triggers
|
|
114
117
|
|
|
@@ -28,7 +28,7 @@ description: How to prompt ElevenLabs Sound Effects v2 in Slates. Read before ca
|
|
|
28
28
|
- `A heavy oak door slams shut in a stone hallway, brief reverberant tail.` (1.5s)
|
|
29
29
|
- `Steady rain on a tin awning, no thunder, no wind gusts, seamless loop.` (18s)
|
|
30
30
|
|
|
31
|
-
**Hard constraint:** it is billed per second and the duration is never left for the model to pick — that would make the charge non-deterministic. It is NOT a speech surface:
|
|
31
|
+
**Hard constraint:** it is billed per second and the duration is never left for the model to pick — that would make the charge non-deterministic. It is NOT a speech surface: a line in a specific voice is `inworld-tts-2`, and dialogue inside a scene is Seed Audio, which casts and performs the line in the room.
|
|
32
32
|
<!-- @card:end -->
|
|
33
33
|
|
|
34
34
|
<!-- @banned:start -->
|
|
@@ -0,0 +1,183 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: slates-prompting-gpt-image-2-5
|
|
3
|
+
description: Prompting GPT Image 2.5 (Flare and Sunburst) — the readable-text / character-sheet / shot-grid engine AND the photoreal front-runner. Read before calling slates_generate_image with model gpt-image-2-5-flare or gpt-image-2-5-sunburst. Covers picking the variant, the five quality tiers (high is the default and the everyday seat), resolution classes (1k/2k=1080p/3k=1440p/4k), reference-image roles, text-accuracy prompting, panel/grid layout direction, edit constraints, and when to route to the Banana line instead.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# GPT Image 2.5 — sheets, grids, and text that actually reads
|
|
7
|
+
|
|
8
|
+
<!-- @card:start -->
|
|
9
|
+
<!-- slates-only -->
|
|
10
|
+
<!-- MACHINE-READ. Everything between the @card markers is extracted by
|
|
11
|
+
src/prompts/craft-cards.ts and returned on every cost estimate for this
|
|
12
|
+
model, so it is the ONE piece of positive craft guidance the agent cannot
|
|
13
|
+
skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved
|
|
14
|
+
compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.
|
|
15
|
+
Keep it under 2,400 characters (the build fails above that) and keep the
|
|
16
|
+
rationale, the receipts and the worked examples in the body below. -->
|
|
17
|
+
<!-- /slates-only -->
|
|
18
|
+
**Card — GPT Image 2.5.** The readable-text, ordered-panel and exact-placement engine, and the photoreal front-runner for people. Structure: subject and action, then the exact copy in quotes, then layout, then light.
|
|
19
|
+
|
|
20
|
+
**Pick the seat first.** `flare` = the small/FAST seat, quality *comparable to* GPT Image 2 — drafts, exploration, volume. `sunburst` = OpenAI's *most capable*, higher quality, slower, same price — finals, hero frames, photoreal, multi-reference edits. Explore on Flare, finish on Sunburst.
|
|
21
|
+
|
|
22
|
+
**The six levers**
|
|
23
|
+
1. **Quote every string that must render verbatim** — `the sign reads "OPEN 24 HOURS"`. Quoted strings render most reliably.
|
|
24
|
+
2. **Font FEEL, never a font name** — `clean geometric sans, high contrast`, `hand-painted brush lettering`.
|
|
25
|
+
3. **Order dense copy explicitly** — `Line 1: "..." Line 2: "..."`. It respects the ordering.
|
|
26
|
+
4. **Name the layout as a grid** for sheets and panels — `a 3x2 grid of panels, reading left to right, equal gutters`.
|
|
27
|
+
5. **Give every reference image a ROLE** — subject / style / clothing / background. New emphasis in 2.5 and the highest-leverage change for multi-reference work.
|
|
28
|
+
6. **Set `quality` deliberately.** Five rungs — `low`, `medium`, `high` (default), `xhigh`, `max` — spanning ~36× end to end, in UNEVEN steps: `max` is 4× `high`, but `xhigh` only ~1.8× it. `medium` is the draft seat; `high` is the everyday tier; reach past it only when tiny type, dense diagrams or many labelled elements ARE the job. 🚨 **Coming from GPT Image 2, the names moved one rung:** its `medium` is this `high`, its `high` is this `max` — same money, renamed ladder. Carrying an old value over silently buys a cheaper picture.
|
|
29
|
+
|
|
30
|
+
**Examples**
|
|
31
|
+
- `A 2x3 character turnaround sheet on a neutral grey field, equal gutters, reading left to right: front, three-quarter, profile, back, three-quarter back, top. One woman, mid-30s, cropped dark hair, olive field jacket. Flat even studio light, no cast shadows. Small caption under each panel naming the angle.`
|
|
32
|
+
- `Photoreal portrait, natural window light from camera-left, visible skin texture and pores, 85mm compression. A man in his 50s in a charcoal knit, half-smile, looking just past lens.`
|
|
33
|
+
|
|
34
|
+
**Hard constraint:** keep total on-image text under about 30 words for perfect accuracy — beyond that it degrades, gracefully but really. It has its own content filter, distinct from Gemini's.
|
|
35
|
+
<!-- @card:end -->
|
|
36
|
+
|
|
37
|
+
<!-- @banned:start -->
|
|
38
|
+
<!-- slates-only -->
|
|
39
|
+
<!-- MACHINE-READ. Every `backticked` token between the @banned markers is
|
|
40
|
+
extracted by src/prompts/banned-tokens.ts and returned on this model's cost
|
|
41
|
+
estimate, and every submitted prompt is matched against it. Keep entries
|
|
42
|
+
backticked and prose outside the backticks. -->
|
|
43
|
+
<!-- /slates-only -->
|
|
44
|
+
**Never use:**
|
|
45
|
+
- a font NAME — describe the feel (`clean geometric sans, high contrast`) instead
|
|
46
|
+
- a reference role essay (`Reference image 1 is a photograph of a woman. Use that exact woman.`) — name the subject inline instead
|
|
47
|
+
- `8k`, `masterpiece`, `best quality`, `highly detailed` — quality incantations do nothing here either
|
|
48
|
+
<!-- @banned:end -->
|
|
49
|
+
|
|
50
|
+
GPT Image's edge is **character-level text accuracy** (~99% on English), ordered panels, and exact element placement — the jobs where every other model garbles a word or shuffles a layout. 2.5 inherits all of it and is better at each.
|
|
51
|
+
|
|
52
|
+
## Which variant
|
|
53
|
+
|
|
54
|
+
**Speed → Flare. Quality → Sunburst.** That is OpenAI's own routing rule, quoted from its image-prompting guide: *"start with GPT Image 2.5 Flare when speed is the priority, or GPT Image 2.5 Sunburst when demanding quality requirements are the priority."* Same price either way, so the trade is purely latency against quality.
|
|
55
|
+
|
|
56
|
+
🚨 **FLARE IS NOT AN UPGRADE OVER GPT IMAGE 2 — IT IS THE FAST ONE.** OpenAI, verbatim: *"GPT Image 2.5 Flare is the small model, optimized for speed, with image quality **comparable to** GPT Image 2. GPT Image 2.5 Sunburst is the base model, optimized for quality, with **higher image quality than** GPT Image 2."* Their model pages agree: Flare is *"our fastest model for high-quality, everyday image generation"*, Sunburst *"our most capable model for image generation and editing."* **Sunburst is the seat that beats what we had; Flare is the one that holds it at half the latency.** An earlier revision of this file called Flare "better than GPT Image 2" and sent Sunburst only to multi-reference edits — both wrong, corrected 2026-09-09 against the vendor docs.
|
|
57
|
+
|
|
58
|
+
**The production pattern: explore on Flare, finish on Sunburst.** Drafts, layout checks and volume go to Flare. Finals, hero frames, photoreal people and any edit that must preserve identity or geometry go to Sunburst.
|
|
59
|
+
|
|
60
|
+
**Sunburst's widest lead is multi-reference editing** — several references all surviving into one frame, the character-consistency-across-shots problem. Reach for it there first, but that is not the only place it belongs.
|
|
61
|
+
|
|
62
|
+
⚠️ **The LMArena receipt, scoped.** At launch Arena had Sunburst #1 and Flare #2 across text-to-image, single-image edit and multi-image edit, with margins over GPT Image 2 of **+81 / +47** on multi-image edit (Image Edit Arena: Sunburst 1520, Flare 1491, GPT Image 2 1461). Two caveats were missing and both matter: the baseline is **GPT Image 2 at `medium`, which is this model's `high`** — not its top tier — and the boards were **preliminary, a few thousand votes each**. Arena says Flare beats GPT Image 2; OpenAI says comparable. Route on OpenAI's wording and treat the board as a tiebreaker, not a spec.
|
|
63
|
+
|
|
64
|
+
🚨 **The GPT Image line is ALSO the photoreal front-runner, and this file said the opposite until 2026-08-24.** **Receipts:** Eric's direct call, plus a head-to-head on the Higgsfield rail where GPT Image 2 at `quality: high`, 2K beat both Nano Banana rails on skin realism for photoreal people — that result is why the whole AI-influencer ad lane generates its plates here. **Route photoreal to this line, not away from it.**
|
|
65
|
+
|
|
66
|
+
⚠️ **Which SEAT reproduces it follows from the two facts above, and it is not the obvious one.** The receipt was measured on GPT Image 2 at `high`, which is this model's **`max`** — the ladder was renamed, not repriced (see `slates-model-selection`). Flare is only *comparable* to GPT Image 2, so **Flare at `max` is the floor: it holds the measured result rather than beating it.** Sunburst is documented as higher quality than GPT Image 2, which makes **Sunburst at `max` the seat most likely to exceed it** — and a photoreal final is exactly the "quality outranks speed" case OpenAI routes to Sunburst. Nobody has re-run the head-to-head on either seat, so this is reasoning from the vendor's positioning, not a measurement. **Run Flare-max against Sunburst-max on one plate before committing the lane, and write the result here.**
|
|
67
|
+
|
|
68
|
+
**What the Banana line still owns:** edit-heavy work, and holding many subjects coherently in one frame. **Not the reference ceiling any more** — that line was true until 2026-09-09, when GPT Image went to its documented 16 against Banana's 14. Route on which model keeps them all recognisable, not on the count.
|
|
69
|
+
|
|
70
|
+
**What would kill this:** a head-to-head at the intended crop going the other way. Per `slates-model-selection` § The meta-rule, re-run the evidence test when the roster changes — never carry a ranking forward on reputation. That rule is exactly what the 2026-08-24 correction failed, and exactly what the two ⚠️ notes above are honouring.
|
|
71
|
+
|
|
72
|
+
## Quality tiers — always set explicitly
|
|
73
|
+
|
|
74
|
+
All five rungs are exposed, and they span ~36× end to end (2k class: $0.0044 → $0.158), which makes this the single biggest cost lever on the model. **The steps are UNEVEN — do not reason about them as a constant multiplier:** ~2.3× `low`→`medium`, ~3.9× `medium`→`high`, ~1.8× `high`→`xhigh`, ~2.25× `xhigh`→`max`. The same ratios hold at every OFFERED resolution class (2k/3k/4k); unoffered 1k differs slightly.
|
|
75
|
+
|
|
76
|
+
| Tier | Use it for |
|
|
77
|
+
|---|---|
|
|
78
|
+
| `low` | Roughest pass — layout and composition checks, throwaway comps. |
|
|
79
|
+
| `medium` | The draft seat. Cheaper than NB2 Lite and available up to 4K, which is why the draft lane moved here. |
|
|
80
|
+
| `high` | **Default.** The everyday tier. Blind benchmarks on GPT Image 2 put this rung — which it called `medium` — within a hair of `max` (which it called `high`) at a quarter of the cost. Inherited from the old ladder, never re-run on 2.5, and it says nothing about `xhigh`. |
|
|
81
|
+
| `xhigh` | One rung short of the top at about half its price (2k: 4 cr against `max`'s 8). Worth trying before `max`. |
|
|
82
|
+
| `max` | Top of the ladder. Tiny type, dense diagrams, many labelled elements. |
|
|
83
|
+
|
|
84
|
+
⚠️ **A tier label means different things on different models.** OpenAI: *"The same quality label does not imply the same image quality or response time across models."* Flare at `max` and Sunburst at `max` are not the same picture, and neither matches Nano Banana's idea of "high".
|
|
85
|
+
|
|
86
|
+
🚨 **The tier NAMES moved between versions and the strings did not.** GPT Image 2's `medium` is this model's `high`; its `high` is this model's `max` — same money, one rung of renaming. So a recipe, a doc or a memory that says "GPT Image at medium" means **`high` here**. Getting this backwards costs picture quality silently: nothing errors, the bill is correct for what was asked, and the image is just worse.
|
|
87
|
+
|
|
88
|
+
Never rely on the provider default. fal's default is `high`, which is correct today — but it is the third rung of five rather than the top of two, so leaning on it means a fal-side change silently reprices you. The Slates ops send `high` unless you say otherwise.
|
|
89
|
+
|
|
90
|
+
**Find the tier from the top down, then walk back.** OpenAI's own procedure: *"If the output falls short, test a higher quality setting. Once it meets your requirements, test lower settings to see whether they preserve acceptable quality while reducing latency. Use `xhigh` or `max` only when they improve an unmet quality requirement within your latency budget."* A higher rung does **not** guarantee a better result on a given prompt. Compare `medium` against `high` when the job is small or dense text; that is where the rungs separate most visibly.
|
|
91
|
+
|
|
92
|
+
## Resolution classes
|
|
93
|
+
|
|
94
|
+
`1k` = 1024²-class · `2k` = 1920×1080-class · `3k` = 2560×1440-class · `4k` = 3840×2160-class. Pick 2k for most sheets/panels; 4k for print-density grids. 4K exists at every tier and is API-only — even paid ChatGPT can't render it.
|
|
95
|
+
|
|
96
|
+
`1k` is not offered, and the reason is not its price: it is strictly dominated. At 1k you pay more for fewer pixels than at 2k, at **all five tiers**. Don't ask for it.
|
|
97
|
+
|
|
98
|
+
⚠️ **Above 2560×1440 you are on a path OpenAI marks EXPERIMENTAL.** Verbatim: *"Outputs with more than 3,686,400 total pixels ('2560x1440') are experimental."* That is the whole **4k** class (≈8.0 MP) plus 3k at 4:3/3:4 (≈3.70 MP). It bills normally and it works — but prove the shot at 2k or 3k 16:9 first, and do not be surprised by an odd frame at 4k.
|
|
99
|
+
|
|
100
|
+
**Hard size bounds**, from fal's schema verbatim: each edge ≤ 3840 px, both edges multiples of 16, longer:shorter ratio ≤ 3:1, total pixels between 655,360 and 8,294,400. **The pixel ceiling is the one that actually bites** — the multiple-of-16 rule is documented but NOT enforced, and we have the receipt: 1920×1080 fails it (1080 = 67.5 × 16), is one of fal's own six priced canonical sizes, and metered clean. Slates picks sizes that respect the ceiling; these matter only if you hand-build a request.
|
|
101
|
+
|
|
102
|
+
🚨 **THE ASPECT RATIO CHANGES THE PRICE ON THIS MODEL, and on no other image model.** OpenAI bills image OUTPUT TOKENS and the count tracks the frame's SHAPE, so at the same resolution class **`1:1` costs about 1.8× and `4:3`/`3:4` about 1.37× what `16:9` costs**; `9:16` costs the same as `16:9`. Metered 2026-09-09 and priced into the cost key, so the quote you get before generating is the real number — but if you are choosing between shapes and the budget is tight, **16:9 or 9:16 is the cheap one.** Every other image model charges the same whatever the shape.
|
|
103
|
+
|
|
104
|
+
## Reference images — give every one a role
|
|
105
|
+
|
|
106
|
+
**Assign a role to every reference image: subject, style, clothing, or background.** This is new emphasis in 2.5 and the highest-leverage change for the 16-reference character lane. An unroled pile of references makes the model guess what each one is for, and it guesses differently every run — which is the drift people mistake for a consistency failure.
|
|
107
|
+
|
|
108
|
+
Reference images route through the edit endpoint, **up to 16** — fal's documented `maxItems`, and the highest reference ceiling of any image seat in Slates (the Banana line takes 14). It was capped at 10 until 2026-09-09, which was never anybody's limit, just a number nobody had checked. The composed "image N" naming applies as everywhere else. Mask-based inpainting exists at the API level but is not surfaced: a mask is something the user has to paint, and there is no painting surface — describe the change instead.
|
|
109
|
+
|
|
110
|
+
## Editing — separate the change from the constraints
|
|
111
|
+
|
|
112
|
+
**State the change, then list what must survive.** "Change only X," then name the invariants explicitly: identity, geometry, lighting, labels. For precise local edits also pin saturation, contrast, camera angle and surrounding objects — anything you do not pin is fair game for the model to move.
|
|
113
|
+
|
|
114
|
+
**One change per iteration, and restate the constraints every turn.** Cross-turn drift is the named failure mode in OpenAI's own guidance: constraints do not persist across turns by themselves, so a multi-turn refinement that stops restating them will slowly rewrite the frame. This applies directly to multi-turn shot refinement.
|
|
115
|
+
|
|
116
|
+
## Prompting for text accuracy
|
|
117
|
+
|
|
118
|
+
- **Quote every string that must render verbatim**: `the sign reads "OPEN 24 HOURS"` — quoted strings render most reliably.
|
|
119
|
+
- Say the text appears **once**, and give its position and typography.
|
|
120
|
+
- Spell unusual words letter-by-letter.
|
|
121
|
+
- Add `no extra text, no watermarks`.
|
|
122
|
+
- Specify font *feel*, not font names: "clean geometric sans, high contrast", "hand-painted brush lettering".
|
|
123
|
+
- For dense text (posters, UI mocks), list the copy as ordered lines: `Line 1: "..." Line 2: "..."` — it respects ordering.
|
|
124
|
+
- **Don't bundle unrelated instructions into a text-rendering request.** A prompt that also redesigns the scene competes with the text for attention.
|
|
125
|
+
- Keep total on-image text under ~30 words for perfect accuracy; beyond that, accuracy degrades gracefully but degrades.
|
|
126
|
+
|
|
127
|
+
## Transparent backgrounds
|
|
128
|
+
|
|
129
|
+
If you need a cut-out rather than a scene, **ask for it explicitly and check the alpha**. OpenAI: request `background=transparent` and use PNG or WebP, then *"check the decoded image's alpha channel, including hair, glass, shadows, and object edges"* — a painted-white backdrop is the common failure and it is not transparency. Say what must NOT appear: *"no solid backdrop, no checkerboard, no scenery, no watermark"*, and do not let the product get restyled while the background is removed. **On every follow-up edit, repeat the transparency requirement** or it gets dropped. (Slates always requests PNG, so the format half is handled for you. **`background` IS surfaced now** — the Background control on the prompt bar, and `backgroundMode` on `slates_generate_image` / `slates_edit_image`. It is free: fal prices this family on size × quality alone.)
|
|
130
|
+
|
|
131
|
+
## When an edit must not touch a region at all
|
|
132
|
+
|
|
133
|
+
Prompting alone cannot guarantee pixel-identical pixels. OpenAI's own instruction: if a region must stay exactly as it was, **composite the approved edit back into the original image** rather than asking the model to preserve it. Treat "preserve" language as a strong bias, never a lock.
|
|
134
|
+
|
|
135
|
+
## Structure a complex prompt in labeled sections
|
|
136
|
+
|
|
137
|
+
For anything with several requirements, OpenAI recommends organising the prompt as **scene, subject, details, constraints** with labeled sections. Same content, easier to read and to change one part without disturbing the rest — which is what makes the one-change-per-iteration rule practical.
|
|
138
|
+
|
|
139
|
+
**Say "photorealistic" or "real photograph" when that is the goal.** It is not inferred from a detailed description; ask for it directly, then describe framing and texture.
|
|
140
|
+
|
|
141
|
+
## Concrete visuals beat mood words
|
|
142
|
+
|
|
143
|
+
Name materials, lighting, colour and medium. Mood words are cues only — "cinematic", "moody", "epic" tell the model almost nothing on their own. Give scale, atmosphere and colour instead. Camera specs (`85mm`, `f/1.4`) are appearance hints, not a physical simulation; they bias the look, they do not compute optics.
|
|
144
|
+
|
|
145
|
+
**For people, state body framing and scale**: "full body visible, feet included", "hands naturally gripping the handlebars". This is also the safest way to phrase a crop — see the blocked-phrasings section below.
|
|
146
|
+
|
|
147
|
+
**No special syntax is required.** Prose, JSON and tagged blocks all work equally well, so pick whatever stays maintainable in the caller.
|
|
148
|
+
|
|
149
|
+
## Panels, sheets, and grids
|
|
150
|
+
|
|
151
|
+
- State the grid explicitly and number the cells: "a 2×3 grid of panels, numbered 1–6, reading left-to-right, top-to-bottom".
|
|
152
|
+
- Give each cell ONE content clause: "Panel 3: the character mid-jump, side view".
|
|
153
|
+
- Character identity sheets: GPT Image holds both the structured panel layout AND photoreal skin, which is why the influencer-ad lane builds its sheets here. Reach for NB2/NB Pro when it is an edit of an existing sheet, or when many subjects have to stay recognisable at once — not for the reference count, which GPT Image now leads at 16.
|
|
154
|
+
|
|
155
|
+
## 🚨 WHAT GETS YOU BLOCKED — read before writing a prompt with a person in it
|
|
156
|
+
|
|
157
|
+
**Receipt: 24 consecutive attempts on one character, 2026-08-24, same project and same rail.** Eleven were refused with `content_policy_violation` on the fal edit endpoint. The refusals were never about the scene — one of the blocked prompts was a woman standing at a kitchen counter with her hand on it. **Two phrasings were hard blocks, 5 for 5 each, and neither ever passed:**
|
|
158
|
+
|
|
159
|
+
**1. Never describe the reference as a photograph of a real person.**
|
|
160
|
+
|
|
161
|
+
> ❌ `Reference image 1 is a photograph of a woman. Use that exact woman.`
|
|
162
|
+
> ✅ `Reference image 1 is a character identity sheet showing one woman across several panels — the face in the large portrait panel is the authority for her identity. Use that exact woman.`
|
|
163
|
+
|
|
164
|
+
The first reads to the filter as *recreate this real person's likeness*, which is a hard refusal regardless of what the rest of the prompt says. The second signals a fictional character and passes. **This is a wording change only — the reference image can be the same file either way.** One plate flipped from refused to accepted on this single sentence with nothing else altered.
|
|
165
|
+
|
|
166
|
+
**2. Never attach a reference sheet containing a headless body panel.** A sheet whose full-body panels are cropped above the neck is refused every time, even with the correct opener. Regenerate the sheet with the head visible in every panel. Related, and already in this file's sheet guidance: phrase a cropped panel as *framing* (`cropped at the collarbone`), never as *absence* (`the head not shown`).
|
|
167
|
+
|
|
168
|
+
⚠️ **These refusals were measured on GPT Image 2, not on 2.5.** The classifier belongs to OpenAI rather than to a model version, so the phrasing rules carry — but they are inherited, not re-measured. If Flare or Sunburst accepts one of the blocked phrasings, that is a new receipt to write down here, not a reason to delete this one.
|
|
169
|
+
|
|
170
|
+
**On top of those, ordinary content triggers still apply** and they stack independently — a correct opener does not rescue them:
|
|
171
|
+
|
|
172
|
+
| Refused | Why, and the fix |
|
|
173
|
+
|---|---|
|
|
174
|
+
| A woman sitting on a bed in a bedroom | Domestic + bed reads as intimate. Move her to a chair, a rug, another room. |
|
|
175
|
+
| A knife, even lying flat on a chopping board next to a lemon | The object is the trigger, not the framing. Swap it — a cast-iron pan cleared instantly. |
|
|
176
|
+
|
|
177
|
+
**🚨 Refusals are PROBABILISTIC. Retry once before rewriting a word.** In the same session an identical prompt, identical reference, identical params was refused and then accepted on a straight re-fire. A rejected job returns no file and costs nothing, so a retry is free and a rewrite is not — rewriting first is how you end up changing four variables and learning nothing. **Only redesign after two or three refusals.**
|
|
178
|
+
|
|
179
|
+
**And change ONE thing at a time.** The eleven refusals above took far longer to diagnose than they should have because a reference swap and an opener rewrite shipped in the same call. Isolate on the prompt you actually want, so a pass leaves you with a usable asset instead of a data point.
|
|
180
|
+
|
|
181
|
+
## Filter regime
|
|
182
|
+
|
|
183
|
+
OpenAI moderate — a third regime distinct from Gemini (NB family) and ByteDance (Seedream). Real-face references pass more readily than Gemini; violence/brand rules are similar. `slates-content-policy` applies unchanged.
|