@slatesvideo/shared 0.7.2 → 0.7.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/clients/cloud.d.ts +4 -0
- package/dist/clients/cloud.js +11 -3
- package/dist/index.d.ts +1 -0
- package/dist/index.js +1 -0
- package/dist/manual/content.d.ts +1 -1
- package/dist/manual/content.js +1 -1
- package/dist/operations/index.d.ts +12 -13
- package/dist/operations/index.js +158 -133
- package/dist/operations/surface.d.ts +6 -2
- package/dist/operations/surface.js +29 -5
- package/dist/prompts/agent-doctrine.d.ts +4 -4
- package/dist/prompts/agent-doctrine.js +17 -28
- package/dist/prompts/guide-discovery.d.ts +23 -0
- package/dist/prompts/guide-discovery.js +39 -0
- package/dist/prompts/guide-retrieval.js +1 -1
- package/dist/prompts/model-capabilities.d.ts +8 -9
- package/dist/prompts/model-capabilities.js +11 -51
- package/dist/prompts/model-facts.d.ts +2 -2
- package/dist/prompts/model-facts.js +15 -26
- package/dist/prompts/partials.generated.js +6 -3
- package/dist/prompts/prompting-tips.d.ts +1 -1
- package/dist/prompts/prompting-tips.js +21 -63
- package/dist/prompts/search-terms.d.ts +3 -0
- package/dist/prompts/search-terms.js +24 -0
- package/dist/skills/content.js +36 -37
- package/dist/skills/metadata.d.ts +7 -0
- package/dist/skills/metadata.js +29 -0
- package/exports/slates-chatgpt-images/generated/SKILL.md +7 -1
- package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
- package/exports/slates-prompt-builder/generated/SKILL.md +28 -16
- package/exports/slates-prompt-builder/generated/reference-character.md +12 -13
- package/exports/slates-prompt-builder/generated/reference-content-policy.md +2 -2
- package/exports/slates-prompt-builder/generated/reference-gpt-image-2-5.md +191 -0
- package/exports/slates-prompt-builder/generated/reference-kling.md +32 -11
- package/exports/slates-prompt-builder/generated/reference-nano-banana.md +24 -6
- package/exports/slates-prompt-builder/generated/reference-omni-flash.md +65 -0
- package/exports/slates-prompt-builder/generated/reference-seedance-2-5.md +362 -0
- package/exports/slates-prompt-builder/generated/reference-seedance.md +34 -4
- package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +77 -23
- package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
- package/package.json +2 -1
- package/skills/_partials/blender-action-curves.md +24 -0
- package/skills/_partials/iteration-diagnosis.md +5 -0
- package/skills/_partials/model-routing.md +35 -0
- package/skills/_partials/seedance-25-timestamps.md +2 -2
- package/skills/_partials/still-gate.md +2 -2
- package/skills/_partials/thresholds.md +1 -1
- package/skills/slates-blocking-to-prompt.md +15 -13
- package/skills/slates-camera-language.md +45 -7
- package/skills/slates-character-identity.md +8 -6
- package/skills/slates-chatgpt-images.md +7 -1
- package/skills/slates-cinematic-look.md +1 -1
- package/skills/slates-content-policy.md +4 -6
- package/skills/slates-cost-discipline.md +18 -12
- package/skills/slates-dialogue-blocking.md +6 -6
- package/skills/slates-direct-response-ad.md +1 -1
- package/skills/slates-edit-and-iterate.md +12 -4
- package/skills/slates-model-selection.md +82 -90
- package/skills/slates-one-prompt-film.md +1 -1
- package/skills/slates-previs-blocking.md +44 -13
- package/skills/slates-project-organization.md +2 -2
- package/skills/slates-prompting-elevenlabs.md +4 -4
- package/skills/slates-prompting-flux-2-max.md +2 -3
- package/skills/slates-prompting-gpt-image-2-5.md +2 -2
- package/skills/slates-prompting-inworld-tts.md +174 -174
- package/skills/slates-prompting-kling-v3.md +11 -9
- package/skills/slates-prompting-lip-sync.md +15 -15
- package/skills/slates-prompting-ltx-2-5.md +5 -6
- package/skills/slates-prompting-minimax-h3.md +11 -11
- package/skills/slates-prompting-motion-transfer.md +8 -8
- package/skills/slates-prompting-nano-banana-2.md +8 -4
- package/skills/slates-prompting-omni-flash.md +9 -9
- package/skills/slates-prompting-seed-audio.md +24 -4
- package/skills/slates-prompting-seedance-2-5.md +40 -30
- package/skills/slates-prompting-seedance.md +4 -4
- package/skills/slates-prompting-seedream-5-lite.md +6 -6
- package/skills/slates-restyle-from-blocking.md +2 -2
- package/skills/slates-script-craft.md +1 -1
- package/skills/slates-shot-variety.md +1 -1
- package/skills/slates-storyboard-from-script.md +1 -1
- package/skills/slates-style-prompting.md +56 -54
- package/skills/slates-ugc-influencer-ad.md +1 -1
- package/skills/slates-vision-feedback-loop.md +118 -110
- package/skills/slates-prompting-veo-3.md +0 -224
|
@@ -1,110 +1,118 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: slates-vision-feedback-loop
|
|
3
|
-
description:
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Vision feedback loop — Slates utility skill
|
|
7
|
-
|
|
8
|
-
Slates returns generated images inline as base64. You see the actual pixels. Use that — don't trust prompt-following blindly.
|
|
9
|
-
|
|
10
|
-
## Asset codes are your shared vocabulary with the user
|
|
11
|
-
|
|
12
|
-
Every asset in Slates has a short stable code (e.g. `IMG-A12`, `VID-V3`, `AUD-S1`) and a label derived from its prompt (e.g. `Beach Sunset`). These are visible in the gallery as a corner badge on each thumbnail. **Always refer to assets by their code in chat** so the user can match what you're saying to a specific card in their gallery.
|
|
13
|
-
|
|
14
|
-
- ✅ "I'm using **IMG-A12 — Beach Sunset** as the first frame. The second-frame candidate **IMG-A15** has the right composition but warmer light — want me to use that one instead?"
|
|
15
|
-
- ❌ "I'm using the beach sunset image..." (user has four beach sunset variants — which one?)
|
|
16
|
-
- ❌ "I'm using asset `7a3f9e4b-...`" (UUIDs aren't readable; user can't match to a badge)
|
|
17
|
-
|
|
18
|
-
The code is the FORMAL reference. The label is human texture. Use both: `IMG-A12 — Beach Sunset`.
|
|
19
|
-
|
|
20
|
-
## Vision tools at your disposal
|
|
21
|
-
|
|
22
|
-
- `slates_get_asset_image` — pull one image into context. Returns its code+label.
|
|
23
|
-
- `slates_get_assets_batch` — pull up to 8 images in one call. Use when picking from a candidate set; cheaper than N individual fetches.
|
|
24
|
-
- `slates_get_asset_video_frames` — extract N keyframes (default 3) from a video and inline them as JPEGs.
|
|
25
|
-
|
|
26
|
-
## Pre-flight is automatic on the gen tools
|
|
27
|
-
|
|
28
|
-
`slates_generate_video
|
|
29
|
-
|
|
30
|
-
## 🔴 The still-gate — never animate a bad frame
|
|
31
|
-
|
|
32
|
-
<!-- @inject:still-gate -->
|
|
33
|
-
**
|
|
34
|
-
|
|
35
|
-
This is a
|
|
36
|
-
<!-- @end:still-gate -->
|
|
37
|
-
|
|
38
|
-
## The pattern
|
|
39
|
-
|
|
40
|
-
1. **Generate.** Call `slates_generate_image` with a prompt. The result is in your context as an image content block.
|
|
41
|
-
2. **Evaluate on TWO axes — they are different questions:**
|
|
42
|
-
- **Brief-conformance** — what did the user actually want? Are the elements right? Composition? Lighting? Subject identity?
|
|
43
|
-
- **Defects** — run the slop rubric below. *A frame can match the brief perfectly and still be slop that mushes the moment it moves.* Checking only the first axis is how a bad frame reaches an expensive video call.
|
|
44
|
-
3. **One of three outcomes:**
|
|
45
|
-
- **Right** → save it (bind to a frame, character slot, etc.) and move on.
|
|
46
|
-
- **Close, but adjustable** → refine with a specific delta, regenerate **once**.
|
|
47
|
-
- **Wrong direction** → ask the
|
|
48
|
-
|
|
49
|
-
## The defect rubric — five slop tells
|
|
50
|
-
|
|
51
|
-
| Tell | What it looks like | Why it matters downstream |
|
|
52
|
-
|---|---|---|
|
|
53
|
-
| **Light with no transitions** | Flat-black pits instead of a shadow ramp; light that stops rather than falls off | Transfers onto every character or object added into that plate later |
|
|
54
|
-
| **Broken-but-plausible objects** | Crates, railings, hardware, mechanisms you can *almost* read but that don't resolve | Turn to mush in motion, and the model multiplies them |
|
|
55
|
-
| **Local logic breaks** | An effect present in only part of the frame — rain scratching one corner, wet ground under one figure | The video model's physical logic breaks along with it |
|
|
56
|
-
| **Oily textures** | Soapy, licked-smooth surfaces that have lost their material identity | Reflections crawl in motion; the plate can't hold continuity |
|
|
57
|
-
| **Too perfect** | A soft light on the face that nothing in the scene could cast, the subject sharper and cleaner than everything around them, every region exposed to be readable, colour pushed warm and saturated | It reads as a subject pasted onto a location, and every shot built from the plate inherits the studio look. Fix it in words: `slates-cinematic-look` |
|
|
58
|
-
|
|
59
|
-
### Per-model accents — check the one you actually used
|
|
60
|
-
|
|
61
|
-
- **Nano Banana Pro** (`nano-banana-pro`) — ruler-straight symmetry, everything parallel and square, flat even light, pretty but staged/stock, textures reading as 3D render rather than photograph. **It hyperbolizes every edit**: ask for graffiti on one wall and the whole location gets tagged.
|
|
62
|
-
- **GPT Image** (`gpt-image-2-5-flare`, `gpt-image-2-5-sunburst`) — microcontrast to the ceiling, hard halos on every edge, no depth or bokeh, white balance pulled warm until the frame yellows, plastic licked-smooth materials. Worst tell: **one sickly texture pattern laid over the entire frame**. ⚠️ Catalogued on `gpt-image-2`, which 2.5 replaced on 2026-09-09 — an accent is a per-model observation, so treat this as a prior to check rather than a finding, and correct it here the first time a 2.5 frame disagrees.
|
|
63
|
-
|
|
64
|
-
> ⚠️ These are accents for **`nano-banana-pro`** and the **GPT Image** line specifically. `nano-banana-2` is a **different model** (Gemini 3.1 Flash Image vs NB Pro's Gemini 3 Pro Image) and we have **no evidence** about its accent. Do not inherit one — say nothing rather than warn about a failure mode you can't substantiate. That caution applies to the GPT Image entry above too: it was measured on `gpt-image-2`, not on either 2.5 seat.
|
|
65
|
-
|
|
66
|
-
## Where the fault lives — triage before you change anything
|
|
67
|
-
|
|
68
|
-
We say "one specific delta per regeneration" but that only helps once you know *which* variable to move. Diagnose first:
|
|
69
|
-
|
|
70
|
-
| Visible pattern | Diagnosis | Fix |
|
|
71
|
-
|---|---|---|
|
|
72
|
-
| The defect exists in the source asset, or stays tied to the same feature when the direction changes | **Source asset** | Fix the sheet / plate, not the prompt |
|
|
73
|
-
| Source is clean, and the defect changes when only the suspect motion clause changes | **Motion direction** | Fix the prompt |
|
|
74
|
-
| Controls conflict, or the failure follows neither variable | **Inconclusive** | Narrow the test — change less, not more |
|
|
75
|
-
|
|
76
|
-
**Review routes; it is not pass/fail.** Geography melts → fix the location. Identity drifts → fix the character sheet. Assets are sound but the action is wrong → fix the video direction. Wrong idea entirely → reopen the brief with the user.
|
|
77
|
-
|
|
78
|
-
**Correct the earliest broken handoff.** Polishing a downstream symptom hides the source and guarantees it resurfaces in the next shot built from the same asset.
|
|
79
|
-
|
|
80
|
-
## Baseline hygiene — isolate the variable you're testing
|
|
81
|
-
|
|
82
|
-
When the **character** is the question, keep the location out of it: test on a plate that already holds its own geometry, depth, materials, and light. **A broken plate gives every character failure a second plausible cause**, and you will spend re-rolls deciding which one you're looking at. The same applies in reverse — test a plate empty before you populate it.
|
|
83
|
-
|
|
84
|
-
## Refinement rules
|
|
85
|
-
|
|
86
|
-
- **One specific delta per regeneration.** Don't change five things at once — you won't know what helped.
|
|
87
|
-
- **
|
|
88
|
-
- On **Seedance**,
|
|
89
|
-
- **Exception — Omni Flash Edit.** Long prompts documentedly destroy its fidelity. There the rule inverts: one short instruction plus *"Keep everything else the same."*
|
|
90
|
-
- **Anchor with references.** If the result drifted from the user's intent, attach the *previous best* generation as a reference image alongside the original brief.
|
|
91
|
-
- **Use `slates_get_asset_image`** to pull a previously-generated image back into context if you need to compare against a fresh generation.
|
|
92
|
-
- **Use `slates_edit_image`** for surgical tweaks instead of full regeneration when ~90% of the image is right — `sourceAssetId` = the asset, `prompt` = the change only. Edits preserve composition and identity; full regen rolls the dice. Recipe: `slates-edit-and-iterate`.
|
|
93
|
-
|
|
94
|
-
## Cost discipline
|
|
95
|
-
|
|
96
|
-
- Track total credits spent across the loop. Surface to the user every 3 iterations.
|
|
97
|
-
-
|
|
98
|
-
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
-
|
|
1
|
+
---
|
|
2
|
+
name: slates-vision-feedback-loop
|
|
3
|
+
description: "Inspect generated media against the brief, diagnose defects and choose a targeted correction. Use during production or iteration; covers reference review, image defects and model-specific failure receipts."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Vision feedback loop — Slates utility skill
|
|
7
|
+
|
|
8
|
+
Slates returns generated images inline as base64. You see the actual pixels. Use that — don't trust prompt-following blindly.
|
|
9
|
+
|
|
10
|
+
## Asset codes are your shared vocabulary with the user
|
|
11
|
+
|
|
12
|
+
Every asset in Slates has a short stable code (e.g. `IMG-A12`, `VID-V3`, `AUD-S1`) and a label derived from its prompt (e.g. `Beach Sunset`). These are visible in the gallery as a corner badge on each thumbnail. **Always refer to assets by their code in chat** so the user can match what you're saying to a specific card in their gallery.
|
|
13
|
+
|
|
14
|
+
- ✅ "I'm using **IMG-A12 — Beach Sunset** as the first frame. The second-frame candidate **IMG-A15** has the right composition but warmer light — want me to use that one instead?"
|
|
15
|
+
- ❌ "I'm using the beach sunset image..." (user has four beach sunset variants — which one?)
|
|
16
|
+
- ❌ "I'm using asset `7a3f9e4b-...`" (UUIDs aren't readable; user can't match to a badge)
|
|
17
|
+
|
|
18
|
+
The code is the FORMAL reference. The label is human texture. Use both: `IMG-A12 — Beach Sunset`.
|
|
19
|
+
|
|
20
|
+
## Vision tools at your disposal
|
|
21
|
+
|
|
22
|
+
- `slates_get_asset_image` — pull one image into context. Returns its code+label.
|
|
23
|
+
- `slates_get_assets_batch` — pull up to 8 images in one call. Use when picking from a candidate set; cheaper than N individual fetches.
|
|
24
|
+
- `slates_get_asset_video_frames` — extract N keyframes (default 3) from a video and inline them as JPEGs. These sampled stills support appearance, framing and identity checks. They do not verify continuous motion, lip sync or sound. Use actual playback or an audio-capable host for those claims when available; otherwise report them unreviewed and retain the saved asset.
|
|
25
|
+
|
|
26
|
+
## Pre-flight is automatic on the gen tools
|
|
27
|
+
|
|
28
|
+
`slates_generate_video` and `slates_generate_image` show you their reference assets **inline** on the confirm response. You don't need to fetch them yourself, but you DO need to look at what comes back, revise the prompt if the references suggest a different motion/framing, and only then re-call with `confirm=true`.
|
|
29
|
+
|
|
30
|
+
## 🔴 The still-gate — never animate a bad frame
|
|
31
|
+
|
|
32
|
+
<!-- @inject:still-gate -->
|
|
33
|
+
**Inspect a start frame before animating it.** Repair a visible defect that would make the intended crop or performance unusable before spending on motion. A clean frame can be animated whenever the brief calls for movement; this check does not require an image stage for text-to-video.
|
|
34
|
+
|
|
35
|
+
This is a cost rule as well as craft: a premium video call can cost many times an image correction. Broken geometry can turn to mush, oily textures can crawl and malformed objects can fall apart in motion. Fix a known source defect at the source instead of buying a more expensive copy. Judge intentional stylisation against the brief, not a universal photoreal standard. Additional image or video requests still follow the existing generation authorization.
|
|
36
|
+
<!-- @end:still-gate -->
|
|
37
|
+
|
|
38
|
+
## The pattern
|
|
39
|
+
|
|
40
|
+
1. **Generate.** Call `slates_generate_image` with a prompt. The result is in your context as an image content block.
|
|
41
|
+
2. **Evaluate on TWO axes — they are different questions:**
|
|
42
|
+
- **Brief-conformance** — what did the user actually want? Are the elements right? Composition? Lighting? Subject identity?
|
|
43
|
+
- **Defects** — run the slop rubric below. *A frame can match the brief perfectly and still be slop that mushes the moment it moves.* Checking only the first axis is how a bad frame reaches an expensive video call.
|
|
44
|
+
3. **One of three outcomes:**
|
|
45
|
+
- **Right** → save it (bind to a frame, character slot, etc.) and move on.
|
|
46
|
+
- **Close, but adjustable** → refine with a specific delta, regenerate **once**.
|
|
47
|
+
- **Wrong direction** → diagnose the failed requirement. Refine within the supplied brief and existing authorization; ask only when the creative intent is unresolved or the next request needs fresh consent.
|
|
48
|
+
|
|
49
|
+
## The defect rubric — five slop tells
|
|
50
|
+
|
|
51
|
+
| Tell | What it looks like | Why it matters downstream |
|
|
52
|
+
|---|---|---|
|
|
53
|
+
| **Light with no transitions** | Flat-black pits instead of a shadow ramp; light that stops rather than falls off | Transfers onto every character or object added into that plate later |
|
|
54
|
+
| **Broken-but-plausible objects** | Crates, railings, hardware, mechanisms you can *almost* read but that don't resolve | Turn to mush in motion, and the model multiplies them |
|
|
55
|
+
| **Local logic breaks** | An effect present in only part of the frame — rain scratching one corner, wet ground under one figure | The video model's physical logic breaks along with it |
|
|
56
|
+
| **Oily textures** | Soapy, licked-smooth surfaces that have lost their material identity | Reflections crawl in motion; the plate can't hold continuity |
|
|
57
|
+
| **Too perfect** | A soft light on the face that nothing in the scene could cast, the subject sharper and cleaner than everything around them, every region exposed to be readable, colour pushed warm and saturated | It reads as a subject pasted onto a location, and every shot built from the plate inherits the studio look. Fix it in words: `slates-cinematic-look` |
|
|
58
|
+
|
|
59
|
+
### Per-model accents — check the one you actually used
|
|
60
|
+
|
|
61
|
+
- **Nano Banana Pro** (`nano-banana-pro`) — ruler-straight symmetry, everything parallel and square, flat even light, pretty but staged/stock, textures reading as 3D render rather than photograph. **It hyperbolizes every edit**: ask for graffiti on one wall and the whole location gets tagged.
|
|
62
|
+
- **GPT Image** (`gpt-image-2-5-flare`, `gpt-image-2-5-sunburst`) — microcontrast to the ceiling, hard halos on every edge, no depth or bokeh, white balance pulled warm until the frame yellows, plastic licked-smooth materials. Worst tell: **one sickly texture pattern laid over the entire frame**. ⚠️ Catalogued on `gpt-image-2`, which 2.5 replaced on 2026-09-09 — an accent is a per-model observation, so treat this as a prior to check rather than a finding, and correct it here the first time a 2.5 frame disagrees.
|
|
63
|
+
|
|
64
|
+
> ⚠️ These are accents for **`nano-banana-pro`** and the **GPT Image** line specifically. `nano-banana-2` is a **different model** (Gemini 3.1 Flash Image vs NB Pro's Gemini 3 Pro Image) and we have **no evidence** about its accent. Do not inherit one — say nothing rather than warn about a failure mode you can't substantiate. That caution applies to the GPT Image entry above too: it was measured on `gpt-image-2`, not on either 2.5 seat.
|
|
65
|
+
|
|
66
|
+
## Where the fault lives — triage before you change anything
|
|
67
|
+
|
|
68
|
+
We say "one specific delta per regeneration" but that only helps once you know *which* variable to move. Diagnose first:
|
|
69
|
+
|
|
70
|
+
| Visible pattern | Diagnosis | Fix |
|
|
71
|
+
|---|---|---|
|
|
72
|
+
| The defect exists in the source asset, or stays tied to the same feature when the direction changes | **Source asset** | Fix the sheet / plate, not the prompt |
|
|
73
|
+
| Source is clean, and the defect changes when only the suspect motion clause changes | **Motion direction** | Fix the prompt |
|
|
74
|
+
| Controls conflict, or the failure follows neither variable | **Inconclusive** | Narrow the test — change less, not more |
|
|
75
|
+
|
|
76
|
+
**Review routes; it is not pass/fail.** Geography melts → fix the location. Identity drifts → fix the character sheet. Assets are sound but the action is wrong → fix the video direction. Wrong idea entirely → reopen the brief with the user.
|
|
77
|
+
|
|
78
|
+
**Correct the earliest broken handoff.** Polishing a downstream symptom hides the source and guarantees it resurfaces in the next shot built from the same asset.
|
|
79
|
+
|
|
80
|
+
## Baseline hygiene — isolate the variable you're testing
|
|
81
|
+
|
|
82
|
+
When the **character** is the question, keep the location out of it: test on a plate that already holds its own geometry, depth, materials, and light. **A broken plate gives every character failure a second plausible cause**, and you will spend re-rolls deciding which one you're looking at. The same applies in reverse — test a plate empty before you populate it.
|
|
83
|
+
|
|
84
|
+
## Refinement rules
|
|
85
|
+
|
|
86
|
+
- **One specific delta per regeneration.** Don't change five things at once — you won't know what helped.
|
|
87
|
+
- **A fresh generation needs a complete coherent prompt.** Change one decision and retain the unchanged requirements so old and new clauses do not conflict. An edit request uses its model's change-only grammar instead; do not turn a surgical edit into a full scene re-description.
|
|
88
|
+
- On **Seedance**, retain the selected model's structure: shot numbers for 2.0, whole-second timing where used for 2.5. See the matching model guide.
|
|
89
|
+
- **Exception — Omni Flash Edit.** Long prompts documentedly destroy its fidelity. There the rule inverts: one short instruction plus *"Keep everything else the same."*
|
|
90
|
+
- **Anchor with references.** If the result drifted from the user's intent, attach the *previous best* generation as a reference image alongside the original brief.
|
|
91
|
+
- **Use `slates_get_asset_image`** to pull a previously-generated image back into context if you need to compare against a fresh generation.
|
|
92
|
+
- **Use `slates_edit_image`** for surgical tweaks instead of full regeneration when ~90% of the image is right — `sourceAssetId` = the asset, `prompt` = the change only. Edits preserve composition and identity; full regen rolls the dice. Recipe: `slates-edit-and-iterate`.
|
|
93
|
+
|
|
94
|
+
## Cost discipline
|
|
95
|
+
|
|
96
|
+
- Track total credits spent across the loop. Surface to the user every 3 iterations.
|
|
97
|
+
- Use the repeated-failure checkpoint below; do not keep submitting an unchanged failed request.
|
|
98
|
+
- **Follow the existing consent for every attempt.** An approved enumerated batch covers its listed calls. A retry or changed input outside that batch needs a new quote and the applicable confirmation; a timeout requires a status check before another submission.
|
|
99
|
+
|
|
100
|
+
<!-- @inject:iteration-diagnosis -->
|
|
101
|
+
## Diagnose repeated failures
|
|
102
|
+
|
|
103
|
+
After three failed attempts at the same requirement, pause unchanged re-rolls and diagnose the source reference, prompt structure, model fit and tool result. Three is a review checkpoint, not a universal limit or proof that the seed cannot matter. Preserve the attempts and name what each test changed.
|
|
104
|
+
|
|
105
|
+
Continue autonomously when the brief is clear, a specific correction is supported and the next request is already authorized. Hand control back when taste or intent cannot be inferred, the next request needs fresh consent, or the available tool cannot meet the requirement. A failed roll never authorizes an additional charge. Follow the existing batch and per-request cost policy.
|
|
106
|
+
<!-- @end:iteration-diagnosis -->
|
|
107
|
+
|
|
108
|
+
## When to break the loop
|
|
109
|
+
|
|
110
|
+
- The user said "good enough" or "ship it." Stop iterating.
|
|
111
|
+
- Repeated attempts show no progress. Diagnose the source, prompt or model before spending again; apply the scoped retry rule above.
|
|
112
|
+
- The user changes brief mid-loop. Treat it as a new brief, not a continuation.
|
|
113
|
+
|
|
114
|
+
## Voice when narrating to the user
|
|
115
|
+
|
|
116
|
+
Tight, observational, no editorializing.
|
|
117
|
+
- ✅ "Frame 2 has the wrong lighting direction — back-lit instead of side. Regenerating with side light."
|
|
118
|
+
- ❌ "I notice that the lighting in frame 2 isn't quite what we were going for. I'll go ahead and try again with a different approach."
|
|
@@ -1,224 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: slates-prompting-veo-3
|
|
3
|
-
description: How to prompt Veo 3.1 (Google). Read before calling slates_generate_video with veo-3.1-fast or veo-3.1-standard. Veo is a NICHE pick, never the default (route per slates-model-selection — Kling is the general default, Seedance the premium tier) — reach for it only when native synchronized audio must generate WITH the video in one gen. 16:9 or 9:16, 4/6/8s. Different cinematography formula than Seedance/Kling. (no subtitles) is mandatory after every dialogue line.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Veo 3.1 — prompting
|
|
7
|
-
|
|
8
|
-
<!-- @card:start -->
|
|
9
|
-
<!-- slates-only -->
|
|
10
|
-
<!-- MACHINE-READ. Everything between the @card markers is extracted by
|
|
11
|
-
src/prompts/craft-cards.ts and returned on every cost estimate for this
|
|
12
|
-
model, so it is the ONE piece of positive craft guidance the agent cannot
|
|
13
|
-
skip. Measured 2026-08-30: a fact inlined where it cannot be skipped moved
|
|
14
|
-
compliance 0/8 to 30/32; the same guidance behind a fetch moved nothing.
|
|
15
|
-
Keep it under 2,400 characters (the build fails above that) and keep the
|
|
16
|
-
rationale, the receipts and the worked examples in the body below. -->
|
|
17
|
-
<!-- /slates-only -->
|
|
18
|
-
**Card — Veo 3.1.** Google's formula, in order: `[Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]`. Sweet spot 50-150 words; the official benchmark is about 50.
|
|
19
|
-
|
|
20
|
-
**The five levers**
|
|
21
|
-
1. **Open with the cinematography** — `Medium shot`, `low angle`, `aerial`, `dolly in`, `rack focus`, `vertigo effect`. Veo reads the first clause as the camera.
|
|
22
|
-
2. **Texture words counter the AI-plastic look** — `fine skin pores`, `visible fabric weave`, `subtle contrast, no gloss or sharpening`. Name materials concretely: `charcoal cotton hoodie`, `matte concrete`, `silk lapel`.
|
|
23
|
-
3. **Weight verbs stop floaty motion** — `trudges`, `drops heavily`, and a ground contact: `boots crunch on gravel`.
|
|
24
|
-
4. **Terse voice direction only** — `says in a weary voice`, `whispers`, `mutters`. Veo is far less responsive to long voice blocks than Kling.
|
|
25
|
-
5. **Always include an ambience line.** Without one the mix feels dead. `Soft office ambience.` `Wind on the open ridge.` And SFX always carries a cause: `SFX: thunder cracks in the distance`, never `SFX: thunder`.
|
|
26
|
-
|
|
27
|
-
**Examples**
|
|
28
|
-
- `Medium shot, a tired founder rubbing her temples in front of a bulky monitor in a cluttered office late at night. Harsh fluorescent overheads and the green glow of the screen. Fine skin pores, visible fabric weave on a charcoal cotton hoodie. Soft office ambience, a fan hum. Retro, slightly grainy.`
|
|
29
|
-
- `Low angle, a farrier trudges across a wet yard carrying a shoeing box, boots crunching on gravel. Overcast north light, matte concrete and wet steel. Wind and distant livestock. He says in a weary voice, "One more and we're done." (no subtitles).`
|
|
30
|
-
|
|
31
|
-
**Hard constraint:** `(no subtitles)` after EVERY dialogue line you do not want burned in as text. Negatives are NOUNS, not instructions — `wall, frame`, never `no walls`. And do not cross syntaxes: Seedance's `single continuous take` suppresses Veo's cuts, and Veo timestamps in a Seedance prompt cause drift.
|
|
32
|
-
<!-- @card:end -->
|
|
33
|
-
|
|
34
|
-
<!-- @banned:start -->
|
|
35
|
-
<!-- slates-only -->
|
|
36
|
-
<!-- MACHINE-READ. Every `backticked` token between the @banned markers is
|
|
37
|
-
extracted by src/prompts/banned-tokens.ts and returned on this model's cost
|
|
38
|
-
estimate, and every submitted prompt is matched against it. Keep entries
|
|
39
|
-
backticked and prose outside the backticks. -->
|
|
40
|
-
<!-- /slates-only -->
|
|
41
|
-
**Never use** (each one is spelled out above):
|
|
42
|
-
- `single continuous take` — Seedance's phrase; in a Veo prompt it suppresses the cuts you asked for
|
|
43
|
-
- `no subtitles` is REQUIRED after dialogue, but instruction-shaped negatives are not: `no walls`, `no man-made structures`, `don't show` — negatives are NOUNS here
|
|
44
|
-
- `SFX: thunder` and any label-only effect — every effect carries a cause and a distance
|
|
45
|
-
<!-- @banned:end -->
|
|
46
|
-
|
|
47
|
-
Google DeepMind's video model. Two tiers: `veo-3.1-fast` (cheaper, quick) and `veo-3.1-standard` (higher quality). 4k variants exist for both (4K video requires Slates Pro).
|
|
48
|
-
|
|
49
|
-
**Native single-shot duration: 4, 6, or 8 seconds** — and **8s only** at 1080p or 4K, or whenever you attach reference images (that endpoint is 8s-fixed). 4s and 6s exist at 720p, text-to-video or single-start-frame only. Longer durations require chaining clips via Extend / last-frame reuse — quality degrades if naively requested past 8s in a single generation. Aspect ratio: **16:9 or 9:16** on the route Slates uses. `slates_generate_video` REFUSES anything outside these before submit and names the legal set — nothing is silently ignored or downgraded.
|
|
50
|
-
|
|
51
|
-
Native synchronized audio at 48kHz: dialogue, SFX, ambient — generated WITH video, not added after.
|
|
52
|
-
|
|
53
|
-
## Official Google formula
|
|
54
|
-
|
|
55
|
-
```
|
|
56
|
-
[Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]
|
|
57
|
-
```
|
|
58
|
-
|
|
59
|
-
Sweet spot length: 50-150 words. Cloud's official benchmark is ~50 words.
|
|
60
|
-
|
|
61
|
-
Verbatim official benchmark:
|
|
62
|
-
> "Medium shot, a tired corporate worker, rubbing his temples in exhaustion, in front of a bulky 1980s computer in a cluttered office late at night. The scene is lit by the harsh fluorescent overhead lights and the green glow of the monochrome monitor. Retro aesthetic, shot as if on 1980s color film, slightly grainy."
|
|
63
|
-
|
|
64
|
-
## Cinematography vocabulary (Vertex AI docs)
|
|
65
|
-
|
|
66
|
-
**Lenses:** wide-angle, telephoto, fisheye, anamorphic, 35mm, 85mm, shallow/deep depth of field
|
|
67
|
-
|
|
68
|
-
**Lighting:** Rembrandt lighting, volumetric lighting, backlighting, golden hour glow, lens flare, rack focus, **vertigo effect** (dolly zoom)
|
|
69
|
-
|
|
70
|
-
**Camera moves:** dolly (in/out), truck (left/right), pan, tilt, crane, aerial/drone, handheld, whip pan, arc shot, zoom
|
|
71
|
-
|
|
72
|
-
## Texture-realism phrases (counter the AI-plastic look)
|
|
73
|
-
|
|
74
|
-
```
|
|
75
|
-
fine skin pores · visible fabric weave · subtle contrast, no gloss or sharpening
|
|
76
|
-
```
|
|
77
|
-
|
|
78
|
-
Specify materials concretely: `charcoal cotton hoodie`, `matte concrete`, `silk lapel`. Generic "smooth, beautiful" rendering is the failure mode you're avoiding.
|
|
79
|
-
|
|
80
|
-
## Dialogue — `(no subtitles)` is mandatory
|
|
81
|
-
|
|
82
|
-
Every dialogue line you don't want burned in as text overlay needs `(no subtitles)`. Verbatim from the founder talking-head benchmark:
|
|
83
|
-
|
|
84
|
-
```
|
|
85
|
-
The founder says, "This update cuts setup time in half, helping teams get started faster." (no subtitles).
|
|
86
|
-
```
|
|
87
|
-
|
|
88
|
-
Without this, Veo will overlay subtitle text on top of your generation.
|
|
89
|
-
|
|
90
|
-
## Voice direction — keep it terse
|
|
91
|
-
|
|
92
|
-
Veo is less responsive to long voice-direction blocks than Kling. Use brief modifiers:
|
|
93
|
-
|
|
94
|
-
```
|
|
95
|
-
says in a weary voice
|
|
96
|
-
whispers
|
|
97
|
-
shouts
|
|
98
|
-
mutters
|
|
99
|
-
```
|
|
100
|
-
|
|
101
|
-
Multi-character: handles 2-3 speakers natively. Past 3, sync degrades — use first-frame/last-frame chaining for 4+.
|
|
102
|
-
|
|
103
|
-
## SFX with cause
|
|
104
|
-
|
|
105
|
-
```
|
|
106
|
-
✅ SFX: thunder cracks in the distance
|
|
107
|
-
❌ SFX: thunder
|
|
108
|
-
```
|
|
109
|
-
|
|
110
|
-
Always specify direction or distance.
|
|
111
|
-
|
|
112
|
-
## Ambient is mandatory
|
|
113
|
-
|
|
114
|
-
Always include an ambience line per scene. Without it, the audio mix feels dead.
|
|
115
|
-
|
|
116
|
-
```
|
|
117
|
-
Soft office ambience.
|
|
118
|
-
Wind on the open ridge.
|
|
119
|
-
Distant city hum.
|
|
120
|
-
```
|
|
121
|
-
|
|
122
|
-
## First-frame + last-frame workflow (Veo's strength)
|
|
123
|
-
|
|
124
|
-
1. Generate start frame (Gemini 2.5 Flash Image is the recommended pair — Slates' Nano Banana 2 works)
|
|
125
|
-
2. Generate end frame
|
|
126
|
-
3. Animate with both frames as anchors
|
|
127
|
-
|
|
128
|
-
**Motion-Lock hack:** Keep ~60% of the same background pixels between start and end frames. Prevents latent drift across the clip.
|
|
129
|
-
|
|
130
|
-
Verbatim arc-shot example:
|
|
131
|
-
> "The camera performs a smooth 180-degree arc shot, starting with the front-facing view of the singer and circling around her to seamlessly end on the POV shot from behind her on stage. The singer sings 'when you look me in the eyes, I can see a million stars.'"
|
|
132
|
-
|
|
133
|
-
## Ingredients-to-Video (multiple references)
|
|
134
|
-
|
|
135
|
-
Verbatim example:
|
|
136
|
-
> "Using the provided images for the detective, the woman, and the office setting, create a medium shot of the detective behind his desk. He looks up at the woman and says in a weary voice, 'Of all the offices in this town, you had to walk into mine.'"
|
|
137
|
-
|
|
138
|
-
## Reference discipline (character / environment refs)
|
|
139
|
-
|
|
140
|
-
<!-- @inject:references-read-literally -->
|
|
141
|
-
> **The general law: the model reads a reference literally.**
|
|
142
|
-
> A reference image is not a suggestion. Whatever is baked into it — lighting, medium, texture, symmetry, competing identities — is read as a **property of the subject** and reproduced downstream. A baked rim light tints every shot made from that sheet. A sheet that looks like a 3D game render gets animated like game footage. Two competing renderings of one face get averaged into a third face.
|
|
143
|
-
|
|
144
|
-
Every reference rule below is a corollary of that one sentence, which is why "prep the reference" beats "prompt around the reference" every time:
|
|
145
|
-
|
|
146
|
-
- **Flat, plain identity refs** — because scene lighting in the sheet becomes scene lighting in the output (Slates' own receipt: a studio-lit sheet produced a subject that looked green-screen-pasted in front of mountains).
|
|
147
|
-
- **One authoritative rendering per subject** — because the model cannot tell which panel is the real one. ByteDance documents this failure directly: multi-view character assets "confuse the model's character recognition, causing it to generate duplicate characters of the same appearance."
|
|
148
|
-
- **No 3D-game-render look in a reference** — the model recognizes the render mood and inherits its motion character, so the *animation* comes out looking like game footage. This is not a taste rule; it is the same literal-reading mechanism applied to the temporal layer.
|
|
149
|
-
- **Break perfect symmetry** — mirrored faces and dead-square framing read as synthetic, and the model preserves that reading rather than correcting it.
|
|
150
|
-
|
|
151
|
-
**What this means in practice:** when output is wrong in a way that tracks the *subject* rather than the *scene* — the lighting is wrong the same way in every shot, the face drifts, the material looks synthetic everywhere — fix the reference, not the prompt. Prompting around a baked-in property is the expensive way to lose.
|
|
152
|
-
<!-- @end:references-read-literally -->
|
|
153
|
-
|
|
154
|
-
<!-- @inject:reference-rules-core -->
|
|
155
|
-
Identity = a few flat-lit neutral angles; one reference per role, named inline; 2-4 refs not 12; describe environments instead of feeding a grid.
|
|
156
|
-
|
|
157
|
-
1. **2-4 strong references beat both extremes.** Not 1 (warps toward itself), not 12 (averages worse). Start with 2-3 focused refs — each one adds context AND another variable to balance.
|
|
158
|
-
2. **One reference per ROLE, named in the prompt** — identity / style-grade / environment. The model does **not** infer a reference's role from its position in the list; the inline name carries it. Same-role competitors drift (two "identity" refs of different people blend into a third face). Slates resolves `@mentions` / `#tags` into numbered citations. You can also bind references directly in scene prose, naming what each image supplies.
|
|
159
|
-
3. **One identity sheet per character, named inline.** A character's identity is a single asset (dominant portrait + body panels), so attach that one asset rather than a pile of views: **fewer competing renderings of a face is better, because the model cannot tell which one is authoritative and averages them.** Slates cites it as `Marcus (image 1)`. **Do NOT hand-write a "Reference Image Instructions" block or role essays** ("use for identity, ignore the outfit, render a neutral expression") — that drags the sheet's studio lighting and wardrobe into a scene that asked for neither. The prompt leads; the user's words own wardrobe, expression, lighting, and action.
|
|
160
|
-
4. **Flat-light identity refs.** Prep identity references with flat, even, shadowless lighting on a plain neutral background. A studio-lit or scene-lit character sheet bleeds its lighting into every generation — the failure looks like the subject was green-screen-pasted in front of the location. Reference prep beats prompting here.
|
|
161
|
-
5. **Environment: describe it, don't feed a grid.** Default to describing the location in words and let the model build a space that fits the shot. Reserve an environment reference for a mandatory exact-match, and then use ONE clean establishing image with natural ambient light that reads as the location's real light — never a multi-panel grid fed whole.
|
|
162
|
-
6. **Grids: explore, don't input.** Use grids to explore compositions cheaply, then pick a cell. Never feed a grid back in as a reference — the cells share a split detail budget and were generated jointly, so their flaws propagate.
|
|
163
|
-
7. **Reuse the same refs across every shot** in a sequence. Lock a set and keep it; swapping references mid-sequence causes drift, because the model adapts each reference to the current prompt rather than copying it.
|
|
164
|
-
8. **Legible in-shot text → bake it into a still start frame, never trust text-to-video.** Have an image model render the text, then animate from that locked frame. Video models smear type.
|
|
165
|
-
9. **Working from existing media — describe ONLY what changes.** The source already carries its composition, motion, timing, and performance; re-describing them fights the model. Narrate the delta. (Video lane: restyle your own clip while keeping the performance; delayed-VFX on "video one"; marker-object insertion; video-as-reference for a series.)
|
|
166
|
-
10. **Style transforms happen in natural language.** By default the source's artistic medium and visual style are inherited. To change it, add a plain-text instruction ("anime → real person"). There are no preset pickers, and there is no style slider.
|
|
167
|
-
<!-- @end:reference-rules-core -->
|
|
168
|
-
|
|
169
|
-
### For Veo specifically
|
|
170
|
-
|
|
171
|
-
- **Veo's idiom for rule 2 is plain-English role naming in the sentence itself** — *"Using the provided images for the detective, the woman, and the office setting, create a medium shot of…"* (see Ingredients-to-Video above). The role rides in the noun phrase, not in a separate label block.
|
|
172
|
-
- **Rule 8 has a second reason to matter here:** Veo bakes subtitle text into the frame unless every dialogue line carries `(no subtitles)`. Text you did not ask for is the failure mode, not just text you did.
|
|
173
|
-
|
|
174
|
-
## Negative prompting — nouns, not instructions
|
|
175
|
-
|
|
176
|
-
Veo has a `negativePrompt` field. **Verbatim Vertex AI rule:**
|
|
177
|
-
> "Describe unwanted elements as nouns rather than instructions. Use 'wall, frame' instead of 'no walls' or 'don't show walls.'"
|
|
178
|
-
|
|
179
|
-
Inline: positive reframing in the body too.
|
|
180
|
-
- ✅ `"a desolate landscape with no buildings or roads"`
|
|
181
|
-
- ❌ `"no man-made structures"`
|
|
182
|
-
|
|
183
|
-
## Common failure modes + fixes
|
|
184
|
-
|
|
185
|
-
| Failure | Fix |
|
|
186
|
-
|---|---|
|
|
187
|
-
| Subject identity shifts mid-clip | Front-load identity at prompt start; use material cues (`charcoal canvas`, `cotton`, `silk`) to stabilize |
|
|
188
|
-
| Floaty / weightless motion | Weight verbs (`trudges`, `drops heavily`), ground contact (`boots crunch on gravel`) |
|
|
189
|
-
| AI-plastic look | `fine skin pores`, `visible fabric weave`, `subtle contrast` |
|
|
190
|
-
| Subtitles baked into video | `(no subtitles)` after every dialogue line |
|
|
191
|
-
| Rushed dialogue | Lines fit one natural breath in 8s |
|
|
192
|
-
| Mismatched ambience | Always include an ambience line |
|
|
193
|
-
| Warped geometry | `photorealistic stability` |
|
|
194
|
-
|
|
195
|
-
## Timestamp shot syntax (for chained / multi-beat scenes)
|
|
196
|
-
|
|
197
|
-
Veo accepts `[00:00-00:02]` brackets for timed sequences within an 8s clip. **Do NOT cross syntaxes** — Veo timestamps in a Seedance prompt cause subject drift; Seedance "single continuous take" in a Veo prompt suppresses cuts.
|
|
198
|
-
|
|
199
|
-
Verbatim multi-beat:
|
|
200
|
-
> "[00:00-00:02] Medium shot from behind a young female explorer with a leather satchel and messy brown hair in a ponytail, as she pushes aside a large jungle vine to reveal a hidden path.
|
|
201
|
-
> [00:02-00:04] Reverse shot of the explorer's freckled face, her expression filled with awe as she gazes upon ancient, moss-covered ruins. SFX: The rustle of dense leaves, distant exotic bird calls.
|
|
202
|
-
> [00:04-00:06] Tracking shot following the explorer as she steps into the clearing and runs her hand over the intricate carvings on a crumbling stone wall.
|
|
203
|
-
> [00:06-00:08] Wide, high-angle crane shot, revealing the lone explorer standing small in the center of the vast, forgotten temple complex, half-swallowed by the jungle. SFX: A swelling, gentle orchestral score begins to play."
|
|
204
|
-
|
|
205
|
-
## Benchmark prompt — founder talking head (full)
|
|
206
|
-
|
|
207
|
-
> "Camera locked at eye level, medium close-up on a 35mm lens: a startup founder in his late 30s with short black hair and light stubble, wearing a charcoal cotton hoodie, speaking directly to camera, leaning slightly forward as he speaks, lifting one hand to emphasize a point, then relaxing back to neutral, in a quiet office during late afternoon, with blurred monitors glowing faintly in the background, lit by soft daylight from a side window with gentle fill on the opposite side and natural falloff across his face. Style: fine skin pores, visible fabric weave, subtle contrast, no gloss or sharpening. Audio: The founder says, 'This update cuts setup time in half, helping teams get started faster.' (no subtitles). Soft office ambience."
|
|
208
|
-
|
|
209
|
-
## Pre-flight: references arrive inline, refer by code
|
|
210
|
-
|
|
211
|
-
When you call `slates_generate_video` with `firstFrameAssetId` / `lastFrameAssetId` / `ingredientAssetIds`, the first call returns those references **inline as image content blocks** alongside a cost estimate and `requires_confirm: true`. Veo's strongest move is first-frame + last-frame; the pre-flight is where you confirm the two frames actually anchor the motion you wrote. Revise the prompt before `confirm=true` if needed.
|
|
212
|
-
|
|
213
|
-
When talking to the user about the gen, refer to each reference by its short code: `IMG-A12 — Founder Headshot`. The user sees that code as a badge on the gallery thumbnail, so they can match what you're saying to what they're looking at.
|
|
214
|
-
|
|
215
|
-
- ✅ "I'm anchoring on **IMG-A12** as the open shot and **IMG-A18** as the close — the 180° arc lands on her looking offscreen left."
|
|
216
|
-
- ❌ "I'm using two of the founder shots..." (which two? They have six.)
|
|
217
|
-
|
|
218
|
-
## Sources
|
|
219
|
-
|
|
220
|
-
- [Google Cloud — Ultimate Prompting Guide for Veo 3.1](https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1)
|
|
221
|
-
- [Google DeepMind — Veo Prompt Guide](https://deepmind.google/models/veo/prompt-guide/)
|
|
222
|
-
- [Google Cloud Docs — Vertex AI Video Generation Prompt Guide](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide)
|
|
223
|
-
- [Atlas Cloud — Veo 3.1 Master Guide](https://www.atlascloud.ai/blog/guides/google-veo-3-1-guide-master-image-to-video-ai-with-native-sound-and-4k-realism)
|
|
224
|
-
- [Invideo — Veo 3.1 Prompt Guide](https://invideo.io/blog/google-veo-prompt-guide/)
|