@slatesvideo/shared 0.6.10 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (85) hide show
  1. package/README.md +1 -1
  2. package/dist/auth.js +2 -2
  3. package/dist/clients/cloud.js +1 -1
  4. package/dist/index.d.ts +2 -1
  5. package/dist/index.js +4 -1
  6. package/dist/manual/content.d.ts +1 -1
  7. package/dist/manual/content.js +1 -1
  8. package/dist/operations/index.d.ts +817 -16
  9. package/dist/operations/index.js +1423 -372
  10. package/dist/operations/surface.d.ts +3 -1
  11. package/dist/operations/surface.js +37 -10
  12. package/dist/prompts/ad-presets.d.ts +77 -0
  13. package/dist/prompts/ad-presets.js +43 -0
  14. package/dist/prompts/agent-doctrine.js +5 -4
  15. package/dist/prompts/banned-tokens.d.ts +4 -29
  16. package/dist/prompts/banned-tokens.js +29 -204
  17. package/dist/prompts/craft-cards.js +2 -2
  18. package/dist/prompts/generation-policy.d.ts +41 -0
  19. package/dist/prompts/generation-policy.js +53 -0
  20. package/dist/prompts/guide-retrieval.d.ts +9 -0
  21. package/dist/prompts/guide-retrieval.js +53 -0
  22. package/dist/prompts/index.d.ts +1 -0
  23. package/dist/prompts/index.js +1 -0
  24. package/dist/prompts/model-capabilities.d.ts +18 -1
  25. package/dist/prompts/model-capabilities.js +72 -19
  26. package/dist/prompts/model-facts.d.ts +59 -0
  27. package/dist/prompts/model-facts.js +121 -15
  28. package/dist/prompts/partials.generated.js +8 -2
  29. package/dist/prompts/prompting-tips.d.ts +1 -1
  30. package/dist/prompts/prompting-tips.js +63 -18
  31. package/dist/prompts/reference-composer.d.ts +2 -0
  32. package/dist/prompts/reference-composer.js +51 -50
  33. package/dist/prompts/script-document.d.ts +165 -0
  34. package/dist/prompts/script-document.js +11 -0
  35. package/dist/prompts/shot-grammar.d.ts +4 -4
  36. package/dist/prompts/shot-grammar.js +3 -3
  37. package/dist/prompts/shot-spec.d.ts +13 -0
  38. package/dist/prompts/shot-spec.js +23 -5
  39. package/dist/skills/content.js +26 -23
  40. package/dist/update-check.d.ts +22 -0
  41. package/dist/update-check.js +109 -0
  42. package/exports/slates-chatgpt-images/generated/SKILL.md +107 -0
  43. package/exports/slates-chatgpt-images/generated/slates-chatgpt-images.skill +0 -0
  44. package/exports/slates-prompt-builder/generated/SKILL.md +3 -3
  45. package/exports/slates-prompt-builder/generated/reference-character.md +9 -1
  46. package/exports/slates-prompt-builder/generated/reference-kling.md +3 -3
  47. package/exports/slates-prompt-builder/generated/reference-nano-banana.md +22 -10
  48. package/exports/slates-prompt-builder/generated/reference-seedance.md +4 -4
  49. package/exports/slates-prompt-builder/generated/slates-prompt-builder-manifest.json +17 -17
  50. package/exports/slates-prompt-builder/generated/slates-prompt-builder.skill +0 -0
  51. package/package.json +10 -4
  52. package/skills/_partials/cinematic-card.md +8 -0
  53. package/skills/_partials/cinematic-routes-short.md +2 -0
  54. package/skills/_partials/cinematic-tips-short.md +2 -0
  55. package/skills/_partials/decision-log.md +1 -13
  56. package/skills/_partials/image-defaults.md +11 -0
  57. package/skills/_partials/lens-video-split.md +1 -0
  58. package/skills/_partials/reference-rules-core.md +1 -1
  59. package/skills/_partials/sheet-tool-defaults.md +6 -0
  60. package/skills/slates-character-identity.md +9 -1
  61. package/skills/slates-chatgpt-images.md +107 -0
  62. package/skills/slates-cinematic-look.md +237 -0
  63. package/skills/slates-cost-discipline.md +18 -12
  64. package/skills/slates-direct-response-ad.md +13 -53
  65. package/skills/slates-edit-and-iterate.md +1 -1
  66. package/skills/slates-model-selection.md +139 -133
  67. package/skills/slates-one-prompt-film.md +38 -95
  68. package/skills/slates-project-organization.md +7 -3
  69. package/skills/slates-prompting-flux-2-max.md +15 -4
  70. package/skills/slates-prompting-gpt-image-2-5.md +41 -28
  71. package/skills/slates-prompting-kling-v3.md +3 -3
  72. package/skills/slates-prompting-lip-sync.md +1 -1
  73. package/skills/slates-prompting-minimax-h3.md +30 -17
  74. package/skills/slates-prompting-motion-transfer.md +1 -1
  75. package/skills/slates-prompting-nano-banana-2.md +24 -11
  76. package/skills/slates-prompting-seedance-2-5.md +12 -12
  77. package/skills/slates-prompting-seedance.md +5 -5
  78. package/skills/slates-prompting-seedream-5-lite.md +14 -3
  79. package/skills/slates-prompting-veo-3.md +1 -1
  80. package/skills/slates-script-craft.md +45 -0
  81. package/skills/slates-shot-variety.md +11 -40
  82. package/skills/slates-storyboard-from-script.md +14 -66
  83. package/skills/slates-style-prompting.md +54 -54
  84. package/skills/slates-ugc-influencer-ad.md +32 -309
  85. package/skills/slates-vision-feedback-loop.md +2 -1
@@ -1,133 +1,139 @@
1
- ---
2
- name: slates-model-selection
3
- description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Kling 3.0 is the general-purpose video default; Seedance 2.0 is the premium tier for anything where physics, effects, or scale remotely matter; Seedance 2.5 is a SECOND SEAT beside 2.0 (30s takes, 30 references and timestamp control, but no 4K and dearer at every shared resolution — never an upgrade); MiniMax H3 is the AUTHORED-AUDIO seat (three directable sound layers in one pass, declared reference relationships, 480p-4K) with MiniMax H3 Max beside it as a faster 768p-capped premium with omni-references; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 or 9:16, 4/6/8s) and never the default.
4
- ---
5
-
6
- # Model selection — the routing doctrine
7
-
8
- Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.
9
-
10
- ## 🔑 The meta-rule — above the table
11
-
12
- The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2.5 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:
13
-
14
- > **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.
15
- >
16
- > **Inspect the output at its intended crop.** A frame that holds up as a thumbnail can fall apart at the size it will actually be watched. For a location, look at atmosphere, material texture, and anchor objects; for a character, identity, skin, pose, and gradients.
17
- >
18
- > **Choose the model that PROVES that requirement** and leaves only failures you can afford to rerun or mask.
19
- >
20
- > **When the roster changes, repeat the evidence test.** Do not carry today's ranking forward on reputation.
21
-
22
- ## Video routing
23
-
24
- | Job | Model | Why |
25
- |---|---|---|
26
- | **General-purpose — the default for most shots** | **Kling 3.0 std** | Cost-effective workhorse. Strong image-to-video: preserves identity, layout, and text from the start frame. 16:9 / 9:16 / 1:1, 3–15s. |
27
- | Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |
28
- | Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
29
- | **Anything with remotely important physics** — effects, destruction, water/fire/smoke/cloth, creature motion, scale, complex simultaneous action | **Seedance 2.0** | The premium tier. Physics and effects are its whole edge; up to 9 ingredient refs, first+last frame, native 4K (4K video is Pro-only). |
30
- | The premium hero shot a piece hangs on | Seedance 2.0 | Spend where it shows. |
31
- | **One take longer than 15 seconds**, or a shot needing more than 9 image references, or an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | A SECOND SEAT beside 2.0, never an upgrade: 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs, and the only Seedance seat that **acts on timestamps** (rules in `slates-prompting-seedance-2-5` § Timestamps) — 480p / 720p / 1080p, **no 4K**, and **dearer than 2.0 at every shared resolution** (720p $0.231/s vs $0.15/s, +54%). If you want 4K, or the same resolution cheaper, stay on 2.0. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) LENGTH is the price dial, not resolution — a 30s 720p face gen is 489 credits and a 30s 1080p faceless gen is 614, against a 1,000-credit welcome grant. Quote before any take over ~10s. |
32
- | **The SOUND has to be directed, not just present** — a specific line delivered a specific way, scene sound that has to sit under it, and score that must stay out of the characters' world | **MiniMax H3** | The only seat where audio is authored in three separate layers in ONE pass (synchronised events in the body, ambience in a soundscape section, audience-only score in its own) rather than toggled on. 5–15s, 480p / 768p / 2K / 4K, 24fps, 32kHz stereo, 11 languages. Rules in `slates-prompting-minimax-h3`. |
33
- | **A reference has to keep a DECLARED amount of itself** — especially moving one subject's characteristic onto a *different* subject | **MiniMax H3** | The only seat that understands a stated retention relationship (kept whole / kept in part / transferred onto another subject / loose echo). 9 images + 3 video + 3 audio, 12 files total. 🚨 The first 5 reference images are free and every one after that costs 4 credits — pass `referenceImages` to `slates_estimate_generation_cost` before a reference-heavy job. |
34
- | **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p/768p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at 768p. It still animates a start frame and an end frame — image-to-video is one of the two things it is for — and since 2026-09-09 it takes the full omni-reference set too (9 images + 3 video + 3 audio), so the seats now differ on ladder and price rather than on what they accept. Never the default; never reach for it to save money. |
35
- | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | Narrow, and now narrower: if the sound needs DIRECTING rather than merely existing, MiniMax H3 is the better seat. |
36
-
37
- ### Named Seedance escalation triggers
38
-
39
- "Physics matter" is an abstract category and it under-fires. These are the beats Seedance is **observably** good at — if the shot contains one, escalate without deliberating:
40
-
41
- - **Real-time → slow-motion contrast.** The signature beat; nearly every strong clip rides it.
42
- - **The camera moving while debris, meteors, sparks or particles crash around the subject.** Distinctly a feature of this model, not just a thing it survives.
43
- - **Massive scale that has to read as genuinely huge** — not "a big thing", a thing whose size is the point of the shot.
44
- - **One continuous unbroken take.**
45
-
46
- Concrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).
47
-
48
- ## Video EDIT routing (changing an existing clip)
49
-
50
- | Job | Tool | Why |
51
- |---|---|---|
52
- | **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + "Keep everything else the same"; long prompts destroy it (see below). |
53
- | **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
54
- | **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
55
- | Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. (2.5's relocate lane reaches 1080p too as of 2026-08-24, at $0.2457/s of combined input+output.) Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
56
- | **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p/1080p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |
57
- | AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
58
-
59
- - **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
60
- - **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.
61
- - **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
62
- - Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
63
-
64
- ## Motion Transfer & Lip Sync routing (Kling-only tools)
65
-
66
- Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that bolts motion or lip movement onto a finished source as a dedicated post-process.
67
-
68
- | Job | Tool | Why |
69
- |---|---|---|
70
- | Motion retarget onto a still character | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |
71
- | Re-voice a clip, or animate a still portrait | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |
72
-
73
- **Want the Seedance version of either?** It is not a switch on these tools — it is a normal `slates_generate_video` on `seedance-2` with the clip attached as a **video reference** and the motion or dialogue written into the prompt ("the character from image 1 performs the exact motion from video 1"). That routes to the same endpoint the tool would have called, with the prompt visible and editable instead of ghost-written. Single-pass conditioning genuinely beats post-hoc retargeting on fast choreography, contact, cloth and hair — and it carries native audio — so escalate there whenever fidelity matters.
74
-
75
- - Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane.
76
- - Faces on that route go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person (premium realface pricing).
77
-
78
- **Rules:**
79
-
80
- - **Default video = Kling 3.0 std.** Escalate to Seedance the moment the shot has physics/effects weight or is the hero moment — and say why in the plan ("physics-heavy, routing to Seedance").
81
- - **Veo is never the default.** 16:9 or 9:16 only, 4/6/8s only (and 8s only at 1080p/4K, or with reference images), and it is not the quality pick — treat it as a single-purpose tool for native-synced-audio shots. If audio can be added after (Kling lip-sync, edit stage), prefer Kling or Seedance + audio in post.
82
- - **9:16 vertical → Kling or Seedance by preference**, not by necessity: Veo does take 9:16 on the route Slates uses. Route away from it because it is the niche seat, not because it can't.
83
- - **Ratios and durations are enforced before submit.** `slates_generate_video` validates the aspect ratio, resolution and duration against the model you picked and refuses out-of-set values with the legal list — it will not silently ignore or downgrade them. The authoritative per-model sets are in the op's own param descriptions, which are generated from the capability SSOT; prefer those over any list written in prose here.
84
- - **Image-to-video from an NB2 start frame** (the standard pipeline) → Kling by default, Seedance when the motion is physics-heavy. Not Veo.
85
- - **User names a model explicitly → use it.** But if it's a mismatch for the job (crazy physics on Kling std, a 30s take on anything but Seedance 2.5, 4K on Seedance 2.5 which has none), say so in one line and offer the right route before generating.
86
-
87
- ## Image routing
88
-
89
- **Video models (Kling, Seedance, Veo) cannot generate standalone images — ever.** A "premium hero reference image" is still an image job: it routes to an image model below, never to Seedance.
90
-
91
- - **Default: Nano Banana 2** — strongest reference HANDLING (14 refs; GPT Image now takes more, at 16, but Banana is still the one that holds many subjects coherently), best legible text, the standard start-frame generator.
92
- - **NB2 Lite** — the fast/draft seat: ~half NB2's price, ~2.7× faster, 1K only. Route iteration volume and drafts here; finals go back to NB2 full (2K/4K).
93
- - **Nano Banana Pro** — the hero-frame/typography ceiling (~2× NB2). NB2 ≈ 95% of Pro; escalate only when spatial composition, cinematic lighting/skin, fine typography-in-scene, or deep multi-element frames must be perfect. Up to 14 refs — feed it a full subject library.
94
- - **GPT Image 2.5** — two seats, `gpt-image-2-5-flare` and `gpt-image-2-5-sunburst`, **same price**. Readable text / panels / UI king: character sheets, shot grids, diagrams, text-bearing panels. **Also the photoreal front-runner (Eric, 2026-08-24)** — it beat both Nano Banana rails head-to-head on skin realism, which is why the AI-influencer ad lane generates every plate on this line. **The seat split is SPEED vs QUALITY, not generate vs edit** (OpenAI's own rule): Flare is the small, fast model with quality *comparable to* GPT Image 2 — drafts, exploration, volume; Sunburst is OpenAI's *most capable* image model, higher quality than GPT Image 2, deliberately slower — finals, hero frames, photoreal, and multi-reference edits, where its lead is widest. **Explore on Flare, finish on Sunburst.** Five quality tiers, cheapest first — `low` (layout checks only), `medium` (drafts), **`high` (the default)**, `xhigh`, `max` (the top). Uneven: `max` is 4× `high`, `xhigh` only ~1.8× it. **16 reference images**, the schema ceiling. **Transparent backgrounds** via `backgroundMode` — free, and the only image family that offers them.
95
-
96
- 🚨 **The tier names moved when 2.5 replaced GPT Image 2, and the strings did not.** GPT Image 2's `medium` is 2.5's `high`; its `high` is 2.5's `max` — same money, one rung of renaming. The 2026-08-24 photoreal result was measured at GPT Image 2 `high`, so **the tier that reproduces it is `max`**. Nobody has re-run it on 2.5; the ranking is inherited, not re-measured.
97
- - **FLUX.2 Max** — photoreal texture, hex-color binding, typography, less censored.
98
- - **Seedream 5 Lite** — uncensored + any-resolution flat price; volume exploration when the Gemini filter is in the way.
99
-
100
- **Split rule of thumb:** readable text / panels / UI → GPT Image 2.5 (Flare to explore, Sunburst to finish); **photoreal people, finals and hero frames → Sunburst at `max`** — the 2026-08-24 result was measured at GPT Image 2's `high`, which is `max` here, and Flare only *matches* GPT Image 2 while Sunburst exceeds it; multi-reference edits where several references must all survive into one frame → Sunburst; edit-heavy work → the Banana line; drafts → GPT Image 2.5 Flare at `medium`, which now undercuts NB2 Lite on both price and resolution; uncensored or odd resolutions → Seedream/FLUX.
101
-
102
- ⚠️ **This line said the opposite until 2026-08-24** — it sent photoreal *away* from GPT Image on reputation, which is the exact failure § The meta-rule above warns about. Re-run the evidence test when the roster moves. It moved again on 2026-09-09, and the ranking was carried across rather than re-measured — exactly what the meta-rule says not to trust. Treat it as a starting hypothesis for 2.5, not a receipt. **The seat choice above is likewise reasoned from OpenAI's positioning, not measured:** run Flare-`max` against Sunburst-`max` on one plate and write the answer into `slates-prompting-gpt-image-2-5`.
103
-
104
- ## Audio routing
105
-
106
- **Image and video models cannot generate standalone audio, and neither audio model can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.
107
-
108
- | Job | Model | Why |
109
- |---|---|---|
110
- | **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines inside a scene | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. The continuity-bed workhorse; dialogue is performed inside the room, not cast. |
111
- | **One named voice saying one line** — a character's own voice, a narrator, a clean VO to lip-sync against | **Inworld TTS-2** (`inworld-tts-2`) | The prompt IS the words, spoken verbatim and billed per character. Voice = the character's clip (cloned for the take), a description, or a preset. No room tone — mix it on the timeline. |
112
- | **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control and a real loop mode. |
113
-
114
- **There is no music model.** A song is imported (Slates reads audio files and puts them on the timeline), not generated. A line that has to be spoken in a SPECIFIC voice is generated on Inworld TTS-2 and lip-synced against; a line that belongs to a scene is performed by Seed Audio inside it.
115
-
116
- ### Named audio escalation triggers
117
-
118
- - **"It needs to sound like a place"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.
119
- - **"Read this line"** → Seed Audio, with the line in quotes inside the scene sentence. Re-roll until the take is right, then lip-sync against it.
120
- - **"That needs a thump right there"** → Sound Effects, with the duration set to roughly the length of the event.
121
- - **"Give it a track"** → there is no music generation. Say so and offer to lay an imported track on an audio track.
122
-
123
- **Rules:**
124
-
125
- - **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.
126
- - **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.
127
- - **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — and remember those extra seconds are billed on both surfaces.
128
- - **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.
129
- - Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`.
130
-
131
- ## Cost is a tiebreaker, not the router
132
-
133
- Route by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.
1
+ ---
2
+ name: slates-model-selection
3
+ description: Which model to pick for a given job — the routing doctrine. Read BEFORE choosing any video or image model, before quoting a plan, and before defaulting anywhere. Seedance 2.5 is the DEFAULT video model (Eric, 2026-09-13: the best in the world — physics, effects, scale, 30s takes, 30 references, timestamps); Seedance 2.0 is the 4K seat and the cheaper one at every shared resolution; Kling 3.0 is the cost-effective seat for performances, start-frame animation and lip-sync; MiniMax H3 is the AUTHORED-AUDIO seat (three directable sound layers in one pass, declared reference relationships, 480p-4K) with MiniMax H3 Max beside it as a faster premium with omni-references and MiniMax H3 Max Turbo as its half-price, frames-only sibling; Veo 3.1 is a narrow niche (native synced audio in one gen, 16:9 or 9:16, 4/6/8s) and never the default.
4
+ ---
5
+
6
+ # Model selection — the routing doctrine
7
+
8
+ Pick the model FIRST, deliberately, before writing a prompt or quoting a plan. Model routing is a core part of the intelligence users are paying for: the agent knows what each model is good at and which ones underperform for a job — defaulting to the wrong model burns the user's credits on a weaker result.
9
+
10
+ ## 🔑 The meta-rule — above the table
11
+
12
+ The tables below are a snapshot. This roster churns constantly (NB2 Lite, Omni Flash, Seedream 5 Lite, GPT Image 2.5 all landed recently) — **a table rots; a rule doesn't.** When the tables and this rule disagree, or when a model appears that the tables don't cover, run the rule:
13
+
14
+ > **Name ONE must-preserve requirement for the shot.** Not a vibe — the single thing that, if it breaks, makes the shot unusable: this face stays this face · the fluid behaves like fluid · the text stays legible · the take stays one unbroken move.
15
+ >
16
+ > **Inspect the output at its intended crop.** A frame that holds up as a thumbnail can fall apart at the size it will actually be watched. For a location, look at atmosphere, material texture, and anchor objects; for a character, identity, skin, pose, and gradients.
17
+ >
18
+ > **Choose the model that PROVES that requirement** and leaves only failures you can afford to rerun or mask.
19
+ >
20
+ > **When the roster changes, repeat the evidence test.** Do not carry today's ranking forward on reputation.
21
+
22
+ ## Video routing
23
+
24
+ | Job | Model | Why |
25
+ |---|---|---|
26
+ | **General-purpose — the default for most shots** | **Seedance 2.5** | The strongest seat in the catalogue: physics, effects, scale and hero shots, 4–30s in one take, 30 image + 10 video + 10 audio references, audio-only refs, and the only Seedance seat that acts on timestamps. 480p / 720p / 1080p, no 4K. LENGTH is the price dial — quote any take over ~10s. |
27
+ | **Cost matters and the shot is a performance or a start-frame animation** | **Kling 3.0 std** | Cost-effective workhorse. Strong image-to-video: preserves identity, layout, and text from the start frame. 16:9 / 9:16 / 1:1, 3–15s. |
28
+ | Higher visual polish, no physics demands | Kling 3.0 pro | Mid-price fidelity bump on the same strengths. |
29
+ | Multi-character dialogue / audio co-generation | Kling 3.0 omni | Dialogue syntax, voice direction, language codes, `@element` refs. |
30
+ | **4K delivery**, or the same resolution cheaper than 2.5 | **Seedance 2.0** | The only Seedance with native 4K (4K video is Pro-only) and cheaper than 2.5 at every shared resolution (720p $0.15/s vs $0.231/s). Same physics and effects strengths; 15s takes, 15 references, no timestamps. |
31
+ | **One take longer than 15 seconds**, more than 15 references, an AUDIO-ONLY reference, or **beats that have to land at a named second** | **Seedance 2.5** | Only 2.5 does these (rules in `slates-prompting-seedance-2-5` § Timestamps); it is the default anyway. 🚨 Two live hazards: (a) with references attached, the words *add / remove / replace / change / extend / continue* make it reclassify the request as a video EDIT and fail AFTER the job queues — describe the finished frame, or use `seedance-2.5-edit`; (b) LENGTH is the price dial, not resolution — a 30s 720p face gen is 489 credits and a 30s 1080p faceless gen is 853, against a 1,000-credit welcome grant. Quote before any take over ~10s. |
32
+ | **The SOUND has to be directed, not just present** — a specific line delivered a specific way, scene sound that has to sit under it, and score that must stay out of the characters' world | **MiniMax H3** | The only seat where audio is authored in three separate layers in ONE pass (synchronised events in the body, ambience in a soundscape section, audience-only score in its own) rather than toggled on. 5–15s, 480p / 768p / 2K / 4K, 24fps, 32kHz stereo, 11 languages. Rules in `slates-prompting-minimax-h3`. |
33
+ | **A reference has to keep a DECLARED amount of itself** — especially moving one subject's characteristic onto a *different* subject | **MiniMax H3** | The only seat that understands a stated retention relationship (kept whole / kept in part / transferred onto another subject / loose echo). 9 images + 3 video + 3 audio, 12 files total. 🚨 The first 5 reference images are free and every one after that costs 4 credits — pass `referenceImages` to `slates_estimate_generation_cost` before a reference-heavy job. |
34
+ | **Turnaround is the requirement** on a text-to-video or start-frame shot at 480p to 1080p | **MiniMax H3 Max** | fal's self-hosted post-train of H3. **Measured 2026-08-27: a 5s 768p clip finished in 4.8s against 57s on base H3 — about 12x faster**, same prompt, queue to file. When turnaround is the requirement this is not a marginal win. 🚨 It is the PREMIUM seat, not a cheap H3 — $0.080/s at 768p against base H3's $0.060/s, 33% more, and it tops out at a 1080p refinement of its 768p render. It still animates a start frame and an end frame — image-to-video is one of the two things it is for — and since 2026-09-09 it takes the full omni-reference set too (9 images + 3 video + 3 audio), so the seats now differ on ladder and price rather than on what they accept. Never the default; never reach for it to save money. |
35
+ | **Drafts and volume** on a text-to-video or start-frame shot, where the credit budget binds and no reference is needed | **MiniMax H3 Max Turbo** | A second fal post-train of H3 with Max's ladder at **half Max's rate at every tier** ($0.040/s at 768p). It takes a start frame and an end frame but has **no reference endpoint**: a shot that needs references goes to H3 Max or base H3. Its 1080p, like Max's, is a refinement of the native 768p render. Re-run the keeper on a hero seat. |
36
+ | Native synchronized audio (dialogue + SFX generated WITH the video in one gen), 16:9, ≤8s | Veo 3.1 | Narrow, and now narrower: if the sound needs DIRECTING rather than merely existing, MiniMax H3 is the better seat. |
37
+
38
+ ### Named Seedance escalation triggers
39
+
40
+ "Physics matter" is an abstract category and it under-fires. These are the beats Seedance is **observably** good at — if the shot contains one, escalate without deliberating:
41
+
42
+ - **Real-time → slow-motion contrast.** The signature beat; nearly every strong clip rides it.
43
+ - **The camera moving while debris, meteors, sparks or particles crash around the subject.** Distinctly a feature of this model, not just a thing it survives.
44
+ - **Massive scale that has to read as genuinely huge** — not "a big thing", a thing whose size is the point of the shot.
45
+ - **One continuous unbroken take.**
46
+
47
+ Concrete beats route better than an abstract category. Cost stays a tiebreaker, never the router (see below).
48
+
49
+ ## Video EDIT routing (changing an existing clip)
50
+
51
+ | Job | Tool | Why |
52
+ |---|---|---|
53
+ | **Footage-synced VFX on real footage** — add/remove an effect, prop, or lighting change while the take stays the take (incl. talking heads) | **Omni Flash Edit** (`slates_edit_video`, `omni-flash-edit`) | **The edit-fidelity winner** (head-to-head receipt 2026-07-09, WITH a short prompt): lip movement held perfectly, audio near-identical, effect landed and released on cue — where Kling missed an action beat and drifted lips. Prompt-only, 3–10s clips, 720p out, ~6.4 cr/s (cheapest). Quirk: occasional tail jitter / doubled final speech beat — trim the tail on the timeline. Fidelity is EARNED by prompt discipline: one short line + "Keep everything else the same"; long prompts destroy it (see below). |
54
+ | **Identity swap needing reference images** — put @marcus into the clip, lock a style from refs | **Kling O3 Edit** (`slates_edit_video`) | The only edit engine that takes element/style reference images (frontal + angles lock identity). ~19¢/s. |
55
+ | **Spoken words must be bit-exact** (VO, legal copy, music) | **Kling O3 Edit** with `keepAudio` (default true) — or segment-splice | Kling keeps the ORIGINAL audio track verbatim — but re-synthesizes the video, so lips can drift slightly against it (7/09 receipt). Omni Flash regenerates audio (voice editing unsupported): on the 7/09 receipt it came back near-identical with perfect lips, but "near-identical" is not a guarantee. Zero-risk path for critical audio: segment-splice — edit only the non-talking seconds and keep the original track under the cut. |
56
+ | Style-transfer-heavy re-imagining, full relocate of the scene, or edit quality worth a premium at 1080p+ | Seedance edit/relocate (`videoReferenceAssetId` on `slates_generate_video`) | Seedance's strength is transfer intensity; it re-generates rather than surgically edits. Head-to-head receipt 2026-07-09 (photoreal-insert job, same clip): at 720p it LOST to Omni Flash edit on result while costing ~3× (vref bills input+output seconds; face-lane rates when people are in frame). Route here for its strengths or at 1080p/4K where its ceiling is higher — never as the cheap default. (2.5's relocate lane reaches 1080p too as of 2026-08-24, at $0.3412/s of combined input+output.) Takes long descriptive prompts fine (no Omni-style hard-fail on timing phrasing). |
57
+ | **A clip LONGER THAN 15 SECONDS** | **Seedance 2.5 Edit** (`slates_edit_video`, `seedance-2.5-edit`) | The only edit engine that takes a 4–30s clip — length is the whole reason to route here. 480p/720p/1080p out, native audio, prompt + clip only (no reference images). Output length AND aspect ratio follow the source, so the billed key is the ceiled source length; an edit bills roughly DOUBLE a plain 2.5 generation of the same length because every provider charges an edit on input + output seconds. Set `seedanceFace: true` when a face is visible — the faceless provider blocks faces outright. No consented-real-face route for editing. Inside 15s, choose on fidelity instead. |
58
+ | AI-edit the user's OWN footage | Omni Flash Edit (3–10s), Kling O3 Edit (3–15s, 720–3840px) or Seedance 2.5 Edit (4–30s) | Both take any MP4/MOV — not just Slates gens. Phone footage MUST be rotation-normalized first (players honor the rotation flag; models don't — raw portrait phone clips come back SIDEWAYS). |
59
+
60
+ - **Edit before re-roll.** A re-roll gambles away the parts the user already likes; an edit changes only what the prompt names. Quote the edit first when a clip is mostly right.
61
+ - **Ship via segment-splice.** Every edit model re-synthesizes the whole clip, so fidelity risk scales with clip length. For real deliverables: trim out ONLY the seconds where the change happens, edit that segment, splice it back over the original on the timeline with the ORIGINAL audio underneath. Most of the final video stays the untouched original — that's how the polished split-screen demos going around actually work, plus gesture-only beats with voiceover laid over in post.
62
+ - **One change per pass, short prompts.** On Omni Flash this is documented law ("overly descriptive prompts can lead to unintended changes" — long identity-lock preambles make drift WORSE, receipt 7/09); on Kling multi-beat instructions get dropped. Chain passes instead.
63
+ - Edited clips are themselves editable clips — chain passes; lineage links each output to its parent.
64
+
65
+ ## Motion Transfer & Lip Sync routing (Kling-only tools)
66
+
67
+ Both tools are **Kling-only**. Every entry in them is a real Kling endpoint that bolts motion or lip movement onto a finished source as a dedicated post-process.
68
+
69
+ | Job | Tool | Why |
70
+ |---|---|---|
71
+ | Motion retarget onto a still character | Kling MC std/pro (`slates_generate_motion_transfer`) | Structured skeleton/depth retarget, ~32–42 credits / 5s, takes up to 30s driving clips. |
72
+ | Re-voice a clip, or animate a still portrait | Kling lip-sync / avatar (`slates_generate_lip_sync`) | ~4–29 credits / 5s blocks. |
73
+
74
+ **Want the Seedance version of either?** It is not a switch on these tools — it is a normal `slates_generate_video` on `seedance-2` with the clip attached as a **video reference** and the motion or dialogue written into the prompt ("the character from image 1 performs the exact motion from video 1"). That routes to the same endpoint the tool would have called, with the prompt visible and editable instead of ghost-written. Single-pass conditioning genuinely beats post-hoc retargeting on fast choreography, contact, cloth and hair — and it carries native audio — so escalate there whenever fidelity matters.
75
+
76
+ - Seedance video-reference gens bill COMBINED input+output seconds (`seedance-2*-vref-*` keys) — pass the clip duration and quote before confirming. Driving clips must be 2–15s on Seedance 2.0 and up to 30s on 2.5; past that it is Kling MC's lane. On Seedance 2.5's AI-face route (EvoLink) the input side counts as at least the output's length: max(input, output) + output.
77
+ - Faces on that route go through the normal cascade: `seedanceFace` for a character, `[REAL_FACE_DETECTED]` → `seedanceRealFace` + `realFaceConsent` for a real person (premium realface pricing).
78
+
79
+ **Rules:**
80
+
81
+ - **Default video = Seedance 2.5.** Route to Seedance 2.0 for 4K or when the same resolution must be cheaper, and to Kling 3.0 std when the budget matters and the shot is a performance or a start-frame animation — and say why in the plan ("4K delivery, routing to 2.0"; "budget dialogue shot, routing to Kling").
82
+ - **Veo is never the default.** 16:9 or 9:16 only, 4/6/8s only (and 8s only at 1080p/4K, or with reference images), and it is not the quality pick — treat it as a single-purpose tool for native-synced-audio shots. If audio can be added after (Kling lip-sync, edit stage), prefer Kling or Seedance + audio in post.
83
+ - **9:16 vertical → Kling or Seedance by preference**, not by necessity: Veo does take 9:16 on the route Slates uses. Route away from it because it is the niche seat, not because it can't.
84
+ - **Ratios and durations are enforced before submit.** `slates_generate_video` validates the aspect ratio, resolution and duration against the model you picked and refuses out-of-set values with the legal list — it will not silently ignore or downgrade them. The authoritative per-model sets are in the op's own param descriptions, which are generated from the capability SSOT; prefer those over any list written in prose here.
85
+ - **Image-to-video from an NB2 start frame** (the standard pipeline) → Seedance 2.5 by default, Kling when the budget matters and the motion is a performance. Not Veo.
86
+ - **User names a model explicitly → use it.** But if it's a mismatch for the job (crazy physics on Kling std, a 30s take on anything but Seedance 2.5, 4K on Seedance 2.5 which has none), say so in one line and offer the right route before generating.
87
+
88
+ ## Image routing
89
+
90
+ **Video models (Kling, Seedance, Veo) cannot generate standalone images — ever.** A "premium hero reference image" is still an image job: it routes to an image model below, never to Seedance.
91
+
92
+ <!-- @inject:image-defaults -->
93
+ **Image default:** gpt-image-2-5-sunburst, quality `high`, 3k. User overrides take priority. Without a project, generation uses the headless Nano Banana 2 seat.
94
+
95
+ | Model | Default resolution |
96
+ |---|---|
97
+ | nano-banana-2 | 2k |
98
+ | nano-banana-2-lite | 1k |
99
+ | nano-banana-pro | 2k |
100
+ | gpt-image-2-5-flare | 2k |
101
+ | gpt-image-2-5-sunburst | 3k |
102
+ | flux-2-max | 1k |
103
+ | seedream-5-lite | 2k |
104
+ <!-- @end:image-defaults -->
105
+
106
+ Use `slates_estimate_generation_cost` for the selected model's current price and craft card. Routing reasons live in the model facts returned by `slates_list_available_models`; use the model's guide for its particular strengths and limits. Choose a different seat when the brief supplies a reason, such as speed, supported output shape, or an edit that failed on the default.
107
+
108
+ **Historical photoreal receipt:** the 2026-08-24 comparison favored GPT Image 2 on one skin-realism task at its old high tier. That is evidence about that comparison, not proof that 2.5 requires its most expensive tier. Raise quality only to address a specific observed shortfall and compare at the delivery crop.
109
+
110
+ ## Audio routing
111
+
112
+ **Image and video models cannot generate standalone audio, and neither audio model can generate images or video.** A shot that needs synced audio generated WITH the picture is still a video job (Kling omni / Veo / Omni Flash / Seedance all carry native audio); the models below produce audio *as its own asset*, to lay on the timeline.
113
+
114
+ | Job | Model | Why |
115
+ |---|---|---|
116
+ | **Default — a whole audio scene in one pass**: room tone, ambience beds, crowds, nature, layered dialogue + effects, spoken lines inside a scene | **Seed Audio 1.0** (`seed-audio`) | One plain sentence in, a complete scene out. The continuity-bed workhorse; dialogue is performed inside the room, not cast. |
117
+ | **One named voice saying one line** — a character's own voice, a narrator, a clean VO to lip-sync against | **Inworld TTS-2** (`inworld-tts-2`) | The prompt IS the words, spoken verbatim and billed per character. Voice = the character's clip (cloned for the take), a description, or a preset. No room tone — mix it on the timeline. |
118
+ | **One effect that lands on a known frame**, or a seamless loop | **Sound Effects v2** (`eleven-sfx`) | The only surface with an exact duration control and a real loop mode. |
119
+
120
+ **There is no music model.** A song is imported (Slates reads audio files and puts them on the timeline), not generated. A line that has to be spoken in a SPECIFIC voice is generated on Inworld TTS-2 and lip-synced against; a line that belongs to a scene is performed by Seed Audio inside it.
121
+
122
+ ### Named audio escalation triggers
123
+
124
+ - **"It needs to sound like a place"** → Seed Audio. Three separate SFX generations layered on the timeline is the wrong shape and costs more.
125
+ - **"Read this line"** → Seed Audio, with the line in quotes inside the scene sentence. Re-roll until the take is right, then lip-sync against it.
126
+ - **"That needs a thump right there"** → Sound Effects, with the duration set to roughly the length of the event.
127
+ - **"Give it a track"** → there is no music generation. Say so and offer to lay an imported track on an audio track.
128
+
129
+ **Rules:**
130
+
131
+ - **🚨 Seed Audio has NO duration parameter.** Length comes from the prompt text, so Slates writes the requested duration into the prompt and **bills what you asked for**. Choose the duration deliberately and never write a second, different length into the sentence. Full doctrine: `slates-prompting-seed-audio`.
132
+ - **Kling's audio syntax does not transfer.** `SFX:` / `Ambient noise:` / `Background music:` prefixes are Kling 3.0 *video* prompt syntax. Seed Audio reads them as literal words and the result degrades.
133
+ - **Beds outlast the cut.** Always ask for more seconds than the clip needs so the edit has fade handles — and remember those extra seconds are billed on both surfaces.
134
+ - **Audio inside the video vs audio as an asset.** If the sound must be locked to what happens on screen, generate it with the video (Kling omni / Seedance / Omni Flash / Veo). If it needs to be moved, trimmed, re-used, or layered, generate it here and drop it on an audio track.
135
+ - Per-model prompting: `slates-prompting-seed-audio`, `slates-prompting-elevenlabs`.
136
+
137
+ ## Cost is a tiebreaker, not the router
138
+
139
+ Route by capability first, then pick the cheapest tier that serves the job (per `slates-cost-discipline`). Never pick a model because its per-second price looked lowest — a cheap clip that has to be regenerated on the right model costs more than routing correctly once.
@@ -1,95 +1,38 @@
1
- ---
2
- name: slates-one-prompt-film
3
- description: Use when the user gives ONE idea and wants a finished video out the other end — "make me a video about X", "turn this idea into an ad", "make a short film from this". The full pipeline: script, project, characters, storyboard, frame images, video generation, timeline assembly, MP4 export. This is the master recipe; the other Slates skills are its sub-steps.
4
- ---
5
-
6
- # One prompt → finished film — Slates master pipeline
7
-
8
- The user gives an idea. You hand back an MP4 on disk. Everything in between is yours, with exactly TWO mandatory user checkpoints: the creative plan, and ONE aggregated cost approval.
9
-
10
- ## The pipeline
11
-
12
- ### 1. Script the beats
13
- Turn the idea into a beat-level script: 4-10 shots, each with subject, action, setting, camera, and duration (4-8s per shot). Surface it as a tight table. Get the user's nod on the plan, format (aspect ratio — 16:9 vs 9:16 decides everything downstream), and rough budget appetite before touching any op.
14
-
15
- 🚨 **Before you fire the set, read its variety counts.** `slates_list_shots` returns the distribution with every listing — shot sizes, camera moves, durations, and any bucket repeating three or more times in a row. Read the table as a COLUMN, not as rows: if push-in is the plurality or every row says wide, the batch is wrong before a credit is spent. The craft is `slates-shot-variety`.
16
-
17
- **Surface a decision log with the plan.**
18
-
19
- <!-- @inject:decision-log -->
20
- When you surface the plan, include a short **decision log** — one line per decision *you* made that the user did not specify **and that no row already records**:
21
-
22
- ```
23
- source phrase or declared default → what you wrote → what it resolves
24
- "in a diner" → warm, and the light is the reason → why the anchor was chosen, not what it is
25
- (no time of day) → late afternoon, low warm key → default; say the word and it changes
26
- ```
27
-
28
- 🚨 **Keep it to what is NOT already data — and almost everything now IS.** A Shot holds the references and their roles, the model, every param, the shot size, the camera, the prop, the action and the spoken line, and `slates_list_shots` reads the whole board back in order with its variety counts. Narrating any of those is retelling a row the user can open. **Write the Shot, and let the log carry only the judgement no field holds** — why this world, why this light, why this register.
29
-
30
- **Hard rule: never silently add weather, props, style, or camera movement.** Four of those are now FIELDS: put the value on the Shot (`prop`, `camera`, `shotSize`, `action`) so the user can read and change it, and put the *reason* in the log only when you invented it rather than being told it. The rule has not softened — it moved from narration into data, which is stronger, because a field can be corrected and a sentence in chat cannot.
31
-
32
- > ❌ **Do NOT turn this into a question gate.** Clarifying questions before optimizing directly fight the locked fast-path rule: *if intent is clear, generate immediately with sane defaults, don't ask questions; only ask for production intent, and batch every question into one message.* Log the decisions, then go. The log is an **output**, not an interrogation — surfaced alongside the plan, never as a separate ceremony, and never as a reason to wait.
33
- <!-- @end:decision-log -->
34
-
35
- A 4-10 shot script is where you invent the most on the user's behalf — time of day, wardrobe, weather, lens feel, camera moves the brief never mentioned. The log is what makes those visible while they are still free to change.
36
-
37
- ### 2. Set up the project
38
- - `slates_create_project` named for the piece.
39
- - Recurring character? Build it properly — `slates_create_character` + the `slates-character-identity` recipe — so every frame references the same identity.
40
- - Recurring location? `slates_create_environment`.
41
- - One-off shots don't need character/environment records; skip the ceremony.
42
-
43
- ### 3. Storyboard skeleton and the Shots (no generation yet)
44
- - `slates_create_storyboard`, `slates_add_scene` per script scene.
45
- - `slates_create_shot` per beat — the prompt, the model, the params and the references, with the roles they carry. **A Shot needs no image**, so the entire film exists as rows before anything is paid for.
46
- - `slates_get_shot` reads one back COMPOSED: the prompt the model will actually receive, its numbered references, and its exact quote. Audit your own work there — you cannot approve something the request will not contain.
47
- - Structure first, spend second — the user catches script problems on the free skeleton, not on burned credits.
48
-
49
- ### 4. ONE aggregated cost approval — then hands-off
50
- The Shots ARE the quote. `slates_generate_from_shots` without `confirm` returns one itemised total for the set plus the largest single item — no hand arithmetic, no `slates_estimate_generation_cost` per call:
51
-
52
- > Plan: 6 frames at 1k 16:9 + 5 × 8s Kling 3.0 std + 1 × 8s Seedance 2 hero shot ≈ N credits total, largest single N. Proceed with the batch?
53
-
54
- Per `slates-cost-discipline` 3b: that single OK authorizes `confirm=true` for **every enumerated call in the batch** — no per-call re-asking. Re-confirm only if a call's price overruns the plan >25% or new calls get added (extra retakes, new shots).
55
-
56
- ### 5. Generate frame images
57
- Fire the image Shots with `slates_generate_from_shots` (`confirm: true` — step 4 authorized it). Slates names each reference inline as "image N"; you never hand-write a role label or a number. Evaluate every result inline against the beat. Bind keepers via `slates_add_frame`, then `slates_update_shot` with `attachFrameId` so the recipe travels with the picture.
58
-
59
- **Multi-take where it matters:** for the hook shot and any shot the whole film hangs on, generate 2-4 variants (cheap model or 1k), pull them back with `slates_get_assets_batch`, pick the strongest on composition + identity, discard the rest. Don't multi-take filler shots.
60
-
61
- ### 6. Generate video per Shot
62
- Fork each bound frame's image Shot with `slates_duplicate_shot` (`model:` the video model — that is the A/B lever the op takes inline), then `slates_update_shot` the copy with `firstFrameAssetId` = the bound frame. Two calls, because `slates_duplicate_shot` forks the prompt, the model and the params; **attachments are changed with `slates_update_shot`.** Then fire the set with `slates_generate_from_shots`.
63
-
64
- ⚠️ **It runs SEQUENTIALLY and blocks until the last clip lands** — a 6-shot film is one long wait, and it will usually outlast the HTTP timeout while the run keeps going. When that happens, poll `slates_get_shot` for each Shot's `generationIds` and then `slates_get_generation_status`; **never re-fire, that double-spends.** (Concurrent batch firing needs a real queue — concurrency limiting, per-item failure isolation, partial-billing semantics — and is deliberately not built yet.)
65
-
66
- **Model mixing — route per `slates-model-selection`** (details in the per-model guides):
67
- - **Kling V3** (`slates-prompting-kling-v3`): the DEFAULT for most shots — 16:9 / 9:16 / 1:1, 3-15s, strong start-frame adherence; std is the workhorse, Omni for multi-character dialogue.
68
- - **Seedance 2** (`slates-prompting-seedance`): the PREMIUM tier — any shot where physics/effects/scale remotely matter, plus the hero shot; audio included, first+last frame guidance, native 4K (4K video is Pro-only).
69
- - **MiniMax H3** (`slates-prompting-minimax-h3`): route here when a shot's SOUND is part of the writing — a line delivered a particular way, scene sound under it, score that must stay outside the characters' world. It authors all three in one pass, which **collapses a shot's audio pass into its video pass** and removes the separate `slates_generate_audio` step for that shot. 5-15s, 480p/768p/2K/4K. Its sibling `minimax-h3-max` is faster, tops out at 768p, takes the same references, and costs MORE at 768p — a deliberate speed pick, never a saving.
70
- - **Veo 3.1** (`slates-prompting-veo-3`): niche, never the default — only when native synced audio must generate WITH the video in one gen; 16:9 or 9:16, 4/6/8s (8s only at 1080p/4K or with reference images).
71
-
72
- Failed gen? The run continues past it and **nothing is retried automatically**. Read the per-Shot error in the result, fix that Shot with `slates_update_shot`, and re-fire only it (a retry beyond the plan = announce the delta cost).
73
-
74
- ### 7. Assemble the timeline
75
- - `slates_get_timeline` once to get the lay of the land.
76
- - `slates_add_clip_to_timeline` for each completed video asset **in story order** — defaults append back-to-back on the first video track, which is exactly an assembly cut.
77
- - Order wrong? `slates_reorder_clips` with the full clip-id list. Dropped a shot? `slates_remove_clip`, then reorder to close the gap.
78
-
79
- ### 8. Export + deliver
80
- - Output path: ask the user, or default to `<slates_get_project_directory>/exports/<name>.mp4`.
81
- - `slates_export_video` (absolute path, `.mp4`; blocks while ffmpeg renders — minutes for long timelines).
82
- - `slates_reveal_file` so the file is literally in front of them.
83
- - Offer the finishing path: `slates_export_timeline_xml` → DaVinci Resolve (File → Import → Timeline) for grading, sound, and titles.
84
-
85
- ### 9. Report
86
- Shots delivered, total spent vs. approved plan, the export path, and the single best next lever ("re-take shot 3 with a tighter prompt" / "add a CTA end-card").
87
-
88
- ## Hard rules
89
-
90
- - **Two checkpoints only.** Creative plan (step 1) and total cost (step 4). Everything else runs without asking — that's the product promise.
91
- - **Skeleton before spend.** Project + storyboard structure are free; generation isn't.
92
- - **Look at everything.** Every image inline, every video via `slates_get_asset_video_frames` if a clip seems off. Never assemble a timeline from clips you haven't evaluated.
93
- - **3-strike rule per shot.** Three failed takes on one shot = stop, show the user what you tried, ask.
94
- - **Consistency comes from references, not luck.** Same identity asset on every character frame; same environment refs across a location's shots.
95
- - **Plan in Shots, not in chat.** Every decision that ends up in a sentence you have to remember is a decision the user cannot see, price, fork or re-fire. A Shot is a row: it survives the conversation, and the user can open it in the app and fix one reference without you.
1
+ ---
2
+ name: slates-one-prompt-film
3
+ description: Deliver a finished video when the user explicitly asks for one, coordinating editable writing, selected media production, a named cut and a verified export. Preserve existing work and follow generation authorization.
4
+ ---
5
+
6
+ # Idea to finished video
7
+
8
+ Carry the requested piece through to an exported file. Use existing work whenever it serves the brief. The creator can enter at any point: writing, importing footage, comparing takes or changing an edit. No mandatory stage sequence, shot count or number of approval checkpoints follows from this guide.
9
+
10
+ ## Make the intended piece visible
11
+
12
+ Use the current project and document unless another destination is requested. Save words through the revision-checked document tools; `slates-script-craft` covers writing and versions. Add production bindings only where needed. A shot needs no image, and a cut can use imported footage with no script.
13
+
14
+ Preserve fixed passages, explicit creative choices and custom prompt bytes. Record production choices in editable shots. Explain only consequential judgments not already visible there. Recurring cast, repeated framing, silence and dependent scenes are valid when they serve the piece.
15
+
16
+ ## Inspect the actual requests and estimate
17
+
18
+ Use `slates_get_shot` to inspect the composed prompt, settings and references. Model choices and supported settings come from `slates-model-selection`, the current capability surface and the selected model's guide. Do not carry limits or prices from an old example.
19
+
20
+ Follow `slates-cost-discipline` and the user's generation policy. Quote the exact requested set with `slates_generate_from_shots` before confirming it. Existing authorization covers its enumerated requests, not extra takes or changed inputs. Editing, choosing versions, importing and building a cut do not spend generation credits.
21
+
22
+ Keep reusable historical media separate from new requests. Matching words alone do not prove matching voice, references or settings. An explicitly requested extra take is never deduplicated away.
23
+
24
+ ## Generate only the authorized material
25
+
26
+ Submit the chosen requests and inspect each returned state. On an uncertain timeout, read the shot's generation IDs and job status before any retry. Diagnose a failure and follow the existing consent policy for added requests. Never discard other takes merely because a new one was selected.
27
+
28
+ Inspect image composition and reference fidelity. Inspect video performance, motion and sound across playback; frame samples alone cannot establish speech or motion quality. If an edit can resolve dead air or order, use the existing take rather than assuming another generation is needed.
29
+
30
+ ## Arrange and deliver
31
+
32
+ Read the available timelines. Name the destination cut explicitly for a variation; independent comparisons use independent cuts. Add selected media in the intended order, preserve trim/level/transform choices and inspect the actual timeline.
33
+
34
+ Use supported video/XML exports and their stated fidelity limits. For selected named cuts, `slates_export_cuts` records distinct outputs and a manifest; retry unfinished outputs with the same manifest identity. Verify the returned files and playback before reporting success. Describe the completed piece, actual spend where available and output paths. Do not label a render complete merely because its submission succeeded.
35
+
36
+ <!-- @inject:decision-log -->
37
+ Record production choices in the editable shot fields. Explain only consequential judgments the user did not specify and no field already records: for example, why a particular light or performance register supports the brief. Do not repeat the shot list in prose or turn this explanation into an approval gate. Follow the separate generation authorization policy before spending.
38
+ <!-- @end:decision-log -->
@@ -5,21 +5,25 @@ description: Use when the user names an asset by code ("use IMG-A36"), asks what
5
5
 
6
6
  # Organizing a Slates project
7
7
 
8
- Slates already gives each REUSABLE reference type its own home — the **Characters**, **Environments**, and **Styles** tabs, each with its own generation + `@mention`/`#ref` behavior. Do NOT recreate those as folders. Folders are for **structure**, never type.
8
+ Slates already gives every REUSABLE reference a home — the **Library**, in categories the user names (Characters, Locations, Products, Looks…; `slates_list_library`), each item used in a prompt as `@name`, or `#name` for a look. Do NOT recreate those as folders, and never invent a category the user did not ask for. Folders are for **structure**, never type.
9
9
 
10
10
  **Folders = where an asset sits in the FILM**, and they mirror to real subfolders on disk (`projects/<id>/…`), so a human can open the project in Resolve/Finder and navigate it like an edit. Use them for work product, not references.
11
11
 
12
12
  Create with `slates_create_folder`; file assets with `slates_move_assets_to_folder`. Generations land in the project's active folder, so set it before a batch.
13
13
 
14
+ A favorite is a keeper, not a folder: `slates_set_asset_favorite` marks one asset (the heart on its card, `isFavorite` in `slates_list_assets`) without moving it. To hand files out of Slates, `slates_export_assets` copies the ORIGINALS of the assets you name into a directory (named by code, never overwriting) — the way to deliver an image-only job that never needed a shot or a timeline. Use it to flag the takes worth a second look; file with folders once the pick is made.
15
+
16
+ To reuse a whole piece rather than one reference, use a TEMPLATE: `slates_export_template` saves a board, a scene or one Shot (recipes, script words, references and the Library items they mention; never takes) as a file, `slates_get_template` reads what a file holds and its swap slots, and `slates_import_template` adds it to a project, optionally swapping a slot for one of that project's own assets. An import generates nothing; quote and fire the returned Shots as usual.
17
+
14
18
  Conventions by project type:
15
19
  - **Short film / narrative:** `Shots` (scene stills) · `Clips` (generated video) · `Final` (the export). Use one folder per scene (`Scene 1`, `Scene 2`, …) instead when the piece has distinct locations/beats.
16
20
  - **Ad / UGC:** `Hooks` · `B-roll` · `Talking-head` · `Final`.
17
21
 
18
22
  Rules of thumb:
19
- - Reusable cast / sets / look → leave in the Characters/Environments/Styles tabs. Don't fold them.
23
+ - Reusable cast / sets / products / look → leave in the Library. Don't fold them.
20
24
  - Scene stills, clips, and the final cut → file into the structural folder they belong to, as you make them.
21
25
  - One folder per asset (folders are structure). Cross-cutting status (hero take, reject, variant) is a tag concern, not a folder.
22
- - Keep the gallery legible: work product lives in folders; the reference scaffolding (sheets, plates, style images) stays in its tabs.
26
+ - Keep the gallery legible: work product lives in folders; the reference scaffolding (sheets, plates, style images) stays in the Library.
23
27
 
24
28
  ## Asset codes — the shared vocabulary (IMG-A12 / VID-V3 / AUD-S1)
25
29