videodraft 0.25.2 → 0.26.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "videodraft",
3
- "version": "0.25.2",
3
+ "version": "0.26.0",
4
4
  "description": "Official VideoDraft CLI for AI videos, images, audio and 3D assets. Agent-friendly: --json everywhere, stable exit codes, async job polling.",
5
5
  "license": "MIT",
6
6
  "type": "module",
package/skills/index.json CHANGED
@@ -22,18 +22,18 @@
22
22
  },
23
23
  {
24
24
  "path": "references/models.md",
25
- "sha256": "1a2f7feb34e5af588c180e64bb19ddf169c1aedc0bd0511c4293378a7fb7cf26",
26
- "bytes": 45577
25
+ "sha256": "031b8d602a17cdf16e8182ada4a696f6d39825abbc9eadbbf0f735b123c3cf91",
26
+ "bytes": 46866
27
27
  },
28
28
  {
29
29
  "path": "references/pipeline.md",
30
- "sha256": "1d8d297f9c97042ab7c13c4fed4a190dad521fe70182d91055c6065f2e7b347f",
31
- "bytes": 14581
30
+ "sha256": "88c6d531e20181dde1e99a0db633b113ac70e559e8b7542ca61c3f509b63f375",
31
+ "bytes": 15112
32
32
  },
33
33
  {
34
34
  "path": "SKILL.md",
35
- "sha256": "c9ce31ecb4cb216108d33e62b2786a9e1cfedef89a399292fc1c3966e3b26903",
36
- "bytes": 43425
35
+ "sha256": "305f3644e63bb94a3676079143473ed68e8624593cf60ca0496f52ce1bd789f7",
36
+ "bytes": 47481
37
37
  }
38
38
  ]
39
39
  }
@@ -15,6 +15,19 @@ VideoDraft is an AI video creation platform where asset generation is the priori
15
15
 
16
16
  ## How to connect
17
17
 
18
+ Generation billing is selected in VideoDraft Settings → API keys. Where enabled,
19
+ **VideoDraft routing** chooses an equivalent provider internally without changing
20
+ the quoted customer credit price. A selected personal key is exclusive: never
21
+ switch to VideoDraft credits or another key after an unsupported-model or
22
+ provider error. Keep the canonical VideoDraft model ID in CLI/MCP calls. Poll
23
+ the original generation after a timeout instead of starting another paid job.
24
+ New provider connections can support only a subset of models and settings;
25
+ follow the server's availability response. ElevenLabs remains audio-only.
26
+ For Pika GPT Image 2.5, supply an explicit quality (`low`, `medium`, `high`,
27
+ `xhigh`, or `max`); `auto` is not supported by that connection. Nano Banana 2
28
+ Lite is 1K only. Never remove requested controls or substitute a different
29
+ model merely to make a personal connection accept the job.
30
+
18
31
  Cloud generation has two equivalent surfaces (same backend, credits, and hosted projects). Native timeline editing is a separate local surface:
19
32
 
20
33
  1. **CLI** (preferred when you have a shell): run `videodraft` if it's on PATH; otherwise `npx -y videodraft@latest` runs it with no install (needs Node ≥20; the `-y` skips npx's install prompt so it runs non-interactively; the package is fetched on first use and cached). For heavy use, `npm install -g videodraft`. If there's no Node/shell here but the MCP connector below is available, use that instead; if neither works, tell the user how to install (https://videodraft.ai/cli).
@@ -102,6 +115,7 @@ Every `videodraft models image|video|audio --json` entry carries `tier` (1 or 2)
102
115
 
103
116
  - Use Seed Audio 1.0 for open-ended text-to-audio, speech/music/sound synthesis, voice conditioning, or prompt-driven editing with up to three audio references or one image. Use `videodraft generate audio`. Reference clips are `@Audio1`, `@Audio2`, and `@Audio3` in array order. There is no duration input. Output is up to two minutes and settles at 19 credits per actual minute, with up to 38 credits reserved during generation. The CLI automatically retries transient responses with one operation key. To recover after the CLI process itself is interrupted, set `--idempotency-key <uuid>` on the original command and reuse it.
104
117
  - Prefer ElevenLabs for voiceover, dialogue, voice changing, dubbing, and sound effects. Honor an explicitly selected supported TTS voice/provider. Use `lyria-3.5` for all music, short or long (up to ~3 minutes, 10 credits flat), with vocals/lyrics or instrumental arrangements. When `--model` is omitted the CLI sends no model and the server default applies, which is `lyria-3.5` wherever the backend supports it. If the server answers that `lyria-3.5` is an unknown model, that backend predates it: rerun without `--model` (the server default then applies), or use `lyria-3-pro-preview` for a track longer than 30 seconds. Ask for the length in the prompt. Keep `lyria-3-clip-preview` (fixed ~30s, 4 credits) and `lyria-3-pro-preview` (8 credits) for explicit requests. For Lyria, put desired length, lyrics and "instrumental only, no vocals" in the prompt; `--length` and `--instrumental` are ElevenLabs-only. Use ElevenLabs Music for exact timing, composition plans, or a style reference track. Lyria allows 10 reference images on Google, 1 on Fal BYOK; Fal 3.5 prompts are limited to 5000 characters. BYOK uses zero VideoDraft credits. ElevenLabs Music means `elevenlabs-music-v2.5` (`elevenlabs-music` is an alias for it); use `elevenlabs-music-v1` only when the user asks for v1 by name.
118
+ - ElevenLabs voiceover runs on Eleven v4 in two modes. Standard is the default and the best quality, at 10 credits per 1000 characters. Turbo (Eleven v4 Turbo) is faster, at 5 credits per 1000 characters. Keep Standard unless the user wants faster or cheaper speech. Pick it with `generate voiceover --mode standard|turbo`, `produce --voice-mode standard|turbo`, or MCP `generate_voiceover.mode` / `produce_project.voice_mode`, and quote Turbo with `costs voiceover --chars <n> --mode turbo`. Google, OpenAI and cloned `custom-*` voices ignore the mode (cloned voices cost 30 per 1000). v4 has no style or speed settings and no SSML `<break>` tags; audio tags such as `[whispers]` work. Dialogue stays on Eleven v3.
105
119
  - For ElevenLabs voiceover, dialogue, and voice changing, accept a supplied raw voice ID (16-64 alphanumeric characters) or `elevenlabs-<id>`. The voice catalog is for discovery, not an allowlist. Do not reject or substitute a supplied ID because it is absent from `videodraft models voices` / `list_available_voices`. Use `--voice <id>` for voiceover and voice changing, or repeat `--line "<id>:Text"` for dialogue; for example, `--voice kPzsL2i3teMYv0FxEYQ6` and `--line "elevenlabs-kPzsL2i3teMYv0FxEYQ6:Hello."` use the same voice. The voice must be accessible to the provider account used for generation. Private or cloned voices may require the user's connected ElevenLabs key. Kling video-control IDs and MiniMax `custom-*` IDs are separate voice systems.
106
120
  - A character who needs to TALK:
107
121
  - **Speaking inside a scene**, or any ordinary clip with dialogue: `generate video` with the line in the prompt, in quotes, with who says it and how. `gemini-omni-1.1-flash`, `seedance-2.5`, `seedance-2`, `kling-3.0` and `kling-o3` all voice it natively. For a specific voice use a Seedance `--ref-audio` clip or a Kling voice bound per element. This is the default path; do not generate speech separately and lip-sync it on.
@@ -111,6 +125,18 @@ Every `videodraft models image|video|audio --json` entry carries `tier` (1 or 2)
111
125
 
112
126
  See [references/models.md](references/models.md) for the detailed routing table and exact capability limits.
113
127
 
128
+ Provider routing uses a maintained local price catalog, with exact setting compatibility and account rules. Platform requests prefer Google's direct image/video APIs and BytePlus for Seedance 2.0/2.5 before comparing public prices. Seedance 1.5 retains its existing Replicate primary; other models remain price-based. Agents should keep using normal model IDs and `get_model_costs` / `--estimate` for customer credits. Do not choose a vendor from a marketing starting price or retry an accepted generation to chase a lower price. Unknown or expired comparable costs retain the existing route; a selected personal provider key stays exclusive and overrides platform-account preferences.
129
+
130
+ Temporary provider promotions are excluded from routing prices. With an Atlas personal key, H3 Max supports text or first/last-frame generation at 480P/768P only with `prompt_expansion_mode: "disabled"` and no seed. Balanced/quality expansion, reference generation and H3 Max Lip Sync require the existing Fal path. Atlas H3 Max is not selected automatically while its published price units remain inconsistent.
131
+
132
+ ## Matching AI Studio controls
133
+
134
+ Use `models image --json` / `models video --json` for the accepted options. Nano Banana Pro/2 expose `--temperature 0..2` and `--google-search-grounding true|false`. Qwen uses one `--ref` plus `--horizontal-angle 0..360`, `--vertical-angle -30..90`, and `--zoom 0..10`; its text prompt is optional. Recraft generates SVG and accepts `--quality Normal|Pro`, repeatable `--recraft-color '#RRGGBB'`, and `--recraft-background '#RRGGBB'`. Original Nano Banana and VideoDraft Image use fixed 1K output.
135
+
136
+ Kling 3.0 / 2.5 Pro accept `--cfg-scale 0..1` and `--negative ""` to clear a negative prompt. Seedance 2/2.5, Wan 3.0, FLUX 3 text/first-frame and Gemini Omni generation accept `--auto-duration`; do not combine it with `--duration`. FLUX 3 auto reserves 20 seconds, so its default estimate includes that ceiling. `--model sora-2 --quality pro --resolution 1080p` selects Sora Pro. Kling O3 reference/edit modes accept image-only `--element` objects, at most four combined with `--ref`. Grok v1 accepts image references for clips up to 10 seconds.
137
+
138
+ Topaz image results may be immediate or queued. The CLI waits by default and downloads either result with `--download`; `--no-wait` returns a queued job ID. Raw MCP callers must poll `check_generation_status` when `upscale_image` returns `job_id`.
139
+
114
140
  ## Prefer references when continuity matters
115
141
 
116
142
  Pure text-to-image or text-to-video is fine for a generic one-off asset. When a specific character, product, location, style, composition, or brand identity must survive generation, use references instead of hoping the prompt recreates it.
@@ -35,7 +35,7 @@ Tier 1 is the rows marked **T1** below. Every other catalog model is Tier 2 and
35
35
 
36
36
  GPT Image 2.5 has two IDs: `gpt-image-2.5-flare` replaces the former GPT Image 2 recommendation; `gpt-image-2.5-sunburst` is the precision/detail alternative. Other model recommendations are unchanged. Explicit `gpt-image-2` selections continue to use the previous model.
37
37
 
38
- Use the same basic image controls as GPT Image 2: `--ref`, `--ar`, `--resolution 1K|2K|4K`, `--quality`, and `--num 1..4`. GPT Image 2.5 quality supports auto/low/medium/high/xhigh/max. Auto is priced at Max. Generation runs directly on OpenAI, or exclusively on the user's Fal key when active.
38
+ Use the same basic image controls as GPT Image 2: `--ref`, `--ar`, `--resolution 1K|2K|4K`, `--quality`, and `--num 1..4`. GPT Image 2.5 quality supports auto/low/medium/high/xhigh/max. Auto is priced at Max in customer-credit estimates. The default platform path uses OpenAI; a selected personal connection stays exclusive. A Pika personal connection requires an explicit quality from low/medium/high/xhigh/max because its upstream API has no auto value. Other unsupported connection settings return an error instead of switching payer or changing the request.
39
39
 
40
40
  ```bash
41
41
  videodraft generate image 'A blue glass bird' --model gpt-image-2.5-flare --ar 1:1 --resolution 2K --quality high --estimate
@@ -119,7 +119,7 @@ Kling O3 is also exposed for reference generation. `videodraft generate video --
119
119
  ### Audio
120
120
 
121
121
  - **Seed Audio 1.0**: use `videodraft generate audio` for open-ended speech, sound, music, or prompt-driven audio editing. It accepts up to three audio references or one image. Address audio references as `@Audio1`, `@Audio2`, and `@Audio3`. Preset and custom cloned voice IDs are supported. Output is up to 120 seconds. There is no requested-duration input. The CLI automatically retries transient responses with one idempotency key. To recover after the CLI process itself is interrupted, set `--idempotency-key <uuid>` on the original command and reuse it.
122
- - **Voiceover/TTS**: prefer ElevenLabs. Brittney is the platform default voice; under ElevenLabs BYOK, use a compatible voice from the user's account. Honor another supported voice/provider when the user explicitly selects it.
122
+ - **Voiceover/TTS**: prefer ElevenLabs. Brittney is the platform default voice; under ElevenLabs BYOK, use a compatible voice from the user's account. Honor another supported voice/provider when the user explicitly selects it. ElevenLabs voices run on Eleven v4 in one of two modes: Standard (the default, best quality) or Turbo (Eleven v4 Turbo, faster and half the price). Keep Standard unless the user wants faster or cheaper speech. Pick the mode with `generate voiceover --mode standard|turbo` or `produce --voice-mode standard|turbo`; over MCP, pass `mode` to `generate_voiceover` or `voice_mode` to `produce_project`. Google, OpenAI and cloned `custom-*` voices ignore it. v4 takes no style or speed settings and no SSML `<break>` tags; audio tags such as `[whispers]` work. Dialogue stays on Eleven v3.
123
123
  - **Dialogue, voice changing, and dubbing**: ElevenLabs only.
124
124
  - **Sound effects**: ElevenLabs Sound Effects only.
125
125
  - **Music**: use `lyria-3.5` for all music, short or long (up to ~3 minutes), with vocals/lyrics or instrumental music. When `--model` is omitted the CLI sends no model and the server default applies, which is `lyria-3.5` wherever the backend supports it. If the server answers that `lyria-3.5` is an unknown model, that backend predates it: rerun without `--model` (the server default then applies), or use `lyria-3-pro-preview` for a track longer than 30 seconds. `lyria-3-clip-preview` (fixed ~30s) and `lyria-3-pro-preview` remain available for explicit requests. Lyria length and structure are prompt-guided, not exact; for an instrumental, include "instrumental only, no vocals" in the prompt. `--length` and `--instrumental` only apply to ElevenLabs. Use `elevenlabs-music-v2.5` when a specified 3-300 second length, composition plan, or style reference track matters. Lyria references: up to 10 images on Google, 1 on Fal BYOK; 3.5 Fal prompts are 1-5000 characters (an image-only request gets a neutral prompt). `elevenlabs-music` is an alias for v2.5. Use `elevenlabs-music-v1` only when the user asks for v1 by name; it takes a prompt, `--length` and `--instrumental` only.
@@ -166,7 +166,7 @@ Direct Fabric text/audio, H3 Max Lip Sync, and Sync Labs do not use the managed
166
166
 
167
167
  ### Upscaling / enhancement
168
168
 
169
- - **Images**: Topaz via `videodraft upscale image <url-or-file> --scale 1x|2x|4x [--mode precision|generative|creative] [--model <name>]`. Modes: `generative` (default; Wonder 3.5, Topaz's recommended model for AI-generated images; `Redefine` takes `--prompt`, `--creativity 1-6`, `--texture 1-5`), `precision` (faithful and cheapest; Standard V2, High Fidelity V3, Low Resolution V2, CGI, Text Refine; use for clean real photos), `creative` (Bloom 2; artistic, `--prompt`, `--creativity 1-9`). Extra knobs: `--no-face-enhance`, `--face-strength`, `--sharpen`, `--denoise`, `--fix-compression`, `--format png`. Use 1x for light enhancement without enlargement, 2x as the general default, and 4x only when the source quality and target size justify it. The server must verify dimensions from a readable image of at most 50 MB; `--width` and `--height` are compatibility hints and cannot bypass a failed probe. The result is synchronous. Cost: 8 credits per started 24 MP of output in precision, per 8 MP (Wonder 3/3.5) or 4 MP in generative, per 2 MP in creative.
169
+ - **Images**: Topaz via `videodraft upscale image <url-or-file> --scale 1x|2x|4x [--mode precision|generative|creative] [--model <name>]`. Modes: `generative` (default; Wonder 3.5, Topaz's recommended model for AI-generated images; `Redefine` takes `--prompt`, `--creativity 1-6`, `--texture 1-5`), `precision` (faithful and cheapest; Standard V2, High Fidelity V3, Low Resolution V2, CGI, Text Refine; use for clean real photos), `creative` (Bloom 2; artistic, `--prompt`, `--creativity 1-9`). Extra knobs: `--no-face-enhance`, `--face-strength`, `--sharpen`, `--denoise`, `--fix-compression`, `--format png`. Use 1x for light enhancement without enlargement, 2x as the general default, and 4x only when the source quality and target size justify it. The server must verify dimensions from a readable image of at most 50 MB; `--width` and `--height` are compatibility hints and cannot bypass a failed probe. The result may be immediate or queued; the CLI waits by default. Use `--no-wait` for an async job ID, and poll it with `status` or `wait`. Cost: 8 credits per started 24 MP of output in precision, per 8 MP (Wonder 3/3.5) or 4 MP in generative, per 2 MP in creative.
170
170
  - **Videos**: Topaz via `videodraft upscale video <url-or-file> --resolution 720p|1080p|4k` (preferred) or `--scale 2x`, plus `--mode precision|generative|creative` and `--model <name>`. `generative` (default; Starlight Precise 2.6, Topaz's recommended model for AI-generated footage; Starlight Fast 2 at half price) costs 12 credits/s up to 1080p and 26 at 4K. `precision` (Proteus, Proteus Natural, Iris, Gaia 2 for animation, Rhea, Theia, Artemis, Dione) is 6x cheaper at 1 / 2 / 6 credits per second for 720p / 1080p / 4K output; use it for real footage or when cost matters. `creative` (Astra 2, `--prompt`, `--creativity`, `--realism`, `--sharp`) always renders 4K at 50 credits/s. `--fps 60` delivers 60fps on the same pass; any output above 30fps (including a 50/60fps source) doubles every rate; the server must verify duration, dimensions, and frame rate from an MP4/MOV source of at most 100 MB. `--duration`, `--width`, `--height`, and `--source-fps` are compatibility hints and cannot override billing or bypass a failed probe. The same source-verification requirement applies to frame interpolation. `--scale` also accepts intermediate factors such as `1.5x`. Max source length 5 minutes. The job is asynchronous; the CLI waits by default, while MCP callers poll `check_generation_status`. MCP video input must be VideoDraft-hosted, so upload local or external sources first.
171
171
  - **Frame interpolation / slow motion**: `videodraft interpolate <url-or-file> --fps 60 [--model Apollo|Chronos|Aion] [--slowdown 1-8]` (MCP `interpolate_video`). Resolution is unchanged. Apollo (default) for smooth general conversion, Chronos for natural slow motion, Aion for extreme slow motion. Apollo/Chronos cost 3 credits per output second up to 1080p (6 at 4K); Aion 5 / 17. These rates cover targets up to 60fps; above 60fps multiply by target FPS / 60 (120fps doubles the rate). Output seconds = source seconds × slowdown. The final charge rounds up to a whole credit.
172
172
  - Use upscaling to preserve the image/video while improving detail, resolution, or cleanup. It cannot fix the wrong subject, misspelled text, bad framing, unwanted objects, broken continuity, or incorrect motion. Use an edit or regeneration for those problems. Keep `generative` for AI-generated sources; switch to `precision` for clean real photos/footage or a cheap pass, and `creative` only when the user wants an artistic reinterpretation.
@@ -198,7 +198,7 @@ Direct Fabric text/audio, H3 Max Lip Sync, and Sync Labs do not use the managed
198
198
  - Direct VEED Fabric: text or normal audio is 8 credits/sec at 480p and 15/sec at 720p; fast audio is 10/sec at 480p and 20/sec at 720p.
199
199
  - MiniMax H3 Max Lip Sync: 5 / 8 / 16 / 32 credits per output second at 480P / 768P (default) / 1080P / 2K. The video runs as long as the audio (at least 5s; only the first 14.8s is used), billed on the server-measured length rounded up, so at most 15 seconds. Use MP3, WAV, M4A (AAC), or AAC audio; other formats run only on the user's own Fal key.
200
200
  - Sync Labs Lipsync 2: 5 credits per verified audio second.
201
- - Voiceover TTS: 10 credits per 1000 characters for standard voices, 30 per 1000 for cloned `custom-*` voices (min 1, pro-rated); applies to standalone voiceovers AND per-scene narration during `produce`. Silent tracks are free. Voice cloning itself is a flat 150 credits per clone.
201
+ - Voiceover TTS: 10 credits per 1000 characters for standard voices (including ElevenLabs Standard, Eleven v4), 5 per 1000 for ElevenLabs Turbo (Eleven v4 Turbo), and 30 per 1000 for cloned `custom-*` voices (min 1, pro-rated); applies to standalone voiceovers AND per-scene narration during `produce`. Quote Turbo with model id `voiceover-turbo` (CLI: `costs voiceover --mode turbo`) and cloned voices with `voiceover-cloned`. Silent tracks are free. Voice cloning itself is a flat 150 credits per clone.
202
202
  - Lyria music: flat per track, 4 credits (Clip) / 10 credits (3.5) / 8 credits (legacy Pro). Fal BYOK uses zero VideoDraft credits.
203
203
  - Seed Audio 1.0: 19 credits per actual output minute, prorated and rounded up to a whole credit. VideoDraft reserves the 120-second maximum of 38 credits and refunds the unused portion after generation. Fal BYOK is free.
204
204
  - ElevenLabs audio: sound effects are per second, dialogue is per character, music/voice-changer/dubbing are per started minute. ElevenLabs Music v2.5 and v1 both cost 60 credits per started output minute; estimate a composition plan with its total length. Voice changer and dubbing reject source media above 300s in the current synchronous flow.
@@ -224,6 +224,7 @@ videodraft costs elevenlabs-dubbing --type audio --duration 60
224
224
  videodraft costs seed-audio-1.0 --type audio --duration 60 # scenario only; model controls actual length
225
225
  videodraft costs elevenlabs-dialogue --type audio --chars 350
226
226
  videodraft costs voiceover --type audio --chars 800 # TTS: 10 cr / 1000 chars
227
+ videodraft costs voiceover --type audio --chars 800 --mode turbo # ElevenLabs Turbo: 5 cr / 1000 chars
227
228
  videodraft generate video "..." --model gemini-omni-1.1-flash --estimate # same quote, inline
228
229
  videodraft generate video "..." --model minimax-h3-max --duration 8 --resolution 768p --prompt-expansion-mode balanced
229
230
  ```
@@ -47,6 +47,7 @@ Use direct asset tools for standalone images, clips, audio, upscales, and descri
47
47
  - **Visual consistency**: never generate a storyboard shot in isolation. Shot prompts carry `[[asset:Name]]` / `[[shot:X-Y]]` tags that `generate_shot_images` resolves against the project's visual assets and prior shots. When generating a single shot whose prompt has no tags, pass `--ref` images yourself (the project's visual assets and/or the previous shot's image; `projects get` exposes both). For scenes with multiple shots or recurring characters, prefer `videodraft shots <project> --model <selected-image-model> --grid`: preserve an explicitly requested compatible image model, otherwise use `nano-banana-2`. It creates one coherent scene grid, then decodes it into individual shot images.
48
48
  - **Reference-first video**: when identity, styling, or composition matters, do not generate each motion clip from text alone. Generate or select the shot still first, then pass the decoded shot image as `--start-image` or `--ref` to the selected video model. AI Production already composes scene grids and sends them to Seedance as references. If the user explicitly requests another compatible video model, bypass fixed Seedance full-video mode and generate the per-shot clips with the requested model, using the individual decoded shot images as anchors.
49
49
  - **Seedance full-video real people**: hosted `full_video` allows real people by default, which applies Fal-tier pricing to every submitted segment and permits the Byteplus-to-Fal fallback. For the lower Byteplus-only rate when no scene grid shows a real identifiable person (non-people, anime, clearly synthetic or stylized characters), use `videodraft produce <project> --mode full_video --no-allow-real-people`, or MCP `produce_project` with `mode: "full_video", allow_real_people: false`. If a partial opted-out run returns `SEEDANCE_REAL_PERSON_OPT_IN_REQUIRED`, re-estimate, follow the user's spend-confirmation preference, and rerun the same project once with an explicit `--allow-real-people` / `allow_real_people: true`. The server reconciles asynchronous results first, preserves running/completed jobs, and retries only failed placeholders carrying that exact code. Do not loop when it was already on. VideoDraft refunds a Byteplus task rejected after asynchronous acceptance, but cannot reroute it; rephrase or change the scene references instead.
50
+ - **Narration voice mode**: ElevenLabs narration runs on Eleven v4. Standard is the default and the best quality, at 10 credits per 1000 characters. `videodraft produce <project> --voice-mode turbo` (MCP `produce_project` with `voice_mode: "turbo"`) uses Eleven v4 Turbo, which is faster at 5 credits per 1000 characters. For one scene, `videodraft generate voiceover --project <id> --scene N --mode turbo` (MCP `generate_voiceover` with `mode: "turbo"`) does the same. Google, OpenAI and cloned `custom-*` voices ignore the mode.
50
51
  - **Hold off generating shot images while the user is still iterating** on storyboard structure.
51
52
  - **produce → export ordering**: `export` requires a produced project where every production scene has timeline media. If `produce` returns `generating_shot_images`, poll the job ids it returns, then re-run produce.
52
53
  - **Do not attach motion clips before production exists**: run `produce` successfully first, then attach finished motion clips to the production timeline. Attaching before `production_data` exists cannot place them in the final timeline.