videodraft 0.22.0 → 0.23.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "videodraft",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.23.1",
|
|
4
4
|
"description": "Official VideoDraft CLI for AI videos, images, audio and 3D assets. Agent-friendly: --json everywhere, stable exit codes, async job polling.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
package/skills/index.json
CHANGED
|
@@ -22,18 +22,18 @@
|
|
|
22
22
|
},
|
|
23
23
|
{
|
|
24
24
|
"path": "references/models.md",
|
|
25
|
-
"sha256": "
|
|
26
|
-
"bytes":
|
|
25
|
+
"sha256": "8161b841894de3a0bb81a7630b19b3ecd6e9a61da14d96e08c74cac61e455479",
|
|
26
|
+
"bytes": 43359
|
|
27
27
|
},
|
|
28
28
|
{
|
|
29
29
|
"path": "references/pipeline.md",
|
|
30
|
-
"sha256": "
|
|
31
|
-
"bytes":
|
|
30
|
+
"sha256": "831ac69ffd9122029423150adf3f21c1f2bba7c3f2d42f9c866bb605d5a284bb",
|
|
31
|
+
"bytes": 13966
|
|
32
32
|
},
|
|
33
33
|
{
|
|
34
34
|
"path": "SKILL.md",
|
|
35
|
-
"sha256": "
|
|
36
|
-
"bytes":
|
|
35
|
+
"sha256": "bd8a43c956b1dfc7d25285c4ecd299618d8a485ab830cf989ce9b72d0906ead7",
|
|
36
|
+
"bytes": 37139
|
|
37
37
|
}
|
|
38
38
|
]
|
|
39
39
|
}
|
|
@@ -89,7 +89,7 @@ Every `videodraft models image|video|audio --json` response carries a top-level
|
|
|
89
89
|
**Audio and utilities:**
|
|
90
90
|
|
|
91
91
|
- Use Seed Audio 1.0 for open-ended text-to-audio, speech/music/sound synthesis, voice conditioning, or prompt-driven editing with up to three audio references or one image. Use `videodraft generate audio`. Reference clips are `@Audio1`, `@Audio2`, and `@Audio3` in array order. There is no duration input. Output is up to two minutes and settles at 19 credits per actual minute, with up to 38 credits reserved during generation. The CLI automatically retries transient responses with one operation key. To recover after the CLI process itself is interrupted, set `--idempotency-key <uuid>` on the original command and reuse it.
|
|
92
|
-
- Prefer ElevenLabs for voiceover, dialogue, voice changing, dubbing, and sound effects. Honor an explicitly selected supported TTS voice/provider. Use Lyria for instrumental music and ElevenLabs Music for vocals, lyrics, or
|
|
92
|
+
- Prefer ElevenLabs for voiceover, dialogue, voice changing, dubbing, and sound effects. Honor an explicitly selected supported TTS voice/provider. Use Lyria for instrumental music and ElevenLabs Music for vocals, lyrics, exact timing, song structure, or a style reference track. ElevenLabs Music means `elevenlabs-music-v2.5` (`elevenlabs-music` is an alias for it); use `elevenlabs-music-v1` only when the user asks for v1 by name.
|
|
93
93
|
- For ElevenLabs voiceover, dialogue, and voice changing, accept a supplied raw voice ID (16-64 alphanumeric characters) or `elevenlabs-<id>`. The voice catalog is for discovery, not an allowlist. Do not reject or substitute a supplied ID because it is absent from `videodraft models voices` / `list_available_voices`. Use `--voice <id>` for voiceover and voice changing, or repeat `--line "<id>:Text"` for dialogue; for example, `--voice kPzsL2i3teMYv0FxEYQ6` and `--line "elevenlabs-kPzsL2i3teMYv0FxEYQ6:Hello."` use the same voice. The voice must be accessible to the provider account used for generation. Private or cloned voices may require the user's connected ElevenLabs key. Kling video-control IDs and MiniMax `custom-*` IDs are separate voice systems.
|
|
94
94
|
- A character who needs to TALK, when you have an image of them, splits by FRAMING:
|
|
95
95
|
- **Talking to camera** (presenter, spokesperson, explainer): the avatar lane, and VEED Fabric is preferred. Use managed `avatar create` then `avatar render` for a reusable avatar record with bundled speech, `avatar fabric` for a one-off portrait plus text or existing audio, and `avatar lipsync` when both the source video and replacement audio already exist.
|
|
@@ -143,7 +143,7 @@ For completed Wan 3.0 jobs, MCP `check_generation_status` and CLI `status`/`wait
|
|
|
143
143
|
|
|
144
144
|
Standalone image/video/audio generations are filed into an AI Studio session in the web app. 3D assets use `assets 3d list/get` separately. You do not have to create an AI Studio session:
|
|
145
145
|
|
|
146
|
-
- **MCP hosts** (Claude Code, claude.ai, Codex, VideoDraft ADE): the server mints an `Mcp-Session-Id` on `initialize`; your host echoes it, and this conversation's generations land in their own session. Tool results echo
|
|
146
|
+
- **MCP hosts** (Claude Code, claude.ai, Codex, VideoDraft ADE): on 2025-era MCP the server mints an `Mcp-Session-Id` on `initialize`; your host echoes it, and this conversation's generations land in their own session. On stateless MCP (2026-07-28) hosts that send a conversation id (ChatGPT) get the same. Otherwise `name_current_ai_studio_session` returns `pass_session_id: true` with a `session_id`: pass that `session_id` to every later standalone generation in the conversation. Tool results echo the session as `ai_studio_session_id`.
|
|
147
147
|
- **CLI**: the same handshake runs once per (profile, server, working directory) and is cached for 12 idle hours, so everything generated from one directory shares one session. `videodraft sessions current` shows it; `videodraft sessions reset` starts a new one.
|
|
148
148
|
- Project generations (`--project <id>` / `project_id`) always go to that project's session.
|
|
149
149
|
|
|
@@ -155,7 +155,7 @@ videodraft sessions name "Purple Seal Rescue Short"
|
|
|
155
155
|
|
|
156
156
|
Choose a concise, specific 3-6 word title for the intended work. Do not copy the client name, date, or exact chat title. Name it once: the operation creates the session with that title. If generation, a user, or an earlier agent created the session first, its existing name is preserved.
|
|
157
157
|
|
|
158
|
-
|
|
158
|
+
Apart from the `pass_session_id: true` case above, pass `--session <id>` / `session_id` only to **continue earlier work** or create an explicit separate group:
|
|
159
159
|
|
|
160
160
|
```bash
|
|
161
161
|
SESSION=$(videodraft sessions create "Fox brand explorations" --json | jq -r '.session.id')
|
|
@@ -118,7 +118,8 @@ Kling O3 is also exposed for reference generation. `videodraft generate video --
|
|
|
118
118
|
- **Voiceover/TTS**: prefer ElevenLabs. Brittney is the platform default voice; under ElevenLabs BYOK, use a compatible voice from the user's account. Honor another supported voice/provider when the user explicitly selects it.
|
|
119
119
|
- **Dialogue, voice changing, and dubbing**: ElevenLabs only.
|
|
120
120
|
- **Sound effects**: ElevenLabs Sound Effects only.
|
|
121
|
-
- **Music**: use `lyria-3-clip-preview` for a short instrumental/background score, `lyria-3-pro-preview` for a longer or higher-quality instrumental score, and `elevenlabs-music` when vocals/lyrics
|
|
121
|
+
- **Music**: use `lyria-3-clip-preview` for a short instrumental/background score, `lyria-3-pro-preview` for a longer or higher-quality instrumental score, and `elevenlabs-music-v2.5` when vocals/lyrics, a specified 3-300 second length, song structure, or a style reference track matter. `elevenlabs-music` is an alias for v2.5. Use `elevenlabs-music-v1` only when the user asks for v1 by name; it takes a prompt, `--length` and `--instrumental` only.
|
|
122
|
+
- **ElevenLabs Music v2.5 inputs**: either a prompt (`--length`, `--instrumental`) or a composition plan, never both. A plan is 1-30 sections of 3-120 seconds each, up to 300 seconds in total, and `--plan`/`--section` pick v2.5 when `--model` is omitted. Empty section parts default to 20 seconds and an instrumental part. Build one with repeatable `--section "<seconds>|<style, style>|<text>"` (use `\n` for line breaks) or pass `--plan plan.json` with `{"chunks":[{"text","duration_ms","positive_styles","negative_styles","context_adherence","audio_reference"}]}`. Section text is an optional `[Section name]`, lyric lines, and `{inline directions}`. Put 6-7 specific English styles on the first section; it sets the genre. `--ref-audio <url|file>` adds a style reference clip to the first section (window up to 30 seconds via `--ref-start`/`--ref-end` in ms, `--ref-strength low|medium|high|xhigh`); local files are uploaded, including relative `audio_reference.audio_url` paths in a plan file. `--seed` works only with a plan. `--format` picks the output (`mp3_48000_192` default on v2.5; `pcm_*`, `ulaw_8000` and `alaw_8000` arrive as stereo WAV files). Style references don't run on the user's own ElevenLabs key; with that key connected, drop the reference or ask the user to turn the key off. The CLI retries transient responses with one idempotency key; set `--idempotency-key <uuid>` to recover after an interruption.
|
|
122
123
|
- Voice Changer and Dubbing require the source media duration and currently accept source media up to 300 seconds.
|
|
123
124
|
|
|
124
125
|
**User-supplied ElevenLabs voice IDs:** voiceover, dialogue, and voice changing accept raw IDs (16-64 alphanumeric characters) or `elevenlabs-<id>`. `videodraft models voices` / MCP `list_available_voices` is for discovery, not an allowlist. Pass a supplied ID directly even when it is absent from the catalog; do not reject it or substitute a catalog voice. The voice must be accessible to the provider account used for generation. Private or cloned voices may require the user's connected ElevenLabs key. A failed voice listing does not prove that a supplied ID cannot generate; a connected key with generation access may still work. If the provider rejects generation access, report that error rather than silently changing the voice.
|
|
@@ -194,7 +195,7 @@ Direct Fabric text/audio and Sync Labs do not use the managed avatar record. The
|
|
|
194
195
|
- Voiceover TTS: 10 credits per 1000 characters for standard voices, 30 per 1000 for cloned `custom-*` voices (min 1, pro-rated); applies to standalone voiceovers AND per-scene narration during `produce`. Silent tracks are free. Voice cloning itself is a flat 150 credits per clone.
|
|
195
196
|
- Lyria music: flat per track, 4 credits (clip) / 8 credits (pro).
|
|
196
197
|
- Seed Audio 1.0: 19 credits per actual output minute, prorated and rounded up to a whole credit. VideoDraft reserves the 120-second maximum of 38 credits and refunds the unused portion after generation. Fal BYOK is free.
|
|
197
|
-
- ElevenLabs audio: sound effects are per second, dialogue is per character, music/voice-changer/dubbing are per started minute. Voice changer and dubbing reject source media above 300s in the current synchronous flow.
|
|
198
|
+
- ElevenLabs audio: sound effects are per second, dialogue is per character, music/voice-changer/dubbing are per started minute. ElevenLabs Music v2.5 and v1 both cost 60 credits per started output minute; estimate a composition plan with its total length. Voice changer and dubbing reject source media above 300s in the current synchronous flow.
|
|
198
199
|
- Seedance 2.0 / 2.5 real people: every listed Seedance 2.x rate assumes `--allow-real-people` is OFF, which uses the Byteplus-priced path (2.0 Mini 4/8 cr/s, Fast 6/13, Standard 7/16/38/78, 2.5 11/24/57 for 480p/720p/1080p). Byteplus refuses real-person likenesses, so a likeness-policy failure does not fall back by default. Passing `--allow-real-people` keeps Byteplus first but permits a submit-time Fal fallback, which allows them, and prices at Fal's rate for that tier: 2.0 Mini 8/16, Fast 11/25, Standard 14/31/69/156, 2.5 23/48/114. That is roughly 2x but not exactly 2x: the Seedance 2.0 1080p pair is 38/69, or about 1.82x. If Byteplus accepts the task and later rejects the generated output, VideoDraft refunds the failure but does not resubmit it to Fal. Pass the option proactively only when supplied visual input media visibly contains a real identifiable person. Otherwise retry once only after the exact opt-in code.
|
|
199
200
|
- Grok Imagine images: `grok-imagine` is a flat 2 cr (3 with a reference). `grok-imagine-2.0` is a separate, newer model priced by resolution and quality: 1K 4 (low) / 6 (medium), 2K 6 / 8, plus 1 cr per reference image (up to 3). v1 is NOT superseded — pick it when cost matters more than 2K.
|
|
200
201
|
- xAI bills refused requests, so failed Grok generations are not refunded.
|
|
@@ -19,6 +19,7 @@ Use direct asset tools for standalone images, clips, audio, upscales, and descri
|
|
|
19
19
|
| Motion clip for a shot | `videodraft generate video --project <id>` | `generate_video` |
|
|
20
20
|
| Attach a finished clip to the timeline | `videodraft attach <project> --scene N --shot M --media <url> --type video` | `attach_media_to_shot` |
|
|
21
21
|
| Background music | `videodraft generate music --attach <project>` | `generate_music` / `set_background_music` |
|
|
22
|
+
| Song with vocals, lyrics or sections | `videodraft generate music --model elevenlabs-music-v2.5 --section "..."` | `generate_music` with `composition_plan` |
|
|
22
23
|
| General or reference-driven audio | `videodraft generate audio "..."` | `generate_audio` |
|
|
23
24
|
| Sound effect | `videodraft generate sound-effect "..."` | `generate_sound_effect` |
|
|
24
25
|
| Dialogue audio | `videodraft generate dialogue --line "voice:text"` | `generate_dialogue` |
|