videodraft 0.4.1 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "videodraft",
3
- "version": "0.4.1",
3
+ "version": "0.6.0",
4
4
  "description": "Official VideoDraft CLI — create AI videos, images and audio from your terminal. Agent-friendly: --json everywhere, stable exit codes, async job polling.",
5
5
  "license": "MIT",
6
6
  "type": "module",
package/skills/index.json CHANGED
@@ -3,27 +3,27 @@
3
3
  "skills": [
4
4
  {
5
5
  "name": "videodraft",
6
- "description": "Create AI videos, images, voiceovers, music, sound effects, dialogue, dubbing, storyboards, avatar videos, media upscales and product/ad videos with VideoDraft. Use when the user mentions VideoDraft, or asks to generate/make a video, video ad, explainer, storyboard, talking-head/avatar video, AI image, voiceover/TTS, background music, sound effects, dialogue audio, voice changing, dubbing, or image/video enhancement and upscaling, including batch/programmatic video generation in scripts or CI. Works via the `videodraft` CLI (preferred in terminals) or the VideoDraft MCP connector.",
6
+ "description": "Create AI videos, images, Seed Audio, voiceovers, music, sound effects, dialogue, dubbing, storyboards, avatar videos, media upscales and product/ad videos with VideoDraft. Use when the user mentions VideoDraft, or asks to generate/make a video, video ad, explainer, storyboard, talking-head/avatar video, AI image, prompt-driven or reference-driven audio, voiceover/TTS, background music, sound effects, dialogue audio, voice changing, dubbing, or image/video enhancement and upscaling, including batch/programmatic video generation in scripts or CI. Works via the `videodraft` CLI (preferred in terminals) or the VideoDraft MCP connector.",
7
7
  "files": [
8
8
  {
9
9
  "path": "references/examples.md",
10
- "sha256": "91cc9aa5a49c34e9a73fdfd224eaf5ce46588c22bd72a951810ae73544efed25",
11
- "bytes": 6254
10
+ "sha256": "5d1bfb72d125db97a0e018643b83a1be6dae60d1ad622ff99ede01d4d916f6a7",
11
+ "bytes": 6390
12
12
  },
13
13
  {
14
14
  "path": "references/models.md",
15
- "sha256": "a871eeec1dc1cc164718f1f52ea40cc90510c2f4c3b71fb57bd6a25981ad06a0",
16
- "bytes": 17138
15
+ "sha256": "2c2fe6a0f5235f6cf6aa7e73181bc37e57a7f6acb48bffc8807eba543ac9068a",
16
+ "bytes": 18025
17
17
  },
18
18
  {
19
19
  "path": "references/pipeline.md",
20
- "sha256": "054109a6ab556bcace1ff843b001afe9bc3a0aac389667e979ac7b77a7200a50",
21
- "bytes": 10644
20
+ "sha256": "f584a668497a7e8fb619c42c3450268f9b365dd11c72d5bc2655002fa0151536",
21
+ "bytes": 10810
22
22
  },
23
23
  {
24
24
  "path": "SKILL.md",
25
- "sha256": "8796e157a835f17edcd8c8345c2668042fdcd793781eace573a570825b6440c6",
26
- "bytes": 15094
25
+ "sha256": "cc80f773350858bc4ef430ebd8cd4e6e55c1a98101c759edea39da443118ab43",
26
+ "bytes": 16074
27
27
  }
28
28
  ]
29
29
  }
@@ -1,13 +1,13 @@
1
1
  ---
2
2
  name: videodraft
3
- description: Create AI videos, images, voiceovers, music, sound effects, dialogue, dubbing, storyboards, avatar videos, media upscales and product/ad videos with VideoDraft. Use when the user mentions VideoDraft, or asks to generate/make a video, video ad, explainer, storyboard, talking-head/avatar video, AI image, voiceover/TTS, background music, sound effects, dialogue audio, voice changing, dubbing, or image/video enhancement and upscaling, including batch/programmatic video generation in scripts or CI. Works via the `videodraft` CLI (preferred in terminals) or the VideoDraft MCP connector.
3
+ description: Create AI videos, images, Seed Audio, voiceovers, music, sound effects, dialogue, dubbing, storyboards, avatar videos, media upscales and product/ad videos with VideoDraft. Use when the user mentions VideoDraft, or asks to generate/make a video, video ad, explainer, storyboard, talking-head/avatar video, AI image, prompt-driven or reference-driven audio, voiceover/TTS, background music, sound effects, dialogue audio, voice changing, dubbing, or image/video enhancement and upscaling, including batch/programmatic video generation in scripts or CI. Works via the `videodraft` CLI (preferred in terminals) or the VideoDraft MCP connector.
4
4
  ---
5
5
 
6
6
  # VideoDraft
7
7
 
8
8
  VideoDraft is an AI video creation platform where asset generation is the priority lane:
9
9
 
10
- - **Asset generation**: standalone images, video clips, voiceovers, music, sound effects, dialogue, voice-changed audio, dubbed media, upscales, and image descriptions. This is the fastest and most important lane. Treat these as complete deliverables when the user asks for assets.
10
+ - **Asset generation**: standalone images, video clips, Seed Audio, voiceovers, music, sound effects, dialogue, voice-changed audio, dubbed media, upscales, and image descriptions. This is the fastest and most important lane. Treat these as complete deliverables when the user asks for assets.
11
11
  - **Asset I/O**: upload local files, download outputs, auto-upload local references, and save generated media where the user can see it.
12
12
  - **Project production**: idea → script → storyboard (scenes + shot images) → project data → production timeline → exported MP4. Use it for a multi-scene video, story, ad, explainer, storyboard, editable timeline, or final export, even when the user does not say "project." A script-only request also creates a script-stage project but stops at the script.
13
13
 
@@ -59,6 +59,7 @@ If the user names a model, use it when compatible. If it cannot handle the reque
59
59
 
60
60
  **Audio and utilities:**
61
61
 
62
+ - Use Seed Audio 1.0 for open-ended text-to-audio, speech/music/sound synthesis, voice conditioning, or prompt-driven editing with up to three audio references or one image. Use `videodraft generate audio`. Reference clips are `@Audio1`, `@Audio2`, and `@Audio3` in array order. There is no duration input. Output is up to two minutes and settles at 19 credits per actual minute, with up to 38 credits reserved during generation. The CLI automatically retries transient responses with one operation key. To recover after the CLI process itself is interrupted, set `--idempotency-key <uuid>` on the original command and reuse it.
62
63
  - Prefer ElevenLabs for voiceover, dialogue, voice changing, dubbing, and sound effects. Honor an explicitly selected supported TTS voice/provider. Use Lyria for instrumental music and ElevenLabs Music for vocals, lyrics, or exact timing.
63
64
  - Talking head/presenter: choose by source. Use managed `avatar create` then `avatar render` when the user wants a reusable avatar record and bundled speech. Use `avatar fabric` for a one-off portrait plus text or existing audio. Use `avatar lipsync` when both the source video and replacement audio already exist.
64
65
  - Enhancement: use Topaz image/video upscaling only when the content is already correct. Use image 1x for cleanup, 2x by default, 4x when justified; use video 2x by default. Edit or regenerate creative errors.
@@ -79,11 +80,11 @@ Do not call `videodraft credits` before routine generations. Paid endpoints vali
79
80
 
80
81
  For expensive work, estimate with `--estimate` or `videodraft costs`, state the selected model/settings/cost, and get a go-ahead. This matters most for shot-image batches, long or high-resolution video, AI Production, and paid audio batches. Honor the user's confirmation preference for the session.
81
82
 
82
- `videodraft models image|video` lists the live image and video catalogs with supported inputs. Video entries are grouped as `generation`, `video_edit`, `motion_control`, `avatar_lipsync`, and `upscale`, and each reports the exact tool. Use `videodraft models video --category video_edit` to narrow the list. `videodraft models audio` lists Google Lyria and ElevenLabs audio/media tools, while `videodraft models voices` lists TTS voices. Consult them instead of guessing capabilities.
83
+ `videodraft models image|video` lists the live image and video catalogs with supported inputs. Video entries are grouped as `generation`, `video_edit`, `motion_control`, `avatar_lipsync`, and `upscale`, and each reports the exact tool. Use `videodraft models video --category video_edit` to narrow the list. `videodraft models audio` lists Seed Audio, Google Lyria, and ElevenLabs audio/media tools, while `videodraft models voices` lists TTS voices. Consult them instead of guessing capabilities.
83
84
 
84
85
  ## Async jobs
85
86
 
86
- Image/video generation is asynchronous: commands submit a job and **wait by default**, printing output URLs (and saving files with `--download`). In scripts/CI prefer explicit control:
87
+ Image/video generation is asynchronous: commands submit a job and **wait by default**, printing output URLs (and saving files with `--download`). Large downloaded images also get a downscaled copy in `previews/` next to them (the `preview` field / "inspect via preview" line in the output) — **look at the preview, deliver the original**; viewing full-resolution images bloats the chat permanently. In scripts/CI prefer explicit control:
87
88
 
88
89
  ```bash
89
90
  JOB=$(videodraft generate image "..." --no-wait --json | jq -r .job_id)
@@ -130,7 +131,7 @@ videodraft produce <project_id> # voiceovers + captions + produc
130
131
  videodraft export <project_id> --download final.mp4
131
132
  ```
132
133
 
133
- Optional between produce and export: per-shot motion clips (`videodraft generate video ... --project <id>` then place it with `videodraft attach <project> --scene N --shot M --media <url|file> --type video --duration <s>`), music (`videodraft generate music "..." --attach <project_id>`), and standalone audio assets (`generate sound-effect`, `generate dialogue`, `generate voice-changer`, `generate dub`). Details, per-step tools and editing rules: [references/pipeline.md](references/pipeline.md).
134
+ Optional between produce and export: per-shot motion clips (`videodraft generate video ... --project <id>` then place it with `videodraft attach <project> --scene N --shot M --media <url|file> --type video --duration <s>`), music (`videodraft generate music "..." --attach <project_id>`), and standalone audio assets (`generate audio`, `generate sound-effect`, `generate dialogue`, `generate voice-changer`, `generate dub`). Details, per-step tools and editing rules: [references/pipeline.md](references/pipeline.md).
134
135
 
135
136
  Avatar/talking-head videos use dedicated commands. For a reusable managed avatar, obtain or generate a clear portrait → `videodraft avatar script` when needed → `videodraft avatar create` → `videodraft avatar render --resolution 720p`. For a one-off portrait, use `videodraft avatar fabric <portrait> --text "..."` or `--audio <file>`. For an existing video plus replacement audio, use `videodraft avatar lipsync <video> --audio <file>`. Managed script/creation is bundled/free; direct Fabric, Sync, the managed Fabric render, and optional portrait generation/upscaling are paid. Confirm expensive steps first.
136
137
 
@@ -39,6 +39,7 @@ videodraft shots "$PROJECT" --grid --estimate # show the user the cost;
39
39
  videodraft shots "$PROJECT" --grid
40
40
  videodraft produce "$PROJECT"
41
41
  videodraft generate music "minimal ambient, warm pads, 60 BPM" --attach "$PROJECT"
42
+ videodraft generate audio "Extend @Audio1 into a 20-second transition" --ref-audio ./intro.wav --format wav --download ./transition.wav
42
43
  videodraft export "$PROJECT" --download solace-launch.mp4
43
44
  ```
44
45
 
@@ -75,6 +75,7 @@ Kling O3 and Wan 2.7 Ref/Edit are dual-mode cards. `videodraft edit video` uses
75
75
 
76
76
  ### Audio
77
77
 
78
+ - **Seed Audio 1.0**: use `videodraft generate audio` for open-ended speech, sound, music, or prompt-driven audio editing. It accepts up to three audio references or one image. Address audio references as `@Audio1`, `@Audio2`, and `@Audio3`. Preset and custom cloned voice IDs are supported. Output is up to 120 seconds. There is no requested-duration input. The CLI automatically retries transient responses with one idempotency key. To recover after the CLI process itself is interrupted, set `--idempotency-key <uuid>` on the original command and reuse it.
78
79
  - **Voiceover/TTS**: prefer ElevenLabs. Brittney is the platform default voice; under ElevenLabs BYOK, use a compatible voice from the user's account. Honor another supported voice/provider when the user explicitly selects it.
79
80
  - **Dialogue, voice changing, and dubbing**: ElevenLabs only.
80
81
  - **Sound effects**: ElevenLabs Sound Effects only.
@@ -132,6 +133,7 @@ Direct Fabric text/audio and Sync Labs do not use the managed avatar record. The
132
133
  - Sync Labs Lipsync 2: 5 credits per verified audio second.
133
134
  - Voiceover TTS: 10 credits per 1000 characters for standard voices, 30 per 1000 for cloned `custom-*` voices (min 1, pro-rated); applies to standalone voiceovers AND per-scene narration during `produce`. Silent tracks are free. Voice cloning itself is a flat 150 credits per clone.
134
135
  - Lyria music: flat per track, 10 credits (clip) / 15 credits (pro).
136
+ - Seed Audio 1.0: 19 credits per actual output minute, prorated and rounded up to a whole credit. VideoDraft reserves the 120-second maximum of 38 credits and refunds the unused portion after generation. Fal BYOK is free.
135
137
  - ElevenLabs audio: sound effects are per second, dialogue is per character, music/voice-changer/dubbing are per started minute. Voice changer and dubbing reject source media above 300s in the current synchronous flow.
136
138
  - Upscales: priced by scale and source size.
137
139
 
@@ -141,6 +143,7 @@ Quote before spending:
141
143
  videodraft costs gemini-omni-flash --type video --duration 8 --resolution 720p --audio
142
144
  videodraft costs seedance-2 --type video --duration 15 --resolution 720p --quality standard --audio
143
145
  videodraft costs elevenlabs-dubbing --type audio --duration 60
146
+ videodraft costs seed-audio-1.0 --type audio --duration 60 # scenario only; model controls actual length
144
147
  videodraft costs elevenlabs-dialogue --type audio --chars 350
145
148
  videodraft costs voiceover --type audio --chars 800 # TTS: 10 cr / 1000 chars
146
149
  videodraft generate video "..." --model gemini-omni-flash --estimate # same quote, inline
@@ -18,6 +18,7 @@ Use direct asset tools for standalone images, clips, audio, upscales, and descri
18
18
  | Motion clip for a shot | `videodraft generate video --project <id>` | `generate_video` |
19
19
  | Attach a finished clip to the timeline | `videodraft attach <project> --scene N --shot M --media <url> --type video` | `attach_media_to_shot` |
20
20
  | Background music | `videodraft generate music --attach <project>` | `generate_music` / `set_background_music` |
21
+ | General or reference-driven audio | `videodraft generate audio "..."` | `generate_audio` |
21
22
  | Sound effect | `videodraft generate sound-effect "..."` | `generate_sound_effect` |
22
23
  | Dialogue audio | `videodraft generate dialogue --line "voice:text"` | `generate_dialogue` |
23
24
  | Voice changer | `videodraft generate voice-changer <audio>` | `change_voice` |