videodraft 0.19.2 → 0.19.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "videodraft",
|
|
3
|
-
"version": "0.19.
|
|
3
|
+
"version": "0.19.4",
|
|
4
4
|
"description": "Official VideoDraft CLI — create AI videos, images and audio from your terminal. Agent-friendly: --json everywhere, stable exit codes, async job polling.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
package/skills/index.json
CHANGED
|
@@ -12,13 +12,13 @@
|
|
|
12
12
|
},
|
|
13
13
|
{
|
|
14
14
|
"path": "references/examples.md",
|
|
15
|
-
"sha256": "
|
|
16
|
-
"bytes":
|
|
15
|
+
"sha256": "32e30ae37f0e9fb95e76421bf8ad245ad89b03e6def57440893d3b237af10f73",
|
|
16
|
+
"bytes": 9200
|
|
17
17
|
},
|
|
18
18
|
{
|
|
19
19
|
"path": "references/models.md",
|
|
20
|
-
"sha256": "
|
|
21
|
-
"bytes":
|
|
20
|
+
"sha256": "3c2ee6f1b4ef3a3b8dea6cf12e494abafb4f0775a495851c0cf0389b105a026e",
|
|
21
|
+
"bytes": 34761
|
|
22
22
|
},
|
|
23
23
|
{
|
|
24
24
|
"path": "references/pipeline.md",
|
|
@@ -27,8 +27,8 @@
|
|
|
27
27
|
},
|
|
28
28
|
{
|
|
29
29
|
"path": "SKILL.md",
|
|
30
|
-
"sha256": "
|
|
31
|
-
"bytes":
|
|
30
|
+
"sha256": "f0e66e0790b3c935f352e8f2e88d2facd6d3cd6fc0246631221c545eaf90c6ed",
|
|
31
|
+
"bytes": 34469
|
|
32
32
|
}
|
|
33
33
|
]
|
|
34
34
|
}
|
|
@@ -70,7 +70,7 @@ Every `videodraft models image|video|audio --json` response carries a top-level
|
|
|
70
70
|
|
|
71
71
|
**Videos:**
|
|
72
72
|
|
|
73
|
-
- `gemini-omni-1.1-flash`: general default for 3-10s generation and uploaded source edits up to 10s. It supports first and last frames, up to 10 total image inputs, up to 3 creative reference videos of at most 3 seconds each, uploaded-video extension, and
|
|
73
|
+
- `gemini-omni-1.1-flash`: general default for 3-10s generation and uploaded source edits up to 10s. It supports first and last frames, up to 10 total image inputs, up to 3 creative reference videos of at most 3 seconds each on `--video-task generate` ONLY, uploaded-video extension, and continuation of an earlier generation through `--previous-interaction-id`. Output is 360p/720p/1080p/4K at 3/10/15/30 cr/s with audio always on. Use `--source-video` for the uploaded edit/extension source; extension sources must be 1-30s. Edit and extend accept EXACTLY ONE input video, so `--ref-video` cannot accompany `--source-video` or `--previous-interaction-id`; only `--ref` images can. One `--ref-video` with no separate source remains a legacy source edit. A previous interaction defaults to conversational edit; add `--video-task extend` or `--extend` with an explicit 3-10s duration to append at the end. A continuation resolves to the prior turn's output and is submitted as an ordinary source, so the same limits apply to it: at most 30s to extend, at most 10s to edit. 40 seconds total is reachable, but only by extending from a source of 30s or less, so a ladder dead-ends once it passes 30s. New dialogue can only be added when the SOURCE video is silent; adding speech on top of a source that already contains speech is refused with "the model is currently unable to process speech edits". The server safely measures creative-reference durations, or you can repeat `--ref-video-duration` when a host blocks metadata probing. Fal BYOK supports its currently callable v1.1 generation and basic-edit endpoints at zero VideoDraft credits, but Fal does not expose continuation/extension or mixed source-edit references as callable endpoints and VideoDraft must never fall back to paid Google.
|
|
74
74
|
- `grok-imagine-video-1.5`: 1-15s text, first-frame, or 1-7 reference-image generation with native audio. Text/first-frame modes support 480p, 720p, and 1080p; reference mode supports 480p/720p. Cite references as `<IMAGE_0>` through `<IMAGE_6>`. It has no last frame, seed, negative prompt, quality tier, reference video, or reference audio.
|
|
75
75
|
- `minimax-h3`: 480p/768p/2K/4K (5/6/13/16 cr/s, 768p default), native stereo audio, and 5-15s text, first/last-frame, or mixed-reference generation. Reference mode accepts up to 9 images, 3 videos, and 3 audio clips, with at most 12 files total. Cite them as `Image 1`, `Video 1`, and `Audio 1` in array order.
|
|
76
76
|
- `minimax-h3-max`: 480p/768p pricing (5/8 cr/s, 768p default), native audio, and 5-15s text or first/last-frame generation. It supports a reproducibility seed, a safety checker, and `disabled` / `balanced` / `quality` prompt expansion. It does not accept reference media.
|
|
@@ -79,7 +79,7 @@ Every `videodraft models image|video|audio --json` response carries a top-level
|
|
|
79
79
|
- `seedance-2`: 11-15s, video/audio/mixed references, wider ratios, selectable audio, or first/last frames. Use `mini` for cost, `fast` for speed, `standard` for quality or 1080p/4K.
|
|
80
80
|
- `seedance-2.5`: 4-30s single takes and up to 50 references (30 image, 10 video, 10 audio). Same modes as 2.0, one quality tier, 480p/720p/1080p. Reach for it when a shot must run past 15s or carry more references than 2.0 allows.
|
|
81
81
|
- `kling-v3-turbo`: fast polished 3-15s with first frame, multi-prompt, and audio, but no elements. `kling-o3`: reference images plus structured image/video elements, first/last frames, multi-prompt, audio control, or 4K. O3 allows 7 combined image references and image-backed elements, reduced to 4 combined items when a video-backed element is present. `kling-3.0`: image-to-video can use structured image/video elements and bind a custom Kling voice ID to either element form. Kling 2.6 Pro uses top-level voice IDs cited as `<<<voice_1>>>` and `<<<voice_2>>>`.
|
|
82
|
-
- Existing-video edits use `videodraft edit video`, not generic generation. `gemini-omni-1.1-flash` is the preferred edit model and is chosen automatically when you omit `--model`: source up to 10s, up to 10 reference images,
|
|
82
|
+
- Existing-video edits use `videodraft edit video`, not generic generation. `gemini-omni-1.1-flash` is the preferred edit model and is chosen automatically when you omit `--model`: source up to 10s, up to 10 reference images, 360p/720p/1080p/4K with audio. Creative `--ref-video` inputs are NOT accepted on an edit, because an edit takes exactly one input video; use `generate video --video-task generate` to guide a new clip with video references instead. Omitting `--model` on a longer or unmeasurable source spends nothing and prints a priced menu (what each model edits, what it drops, what it costs) so you can put the choice to the user. **Truncation:** only Gemini refuses an over-length source. Happy Horse silently edits just the first 15s, Kling O3 the first 10s, and Grok the first 8s. The command warns you when that will happen; always relay it to the user. Gemini regenerates the audio track, so use Happy Horse or Kling O3 with `--preserve-audio` when the source audio must survive. Choose Grok for the cheapest prompt-only edit, Happy Horse for up to 5 image references, or Kling O3 for controlled reference-image edits.
|
|
83
83
|
- To edit a source longer than 10s without losing its tail, cut it into <=10s pieces in the native editor, edit each with Gemini, and reassemble. VideoDraft has no server-side split/concat, so this path needs `videodraft_editor`.
|
|
84
84
|
- Kling O3 also has a reference-generation mode. Use `videodraft generate video --model kling-o3-video-ref-edit` with exactly one `--ref-video` to generate a new guided clip; use `videodraft edit video` when changing the source itself. Use Wan 3.0 for new mixed-reference generation, not source-video editing.
|
|
85
85
|
- Motion transfer uses `videodraft edit motion` with Kling V3 by default, or Kling 2.6 when explicitly requested or lower cost matters. It requires a subject image and a motion-reference video.
|
|
@@ -36,19 +36,18 @@ Submit-then-collect parallelizes server-side generation; the single multi-id `wa
|
|
|
36
36
|
|
|
37
37
|
Extend an uploaded clip or conversationally edit an official Google interaction with Gemini Omni 1.1 Flash:
|
|
38
38
|
|
|
39
|
-
Each extension appends 3-10 seconds at the end
|
|
39
|
+
Each extension appends 3-10 seconds at the end. Edit and extend accept EXACTLY ONE input video, so `--ref-video` cannot accompany `--source-video` or `--previous-interaction-id`; only `--ref` images can. New dialogue is allowed only when the source video is silent, so a chain of spoken beats must carry its speech in the first turn. A continuation is submitted as the prior turn's output, so 40 seconds total is reachable only by extending from a source of 30s or less, and a ladder dead-ends past 30s.
|
|
40
40
|
|
|
41
41
|
```bash
|
|
42
|
-
# Uploaded-video extension
|
|
42
|
+
# Uploaded-video extension. Reference IMAGES are fine here; reference videos are not.
|
|
43
43
|
videodraft generate video --model gemini-omni-1.1-flash \
|
|
44
|
-
--source-video ./ending.mp4 --ref
|
|
45
|
-
--
|
|
44
|
+
--source-video ./ending.mp4 --ref ./wardrobe.png \
|
|
45
|
+
--extend --duration 6 --resolution 1080p \
|
|
46
46
|
--download ./media/extended.mp4
|
|
47
47
|
|
|
48
|
-
# Conversational edit from the interaction_id
|
|
48
|
+
# Conversational edit from the interaction_id of an earlier generation. Add --extend to lengthen instead.
|
|
49
49
|
videodraft generate video --model gemini-omni-1.1-flash \
|
|
50
50
|
--previous-interaction-id "$INTERACTION_ID" \
|
|
51
|
-
--ref-video ./new-performance-reference.mp4 --ref-video-duration 2.5 \
|
|
52
51
|
--resolution 720p \
|
|
53
52
|
--download ./media/continued.mp4
|
|
54
53
|
```
|
|
@@ -34,7 +34,7 @@ Use `--num 1..4` for variations of one prompt in a single call. Never loop separ
|
|
|
34
34
|
|
|
35
35
|
| Need | Choose | Important limits |
|
|
36
36
|
| ---------------------------------------------------------------------------------------------------------- | ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
37
|
-
| Most generation, first/last-frame, mixed-reference, source-edit, or extension requests | `gemini-omni-1.1-flash` | 3-10s output; 360p/720p/1080p/4K at 3/10/15/30 cr/s; audio always; up to 10 image inputs and 3 reference videos <=3s; edit source <=10s;
|
|
37
|
+
| Most generation, first/last-frame, mixed-reference, source-edit, or extension requests | `gemini-omni-1.1-flash` | 3-10s output; 360p/720p/1080p/4K at 3/10/15/30 cr/s; audio always; up to 10 image inputs, and 3 reference videos <=3s on generate only; edit source <=10s; continuation or 3-10s extension of a 1-30s source, 40s total reachable only by extending from <=30s |
|
|
38
38
|
| Grok 1.5 text, first-frame, or 1-7 image-reference clips with native audio and optional 1080p | `grok-imagine-video-1.5` | 1-15s; 480p/720p/1080p for text/first-frame; references are 480p/720p only; no last frame |
|
|
39
39
|
| Unified text, first/last-frame, mixed-media, document, or webpage reference generation | `wan-3.0` | 2-30s or auto; 480p/720p/1080p at 7/14/28 cr/s; 10 image, 5 video, 5 audio refs, 20 media files total; document/web refs require thinking |
|
|
40
40
|
| 480p/768p/2K/4K with native stereo audio, first/last frames, or mixed image/video/audio references | `minimax-h3` | 5-15s; 5/6/13/16 cr/s by resolution; up to 9 image, 3 video, 3 audio refs, 12 files total; first 5 reference images free then 8 cr each; reference video/audio each total <=15s |
|
|
@@ -56,9 +56,9 @@ Routing rules:
|
|
|
56
56
|
- Grok 1.5 always generates native audio. Do not pass `--no-audio`, `--seed`, `--negative`, or `--quality`.
|
|
57
57
|
- Around 11-15 seconds with native audio: use MiniMax H3, Kling, or Seedance, not Gemini.
|
|
58
58
|
- One existing source video that should be edited: use `videodraft edit video`, which auto-selects Gemini Omni 1.1 Flash for a source up to 10s. Generic `generate video --ref-video` retains source-edit behavior when `--video-task` is omitted, but the dedicated edit command is clearer.
|
|
59
|
-
- Video supplied as a creative reference: Gemini Omni 1.1 Flash accepts up to 3 videos of at most 3 seconds each, including mixed image and video input.
|
|
59
|
+
- Video supplied as a creative reference: Gemini Omni 1.1 Flash accepts up to 3 videos of at most 3 seconds each, including mixed image and video input. Creative video references are accepted on `--video-task generate` ONLY: edit and extend take exactly one input video, so `--ref-video` cannot accompany `--source-video` or `--previous-interaction-id`. Use Wan 3.0 for up to 5 ordered video/audio references and 1080p, MiniMax H3 for 2K/4K, or Seedance 2.0 when quality-tier control matters.
|
|
60
60
|
- First and last frame control: Gemini Omni 1.1 Flash, Wan 3.0, MiniMax H3, MiniMax H3 Max, Seedance, Kling O3, and Kling 3.0 support it. For Gemini, every start/end frame counts toward the 10-image total, leaving up to 9 references with a start frame or 8 with both frames.
|
|
61
|
-
- Extension with a 1-30s uploaded source uses Gemini Omni 1.1 Flash plus `--source-video` and `--video-task extend` (or `--extend`). Pass an explicit 3-10s output duration; the model appends it at the end, up to 40 seconds total.
|
|
61
|
+
- Extension with a 1-30s uploaded source uses Gemini Omni 1.1 Flash plus `--source-video` and `--video-task extend` (or `--extend`). Pass an explicit 3-10s output duration; the model appends it at the end, up to 40 seconds total. Creative `--ref-video` clips CANNOT accompany the source: edit and extend take exactly one input video. New dialogue is allowed only when the source video is silent; adding speech on top of a source that already has speech is refused with "the model is currently unable to process speech edits". Continuation uses `--previous-interaction-id <id>` and defaults to a conversational edit; `--video-task extend` lengthens it. A continuation is submitted as the prior turn's output video, so it obeys the same source limits (<=30s to extend, <=10s to edit) and the same silent-source dialogue rule. 40s total is reachable only by extending from a source of 30s or less, so a ladder dead-ends past 30s. Fal returns interaction IDs, but its currently published callable v1.1 endpoints do not expose continuation/extension or mixed source-edit references; never infer a route or silently fall back to paid Google.
|
|
62
62
|
- Wan 3.0 reference mode and frame mode are separate. It accepts 10 images, 5 videos, and 5 audio clips, at most 20 media files total. Video and audio each total at most 15 seconds. `--file-url` and `--web-url` require `--thinking`. `--auto-duration` cannot be combined with `--duration`.
|
|
63
63
|
- MiniMax H3 reference mode and first-plus-last-frame mode are separate. Audio cannot be the only reference. Address references as `Image 1`, `Video 1`, and `Audio 1` in array order.
|
|
64
64
|
- Seedance reference mode and first-plus-last-frame mode are separate. Do not promise reference video/audio plus a last frame in one generation.
|
|
@@ -158,7 +158,7 @@ Direct Fabric text/audio and Sync Labs do not use the managed avatar record. The
|
|
|
158
158
|
|
|
159
159
|
- Images: per image (× `--num`). Matrix-priced models (GPT-Image, Nano Banana Pro, Seedream v5 Pro) vary by resolution/quality.
|
|
160
160
|
- Video: usually credits/second × duration; rate depends on model + resolution + quality + native audio on/off.
|
|
161
|
-
- Gemini Omni 1.1 Flash: 3 / 10 / 15 / 30 credits per output second at 360p / 720p / 1080p / 4K (720p default).
|
|
161
|
+
- Gemini Omni 1.1 Flash: 3 / 10 / 15 / 30 credits per output second at 360p / 720p / 1080p / 4K (720p default). Extend appends an explicit 3-10 seconds to a 1-30s source, whether uploaded or resolved from a prior interaction; 40 seconds total is reachable only while the source stays at or under 30s. Fal BYOK generation and basic source edits cost zero VideoDraft credits; continuation, extension, and separate creative references on a source edit are unavailable under Fal BYOK.
|
|
162
162
|
- MiniMax H3: 5 / 6 / 13 / 16 credits per output second at 480p / 768p / 2K / 4K (768p default). The first 5 reference images are included, then 8 credits for each additional image. Reference video and reference audio are NOT billed.
|
|
163
163
|
- MiniMax H3 Max: 5 credits per output second at 480p or 8 credits per second at 768p (default).
|
|
164
164
|
- Wan 3.0: 7 / 14 / 28 credits per output second at 480p / 720p / 1080p. Auto duration reserves 30 seconds and reconciles unused credits from the provider-reported output duration. Fal BYOK charges zero VideoDraft credits.
|