venice-video-harness 2.18.0 → 2.20.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/venice-agent-guide/SKILL.md +8 -0
- package/.agents/skills/venice-video-model-routing/SKILL.md +75 -0
- package/AGENTS.md +49 -0
- package/CHANGELOG.md +150 -0
- package/README.md +112 -1
- package/capabilities.json +202 -2
- package/dist/agent/guide.d.ts.map +1 -1
- package/dist/agent/guide.js +11 -0
- package/dist/agent/guide.js.map +1 -1
- package/dist/agent/pipeline.d.ts +23 -0
- package/dist/agent/pipeline.d.ts.map +1 -1
- package/dist/agent/pipeline.js +26 -2
- package/dist/agent/pipeline.js.map +1 -1
- package/dist/mini-drama/choices.d.ts.map +1 -1
- package/dist/mini-drama/choices.js +2 -0
- package/dist/mini-drama/choices.js.map +1 -1
- package/dist/mini-drama/cli.d.ts.map +1 -1
- package/dist/mini-drama/cli.js +168 -3
- package/dist/mini-drama/cli.js.map +1 -1
- package/dist/mini-drama/loop-engine.d.ts +216 -0
- package/dist/mini-drama/loop-engine.d.ts.map +1 -0
- package/dist/mini-drama/loop-engine.js +713 -0
- package/dist/mini-drama/loop-engine.js.map +1 -0
- package/dist/mini-drama/montage.d.ts.map +1 -1
- package/dist/mini-drama/montage.js +4 -3
- package/dist/mini-drama/montage.js.map +1 -1
- package/dist/mini-drama/prompt-builder.d.ts +0 -11
- package/dist/mini-drama/prompt-builder.d.ts.map +1 -1
- package/dist/mini-drama/prompt-builder.js +72 -12
- package/dist/mini-drama/prompt-builder.js.map +1 -1
- package/dist/mini-drama/video-generator.d.ts +88 -1
- package/dist/mini-drama/video-generator.d.ts.map +1 -1
- package/dist/mini-drama/video-generator.js +105 -44
- package/dist/mini-drama/video-generator.js.map +1 -1
- package/dist/series/types.d.ts +32 -1
- package/dist/series/types.d.ts.map +1 -1
- package/dist/series/types.js +65 -3
- package/dist/series/types.js.map +1 -1
- package/dist/session/status.d.ts +12 -0
- package/dist/session/status.d.ts.map +1 -1
- package/dist/session/status.js +11 -0
- package/dist/session/status.js.map +1 -1
- package/dist/venice/models.d.ts +35 -0
- package/dist/venice/models.d.ts.map +1 -1
- package/dist/venice/models.js +100 -0
- package/dist/venice/models.js.map +1 -1
- package/dist/web/jobs.js +1 -1
- package/dist/web/jobs.js.map +1 -1
- package/dist/web/server.d.ts +19 -0
- package/dist/web/server.d.ts.map +1 -1
- package/dist/web/server.js +73 -1
- package/dist/web/server.js.map +1 -1
- package/dist/web/state.d.ts +3 -0
- package/dist/web/state.d.ts.map +1 -1
- package/dist/web/state.js +2 -0
- package/dist/web/state.js.map +1 -1
- package/dist/web/ui/dist/assets/index-BM-V1s_N.css +1 -0
- package/dist/web/ui/dist/assets/index-BQPyylPl.js +41 -0
- package/dist/web/ui/dist/assets/index-BbATN-Iz.js +41 -0
- package/dist/web/ui/dist/assets/index-CGIShJAr.js +41 -0
- package/dist/web/ui/dist/assets/index-CbErOQAn.js +41 -0
- package/dist/web/ui/dist/index.html +2 -2
- package/package.json +1 -1
|
@@ -38,6 +38,14 @@ in the other.
|
|
|
38
38
|
- State placement explicitly in every prompt: lock each location's landmark geography (`spatialAnchors`) and give every character shot a `blocking` field (position vs named anchors, screen side, depth, facing/eyeline). Keep screen sides and eyelines constant across a scene unless a movement is scripted.
|
|
39
39
|
- Prefer native model dialogue (Seedance 2.0, HappyHorse 1.1 with voice-donor references) over exact TTS lip-sync.
|
|
40
40
|
|
|
41
|
+
## Modes of operation — the linear pipeline is not the only path
|
|
42
|
+
- Default path: the gated pipeline (`venice-video pipeline`) — aesthetic → cast → episode → script → approve → storyboard → QA → render → assemble. Use it for a finished, identity-locked cut.
|
|
43
|
+
- Loop mode (`venice-video loop -p <project> -e <n> --mode <looping|production>`): once a shot script exists, render the whole plan continuously and watch it as a live browser loop that hot-swaps fresh takes. It SKIPS the storyboard/QA gates and writes only under `episodes/episode-NNN/loop/`, so it never touches the canonical cut. `looping` = creative flow, lower quality (Turbo 480P, disposable); `production` = gather usable, identity-locked takes (Max R2V @768P). `--mode` is REQUIRED non-interactively; `--budget` caps spend (default $2). This is NOT in the `pipeline` stage list — it is under `branches` in `pipeline --json`.
|
|
44
|
+
- **Loop mode plays AND renders at the same time — it does NOT pre-generate everything.** The browser loops the takes that already exist while the worker keeps generating new ones, swapping each shot's newer take in on the loop's **next pass** (never mid-clip). Two things make a running loop *look* pre-generated when it isn't: (a) it **pauses** when it hits `--budget` or you click Pause, then just replays what's on disk — raise `--budget` or pass `--unbounded` to keep it generating; (b) `--max-takes` is a per-shot ring buffer (default 3, older takes pruned), not a total.
|
|
45
|
+
- **Which mode evolves while you watch:** `looping` (Turbo, 480P) renders faster than it plays, so it visibly changes as you watch. `production`/create renders each shot slowly on Max R2V — it's for **gathering keeper takes**, not a continuously-evolving watch, and on a short/few-shot plan it can look static even while running (one take takes longer to render than a full loop cycle takes to play). Want "watch it keep changing"? Use `looping` with a higher/unbounded budget.
|
|
46
|
+
- Three video lanes, chosen per shot by the router (see the `venice-video-model-routing` skill): **t2v** (prompt only), **i2v** (animate a supplied START image via `image_url` — establishing/atmosphere shots, or chaining off a previous last frame), **R2V** (identity anchored to a reference stack — the default for character shots). "i2v" means a supplied first frame; "R2V" means `reference_image_urls`, not a start frame — do not conflate them.
|
|
47
|
+
- To animate a single image you already have (plain i2v, no project), the routing skill's bundled `scripts/venice-video.py --image <file> --model <...-image-to-video>` is the standalone path; the project pipeline is for multi-shot, consistency-first work.
|
|
48
|
+
|
|
41
49
|
## Where the full knowledge lives
|
|
42
50
|
- `AGENTS.md` — 49 rules and 28 production anti-patterns, shipped in the package.
|
|
43
51
|
- `.agents/skills/` — `venice-api`, `venice-video-model-routing`, `character-consistency`, `shot-composition`, `burn-in-subtitles`, `video-editing`, and more.
|
|
@@ -318,6 +318,73 @@ Generate consistent storyboard panels using two Venice models in sequence:
|
|
|
318
318
|
|
|
319
319
|
Construct prompts differently depending on the resolved model's capabilities:
|
|
320
320
|
|
|
321
|
+
### Simple-Prompt Models (MiniMax H3 Max, H3 Max Turbo)
|
|
322
|
+
|
|
323
|
+
Read this before anything else in this section: everything below assumes a model that
|
|
324
|
+
renders what it is told and drifts when it is not. The H3 Max pair is the opposite, and
|
|
325
|
+
the registry marks it with `promptStyle: 'simple'` (`modelWantsSimplePrompt(modelId)`).
|
|
326
|
+
These models compose their own coverage — framing, cutting, beat rhythm — from one plain
|
|
327
|
+
statement of intent, and the directorial stack fights the shot the model would otherwise
|
|
328
|
+
have chosen.
|
|
329
|
+
|
|
330
|
+
`buildVideoPrompt()` and `buildMontagePrompt()` already drop the heavy blocks for them, so
|
|
331
|
+
the adaptation is mostly a matter of not writing them back in by hand:
|
|
332
|
+
|
|
333
|
+
- **Dropped:** the authored `Blocking:` restatement, the location description and
|
|
334
|
+
`spatialAnchors` "Fixed layout (never rearrange)" line, the geography-hold /
|
|
335
|
+
no-mirroring lecture, and the full aesthetic string (the compact one is used instead).
|
|
336
|
+
- **Kept:** `@ImageN` identity declarations and role clauses on the R2V lane, the beat
|
|
337
|
+
description, dialogue with delivery, and the audio-exclusion suffix. Identity and look
|
|
338
|
+
still have to be bound; only the staging instructions go.
|
|
339
|
+
- **Dialogue is IMPROVISED, not scripted.** These models carry natural, continuous speech
|
|
340
|
+
across a whole generation and degrade when handed an exact line to recite. So in
|
|
341
|
+
native-dialogue mode the prompt gives the scripted line as INTENT — `[@Image1, voice,
|
|
342
|
+
delivery] conveys: "…"` plus a one-line "improvise naturally in character, keep the intent
|
|
343
|
+
and tone, don't recite word for word" note — instead of the directorial `[…]: "exact line"`
|
|
344
|
+
quote. `buildVideoPrompt` / `buildMontagePrompt` do this automatically via
|
|
345
|
+
`shouldImproviseDialogue(modelId, series)` (gated on `modelWantsSimplePrompt` +
|
|
346
|
+
`audioStrategy !== 'lip-sync'`). Do NOT hand-write exact quotes for these models expecting
|
|
347
|
+
them verbatim. The exception is **exact-lip-sync**, where the `audio_url` drives the exact
|
|
348
|
+
words, so the line stays verbatim. One consequence: captions must come from transcribing
|
|
349
|
+
the rendered audio, not from `script.json` (the model won't say it word for word).
|
|
350
|
+
- **Don't** add camera terms, shot lists, or `Lens switch.` lines per beat. State the
|
|
351
|
+
sequence in a sentence or two and let the model cut it — that instinct is why this is
|
|
352
|
+
the montage family.
|
|
353
|
+
- Resolution is pinned to **768P** by default (2K is a hard 400 — the inverse of plain
|
|
354
|
+
MiniMax H3, which is 2K-only). **480P is the draft tier** and is only selected via an
|
|
355
|
+
explicit `resolution` override on `renderVideoFile`. Durations run the 5-15s ladder.
|
|
356
|
+
`audio` is not configurable, so the field is omitted from the body entirely.
|
|
357
|
+
- Turbo ships **no R2V lane**; identity and lip-sync shots cross to
|
|
358
|
+
`minimax-h3-max-reference-to-video`.
|
|
359
|
+
|
|
360
|
+
**Loop mode uses this family, with two modes.** `venice-video loop` renders the whole shot
|
|
361
|
+
script continuously and plays it as a live browser loop that hot-swaps in fresh takes:
|
|
362
|
+
|
|
363
|
+
- `--mode watch` (default) — `minimax-h3-max-turbo-text-to-video` (or `-image-to-video` off
|
|
364
|
+
an existing panel) at 480P, ~$0.012/s. The cheapest lane; a disposable draft that does NOT
|
|
365
|
+
lock identity (Turbo has no R2V lane).
|
|
366
|
+
- `--mode create` — the real reference-first routing on the **non-Turbo** H3 Max family at
|
|
367
|
+
768P (~$0.024/s): character shots on `minimax-h3-max-reference-to-video` with the full
|
|
368
|
+
`@Image` reference stack + voice-donor audio (identity locked, takes usable), atmosphere on
|
|
369
|
+
Max i2v/t2v. Shots with missing references degrade to i2v/t2v.
|
|
370
|
+
|
|
371
|
+
Both skip the storyboard/QA gates and write only under `loop/`. See `LoopEngine`
|
|
372
|
+
(`src/mini-drama/loop-engine.ts`) and AGENTS.md rule 58. Watch is a preview, not a production
|
|
373
|
+
path; create is the "keep the good takes" path.
|
|
374
|
+
|
|
375
|
+
**The loop plays AND renders concurrently — it does not pre-generate.** The browser loops the
|
|
376
|
+
takes that already exist while the worker keeps generating new ones, swapping each shot's newer
|
|
377
|
+
take in on the loop's NEXT pass (never mid-clip). It only *looks* pre-generated when (a) it is
|
|
378
|
+
PAUSED — budget hit or Pause clicked — and just replays what is on disk (raise `--budget` or
|
|
379
|
+
pass `--unbounded` to keep it going), or (b) it is in `create` mode, whose slow Max R2V renders
|
|
380
|
+
can't keep up with a short loop cycle, so evolution lags. `watch` (looping/Turbo) renders faster
|
|
381
|
+
than it plays and visibly evolves as you watch; `create` (production) is for gathering keeper
|
|
382
|
+
takes, not a live-evolving watch. `--max-takes` is a per-shot ring buffer, not a total/stop.
|
|
383
|
+
|
|
384
|
+
The Creator app mirrors this: `VideoModelCapabilities.wantsSimplePrompt(id:)` gates the
|
|
385
|
+
same trimming in `ShotPromptBuilder`, and relaxes the `produce_shots` motion/length gate
|
|
386
|
+
so a correctly short H3 Max prompt isn't rejected as thin.
|
|
387
|
+
|
|
321
388
|
### Image-Tag R2V Models (Seedance 2.0 R2V Enhanced — default for all lanes)
|
|
322
389
|
|
|
323
390
|
- Replace character names in descriptions with `@Image1`, `@Image2` tokens via regex
|
|
@@ -454,6 +521,9 @@ Seedance 2.0 (now the default for both atmosphere and character shots) accepts *
|
|
|
454
521
|
- **Omitting `aspect_ratio` from R2V models:** Both Seedance R2V and Kling O3 R2V require `aspect_ratio`. If omitted, it defaults to `16:9` in code. Always pass `aspect_ratio` explicitly.
|
|
455
522
|
- **Sending `image_references`/`image_1` to `nano-banana-pro`:** Returns 400. The generation model does not accept reference payloads at all.
|
|
456
523
|
- **Sending invalid durations:** Seedance 2.0 accepts 4s/5s/8s/10s/12s/15s. Veo 3.1 accepts 4s/6s/8s. Duration auto-snap corrects this.
|
|
524
|
+
- **Sending `2K` to MiniMax H3 Max, or `768P` to plain MiniMax H3:** The two families share a name and invert on resolution. H3 is 2K-only; H3 Max and H3 Max Turbo top out at 768P and reject 2K. `video-generator.ts` pins each family, and the `-max` branch has to stay above the `minimax-h3` substring match or every H3 Max render 400s.
|
|
525
|
+
- **Reaching for `minimax-h3-max-turbo-reference-to-video`:** It doesn't exist ("Specified model not found"). Turbo has no R2V lane; identity shots route to `minimax-h3-max-reference-to-video`.
|
|
526
|
+
- **Sending `audio` to any MiniMax H3 family model:** `audioConfigurable: false` — audio is always generated and the field must be omitted from the body, not set to `true`.
|
|
457
527
|
- **Reference images below 300x300:** R2V models reject `reference_image_urls` and `elements` images smaller than 300x300 pixels. Never downscale character references below this threshold.
|
|
458
528
|
- **Seedance + non-seedream face images (no longer an issue, 2026-07):** Venice removed the restriction that Seedance 2.0 only accepts face-bearing input images from `seedream-v5-lite` / `seedream-v5-lite-edit`. Any image family now works for face-bearing inputs, so there's nothing to pair, reroute, or launder — the pre-flight gate is a no-op.
|
|
459
529
|
|
|
@@ -471,6 +541,11 @@ Seedance 2.0 (now the default for both atmosphere and character shots) accepts *
|
|
|
471
541
|
- **Sequential action in image descriptions:** Causes comic-panel layouts instead of single frames. Separate the single-frame panel description from the full video action description.
|
|
472
542
|
- **Vague body orientation:** Produces twisted poses. Always specify full-body direction explicitly (e.g., "seen entirely from behind", "facing camera directly").
|
|
473
543
|
|
|
544
|
+
### Prompt-Style Mismatches
|
|
545
|
+
|
|
546
|
+
- **Over-directing a simple-prompt model:** Hand-writing blocking, per-beat camera terms, `Lens switch.` lines, or geography-hold clauses into a MiniMax H3 Max prompt flattens the result — it stages its own coverage and the clauses fight it. The prompt builder strips these automatically; don't re-add them via authored shot fields expecting them to help.
|
|
547
|
+
- **Padding an H3 Max prompt to clear the motion gate:** The Creator's `produce_shots` money gate normally demands 12+ words and motion vocabulary. For simple-prompt models it only asks for a stated subject and setting. Inflating a short, correct prompt to satisfy the old bar is the failure, not the fix.
|
|
548
|
+
|
|
474
549
|
### Style Consistency Failures
|
|
475
550
|
|
|
476
551
|
- **Aesthetic description buried at end of prompt:** The model commits to a rendering style before reaching the style instructions, causing inconsistency between angles/shots. Always front-load style with a `STYLE:` prefix and add a `STYLE REMINDER:` suffix.
|
package/AGENTS.md
CHANGED
|
@@ -123,6 +123,44 @@ The full model registry lives in `src/venice/models.ts` with typed specs for eve
|
|
|
123
123
|
- **R2V is pure-reference-only.** Sending `image_url` (or `end_image_url`) alongside `reference_image_urls` is a hard 400: *"image_url and end_image_url cannot be combined with reference media for this model."* `minimax-h3-reference-to-video` is therefore in `MODELS_USING_IMAGE_TAGS`, which is what puts the generator in pure reference mode. It honors `@ImageN` tags — verified by paid render, both tagged characters landed on their assigned `@Image1` / `@Image2` slots.
|
|
124
124
|
- **Reference aspect influences output orientation, so keep a 16:9 plate in the stack.** With the harness's normal slot plan (1:1 character sheets + the 16:9 storyboard blocking plate) and `aspect_ratio: '16:9'`, a paid render returned a true 2560×1440. But a stack of uniformly portrait references returned 1440×1920 *despite* `aspect_ratio: '16:9'` — the requested ratio did not override them. Character-only H3 shots with no blocking plate are the orientation risk; check the first-frame contact sheet before assembling.
|
|
125
125
|
|
|
126
|
+
**MiniMax H3 Max / H3 Max Turbo (added 2026-09-03):**
|
|
127
|
+
- `minimax-h3-max-text-to-video` / `-image-to-video` / `-reference-to-video`, and
|
|
128
|
+
`minimax-h3-max-turbo-text-to-video` / `-image-to-video`.
|
|
129
|
+
- **Related to MiniMax H3 in name only.** Four differences, each of which costs
|
|
130
|
+
a render if you assume H3 behavior:
|
|
131
|
+
- **768P, and 2K is a hard 400** (`Expected '480P' | '768P'`) — the exact
|
|
132
|
+
inverse of base H3. The generator's resolution pin matches
|
|
133
|
+
`minimax-h3-max` *before* `minimax-h3` for this reason; do not reorder
|
|
134
|
+
those branches. 480P is the draft tier, 768P the finish.
|
|
135
|
+
- **They want plain prompts** (`promptStyle: 'simple'` in the registry).
|
|
136
|
+
These models stage their own framing, coverage, and cutting from a stated
|
|
137
|
+
intent, and the directorial stack overrides that instinct. `buildVideoPrompt`
|
|
138
|
+
and `buildMontagePrompt` drop blocking, the locked location description, and
|
|
139
|
+
the geography-hold paragraphs for them; identity (`@ImageN`), the beat, the
|
|
140
|
+
line, the sound, and a compact look survive. Use `modelWantsSimplePrompt(id)`
|
|
141
|
+
rather than an id check when adding new behavior.
|
|
142
|
+
- **`private` tier** (H3 is `anonymized`), and uncensored. Prompt cap 10000
|
|
143
|
+
chars, though the useful prompt is a couple of sentences.
|
|
144
|
+
- **Price.** $0.024/s for H3 Max and $0.012/s for Turbo at 768P, against
|
|
145
|
+
$0.10/s for base H3 — Turbo is the cheapest lane in the registry, cheap
|
|
146
|
+
enough that a 15s take is disposable: render several and pick.
|
|
147
|
+
- **Best used for montages and single-take storytelling.** This is where the
|
|
148
|
+
simple-prompt instinct pays: describe the sequence and let the model cut it.
|
|
149
|
+
Note the montage window now derives from the montage model's own ladder, so
|
|
150
|
+
H3 Max montages plan at 5-15s rather than Seedance's 30s.
|
|
151
|
+
- Shared with H3: the 5-15s ladder (4s is a hard 400) and native audio that is
|
|
152
|
+
**not** toggleable, so the generator omits the `audio` field entirely.
|
|
153
|
+
- **Turbo has no R2V lane.** `minimax-h3-max-turbo-reference-to-video` is
|
|
154
|
+
"Specified model not found", so the `minimax-h3-max-turbo` family routes
|
|
155
|
+
character-consistency and lip-sync shots to `minimax-h3-max-reference-to-video`.
|
|
156
|
+
- R2V is treated as pure-reference (in `MODELS_USING_IMAGE_TAGS`) like H3 R2V.
|
|
157
|
+
Note the difference from H3: `/video/quote` *accepted* `image_url` alongside
|
|
158
|
+
`reference_image_urls` here, but quote validates less than queue, and
|
|
159
|
+
pure-reference is the right mode regardless — it keeps compositional
|
|
160
|
+
authority with the reference stack and is what makes `@ImageN` resolve.
|
|
161
|
+
- i2v lanes inherit aspect from the start image and expose no `aspect_ratios`;
|
|
162
|
+
t2v and R2V accept `16:9 / 21:9 / 4:3 / 1:1 / 3:4 / 9:16`.
|
|
163
|
+
|
|
126
164
|
**Long Duration:**
|
|
127
165
|
- `longcat-image-to-video` / `longcat-distilled-image-to-video` (up to **30s**, no audio)
|
|
128
166
|
- `ltx-2-fast-image-to-video` / `ltx-2-v2-3-fast-image-to-video` (up to **20s**, up to 4K)
|
|
@@ -440,6 +478,10 @@ Use `POST /video/quote` (via `quoteVideo()`) to estimate costs before committing
|
|
|
440
478
|
|
|
441
479
|
57. **The browser UI is the built-in node web app — default to `venice-video web`.** When the operator asks for a browser UI, a dashboard, or to "open the harness in a browser," start the bundled local web app: `venice-video web` (browser dashboard + a whitelisted command runner over the workspace; binds localhost only, `http://127.0.0.1:3000` by default). Do NOT scaffold a separate/ad-hoc UI or point them at anything else. Per the workspace dev-server rule, kill existing node processes first and run on port 3000. The compiled front-end ships inside the npm package (`dist/web/ui/dist`); `npm run web:build` rebuilds it. Command lives in `src/mini-drama/cli.ts` (`web`), server in `src/web/server.ts`.
|
|
442
480
|
|
|
481
|
+
58. **Loop mode has two modes — a disposable Turbo draft (watch) and an identity-locked Max R2V create — both gate-skipping and money-capped (2026-09-04).** `venice-video loop -p <project> -e <n> [--mode watch|create]` boots the web UI (Loop tab) plus an in-process `LoopEngine` (`src/mini-drama/loop-engine.ts`) that renders every shot into `episodes/episode-NNN/loop/` and keeps regenerating fresh takes so the browser can watch the whole plan on repeat while it evolves. Both modes **bypass the references(only for watch) / storyboard / QA gates** — the only hard precondition is a shot script with ≥1 shot — and write ONLY under `loop/` (`shot-NNN--takeK.mp4` + `loop-manifest.json`), never touching canonical `scene-001/` renders or `series.json`, so a loop can run alongside real production. **The mode is the session's first, REQUIRED decision** — the `loop` command asks "is this for LOOPING (creative flow, lower quality) or PRODUCTION (gather usable shots, higher quality)?" It is a deliberate quality-vs-flow tradeoff, never a silent default: interactive `promptChoice` in a TTY, and a **hard error** in a non-interactive run with no `--mode` (agents MUST pass it). `--mode` accepts natural words — `looping`/`loop`/`fun`/`creative` → watch, `production`/`prod`/`gather` → create (`normalizeLoopMode`). Internally the modes are still `watch`/`create`. **The two modes are the point:** (a) **watch** (enjoyment / creative flow) = MiniMax H3 Max **Turbo** at 480P (~$0.012/s): the first generation is **t2v**, every shot after it **chains i2v off the previous last frame**, and it **NEVER uses R2V** (R2V renders are too slow for a loop, and Turbo has none anyway). It ignores panels and the reference stack — identity is NOT locked; it is a fast fun loop, never production-fidelity. (b) **create** (gather good shots for a project) = the real reference-first routing on the **non-Turbo** H3 Max family at 768P (~$0.024/s): character shots render on `minimax-h3-max-reference-to-video` with the full `@Image` reference stack + voice-donor audio (identity **locked**, takes usable), atmosphere shots on Max i2v/t2v, **each shot rendered independently (chaining OFF by default** — R2V and a start frame can't combine on MiniMax). Create degrades a character shot to i2v/t2v only when its references are missing on disk — so for create, generate character/location references first (rule 54). Both share the render primitive: create mode reuses `resolveShotReferenceInputs` + `ensureVoiceReferenceForShot` (exported from `video-generator.ts`, the SAME resolution `renderSingleShotUnit` uses), so its reference stack can't drift from the real pipeline. Other invariants: (c) **chaining default follows the mode:** watch chains (shot 1 t2v, every later shot i2v off the previous shot's current-take LAST frame via `extractLastFrame`, so the loop plays as one piece); create does NOT chain (each shot is independent R2V — chaining and R2V are mutually exclusive on MiniMax). `--no-chain` forces it off. Chaining uses a lean prompt because the start frame, not a reference stack, drives the render. (d) **Every take renders the model's full length (15s default)** — not the shot's scripted duration — for maximum footage/playback per render; override with `--duration`. (e) **The loop regenerates continuously** and does NOT settle after a fixed number of takes: `--max-takes` is a **ring buffer** (candidate takes kept per shot; older non-current takes are pruned and their files deleted so an infinite run can't fill the disk), NOT a stop condition. It stops only on Pause, `--once` (one pass), or the budget. (f) **Money:** billed at queue time; `--budget` (default $2) pauses the loop, and the UI **Resume** button (or a per-shot regenerate) authorizes another budget's worth via `engine.start()` — so "Start/Resume" always does something. `--unbounded` removes the budget cap (spends until stopped) but keeps the ring buffer. (g) Resolution is reachable because `renderVideoFile` honors an optional `resolution` override validated against the model's `resolutions` (it otherwise force-pins `768P` for every `minimax-h3-max*` id); **watch defaults 480P (the infinite-loop tier), create 768P**. (h) The engine shares the web server's `EventHub` and broadcasts `loop-updated` for instant hot-swap; the manifest (with `mode` + `chain`) also feeds `collectEpisodeState` so a plain `venice-video web` shows the last loop state. (i) **A shot that fails `MAX_CONSECUTIVE_SHOT_ERRORS` (3) times in a row is given up on** — marked `failed`, dropped from `pickNext`, no longer re-queued/re-billed (a manual `regenerate` clears it). This caps the money leak where a server-side-doomed shot (a MiniMax i2v face start frame — billed at queue time, then `/video/retrieve` 500s) is otherwise re-selected fewest-takes-first and re-billed every cycle. (j) **Face-continuity prompting (`--face-continuity`, default on):** prompts each chained character shot to END on the character's face so the next i2v continuation is smoother — but it is **auto-suppressed** when the chain i2v model rejects face start frames (`i2vRejectsFaceStartFrame`, i.e. all MiniMax i2v lanes; a face-ending frame becomes the next start frame and would trip the death in anti-pattern 31). So it is dormant on both current loop lanes and activates on a face-accepting i2v lane; for face loops today use create/R2V. When changing loop behavior, keep `loop-engine.ts`, the `loop` command in `cli.ts`, the `/api/projects/:slug/loop/*` endpoints + `WebServerOptions.loop` in `src/web/server.ts`, `resolveShotReferenceInputs`/`extractLastFrame`/`i2vRejectsFaceStartFrame` in `video-generator.ts`/`models.ts`, and the `LoopView` UI in sync.
|
|
482
|
+
|
|
483
|
+
59. **MiniMax H3 Max (simple-prompt) models IMPROVISE dialogue — the scripted line is intent, not a script (2026-09-04).** Unlike Seedance/Wan, which render the exact line you give them, the H3 Max family (`promptStyle: 'simple'`) performs markedly better carrying natural, continuous speech across a whole generation than reciting a verbatim quote — a fixed line fights the model the same way the directorial blocks do. So in **native-dialogue mode** the prompt builders render a speaker's line as INTENT — `[@ImageN, voice, delivery] conveys: "…"` plus a one-line "improvise naturally in character, keep the meaning and tone, don't recite word for word" note — instead of the directorial `[…]: "exact line"` quote. This is automatic via `shouldImproviseDialogue(modelId, series)` in `prompt-builder.ts` (`modelWantsSimplePrompt(modelId)` AND `videoDefaults.audioStrategy !== 'lip-sync'`), applied in both `buildVideoPrompt` (singles) and `buildMontagePrompt` (montage) — the two paths the H3 Max family renders through. Directorial models keep the exact quote; the legacy Seedance-native / Kling multi-shot builders (`buildMultiShotPrompt`) are untouched because those lanes are never simple-prompt. **The exception is exact-lip-sync**: there the `audio_url` drives the exact spoken words, so the line stays verbatim (the improv gate excludes `audioStrategy === 'lip-sync'`). **Consequence:** when a simple-prompt model improvises, burned/exported captions must be derived by transcribing the rendered audio (the editing pipeline's `silencedetect`/whisper path), NOT from `script.json` — the model will not say it word for word. When changing this, keep `shouldImproviseDialogue` / `formatDialogueLine` / `IMPROV_DIALOGUE_NOTE` in `prompt-builder.ts` and this rule in sync.
|
|
484
|
+
|
|
443
485
|
## Learned Anti-Patterns (Production Issues Log)
|
|
444
486
|
|
|
445
487
|
Issues discovered during production and their fixes. The agent should internalize these to avoid repeating them.
|
|
@@ -632,6 +674,13 @@ Issues discovered during production and their fixes. The agent should internaliz
|
|
|
632
674
|
**Root cause:** Seedance montage generations front-load transition frames so the beats can be cut apart; the per-beat cutter slices at the planned timestamps, so when the model's actual transition lands a few frames off the boundary, junk frames survive at a cut's head — most visibly on the first beat of unit 1, which becomes the film's opening frames.
|
|
633
675
|
**Fix:** `qa-videos` detects it programmatically (per-frame luma scan over each unit's first second; a spike that reverts within 3 frames is a flash, not a scene change — ffmpeg `signalstats`, zero API cost). Fix by re-rendering the unit or trimming the flagged frames off the head of that beat's cut before assembly; for the film's first shot, always eyeball frames 0-10 of the final master before delivery.
|
|
634
676
|
|
|
677
|
+
### 31. Loop Watch Mode: t2v Aspect 400, Chain Frame Past Stream End, Face Frames Die Server-Side (2026-09-04)
|
|
678
|
+
**Symptom:** First watch-mode loop run failed at three successive layers. (a) Every t2v queue 400'd with `aspect_ratio: Required`. (b) After that was fixed, chained i2v takes 400'd with `image_url is required` because the extracted last-frame PNG did not exist — yet `extractLastFrame` had not thrown. (c) Once chaining produced real frames, every chained render whose start frame contained a recognizable human face queued successfully, then died server-side: `/video/retrieve` 500s ("An unknown error occurred") forever. The engine's give-up-after-6-polls path cleared the pending record and re-queued fresh, billing ~$0.18 per attempt (4 attempts on one shot).
|
|
679
|
+
**Root cause:** (a) `renderVideoFile` set `aspect_ratio` only for `reference-to-video` and non-i2v Seedance; MiniMax H3 Max (Turbo) t2v REQUIRES it. (b) `extractLastFrame` probed the CONTAINER duration and sought to `duration - 0.05`; MiniMax writes an audio track slightly longer than the video stream, so the seek landed past the last decodable frame — and ffmpeg exits 0 having written no file. (c) MiniMax i2v accepts a face-bearing start frame at queue time (billed) but the render fails server-side, surfacing only as a retrieve 500. Faceless start frames (the robot) render fine; t2v (no input image) renders fine. There is no MiniMax equivalent of the Seedance 409 `needs_consent` handshake.
|
|
680
|
+
**Fix:** (a) `video-generator.ts` now sets `body.aspect_ratio` for every `text-to-video` model as well. (b) `extractLastFrame` probes the `v:0` stream duration and steps back in widening offsets (0.1/0.3/0.6/1.0s) until the PNG actually exists on disk, throwing otherwise. (c) The face-death is Venice-side, but the fallout is now capped: `LoopEngine` counts consecutive render failures per shot and **gives up on a shot after `MAX_CONSECUTIVE_SHOT_ERRORS` (3)** — it is marked `failed`, dropped from `pickNext`, and no longer re-queued/re-billed (a manual `regenerate` revives it). This is the real fix for the money leak — note the re-queue was NOT `isQueueGoneError` (400/404/410, not 500): the loop force-requeues, so a doomed shot was re-selected fewest-takes-first by the *scheduler* every cycle. Watch mode with human faces still hits the death per chained take, so `--no-chain` remains the clean workaround; create mode routes character shots to R2V (face **references**, not a start frame) and, when refs are missing, now degrades to **t2v** instead of a face-bearing i2v (`i2vRejectsFaceStartFrame` in `models.ts`). **Verified 2026-09-04:** MiniMax **R2V** accepts face-bearing reference sheets — a 5s render off a character `front.png` succeeded in ~13s (`scripts/probe-minimax-r2v-face.ts`). So only i2v *start frames* die on a face, not R2V *references*: create-mode character loops are viable, and create/R2V is the working path for smooth character-face loops.
|
|
681
|
+
**Related — face-continuity prompting:** the loop can prompt each chained shot to END on the character's face for smoother i2v continuations (`--face-continuity`, on by default), but it is **auto-suppressed on any i2v model that rejects face start frames** (all MiniMax i2v lanes), because a face-ending frame is the next shot's start frame and would trip exactly this death. So today it is dormant on both loop lanes; it activates on a non-MiniMax i2v lane or once Venice fixes MiniMax i2v. For face continuity now, use create mode (R2V locks the face from the sheets, no i2v chaining).
|
|
682
|
+
**Files:** `src/mini-drama/video-generator.ts`, `src/mini-drama/loop-engine.ts`, `src/venice/models.ts`, `scripts/probe-minimax-r2v-face.ts`
|
|
683
|
+
|
|
635
684
|
## Output
|
|
636
685
|
|
|
637
686
|
Generated project output belongs in:
|
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,155 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 2.20.0 — 2026-09-04
|
|
4
|
+
|
|
5
|
+
### Fixed
|
|
6
|
+
|
|
7
|
+
- **Loop watch mode: t2v `aspect_ratio` 400 + chain-frame past stream end (#25).**
|
|
8
|
+
`renderVideoFile` now sends `aspect_ratio` for every `text-to-video` model
|
|
9
|
+
(MiniMax H3 / H3 Max t2v **require** it — every opening watch take 400'd
|
|
10
|
+
without it). `extractLastFrame` now probes the `v:0` **stream** duration (not
|
|
11
|
+
the container, whose longer audio track pushed the seek past the last
|
|
12
|
+
decodable frame, making ffmpeg exit 0 with no file) and steps back in widening
|
|
13
|
+
offsets until the PNG lands, throwing otherwise so callers fall back to an
|
|
14
|
+
unchained render instead of queueing an `image_url`-less i2v.
|
|
15
|
+
- **Loop money leak: a persistently-failing shot is now given up on.**
|
|
16
|
+
`LoopEngine` tracks consecutive render failures per shot and, after
|
|
17
|
+
`MAX_CONSECUTIVE_SHOT_ERRORS` (3), marks the shot `failed`, drops it from the
|
|
18
|
+
scheduler, and stops re-queueing it. Previously a server-side-doomed shot (a
|
|
19
|
+
MiniMax i2v start frame with a human face — billed at queue time, then
|
|
20
|
+
`/video/retrieve` 500s forever) was re-selected fewest-takes-first and
|
|
21
|
+
re-billed every cycle until the budget paused the loop. A manual `regenerate`
|
|
22
|
+
revives a given-up-on shot. `failed`/`lastError` are surfaced on the manifest
|
|
23
|
+
and `loop-updated` events.
|
|
24
|
+
- **Create mode no longer degrades a character shot to a face-killing i2v.**
|
|
25
|
+
When a create-mode character shot has no R2V references on disk it now degrades
|
|
26
|
+
to **t2v** rather than i2v-off-a-panel on models that reject face start frames
|
|
27
|
+
(MiniMax) — a character panel almost always shows a face, which would 500.
|
|
28
|
+
|
|
29
|
+
### Added
|
|
30
|
+
|
|
31
|
+
- **`venice-video loop --face-continuity` (on by default).** Prompts each chained
|
|
32
|
+
character shot to end on the character's face so the next clip's i2v
|
|
33
|
+
continuation is smoother. **Auto-suppressed** on i2v models that reject face
|
|
34
|
+
start frames (all MiniMax i2v lanes — a face-ending frame is the next start
|
|
35
|
+
frame and would trip the server-side death), so it can't break the watch loop;
|
|
36
|
+
it activates on a face-accepting i2v lane. `--no-face-continuity` to disable.
|
|
37
|
+
New `i2vRejectsFaceStartFrame()` capability in `models.ts`.
|
|
38
|
+
- **`scripts/probe-minimax-r2v-face.ts`** — one paid 5s probe to settle whether
|
|
39
|
+
MiniMax H3 Max **R2V** accepts face-bearing reference sheets (the open question
|
|
40
|
+
behind create-mode character loops). **Verified 2026-09-04: it does** (render
|
|
41
|
+
succeeded in ~13s) — only i2v *start frames* die on a face, so create-mode
|
|
42
|
+
character loops are viable.
|
|
43
|
+
- Injectable `errorBackoffMs` on `LoopEngine` (test seam).
|
|
44
|
+
|
|
45
|
+
See AGENTS.md rule 58 and anti-pattern 31.
|
|
46
|
+
|
|
47
|
+
## 2.19.0 — 2026-09-04
|
|
48
|
+
|
|
49
|
+
### Added
|
|
50
|
+
|
|
51
|
+
- **`venice-video loop` — infinite loop mode (watch + create).** Boots the local
|
|
52
|
+
web UI (a new **Loop** tab) plus an in-process `LoopEngine`
|
|
53
|
+
(`src/mini-drama/loop-engine.ts`) that renders the approved shot script
|
|
54
|
+
continuously and plays the whole plan as a live browser loop, hot-swapping each
|
|
55
|
+
shot in as its take finishes and regenerating fresh takes while it plays. Two
|
|
56
|
+
modes, **chosen by a required decision at session start** — "is this for LOOPING
|
|
57
|
+
(creative flow, lower quality) or PRODUCTION (gather usable shots, higher
|
|
58
|
+
quality)?" — asked interactively in a terminal and a hard error in a
|
|
59
|
+
non-interactive run with no `--mode` (never a silent default). `--mode` accepts
|
|
60
|
+
natural words (`looping`/`loop`/`fun`, `production`/`prod`/`gather`). The modes:
|
|
61
|
+
**watch** (looping) renders
|
|
62
|
+
**MiniMax H3 Max Turbo at 480P** — the first generation is t2v, every later shot
|
|
63
|
+
chains i2v off the previous last frame, and it **never uses R2V** (too slow for
|
|
64
|
+
a loop); identity is NOT locked. **create** renders the **non-Turbo** H3 Max
|
|
65
|
+
family at 768P using the real reference-first routing — character shots on
|
|
66
|
+
`minimax-h3-max-reference-to-video` with the full `@Image` reference stack +
|
|
67
|
+
voice-donor audio (identity locked, takes usable), each shot rendered
|
|
68
|
+
independently, degrading to i2v/t2v when references are missing.
|
|
69
|
+
- **Continuous by default:** the loop regenerates forever (it does not stop
|
|
70
|
+
after N takes); `--max-takes` is a ring buffer (candidate takes kept per shot,
|
|
71
|
+
older ones pruned + deleted), not a stop condition. It stops only on Pause,
|
|
72
|
+
`--once`, or the budget.
|
|
73
|
+
- **Last-frame chaining** defaults to the mode (on for watch, off for create,
|
|
74
|
+
since R2V and a start frame can't combine on MiniMax); `--no-chain` forces it
|
|
75
|
+
off. Shot 1 renders normally, every later chained shot renders i2v off the
|
|
76
|
+
previous shot's last frame so the loop plays as one continuous piece.
|
|
77
|
+
- **Full-length takes:** every generation renders the model max (15s default),
|
|
78
|
+
override with `--duration`.
|
|
79
|
+
- **Budget as pause, not a hard stop:** `--budget` (default $2) pauses the loop;
|
|
80
|
+
the UI **Resume** button (and per-shot regenerate) authorizes another budget's
|
|
81
|
+
worth, so "Start/Resume" always does something. `--unbounded` removes the cap.
|
|
82
|
+
Both modes skip the storyboard/QA gates, write only under
|
|
83
|
+
`episodes/episode-NNN/loop/` (per-take mp4s + `loop-manifest.json`), and never
|
|
84
|
+
touch canonical renders or `series.json`. Resumable across restarts; pins and
|
|
85
|
+
per-shot regenerate from the UI. New
|
|
86
|
+
`/api/projects/:slug/loop/{state,start,stop,pin,regenerate}` endpoints, a shared
|
|
87
|
+
`EventHub` `loop-updated` event, and loop state surfaced through
|
|
88
|
+
`collectEpisodeState`. See AGENTS.md rule 58.
|
|
89
|
+
- **MiniMax H3 Max (simple-prompt) models now improvise dialogue.** In
|
|
90
|
+
native-dialogue mode, `buildVideoPrompt` and `buildMontagePrompt` render a
|
|
91
|
+
speaker's scripted line as INTENT for these models — `[@ImageN, voice,
|
|
92
|
+
delivery] conveys: "…"` plus a "improvise naturally in character, keep the
|
|
93
|
+
meaning and tone, don't recite word for word" note — instead of the
|
|
94
|
+
directorial `[…]: "exact line"` quote. Gated on
|
|
95
|
+
`shouldImproviseDialogue(modelId, series)` (`modelWantsSimplePrompt` AND
|
|
96
|
+
`audioStrategy !== 'lip-sync'`). Seedance/Wan/Kling keep the exact quote, and
|
|
97
|
+
exact-lip-sync keeps the exact line (the `audio_url` drives the words). See
|
|
98
|
+
AGENTS.md rule 59. Note: captions for these shots should come from
|
|
99
|
+
transcribing the rendered audio, not `script.json`.
|
|
100
|
+
- **`resolveShotReferenceInputs` / `ensureVoiceReferenceForShot` exported from
|
|
101
|
+
`video-generator.ts`.** The per-shot reference/scene/voice resolution block was
|
|
102
|
+
extracted from `renderSingleShotUnit` into a shared `resolveShotReferenceInputs`
|
|
103
|
+
helper (behavior-identical; `renderSingleShotUnit` now calls it) so loop create
|
|
104
|
+
mode resolves the exact same reference stack the real pipeline does instead of a
|
|
105
|
+
divergent copy.
|
|
106
|
+
- **`resolution` override on `renderVideoFile`.** Honored only when the model
|
|
107
|
+
lists it (validated against the registry, else the family default applies), so
|
|
108
|
+
loop mode can pin H3 Max Turbo to its 480P draft tier without disturbing the
|
|
109
|
+
`768P`/`2K`/`720p` auto-pins every other path relies on.
|
|
110
|
+
- **MiniMax H3 Max + H3 Max Turbo** (probe-verified 2026-09-03):
|
|
111
|
+
`minimax-h3-max-text-to-video` / `-image-to-video` / `-reference-to-video`
|
|
112
|
+
and `minimax-h3-max-turbo-text-to-video` / `-image-to-video`. 768P/480P,
|
|
113
|
+
5-15s, native non-toggleable audio, `private`, uncensored. $0.024/s and
|
|
114
|
+
$0.012/s at 768P (base H3 is $0.10/s). Turbo ships no R2V lane. Two new
|
|
115
|
+
video families — `minimax-h3-max` and `minimax-h3-max-turbo` — both routing
|
|
116
|
+
identity to `minimax-h3-max-reference-to-video`, which is the only lane in
|
|
117
|
+
the pair with `audio_input: true` and therefore also the family lip-sync model.
|
|
118
|
+
- **`promptStyle` on `VideoModelSpec`, and a simple-prompt path in the prompt
|
|
119
|
+
builder.** H3 Max models are `promptStyle: 'simple'`: they stage their own
|
|
120
|
+
framing, coverage, and cutting from a plain statement of intent, and the
|
|
121
|
+
directorial stack flattens that. `buildVideoPrompt` and `buildMontagePrompt`
|
|
122
|
+
now drop spatial blocking, the locked location description, and the
|
|
123
|
+
geography-hold paragraphs for these models, and use the compact aesthetic
|
|
124
|
+
instead of the full one. Identity declarations, reference role clauses, the
|
|
125
|
+
beat, dialogue, sound, and the hard-cut instruction are kept — those are not
|
|
126
|
+
inferable. Gate new behavior on `modelWantsSimplePrompt(id)`, not id checks.
|
|
127
|
+
Every other family is unchanged (`'directorial'` is the default).
|
|
128
|
+
|
|
129
|
+
### Fixed
|
|
130
|
+
|
|
131
|
+
- **The resolution pin no longer sends H3 Max to 2K.** `renderVideoFile` matched
|
|
132
|
+
`minimax-h3` by substring, so every `minimax-h3-max-*` render would have been
|
|
133
|
+
pinned to `2K` — a hard 400 on those models. The `minimax-h3-max` branch now
|
|
134
|
+
precedes it and pins `768P`.
|
|
135
|
+
- **Montage windows are bounded by the montage model, not a flat 30s.**
|
|
136
|
+
`resolveMontageMaxDurationSec` accepted a `montageModel` and ignored it,
|
|
137
|
+
always returning Seedance 2.5's 30s ceiling. Pointing `montageModel` at any
|
|
138
|
+
shorter-ladder model (H3 Max tops out at 15s) therefore planned 30s units
|
|
139
|
+
that all failed `assertShotDurationsValid` *after* the plan was written. The
|
|
140
|
+
ceiling now comes from the model's `maxDurationSec` and an explicit
|
|
141
|
+
`montageMaxDurationSec` is clamped to it. New `resolveMontageMinDurationSec`
|
|
142
|
+
does the same for the floor, so H3 Max montages respect its 5s ladder start
|
|
143
|
+
instead of Seedance's 4s.
|
|
144
|
+
- **Registry resolution ORDER is now load-bearing, and the Creator app honors
|
|
145
|
+
it.** The app defaults to the first allowed resolution whenever a plan's own
|
|
146
|
+
value isn't offered, and Venice's live `/models` lists H3 Max as
|
|
147
|
+
`["480P", "768P"]` — so the app would have rendered every H3 Max shot at its
|
|
148
|
+
draft tier. H3 Max entries list `['768P', '480P']` deliberately, and the app
|
|
149
|
+
reorders the live list to follow the manifest
|
|
150
|
+
(`VideoModelCapabilities.preferredResolutionOrder`). When editing a
|
|
151
|
+
`resolutions` array, treat position 0 as the default, not as arbitrary.
|
|
152
|
+
|
|
3
153
|
## 2.18.0 — 2026-08-21
|
|
4
154
|
|
|
5
155
|
### Added
|
package/README.md
CHANGED
|
@@ -552,6 +552,8 @@ Live catalog (synced against `GET /api/v1/models?type=video` — 103 entries). F
|
|
|
552
552
|
| **HappyHorse 1.1** | i2v, R2V (up to 9 refs) | t2v | 15s | Yes (joint single-pass, 7-lang phoneme lip-sync) | **#1 blind-preference T2V + I2V** (Alibaba 15B). 3-15s, 720p/1080p, nine aspect ratios. Best for talking characters + multilingual localization; SFW/commercial-leaning. The `happyhorse` video-family now routes here. |
|
|
553
553
|
| **HappyHorse 1.0** | i2v, R2V | t2v | 15s | Yes | Prior line, kept for back-compat. Livelier hand-camera realism / cinematic grain vs Seedance. |
|
|
554
554
|
| **MiniMax H3** | i2v, R2V (up to 9 refs) | t2v | 15s (**5s floor**) | Yes (native stereo, not toggleable) | Open-weight omni-modal model — one net covers T2V/I2V/reference. **2K is the only resolution** (no draft tier) at ~1/3 the per-second cost of other families; 24fps, 2500-char prompts. The `minimax-h3` video-family routes here. Sub-5s durations are a hard 400. |
|
|
555
|
+
| **MiniMax H3 Max** | i2v, R2V (up to 9 refs) | t2v | 15s (**5s floor**) | Yes (native, not toggleable) | **Simple prompts — the model stages its own coverage.** Registry `promptStyle: 'simple'`, so the prompt builder strips blocking, locked location descriptions, and geography-hold clauses; say the intent in a sentence or two. Best for montages and beats where the model telling its own story is the point. **768P max — 2K is a hard 400**, the inverse of base H3 (480P is the draft tier). `private` tier, uncensored, 10000-char prompts. $0.024/s. The `minimax-h3-max` video-family routes here. |
|
|
556
|
+
| **MiniMax H3 Max Turbo** | i2v | t2v | 15s (**5s floor**) | Yes (native, not toggleable) | Same model and constraints at **$0.012/s — the cheapest lane in the registry**, which makes 15s takes cheap enough to render several and pick. **No R2V lane** (`-turbo-reference-to-video` does not exist), so the `minimax-h3-max-turbo` family routes identity shots to `minimax-h3-max-reference-to-video`. |
|
|
555
557
|
| **Wan 3.0** | i2v, R2V (up to 9 refs), Enhanced | t2v | **30s** | Yes (always on, not toggleable) | **Longest shots on Venice** — 5/10/15/20/25/30s at 480p/720p/1080p, five aspect ratios plus adaptive, 5000-char prompts. The `wan-3-0` video-family routes here. No audio input anywhere in the family, so it can't lip-sync to a supplied recording. `*-enhanced-*` variants are beta. |
|
|
556
558
|
| **Wan 2.7** | i2v, R2V, V2V, Spicy | t2v | 15s | Wan i2v has no audio; lip-syncs via `audio_url` input | **The audio-driven fallback for exact lip-sync.** R2V exposes per-element `audio_url` for multi-speaker. Spicy = uncensored i2v variant. Seedance 2.x R2V and MiniMax H3 R2V also accept a top-level `audio_url`, so those families never route here. |
|
|
557
559
|
| **Wan 2.6** | Standard, Flash, R2V | Standard | 15s | Yes (i2v/t2v); R2V capped at 10s | Now has R2V variant with `audio_url` input. 1080p. |
|
|
@@ -930,6 +932,115 @@ a persistent history file; `Ctrl-C` cancels the running command without killing
|
|
|
930
932
|
the session (`Ctrl-D` or `/exit` leaves); `/help`, `/status`, `/jobs`, `/cd`, and
|
|
931
933
|
`/pwd` are shell meta-commands; `!<cmd>` runs something in your system shell.
|
|
932
934
|
|
|
935
|
+
### Loop mode — watch the whole plan while it renders, or iterate on real shots
|
|
936
|
+
|
|
937
|
+
Once a plan exists (an approved shot script), you can play the entire film as a
|
|
938
|
+
live browser loop while the harness renders it, instead of waiting for the full
|
|
939
|
+
gated pipeline:
|
|
940
|
+
|
|
941
|
+
```bash
|
|
942
|
+
venice-video loop -p ~/VeniceVideos/my-film -e 1 # asks the purpose
|
|
943
|
+
venice-video loop -p ~/VeniceVideos/my-film -e 1 --mode looping # or state it
|
|
944
|
+
venice-video loop -p ~/VeniceVideos/my-film -e 1 --mode production
|
|
945
|
+
```
|
|
946
|
+
|
|
947
|
+
Loop mode starts with one **required, deliberate decision** — **is this for
|
|
948
|
+
LOOPING or for PRODUCTION?** — because it is a real quality-vs-flow tradeoff, not
|
|
949
|
+
a default to fall through. In a terminal it asks; non-interactively you must pass
|
|
950
|
+
`--mode` (it errors otherwise). You can state it in plain words —
|
|
951
|
+
`--mode looping` / `loop` / `fun` / `creative`, or `--mode production` / `prod` /
|
|
952
|
+
`gather`:
|
|
953
|
+
|
|
954
|
+
- **Looping — creative flow, lower quality.** The first generation is **t2v**,
|
|
955
|
+
every shot after it **chains i2v off the previous shot's last frame**, and it
|
|
956
|
+
**never uses R2V** (those renders are too slow for a loop). Turbo, 480P, fast.
|
|
957
|
+
Not final-quality; it's for watching and riffing.
|
|
958
|
+
- **Production — gather usable shots, higher quality.** **Max R2V + references**
|
|
959
|
+
at 768P, identity locked, each shot rendered independently. Slower, but the
|
|
960
|
+
takes you pin are keepers.
|
|
961
|
+
|
|
962
|
+
Either way it boots the local web UI, opens the browser to a **Loop** tab, and
|
|
963
|
+
**auto-starts** a background engine that renders each shot into the episode's
|
|
964
|
+
`loop/` directory and **keeps regenerating fresh takes continuously** (it does
|
|
965
|
+
not stop after a fixed number of takes — only a Pause or the budget stops it).
|
|
966
|
+
The plan plays on repeat and each shot hot-swaps in as its take finishes;
|
|
967
|
+
because the render outruns playback, the video keeps evolving. Pin the keepers,
|
|
968
|
+
regenerate the ones you don't, and watch a running spend meter. Both modes
|
|
969
|
+
**skip the storyboard/QA gates** and write **only** under `loop/` — canonical
|
|
970
|
+
`scene-001/shot-NNN.mp4` renders and `series.json` are never touched, so a loop
|
|
971
|
+
can run alongside real production.
|
|
972
|
+
|
|
973
|
+
Two behaviors make the loop play as one continuous piece:
|
|
974
|
+
|
|
975
|
+
- **Last-frame chaining (default on).** Shot 1 renders normally; **every shot
|
|
976
|
+
after the first renders i2v using the previous shot's last frame as its first
|
|
977
|
+
frame**, so the clips flow into each other. Turn it off with `--no-chain` to
|
|
978
|
+
render each shot independently (in create mode that keeps per-shot R2V identity
|
|
979
|
+
locking).
|
|
980
|
+
- **Full-length takes.** Every generation renders the model's full length
|
|
981
|
+
(**15s** by default — MiniMax H3 Max's max), for maximum footage and playback
|
|
982
|
+
per render. Override with `--duration`.
|
|
983
|
+
|
|
984
|
+
Because the engine auto-starts, the **Loop** tab shows **Pause** while it's
|
|
985
|
+
running. It regenerates until you Pause or the budget is reached; when the budget
|
|
986
|
+
is reached it pauses and the button becomes **Resume**, which authorizes another
|
|
987
|
+
budget's worth and continues. (`--max-takes` is a ring buffer — the number of
|
|
988
|
+
candidate takes kept per shot — not a stop condition; older takes are pruned so
|
|
989
|
+
an infinite run can't fill the disk.)
|
|
990
|
+
|
|
991
|
+
The two modes differ in what they render:
|
|
992
|
+
|
|
993
|
+
| Purpose (`--mode`) | Model | Resolution | Identity | Use it to… |
|
|
994
|
+
|---|---|---|---|---|
|
|
995
|
+
| **looping** | MiniMax H3 Max **Turbo** t2v/i2v (~$0.012/s) | 480P | **not** locked (Turbo has no R2V lane) | keep a fast, continuous loop going for creative flow |
|
|
996
|
+
| **production** | MiniMax H3 Max **R2V** for character shots, i2v/t2v otherwise (~$0.024/s) | 768P | **locked** via the project's reference stack | gather real, usable shots and pin keepers |
|
|
997
|
+
|
|
998
|
+
Production mode uses the same reference-first routing as the real pipeline:
|
|
999
|
+
character shots render on `minimax-h3-max-reference-to-video` with the full
|
|
1000
|
+
`@Image` reference stack (character sheets, location angles, blocking plates)
|
|
1001
|
+
plus voice-donor audio, so identity holds. Shots with no references on disk
|
|
1002
|
+
degrade to i2v (off a panel) or t2v, so generate your character/location
|
|
1003
|
+
references first for the full effect.
|
|
1004
|
+
|
|
1005
|
+
Continuous regeneration spends money, so it is capped by default:
|
|
1006
|
+
|
|
1007
|
+
```bash
|
|
1008
|
+
venice-video loop -p <dir> -e 1 \
|
|
1009
|
+
--mode production \ # looping | production (required; also accepts loop/fun, prod/gather)
|
|
1010
|
+
--resolution 768P \ # defaults: 480P (looping) / 768P (production)
|
|
1011
|
+
--duration 15s \ # per-take length, snapped to the 5-15s ladder (default 15s)
|
|
1012
|
+
--budget 2 \ # pause after ~$2; Resume/regenerate authorizes another budget
|
|
1013
|
+
--max-takes 3 \ # candidate takes kept per shot (ring buffer, not a stop)
|
|
1014
|
+
--no-chain \ # render shots independently instead of i2v last-frame chaining
|
|
1015
|
+
--no-face-continuity \ # don't prompt chained shots to end on the character's face (see below)
|
|
1016
|
+
--once # or: render one take per shot, then stop
|
|
1017
|
+
# --unbounded # remove the budget cap (spends until you Ctrl-C)
|
|
1018
|
+
```
|
|
1019
|
+
|
|
1020
|
+
The loop is resumable: takes, pins, and spend are recorded in
|
|
1021
|
+
`loop/loop-manifest.json`, so re-running `loop` picks up where it left off.
|
|
1022
|
+
`Ctrl-C` stops the engine and the server.
|
|
1023
|
+
|
|
1024
|
+
**A shot that keeps failing is given up on, not re-billed forever.** After 3
|
|
1025
|
+
consecutive render failures the engine marks the shot `failed`, stops scheduling
|
|
1026
|
+
it, and moves on — so a server-side-doomed shot (e.g. a MiniMax i2v start frame
|
|
1027
|
+
with a human face, which Venice bills at queue time then 500s on retrieve) can't
|
|
1028
|
+
burn the whole budget one failed take at a time. A manual **regenerate** in the
|
|
1029
|
+
UI revives it.
|
|
1030
|
+
|
|
1031
|
+
**Face continuity (on by default, for smoother i2v transitions).** In a chained
|
|
1032
|
+
loop each shot's last frame becomes the next shot's i2v start frame, so
|
|
1033
|
+
`--face-continuity` (default on) prompts each character shot to **end on the
|
|
1034
|
+
character's face**, giving the next clip a clean anchor to continue from
|
|
1035
|
+
(`--no-face-continuity` turns it off). One important caveat: MiniMax i2v renders
|
|
1036
|
+
**die server-side when the start frame shows a face** (AGENTS.md anti-pattern
|
|
1037
|
+
31), so on the MiniMax loop lanes this prompting is **auto-suppressed** — a
|
|
1038
|
+
face-ending frame would kill the next chained render. It activates on any i2v
|
|
1039
|
+
model that accepts face start frames. For smooth character-face loops **today**,
|
|
1040
|
+
use **production** mode: R2V locks the face from the reference sheets across every
|
|
1041
|
+
shot, with no i2v chaining involved (verified — MiniMax R2V accepts face-bearing
|
|
1042
|
+
reference sheets; only i2v *start frames* die).
|
|
1043
|
+
|
|
933
1044
|
### Interrupted renders are resumable
|
|
934
1045
|
|
|
935
1046
|
The harness records each Venice `queue_id` to disk *before* it starts polling, so
|
|
@@ -1018,7 +1129,7 @@ These defaults are overridable per-project via `series.json` → `videoDefaults`
|
|
|
1018
1129
|
|
|
1019
1130
|
### Picking a family at project creation
|
|
1020
1131
|
|
|
1021
|
-
`venice-video new` asks which family to use, and `venice-video new-series` asks too when it's run on a terminal without `--video-family`. Both write the answer to `series.json` → `videoDefaults.videoFamilyPreference` and swap the action / atmosphere / character-consistency models to match. The wizard orders the families as Automatic, Seedance, Wan 3.0, MiniMax H3, HappyHorse, Grok Imagine, then Kling O3.
|
|
1132
|
+
`venice-video new` asks which family to use, and `venice-video new-series` asks too when it's run on a terminal without `--video-family`. Both write the answer to `series.json` → `videoDefaults.videoFamilyPreference` and swap the action / atmosphere / character-consistency models to match. The wizard orders the families as Automatic, Seedance, Wan 3.0, MiniMax H3, MiniMax H3 Max, MiniMax H3 Max Turbo, HappyHorse, Grok Imagine, then Kling O3.
|
|
1022
1133
|
|
|
1023
1134
|
### Choosing dialogue audio
|
|
1024
1135
|
|