venice-video-harness 2.22.1 → 2.25.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -184,8 +184,8 @@ The full model registry lives in `src/venice/models.ts` with typed specs for eve
184
184
  | `@Image` tags (flat ref prompt syntax) | Seedance 2.0 R2V, Grok Imagine R2V |
185
185
  | Native stereo audio with lip-sync | Seedance 2.0 (8+ languages) |
186
186
  | Native stereo audio, not toggleable | HappyHorse 1.1, MiniMax H3 (omit the `audio` field or the request 400s) |
187
- | 2K output | MiniMax H3 (2K is its ONLY resolution) |
188
- | 4K output | Veo 3.1, LTX 2.0 |
187
+ | 2K output | MiniMax H3, MiniMax Hailuo 03 (2K is their ONLY resolution) |
188
+ | 4K output | **Seedance 2.0** (i2v/t2v/r2v — re-probed live 2026-09-07: `resolution: '4k'` quotes 200), Veo 3.1, LTX 2.5 Fast (2160p), Kling O3/V3 4K. Seedance 2.5 tops out at **1080p** (2K/4K 400). The harness had been capping Seedance at 720p until 2026-09-07; the resolution picker (Stream) now offers each model's true ladder. |
189
189
  | 30s duration | Longcat |
190
190
  | 20s duration | LTX 2.0 Fast, LTX 2.0 v2.3 Fast |
191
191
  | 15s duration | Seedance 2.0, Kling O3/V3, Wan 2.6 |
@@ -222,7 +222,7 @@ Preferred defaults (overridable per-project via `series.json` → `videoDefaults
222
222
 
223
223
  | Role | Default Model | When Used |
224
224
  |------|--------------|-----------|
225
- | Character shots (up to ~6 characters) | `seedance-2-5-reference-to-video` | Default R2V — up to 30 `reference_image_urls` with `@Image` tags (chars + storyboard plate + location angles), pure reference mode (no start image), every integer 4-30s, 480p/720p (harness pins 720p), native stereo audio |
225
+ | Character shots (up to ~6 characters) | `seedance-2-5-reference-to-video` | Default R2V — up to 30 `reference_image_urls` with `@Image` tags (chars + storyboard plate + location angles), pure reference mode (no start image), every integer 4-30s, **480p/720p/1080p** (harness auto-pins 720p; 1080p available on request — quote 2026-09-07), native stereo audio |
226
226
  | Character shots (budget overflow) | `kling-o3-standard-reference-to-video` | Auto-fallback — structured `elements` + `reference_image_urls`. Rare on 2.5's 30-ref budget; kept for extreme character counts / scene-image needs |
227
227
  | Native character dialogue | `seedance-2-5-reference-to-video` | Default. Generates the authored line in-frame; `reference_audio_urls` voice donors preserve timbre/accent/pacing. Native mouth sync is prompt-driven, not deterministic to an exact supplied recording. |
228
228
  | Exact lip-sync, low/medium motion | `resolveLipSyncModel(family)` | Only when `audioStrategy: lip-sync`. Venice speech drives mouth movement through `audio_url`; min 3s audio. Stays in-family on Seedance (`seedance-2-5-reference-to-video`) and MiniMax H3 (`minimax-h3-reference-to-video`), both of which accept a top-level `audio_url`. Other families fall back to `wan-2-7-image-to-video`, which needs a keyframe from a Seedance R2V identity pass (rule 32). |
@@ -482,7 +482,7 @@ Use `POST /video/quote` (via `quoteVideo()`) to estimate costs before committing
482
482
 
483
483
  59. **MiniMax H3 Max (simple-prompt) models IMPROVISE dialogue — the scripted line is intent, not a script (2026-09-04).** Unlike Seedance/Wan, which render the exact line you give them, the H3 Max family (`promptStyle: 'simple'`) performs markedly better carrying natural, continuous speech across a whole generation than reciting a verbatim quote — a fixed line fights the model the same way the directorial blocks do. So in **native-dialogue mode** the prompt builders render a speaker's line as INTENT — `[@ImageN, voice, delivery] conveys: "…"` plus a one-line "improvise naturally in character, keep the meaning and tone, don't recite word for word" note — instead of the directorial `[…]: "exact line"` quote. This is automatic via `shouldImproviseDialogue(modelId, series)` in `prompt-builder.ts` (`modelWantsSimplePrompt(modelId)` AND `videoDefaults.audioStrategy !== 'lip-sync'`), applied in both `buildVideoPrompt` (singles) and `buildMontagePrompt` (montage) — the two paths the H3 Max family renders through. Directorial models keep the exact quote; the legacy Seedance-native / Kling multi-shot builders (`buildMultiShotPrompt`) are untouched because those lanes are never simple-prompt. **The exception is exact-lip-sync**: there the `audio_url` drives the exact spoken words, so the line stays verbatim (the improv gate excludes `audioStrategy === 'lip-sync'`). **Consequence:** when a simple-prompt model improvises, burned/exported captions must be derived by transcribing the rendered audio (the editing pipeline's `silencedetect`/whisper path), NOT from `script.json` — the model will not say it word for word. When changing this, keep `shouldImproviseDialogue` / `formatDialogueLine` / `IMPROV_DIALOGUE_NOTE` in `prompt-builder.ts` and this rule in sync.
484
484
 
485
- 60. **`stream` is not `loop` — "infinite" means the STORY never ends, not that the playback cycles (2026-09-04).** When the operator asks for an infinite / never-ending / always-on story, use `venice-video stream -p <project> [--direction "…"]`, not `loop`. The two are different products: `loop` takes a fixed shot plan and re-renders the same N shots forever so the *playback* cycles (`--max-takes` is a ring buffer of candidate renders per shot); `stream` (`src/mini-drama/stream-engine.ts`) never repeats and never re-renders. It authors the story live: the intelligence model (`series.intelligence`, or `--writer`) writes ONE beat at a time from the series bible + `story-so-far.md` (one summary line per prior beat) + the last `STREAM_RECENT_BEATS` (6) beats verbatim, with a hard continuity rule ("begins EXACTLY where the previous beat ended, the camera does not cut"). Beat 1 renders t2v on MiniMax H3 Max Turbo; every later beat renders i2v off the previous beat's last frame (`extractLastFrame`). Invariants: (a) **no re-anchoring, ever** — every frame descends from the frame before it, identity drifts slowly by design, and there is no t2v reset cadence; (b) **no ring buffer** — every beat stays on disk in order (`stream/beat-NNNNN.mp4` + `.json`), disk is the only limit; (c) **a stream cannot skip a beat** — after `MAX_CONSECUTIVE_ERRORS` (3) failures at write, chain-frame, or render the engine STOPS (a skipped beat would be a hidden cut), where the loop would give up on one shot and move on; (d) it needs only `series.json` — no script, storyboard, QA, or references; a locked aesthetic and a cast (`--skip-images` is fine) make the writer much better; the command registers its episode in `series.json` when missing (the browser builds its episode list from there — an unregistered episode rendered beats the Stream tab could not show, and hid the Start button behind "No episodes yet."); (e) `--direction` is standing direction folded into every writer prompt (the place for "laugh track after every joke"), and `openingBeat` lets a caller pin beat 1 verbatim; (f) budget/resume semantics match the loop (billed at queue time, `--budget` stops, Start/Continue authorizes another, `--unbounded` lifts the cap; `stream-manifest.json` resumes from the longest unbroken prefix of beats whose files exist and chains off the last one); (g) **the writer is the session's first decision — ASK the operator which model writes the beats before starting a new stream.** It is the voice of the whole story and it bills from beat 1. Interactive runs get a `promptChoice` over `selectableTextModels()`; a non-interactive new stream with no `--writer` is a hard error (pass `--writer <model>` or `--writer default`), mirroring `loop --mode`. A resumed stream keeps the project default. The writer and the per-beat cost print before beat 1 bills. Do not treat `series.intelligence` as the operator's choice for the stream — it was chosen for QA and scripting, not latency. The stream's own default is `STREAM_DEFAULT_WRITER` (`deepseek-v4-flash-0731-fast`, 3.8s median, 9/9 valid, thinking off) from the 2026-09-05 bakeoff (`scripts/bakeoff-stream-writer.ts`; results in `src/mini-drama/stream-choices.ts` and the README). `chatJson` takes `disableThinking`; the stream writer sets it per choice — with thinking on, the same models are 3-10x slower and reasoning-only models burn the token budget and return nothing. Both the writer and the video family are dropdowns in the Stream tab (`POST /stream/config`, `engine.configure()`); a change applies to the NEXT beat, the i2v chain survives a family switch (the start frame is a PNG), and a resumed stream keeps the models in its manifest. Every video family other than Turbo renders slower than playback; the UI says so beside the selector. Per-beat cost comes from the family's quoted `usdPer15s` (Turbo is $0.11 at 480P, not the $0.18 the old constant assumed). (h) **A face-ending beat must not kill the stream.** MiniMax i2v dies server-side on a face-filled start frame after billing (anti-pattern 31), and the chain makes one bad frame poison every retry. Three defenses, in order: the writer's system prompt carries a MANDATORY camera rule — end every beat wide, never on a human face close-up; a failed chained render steps the start frame back (`STREAM_CHAIN_STEP_BACK_SEC`); and after `STREAM_CHAIN_FAILURES_BEFORE_RESET` (2) chained failures on one beat, the engine renders that beat **t2v as a soft reset** (`lane: 't2v-reset'`), with the previous beat's summary prepended so the prompt re-establishes the scene. Identity drifts for one beat; a skipped beat or a dead stream is worse. The Stream tab shows `lastError` while a retry is in flight and marks reset beats. (i) **The Stream tab merges, never replaces.** The on-disk manifest is re-read on every `state-changed` (the workspace watcher fires on each beat's files); it is a fallback. `stream-updated` SSE is the truth. The view merges beats by number so a stale disk snapshot can never remove a beat the SSE already delivered — that race is what made new beats appear only after a reload. Only the `<video>` element swaps source (keyed by file); the page never reloads, and playback is kicked explicitly after each swap. (j) **Every beat keeps its exact video prompt.** `StreamBeat.render` = model, prompt, resolution, duration, start frame — byte-for-byte what `renderVideoFile` sent. The Stream tab exposes it per beat ("Full prompt") and as a JSON/Markdown export of the whole stream (`/stream/export.json|md`), so an operator can fine-tune outside the harness. Older beats are backfilled from `.recipe.json` on resume. When the prompt builder changes, the recorded prompt is the evidence of what a given beat actually got. The writer is `AuthorFn` and the renderer `RenderFn`, both injectable (`tests/stream-engine.test.mjs`). Output for the browser is the **Stream** tab (`StreamView.tsx`, `stream-updated` SSE, `/api/projects/:slug/stream/{state,start,stop}`), which plays forward from beat 1 and holds on the newest beat until the next lands. When changing stream behavior, keep `stream-engine.ts`, the `stream` command in `cli.ts`, `server.ts` (`WebServerOptions.stream`), `state.ts`, `StreamView.tsx`, `api.ts` (`names` list), `types.ts`, and this rule in sync. Note the `-e` option on `stream` is a string parsed by hand — passing Commander's `parseInt` collided with the `applyContextDefaults` hook and yielded `episode-NaN`.
485
+ 60. **`stream` is not `loop` — "infinite" means the STORY never ends, not that the playback cycles (2026-09-04).** When the operator asks for an infinite / never-ending / always-on story, use `venice-video stream -p <project> [--direction "…"]`, not `loop`. The two are different products: `loop` takes a fixed shot plan and re-renders the same N shots forever so the *playback* cycles (`--max-takes` is a ring buffer of candidate renders per shot); `stream` (`src/mini-drama/stream-engine.ts`) never repeats and never re-renders. It authors the story live: the intelligence model (`series.intelligence`, or `--writer`) writes ONE beat at a time from the series bible + `story-so-far.md` (one summary line per prior beat) + the last `STREAM_RECENT_BEATS` (6) beats verbatim, with a hard continuity rule ("begins EXACTLY where the previous beat ended, the camera does not cut"). Beat 1 renders t2v on MiniMax H3 Max (the default as of 2026-09-07 — Turbo read too low-quality; the faster/cheaper `minimax-h3-max-turbo` lane is one dropdown away); every later beat renders i2v off the previous beat's last frame (`extractLastFrame`). Invariants: (a) **no re-anchoring, ever** — every frame descends from the frame before it, identity drifts slowly by design, and there is no t2v reset cadence; (b) **no ring buffer** — every beat stays on disk in order (`stream/beat-NNNNN.mp4` + `.json`), disk is the only limit; (c) **a stream cannot skip a beat** — after `MAX_CONSECUTIVE_ERRORS` (3) failures at write, chain-frame, or render the engine STOPS (a skipped beat would be a hidden cut), where the loop would give up on one shot and move on; (d) it needs only `series.json` — no script, storyboard, QA, or references; a locked aesthetic and a cast (`--skip-images` is fine) make the writer much better; the command registers its episode in `series.json` when missing (the browser builds its episode list from there — an unregistered episode rendered beats the Stream tab could not show, and hid the Start button behind "No episodes yet."); (e) `--direction` is standing direction folded into every writer prompt (the place for "laugh track after every joke"), and `openingBeat` lets a caller pin beat 1 verbatim; (f) budget/resume semantics match the loop (billed at queue time, `--budget` stops, Start/Continue authorizes another, `--unbounded` lifts the cap; `stream-manifest.json` resumes from the longest unbroken prefix of beats whose files exist and chains off the last one); (g) **the writer is the session's first decision — ASK the operator which model writes the beats before starting a new stream.** It is the voice of the whole story and it bills from beat 1. Interactive runs get a `promptChoice` over `selectableTextModels()`; a non-interactive new stream with no `--writer` is a hard error (pass `--writer <model>` or `--writer default`), mirroring `loop --mode`. A resumed stream keeps the project default. The writer and the per-beat cost print before beat 1 bills. Do not treat `series.intelligence` as the operator's choice for the stream — it was chosen for QA and scripting, not latency. The stream's own default is `STREAM_DEFAULT_WRITER` (`deepseek-v4-flash-0731-fast`, 3.8s median, 9/9 valid, thinking off) from the 2026-09-05 bakeoff (`scripts/bakeoff-stream-writer.ts`; results in `src/mini-drama/stream-choices.ts` and the README). `chatJson` takes `disableThinking`; the stream writer sets it per choice — with thinking on, the same models are 3-10x slower and reasoning-only models burn the token budget and return nothing. Both the writer and the video family are dropdowns in the Stream tab (`POST /stream/config`, `engine.configure()`); a change applies to the NEXT beat, the i2v chain survives a family switch (the start frame is a PNG), and a resumed stream keeps the models in its manifest. The default `minimax-h3-max` is pinned to 480P for speed (~45s/beat, $0.22) — sharper than Turbo but still slower than playback; the look-ahead buffer removes the writer latency but not the render, and only the cheaper/lower-quality Turbo lane nearly keeps pace (the UI says so beside the selector). Per-beat cost comes from the family's quoted `usdPer15s` (verified via `POST /video/quote`: H3 Max is $0.22 at 480P and $0.36 at 768P; Turbo $0.11 at 480P). (h) **A face-ending beat must not kill the stream.** MiniMax i2v dies server-side on a face-filled start frame after billing (anti-pattern 31), and the chain makes one bad frame poison every retry. Three defenses, in order: the writer's system prompt carries a MANDATORY camera rule — end every beat wide, never on a human face close-up; a failed chained render steps the start frame back (`STREAM_CHAIN_STEP_BACK_SEC`); and after `STREAM_CHAIN_FAILURES_BEFORE_RESET` (2) chained failures on one beat, the engine renders that beat **t2v as a soft reset** (`lane: 't2v-reset'`), with the previous beat's summary prepended so the prompt re-establishes the scene. Identity drifts for one beat; a skipped beat or a dead stream is worse. The Stream tab shows `lastError` while a retry is in flight and marks reset beats. (i) **The Stream tab merges, never replaces.** The on-disk manifest is re-read on every `state-changed` (the workspace watcher fires on each beat's files); it is a fallback. `stream-updated` SSE is the truth. The view merges beats by number so a stale disk snapshot can never remove a beat the SSE already delivered — that race is what made new beats appear only after a reload. Only the `<video>` element swaps source (keyed by file); the page never reloads, and playback is kicked explicitly after each swap. The whole `StreamView` is keyed by project slug in `App.tsx` so switching projects REMOUNTS it with fresh state — without the key, the previous project's beats stayed under the player and the writer/video/resolution selects stayed greyed out because the engine was still bound to the other project (fixed 2026-09-07). (j) **Every beat keeps its exact video prompt.** `StreamBeat.render` = model, prompt, resolution, duration, start frame — byte-for-byte what `renderVideoFile` sent. The Stream tab exposes it per beat ("Full prompt") and as a JSON/Markdown export of the whole stream (`/stream/export.json|md`), so an operator can fine-tune outside the harness. Older beats are backfilled from `.recipe.json` on resume. When the prompt builder changes, the recorded prompt is the evidence of what a given beat actually got. The writer is `AuthorFn` and the renderer `RenderFn`, both injectable (`tests/stream-engine.test.mjs`). Output for the browser is the **Stream** tab (`StreamView.tsx`, `stream-updated` SSE, `/api/projects/:slug/stream/{state,start,stop}`), which plays forward from beat 1 and holds on the newest beat until the next lands. When changing stream behavior, keep `stream-engine.ts`, the `stream` command in `cli.ts`, `server.ts` (`WebServerOptions.stream`), `state.ts`, `StreamView.tsx`, `api.ts` (`names` list), `types.ts`, and this rule in sync. Note the `-e` option on `stream` is a string parsed by hand — passing Commander's `parseInt` collided with the `applyContextDefaults` hook and yielded `episode-NaN`. (k) **Pre-written beats: `stream --beats-file` (2026-09-07).** When the operator wants the beats authored up front — no live writer at all for the scripted span — pass `--beats-file <path>`: a JSON file holding a bare array of `AuthoredBeat` objects or the `{ "beats": [...] }` shape of `/stream/export.json` (entries with an `authored` object are unwrapped, so an exported stream replays as-is). `parseScriptedBeats()` unwraps and type-checks; `normalizeBeat()` runs per entry against the locked cast BEFORE anything bills, so a bad beat fails at load. Engine side, `scriptedBeats` on `StreamEngineOptions` wraps any author (injected override included) in `makeScriptedAuthor()`: beat N of the stream is served from file position N−1 until the file runs out, then the live writer takes over as the fallback — for a new stream that fallback defaults to `STREAM_DEFAULT_WRITER`, so `--beats-file` alone satisfies the writer decision of (g). A writer switch from the Stream tab (`configure`) or a resumed manifest changes ONLY the fallback; scripted beats keep serving, and already-rendered beats are never re-rendered. The continuity rules still bind the file's author: each beat is one continuous shot beginning where the previous ended, and every beat must END wide (anti-pattern 31 — the chain's start frame must not be a human-face close-up). (l) **Look-ahead writer buffer — the writer authors AHEAD of the render by default (2026-09-07, 2.24.0).** The stream runs the writer (producer) and renderer (consumer) concurrently: `runWriter()` keeps up to `lookahead` beats (default `STREAM_DEFAULT_LOOKAHEAD` = 15) authored and waiting in `buffer[]`, and `renderNext()` consumes `buffer[0]` — the in-flight beat, kept there across render retries (this replaced the old single `pendingBeat`) — so a render NEVER blocks on a writer-model call. This is the whole point: previously each beat's wall time was writer latency + render latency in series; now the writer stays ahead so it is just render latency, and a slower/better writer is free as long as it keeps ahead. Priming fills the buffer while the stream is paused, so Start renders back to back immediately. Controls: `--lookahead <n>` (0 = the old serial path, where the render worker authors each beat inline just before rendering it) and `--no-refill` (`autoRefill=false` — fill the buffer once, then author on demand as it drains; default keeps it topped up to the depth). Both are switchable at runtime via `engine.configure()` / `POST /stream/config` and the Stream tab's **Look-ahead buffer** control (depth field + "keep topped up" toggle); a live `buffered/depth` meter rides the `stream-updated` SSE (`buffered`, `lookahead`, `autoRefill`). Invariants: the writer's author context is built in-memory from the union of rendered + buffered beats (`authoredList()` / `buildAuthorContext()`), NOT `story-so-far.md`, so beat N sees the beats already queued ahead of the render (or the writer would repeat itself); the budget still bounds it — `authorTarget()` never authors beats the budget cannot render; the buffer persists as `pendingBeats` in the manifest and is restored on resume ONLY when the beats prefix is intact (a truncated prefix would misalign the chain); switching the writer drops the non-in-flight buffered beats so the new writer takes over from the next beat (preserving the "a switch applies to the next beat" contract); and the writer's idle/backoff timers are `unref`'d so a stopped engine can never hold the process open. Keep `stream-engine.ts`, `cli.ts` (`--lookahead`/`--no-refill`), `server.ts` (`/stream/config` body), `types.ts`, `api.ts`, and `StreamView.tsx` in sync. (m) **Always run and OPEN the local UI when streaming — 100% of the time, unprompted (2026-09-07).** A stream is a live broadcast; a running stream the operator cannot watch is a failure. Whenever the operator asks to start/run/kick off a stream, the agent ALWAYS launches the local web UI and opens it in the browser as soon as the opening beat is ready — the operator must NEVER have to ask for it. Run `venice-video stream -p <project> [-e N] [--direction …]` (it starts the server on port 3000, primes beat 1, and opens the browser by default); never pass `--no-open`. Always surface the `http://127.0.0.1:3000/?project=<slug>&tab=Stream` URL in the reply so the operator can reopen it, and if the browser cannot auto-open (headless/remote/SSH), print that URL prominently and tell them to open it. Free port 3000 first if it is taken.
486
486
 
487
487
  ## Learned Anti-Patterns (Production Issues Log)
488
488
 
package/CHANGELOG.md CHANGED
@@ -1,5 +1,102 @@
1
1
  # Changelog
2
2
 
3
+ ## 2.25.0 — 2026-09-07
4
+
5
+ ### Fixed
6
+
7
+ - **Resolution was silently capped far below what Venice accepts.** The registry
8
+ listed Seedance as `480p/720p` and `renderVideoFile` hard-pinned Seedance to
9
+ `720p`, rejecting any higher override against that stale list. Re-probed live
10
+ via `POST /video/quote` (2026-09-07): `seedance-2-0-*` accept **`4k`**
11
+ ($4.86/5s) and `seedance-2-5-*` accept **`1080p`** ($2.56/5s) on the exact
12
+ ids the harness already sends. Registry corrected; a chosen resolution up to
13
+ each model's real ceiling now reaches the queue body, and an unsupported
14
+ choice (e.g. Seedance 2.5 @ `4k`) safely falls back to the family default
15
+ instead of 400ing. This is why Stream/Loop output looked low-fidelity: the
16
+ MiniMax H3 Max/Turbo speed lanes render at a **480P draft** and top out at
17
+ 768P (2K is a hard 400) — draft lanes for keeping up with live playback, not
18
+ a quality tier.
19
+
20
+ ### Added
21
+
22
+ - **Full resolution control in the Stream tab, per each model's real ladder.**
23
+ The video-family dropdown's resolution selector is now fed by the live-
24
+ verified ladders: Seedance 2.0 → `480p/720p/1080p/4k`, Seedance 2.5 →
25
+ `480p/720p/1080p`. Two true high-res lanes added to the Stream choices:
26
+ **LTX Video 2.5 Fast** (up to 2160p, even-second durations) and **Veo 3.1
27
+ Fast** (up to 4K, 8s max). Defaults stay on the fast 480P draft lanes — you
28
+ opt up when quality matters more than live pacing. No UI rebuild required
29
+ (the selector already renders `choices.video[].resolutions`).
30
+ - **Live catalog sync (2026-09-07).** Registered high-res models the harness
31
+ was missing: `ltx-2-5-fast-*` / `ltx-2-5-pro-*`, `minimax-hailuo-03-*` (2K),
32
+ `wan-3-0-prime-*`, and the live-listed `seedance-2-0-*-basic` (4K). Capability
33
+ sets in `series/types.ts` updated in lockstep (registry-coverage test).
34
+
35
+ ## 2.24.0 — 2026-09-07
36
+
37
+ ### Added
38
+
39
+ - **Look-ahead writer buffer — the writer authors beats ahead of the render by
40
+ default.** The stream now runs the writer and the renderer as a
41
+ producer/consumer pair: the writer keeps up to `--lookahead` beats (default
42
+ **15**) authored and waiting in a buffer, so a render never blocks on a
43
+ writer-model call. Priming fills the buffer while the stream is paused, so
44
+ clicking Start renders back to back with no writer latency. This also lets a
45
+ slower, better writer keep pace as long as it stays ahead.
46
+ - `--lookahead <n>` sets the buffer depth. `0` restores the pre-2.24 serial
47
+ behaviour (author each beat just before it renders).
48
+ - `--no-refill` fills the buffer once, then authors on demand as it drains;
49
+ the default keeps the buffer topped up as the renderer consumes it.
50
+ - Both are switchable at runtime from the Stream tab (a **Look-ahead buffer**
51
+ control: depth input + "keep topped up" toggle) and via
52
+ `POST /stream/config` (`lookahead`, `autoRefill`). The Stream tab shows a
53
+ live `buffered / depth` meter.
54
+ - Switching the writer drops the beats the old writer had buffered (keeping
55
+ only the one on the wire) so the new writer takes over from the next beat.
56
+ - Budget still bounds it: the writer never authors beats the budget cannot
57
+ render. The buffer persists in the manifest (`pendingBeats`) so a resume
58
+ renders the pre-authored beats without paying for them again.
59
+ - Engine: `lookahead` / `autoRefill` on `StreamEngineOptions`, `lookahead` /
60
+ `autoRefill` / `buffered` / `pendingBeats` in the manifest, and
61
+ `STREAM_DEFAULT_LOOKAHEAD` (15), all exported.
62
+
63
+ ### Changed
64
+
65
+ - **Default stream video family is now `minimax-h3-max`, pinned to 480P.** Turbo
66
+ reads noticeably lower quality, so the default is the sharper MiniMax H3 Max
67
+ model — but kept at **480P** (not its 768P draft tier) so it still generates
68
+ fast: ~45 s/beat at $0.22 per 15 s (verified via `POST /video/quote`; 768P is
69
+ $0.36 and selectable). It renders slower than playback, but the look-ahead
70
+ buffer takes the writer latency out of the picture and the Stream tab shows
71
+ the hold honestly. Turbo ($0.11, ~30 s) is still one dropdown away for a
72
+ cheaper live watch. `STREAM_DEFAULT_VIDEO_FAMILY`, the default
73
+ `STREAM_MODEL_T2V` / `STREAM_MODEL_I2V` lanes, the H3 Max draft resolution,
74
+ and the CLI `--video-family` default all move to `minimax-h3-max` @ 480P.
75
+
76
+ ### Fixed
77
+
78
+ - **Stream tab: switching projects no longer leaks the previous project's
79
+ beats or its disabled/attached controls.** The view is keyed by project, so
80
+ it remounts with fresh state on a project switch (the old project's beats
81
+ stayed under the player, and the writer/video/resolution selects stayed
82
+ greyed out because the engine was still bound to the other project).
83
+
84
+ ## 2.23.0 — 2026-09-07
85
+
86
+ ### Added
87
+
88
+ - **Pre-written beats for the stream: `stream --beats-file`.** Author beats up
89
+ front and the stream renders them without ever calling the writer model.
90
+ Accepts a bare JSON array of beats or the `{ "beats": [...] }` shape of
91
+ `/stream/export.json` (entries with an `authored` object are unwrapped, so an
92
+ exported stream replays as-is). Each entry is normalized against the locked
93
+ cast the same way a writer's output would be; a beat with no description
94
+ fails at load, before anything bills. The live writer (defaulting to
95
+ `STREAM_DEFAULT_WRITER` for a new stream) is only the fallback past the last
96
+ scripted beat, and a writer switch from the Stream tab changes only that
97
+ fallback. Engine side: `scriptedBeats` on `StreamEngineOptions`,
98
+ `makeScriptedAuthor()`, and `parseScriptedBeats()`, all exported.
99
+
3
100
  ## 2.22.1 — 2026-09-05
4
101
 
5
102
  ### Added
package/README.md CHANGED
@@ -944,6 +944,58 @@ venice-video loop -p ~/VeniceVideos/my-film -e 1 --mode looping # or state it
944
944
  venice-video loop -p ~/VeniceVideos/my-film -e 1 --mode production
945
945
  ```
946
946
 
947
+ ### Pre-written beats: `stream --beats-file`
948
+
949
+ The stream writes every beat with a live writer model. To author the beats
950
+ yourself — or have an agent write them up front — pass `--beats-file`. The
951
+ first N beats of the stream are then served from the file and the writer model
952
+ is **never called** for them; only if the stream runs past the last scripted
953
+ beat does the live writer take over (defaulting to `STREAM_DEFAULT_WRITER`).
954
+
955
+ ```bash
956
+ venice-video stream -p ~/VeniceVideos/my-film -e 1 \
957
+ --beats-file ~/VeniceVideos/my-film/beats.json \
958
+ --direction "live studio audience laugh track after every joke" \
959
+ --budget 2
960
+ ```
961
+
962
+ With `--beats-file` a new stream needs no `--writer`: the file IS the writer
963
+ decision for the beats it covers. A `--writer` still overrides the fallback
964
+ used past the file. On resume the scripted lane re-attaches the same way —
965
+ beats already rendered are never re-rendered, and a writer switch from the
966
+ Stream tab changes only the fallback.
967
+
968
+ The file is JSON and accepts two shapes:
969
+
970
+ ```jsonc
971
+ // 1. A bare array of beats.
972
+ [
973
+ {
974
+ "description": "The bell jingles as JAKE strides in and takes the couch.",
975
+ "characters": ["JAKE KELLER", "MEL"],
976
+ "dialogue": { "character": "JAKE KELLER", "line": "The usual.", "delivery": "cheerful" },
977
+ "sfx": "door bell, live studio audience applause",
978
+ "cameraMovement": "slow dolly in to a wide of the cafe",
979
+ "summary": "Jake arrives at the cafe."
980
+ }
981
+ ]
982
+ ```
983
+
984
+ ```jsonc
985
+ // 2. The { "beats": [...] } shape of /stream/export.json — entries with an
986
+ // "authored" object are unwrapped, so an exported stream replays as-is.
987
+ { "beats": [ { "n": 1, "authored": { "description": "…", … } } ] }
988
+ ```
989
+
990
+ Beat fields match `AuthoredBeat` in `stream-engine.ts`. Each entry is
991
+ normalized against the locked cast (names snap to the cast's spelling, missing
992
+ fields are completed), and a beat with no `description` fails at load — before
993
+ anything bills. The stream's continuity rules still apply to what you write:
994
+ each beat is one continuous shot that begins where the previous beat ended,
995
+ and every beat should END on a wide or medium-wide frame, never a human-face
996
+ close-up (the next beat chains off that frame, and MiniMax i2v dies on a
997
+ face-filled start frame — anti-pattern 31).
998
+
947
999
  Loop mode starts with one **required, deliberate decision** — **is this for
948
1000
  LOOPING or for PRODUCTION?** — because it is a real quality-vs-flow tradeoff, not
949
1001
  a default to fall through. In a terminal it asks; non-interactively you must pass
@@ -1061,7 +1113,7 @@ How it works:
1061
1113
  writer and the per-beat cost print before beat 1 bills.
1062
1114
  2. The writer writes beat 1 from the series bible: concept, setting, aesthetic,
1063
1115
  and cast.
1064
- 3. Beat 1 renders text-to-video on MiniMax H3 Max Turbo.
1116
+ 3. Beat 1 renders text-to-video on MiniMax H3 Max (the default; the faster, lower-quality Turbo lane is selectable).
1065
1117
  4. The writer reads `story-so-far.md` (one line per prior beat) plus the last
1066
1118
  6 beats verbatim, and writes beat 2 so it begins exactly where beat 1 ended.
1067
1119
  5. Beat 2 renders image-to-video off beat 1's last frame.
@@ -1081,13 +1133,39 @@ venice-video stream -p <dir> \
1081
1133
  -e 1 \ # episode the stream lives under (default 1)
1082
1134
  --direction "<text>" \ # standing direction folded into every beat's writer prompt
1083
1135
  --writer <model> \ # writer; asked for a new stream, required non-interactively (see the bakeoff table)
1084
- --video-family <family> \ # minimax-h3-max-turbo (default) | minimax-h3-max | wan-3-0 | grok-imagine | seedance-2-0 | seedance-2-5 | kling-o3-standard
1136
+ --video-family <family> \ # minimax-h3-max (default) | minimax-h3-max-turbo | wan-3-0 | grok-imagine | seedance-2-0 | seedance-2-5 | kling-o3-standard
1085
1137
  --resolution 480P \ # default: the family's draft tier
1086
1138
  --duration 15s \ # per-beat length, snapped to the 5-15s ladder
1139
+ --lookahead 15 \ # beats authored AHEAD of the render (0 = serial)
1087
1140
  --budget 2 # stop after ~$2; Continue authorizes another budget
1141
+ # --no-refill # fill the look-ahead buffer once, then author on demand
1088
1142
  # --unbounded # no cap (streams until Ctrl-C)
1089
1143
  ```
1090
1144
 
1145
+ #### Look-ahead writer buffer
1146
+
1147
+ By default the writer runs **ahead** of the render. It is a producer/consumer
1148
+ pair: the writer keeps up to `--lookahead` beats (default **15**) authored and
1149
+ waiting in a buffer, and the renderer pulls from it — so a render never blocks
1150
+ on a writer-model call. While the stream is paused after priming, the writer is
1151
+ already filling the buffer, so clicking Start renders back to back with no
1152
+ writer latency between beats. It also lets you run a slower, better writer
1153
+ without stalling playback, as long as the writer stays ahead of the render.
1154
+
1155
+ - `--lookahead <n>` sets the depth. `0` is serial: each beat is authored just
1156
+ before it renders (the pre-2.24 behaviour), so every beat pays the writer
1157
+ latency.
1158
+ - `--no-refill` fills the buffer once and then authors on demand as it drains;
1159
+ the default keeps it topped up to the depth as the renderer consumes it.
1160
+ - Both are switchable live from the Stream tab (the **Look-ahead buffer**
1161
+ control — a depth field and a "keep topped up" toggle) and via
1162
+ `POST /stream/config`. The tab shows a live `buffered / depth` meter.
1163
+ - Switching the writer drops the beats the old writer had queued (keeping only
1164
+ the one on the wire) so the new writer takes over from the next beat.
1165
+ - The budget still bounds it — the writer never authors beats the budget cannot
1166
+ render — and the buffer is saved in `stream-manifest.json` (`pendingBeats`),
1167
+ so a resume renders the pre-authored beats without paying for them again.
1168
+
1091
1169
  The stream is resumable: re-running `stream` continues from the last beat on
1092
1170
  disk and chains off it. After 3 consecutive failures (write, chain, or render)
1093
1171
  the engine stops rather than skip a beat — a stream cannot have a hidden cut.
@@ -1171,16 +1249,18 @@ Not offered, with the reason:
1171
1249
  ##### Video Family Matrix
1172
1250
 
1173
1251
  Speed is the wall time to render one 15 s beat. "Lag" is what the viewer
1174
- feels: with the default writer (~4 s) added, Turbo makes a 15 s beat in ~35 s,
1175
- so the player holds ~20 s between beats once it has caught up. Every other
1176
- family holds for a minute or more. Cost is the quote for 15 s at the family's
1252
+ feels: the default `minimax-h3-max` at 480P renders a 15 s beat in ~45 s, so the
1253
+ player holds ~30 s between beats once it has caught up; the faster Turbo lane
1254
+ cuts that to a ~20 s hold at lower quality. The look-ahead buffer takes the
1255
+ writer's time out of this — only the render remains. Every family other than
1256
+ Turbo holds for a minute or more. Cost is the quote for 15 s at the family's
1177
1257
  draft resolution. Quality is relative to what the harness knows about each
1178
1258
  family (see the model registry and AGENTS.md).
1179
1259
 
1180
1260
  | Family | Privacy | Speed (15 s beat) | Cost / 15 s | Quality | Faces on start frame | Verdict |
1181
1261
  |---|---|---|---|---|---|---|
1182
- | `minimax-h3-max-turbo` **(default)** | ●●● private | ●●● ~30 s | ●●● $0.11 | ●●○ good motion, native audio, improvises dialogue | ✗ dies after billing; engine soft-resets | The only lane that nearly keeps pace. Draft look at 480P; 768P selectable. |
1183
- | `minimax-h3-max` | ●●● private | ●●○ ~60 s | ●●● $0.22 | ●●● sharper than Turbo, same model family | ✗ same limit | Pick when you want the Turbo look at finish quality and will accept a 1-minute hold. |
1262
+ | `minimax-h3-max` **(default)** | ●●● private | ●●○ ~45 s @ 480P | ●●● $0.22 | ●●● sharper than Turbo, same model | ✗ dies after billing; engine soft-resets | The default. H3 Max quality pinned to 480P for speed; ~30 s hold. 768P selectable at $0.36. |
1263
+ | `minimax-h3-max-turbo` | ●●● private | ●●● ~30 s | ●●● $0.11 | ●●○ good motion, native audio, lower quality | ✗ same limit | Fastest and cheapest, the only lane that nearly keeps pace. Draft look at 480P; pick when a live watch matters more than fidelity. |
1184
1264
  | `wan-3-0` | ●○○ anonymized | ●○○ ~120 s | ●●○ $0.68 | ●●● strong, up to 1080p, 30 s ladder | ✓ accepts faces | Best choice if the show is face-heavy and the camera rule is not enough. Slow. |
1185
1265
  | `grok-imagine` | ●○○ anonymized | ●○○ ~90 s | ●○○ $0.95 | ●●○ | ✓ | Faster than Wan, pricier, lower ceiling. |
1186
1266
  | `seedance-2-0` | ●○○ anonymized | ○○○ ~180 s | ●○○ $1.32 | ●●● the harness production look, native lip-synced dialogue | ✓ | Production fidelity. The viewer waits ~3 min per beat. Use for a stream you export, not one you watch. |
@@ -1193,7 +1273,8 @@ family (see the model registry and AGENTS.md).
1193
1273
  |---|---|---|---|
1194
1274
  | Watch it live, cheapest, private | `deepseek-v4-flash-0731-fast` | `minimax-h3-max-turbo` @ 480P | ~35 s per beat, ~20 s hold, ~$0.11/beat, ~$13/hour of story |
1195
1275
  | Watch it live, best sitcom writing | `mistral-small-2603` | `minimax-h3-max-turbo` | Same lag, warmer beats |
1196
- | Sharper picture, still private | `deepseek-v4-flash-0731-fast` | `minimax-h3-max` @ 768P | ~65 s per beat, ~50 s hold, $0.22/beat |
1276
+ | Sharper picture, still fast (default) | `deepseek-v4-flash-0731-fast` | `minimax-h3-max` @ 480P | ~45 s per beat, ~30 s hold, $0.22/beat |
1277
+ | Max fidelity, will accept the wait | `deepseek-v4-flash-0731-fast` | `minimax-h3-max` @ 768P | ~65 s per beat, ~50 s hold, $0.36/beat |
1197
1278
  | Human faces fill the frame often | any fast writer | `wan-3-0` | Faces never kill the chain; ~2 min per beat |
1198
1279
  | Production look to export later | `kimi-k3` | `seedance-2-0` or `-2-5` | ~3.5 min per beat, $1.32-1.93/beat; run it overnight, do not watch it live |
1199
1280
  | Strict privacy for both text and pixels | `deepseek-v4-flash-0731-fast` or `mistral-small-2603` | `minimax-h3-max-turbo` or `minimax-h3-max` | The only fully private pairing; MiniMax is the sole private video family here |