@sogni-ai/sogni-protocol 1.0.0-alpha.39 → 1.0.0-alpha.40

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -236,6 +236,7 @@
236
236
  "minimax-h3-t2v",
237
237
  "minimax-h3-t2v-turbo",
238
238
  "minimax-h3-fasth3-t2v-turbo",
239
+ "minimax-h3-fasth3-t2v-turbo-2stage",
239
240
  "happyhorse-1.1-t2v",
240
241
  "happyhorse-1.1-i2v",
241
242
  "happyhorse-1.1-r2v",
@@ -244,7 +245,7 @@
244
245
  "wan3.0-video",
245
246
  "wan3.0-spicy-video"
246
247
  ],
247
- "description": "\"ltx25\" (default): LTX 2.5 with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Video model. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio. \"minimax-h3-t2v\": standard 20-step MiniMax H3 text-to-video; \"minimax-h3-t2v-turbo\": the existing 4-step LightX2V Turbo text-to-video; \"minimax-h3-fasth3-t2v-turbo\": the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. All use native audio, fixed 24fps, 5.17-15.08s, and a 768p-class 32px-grid canvas; use animate_photo for H3 image-conditioned modes. Base and Turbo T2V/I2V/FLF2V prompts use the exact ordered fields integrated_multimodal_description, overall_soundscape, and non_diegetic_music; I2V/FLF2V prepend the official alignment line. \"minimax-h3-r2v\": standard 20-step MiniMax H3 reference-to-video; \"minimax-h3-r2v-turbo\": the dedicated LightX2V 4-step Ref2VA Turbo workflow using Euler/simple and a 960x544 default. FastH3 has no R2V mode. Both R2V selectors accept up to 9 images, 3 videos, and 3 audios (12 files total); at least one visual reference (image or video) is required and audio alone is invalid. Select references with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and address them with the official <Subject N>/<Picture N>/<Video N>/<Audio N> semantics. Seedance quality is selected only by model: use \"seedance2-mini\" for Seedance 2.0 Mini or faster/lower-cost 720p iteration, and use \"seedance2\" for the full Seedance 2.0 model, explicit full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft or Mini. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p and 720p ONLY (it cannot render 1080p or 4K), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 1080p/4K requests, which Seedance 2.5 cannot satisfy. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. Alibaba HappyHorse 1.1 video models (third-party vendor — requires Premium Spark). Select by mode: \"happyhorse-1.1-t2v\" for text-to-video, \"happyhorse-1.1-i2v\" for image-to-video from one first-frame image, and \"happyhorse-1.1-r2v\" for reference-to-video with up to 9 reference images. Resolutions 720P and 1080P; duration 3-15 seconds at 24 fps; native synchronized audio is always generated (do not set generateAudio or negativePrompt). Supported aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9. HappyHorse 1.1 takes image references only and renders a native synchronized audio track (always on; do not set generateAudio or a negative prompt). Pick the model by mode: happyhorse-1.1-t2v for text-to-video (no reference image), happyhorse-1.1-i2v for image-to-video from a single first frame, and happyhorse-1.1-r2v for reference-to-video with 1 to 9 reference images. For r2v, tag the images in the prompt as [Image 1]…[Image 9] and assign each a clear role. HappyHorse does not accept reference videos or reference audios. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
248
+ "description": "\"ltx25\" (default): LTX 2.5 with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Video model. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio. \"minimax-h3-t2v\": standard 20-step MiniMax H3 text-to-video; \"minimax-h3-t2v-turbo\": the existing 4-step LightX2V Turbo text-to-video; \"minimax-h3-fasth3-t2v-turbo\": the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple; \"minimax-h3-fasth3-t2v-turbo-2stage\": the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution picks the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 14 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 14 Spark per second, and 720 renders a 384px canvas (672x384 is delivered at 1344x768) for about the regular FastH3 price of 4 Spark per second. The estimate prices every request. It takes the same inputs, durations and LoRAs as \"minimax-h3-fasth3-t2v-turbo\". Choose it when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. All use native audio, fixed 24fps, 5.17-15.08s, and a 768p-class 32px-grid canvas; use animate_photo for H3 image-conditioned modes. Base and Turbo T2V/I2V/FLF2V prompts use the exact ordered fields integrated_multimodal_description, overall_soundscape, and non_diegetic_music; I2V/FLF2V prepend the official alignment line. \"minimax-h3-r2v\": standard 20-step MiniMax H3 reference-to-video; \"minimax-h3-r2v-turbo\": the dedicated LightX2V 4-step Ref2VA Turbo workflow using Euler/simple and a 960x544 default. FastH3 has no R2V mode. Both R2V selectors accept up to 9 images, 3 videos, and 3 audios (12 files total); at least one visual reference (image or video) is required and audio alone is invalid. Select references with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and address them with the official <Subject N>/<Picture N>/<Video N>/<Audio N> semantics. Seedance quality is selected only by model: use \"seedance2-mini\" for Seedance 2.0 Mini or faster/lower-cost 720p iteration, and use \"seedance2\" for the full Seedance 2.0 model, explicit full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft or Mini. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p and 720p ONLY (it cannot render 1080p or 4K), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 1080p/4K requests, which Seedance 2.5 cannot satisfy. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. Alibaba HappyHorse 1.1 video models (third-party vendor — requires Premium Spark). Select by mode: \"happyhorse-1.1-t2v\" for text-to-video, \"happyhorse-1.1-i2v\" for image-to-video from one first-frame image, and \"happyhorse-1.1-r2v\" for reference-to-video with up to 9 reference images. Resolutions 720P and 1080P; duration 3-15 seconds at 24 fps; native synchronized audio is always generated (do not set generateAudio or negativePrompt). Supported aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9. HappyHorse 1.1 takes image references only and renders a native synchronized audio track (always on; do not set generateAudio or a negative prompt). Pick the model by mode: happyhorse-1.1-t2v for text-to-video (no reference image), happyhorse-1.1-i2v for image-to-video from a single first frame, and happyhorse-1.1-r2v for reference-to-video with 1 to 9 reference images. For r2v, tag the images in the prompt as [Image 1]…[Image 9] and assign each a clear role. HappyHorse does not accept reference videos or reference audios. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
248
249
  },
249
250
  "generateAudio": {
250
251
  "type": "boolean",
@@ -281,15 +282,7 @@
281
282
  },
282
283
  "targetResolution": {
283
284
  "type": "number",
284
- "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini and Seedance 2.5 support 480p/720p only, so never set 1080p or 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. HappyHorse supports only 720p and 1080p. Never set 4K for Wan 3 or HappyHorse. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for H3 and never 1080p or 4K. Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, Wan 3, HappyHorse, or MiniMax H3, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead. For MiniMax H3 2K requests keep targetResolution at 768 (or omit it) and set outputScale to 2 instead."
285
- },
286
- "outputScale": {
287
- "type": "integer",
288
- "enum": [
289
- 1,
290
- 2
291
- ],
292
- "description": "MiniMax H3 only. 2 delivers 2K output: the clip renders on the normal H3 canvas and is delivered at twice its width and height (1344x768 becomes 2688x1536) with the same length and audio, for +10 Spark per second (+6 at 480p). Set 2 only when the user asks for 2K, 1440p-class or extra-sharp H3 output; leave unset otherwise. Ignored for other models."
285
+ "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini and Seedance 2.5 support 480p/720p only, so never set 1080p or 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. HappyHorse supports only 720p and 1080p. Never set 4K for Wan 3 or HappyHorse. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for the regular H3 selectors and never 1080p or 4K. The two-stage H3 selector \"minimax-h3-fasth3-t2v-turbo-2stage\" delivers twice the canvas, so there targetResolution names the delivered short-edge class: 1080 (544px canvas short edge: 960x544 delivered at 1920x1088), 1440 for 2K (the 1344x768 canvas delivered at 2688x1536), or 720 (384px canvas: 672x384 delivered at 1344x768); omit it for 2K. Never set 4K for H3. Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, Wan 3, HappyHorse, or MiniMax H3, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead."
293
286
  },
294
287
  "numberOfVariations": {
295
288
  "type": "number",
@@ -313,7 +306,7 @@
313
306
  "type": "string",
314
307
  "minLength": 1
315
308
  },
316
- "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-t2v\", \"minimax-h3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo\", \"minimax-h3-r2v\", \"minimax-h3-r2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
309
+ "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-t2v\", \"minimax-h3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo-2stage\", \"minimax-h3-r2v\", \"minimax-h3-r2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
317
310
  },
318
311
  "loraStrengths": {
319
312
  "type": "array",
@@ -833,13 +826,15 @@
833
826
  "minimax-h3-i2v",
834
827
  "minimax-h3-i2v-turbo",
835
828
  "minimax-h3-fasth3-i2v-turbo",
829
+ "minimax-h3-fasth3-i2v-turbo-2stage",
836
830
  "minimax-h3-flf2v",
837
831
  "minimax-h3-flf2v-turbo",
838
832
  "minimax-h3-fasth3-flf2v-turbo",
833
+ "minimax-h3-fasth3-flf2v-turbo-2stage",
839
834
  "wan3.0-video",
840
835
  "wan3.0-spicy-video"
841
836
  ],
842
- "description": "\"ltx25\" (default): LTX 2.5 I2V or first/last-frame video with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Which video model to use. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio, up to 10s. \"minimax-h3-i2v\" is standard MiniMax H3 from one first frame; \"minimax-h3-i2v-turbo\" is the existing 4-step LightX2V Turbo engine; \"minimax-h3-fasth3-i2v-turbo\" is the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. The matching FLF2V selectors provide standard, LightX2V Turbo, and FastH3 Turbo first/last-frame generation; FastH3 has no R2V mode; use frameRole=\"both\" and provide the end frame. H3 generates native audio at fixed 24fps for 5.17-15.08s and has no negative-prompt input. H3 Base and Turbo prompts use the exact three-field contract and the official mode-specific alignment line. Do not set Seedance here; use generate_video with Seedance references. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
837
+ "description": "\"ltx25\" (default): LTX 2.5 I2V or first/last-frame video with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Which video model to use. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio, up to 10s. \"minimax-h3-i2v\" is standard MiniMax H3 from one first frame; \"minimax-h3-i2v-turbo\" is the existing 4-step LightX2V Turbo engine; \"minimax-h3-fasth3-i2v-turbo\" is the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. \"minimax-h3-fasth3-i2v-turbo-2stage\" (first frame) and \"minimax-h3-fasth3-flf2v-turbo-2stage\" (first and last frame) are the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution picks the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 14 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 14 Spark per second, and 720 renders a 384px canvas (672x384 is delivered at 1344x768) for about the regular FastH3 price of 4 Spark per second. The estimate prices every request. They take the same inputs, durations and LoRAs as their FastH3 selectors. Choose them when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. The matching FLF2V selectors provide standard, LightX2V Turbo, FastH3 Turbo, and two-stage FastH3 first/last-frame generation; FastH3 has no R2V mode; use frameRole=\"both\" and provide the end frame. H3 generates native audio at fixed 24fps for 5.17-15.08s and has no negative-prompt input. H3 Base and Turbo prompts use the exact three-field contract and the official mode-specific alignment line. Do not set Seedance here; use generate_video with Seedance references. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
843
838
  },
844
839
  "negativePrompt": {
845
840
  "type": "string",
@@ -871,15 +866,7 @@
871
866
  },
872
867
  "targetResolution": {
873
868
  "type": "number",
874
- "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", or \"1080p\" without exact pixels or an output orientation. This preserves the source image aspect ratio. Wan 3 supports 480p, 720p, and 1080p; HappyHorse supports only 720p and 1080p. Never set 4K for either. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for H3 and never 1080p or 4K. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\" or \"720p landscape\", use exact-pixel aspectRatio instead. For MiniMax H3 2K requests keep targetResolution at 768 (or omit it) and set outputScale to 2 instead."
875
- },
876
- "outputScale": {
877
- "type": "integer",
878
- "enum": [
879
- 1,
880
- 2
881
- ],
882
- "description": "MiniMax H3 only. 2 delivers 2K output: the clip renders on the normal H3 canvas and is delivered at twice its width and height (1344x768 becomes 2688x1536) with the same length and audio, for +10 Spark per second (+6 at 480p). Set 2 only when the user asks for 2K, 1440p-class or extra-sharp H3 output; leave unset otherwise. Ignored for other models."
869
+ "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", or \"1080p\" without exact pixels or an output orientation. This preserves the source image aspect ratio. Wan 3 supports 480p, 720p, and 1080p; HappyHorse supports only 720p and 1080p. Never set 4K for either. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for the regular H3 selectors and never 1080p or 4K. The two-stage H3 selectors \"minimax-h3-fasth3-i2v-turbo-2stage\" and \"minimax-h3-fasth3-flf2v-turbo-2stage\" deliver twice the canvas, so there targetResolution names the delivered short-edge class: 1080 (544px canvas short edge: 960x544 delivered at 1920x1088), 1440 for 2K (the 1344x768 canvas delivered at 2688x1536), or 720 (384px canvas: 672x384 delivered at 1344x768); omit it for 2K. The source aspect is kept: a portrait source at 1080 renders 544x960 and is delivered at 1088x1920. Never set 4K for H3. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\" or \"720p landscape\", use exact-pixel aspectRatio instead."
883
870
  },
884
871
  "sourceImageIndex": {
885
872
  "type": "number",
@@ -947,7 +934,7 @@
947
934
  "type": "string",
948
935
  "minLength": 1
949
936
  },
950
- "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-i2v\", \"minimax-h3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo\", \"minimax-h3-flf2v\", \"minimax-h3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
937
+ "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-i2v\", \"minimax-h3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo-2stage\", \"minimax-h3-flf2v\", \"minimax-h3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo-2stage\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
951
938
  },
952
939
  "loraStrengths": {
953
940
  "type": "array",
@@ -195,6 +195,7 @@
195
195
  "minimax-h3-t2v",
196
196
  "minimax-h3-t2v-turbo",
197
197
  "minimax-h3-fasth3-t2v-turbo",
198
+ "minimax-h3-fasth3-t2v-turbo-2stage",
198
199
  "happyhorse-1.1-t2v",
199
200
  "happyhorse-1.1-i2v",
200
201
  "happyhorse-1.1-r2v",
@@ -203,7 +204,7 @@
203
204
  "wan3.0-video",
204
205
  "wan3.0-spicy-video"
205
206
  ],
206
- "description": "\"ltx25\" (default): LTX 2.5 with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Video model. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio. \"minimax-h3-t2v\": standard 20-step MiniMax H3 text-to-video; \"minimax-h3-t2v-turbo\": the existing 4-step LightX2V Turbo text-to-video; \"minimax-h3-fasth3-t2v-turbo\": the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. All use native audio, fixed 24fps, 5.17-15.08s, and a 768p-class 32px-grid canvas; use animate_photo for H3 image-conditioned modes. Base and Turbo T2V/I2V/FLF2V prompts use the exact ordered fields integrated_multimodal_description, overall_soundscape, and non_diegetic_music; I2V/FLF2V prepend the official alignment line. \"minimax-h3-r2v\": standard 20-step MiniMax H3 reference-to-video; \"minimax-h3-r2v-turbo\": the dedicated LightX2V 4-step Ref2VA Turbo workflow using Euler/simple and a 960x544 default. FastH3 has no R2V mode. Both R2V selectors accept up to 9 images, 3 videos, and 3 audios (12 files total); at least one visual reference (image or video) is required and audio alone is invalid. Select references with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and address them with the official <Subject N>/<Picture N>/<Video N>/<Audio N> semantics. Seedance quality is selected only by model: use \"seedance2-mini\" for Seedance 2.0 Mini or faster/lower-cost 720p iteration, and use \"seedance2\" for the full Seedance 2.0 model, explicit full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft or Mini. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p and 720p ONLY (it cannot render 1080p or 4K), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 1080p/4K requests, which Seedance 2.5 cannot satisfy. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. Alibaba HappyHorse 1.1 video models (third-party vendor — requires Premium Spark). Select by mode: \"happyhorse-1.1-t2v\" for text-to-video, \"happyhorse-1.1-i2v\" for image-to-video from one first-frame image, and \"happyhorse-1.1-r2v\" for reference-to-video with up to 9 reference images. Resolutions 720P and 1080P; duration 3-15 seconds at 24 fps; native synchronized audio is always generated (do not set generateAudio or negativePrompt). Supported aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9. HappyHorse 1.1 takes image references only and renders a native synchronized audio track (always on; do not set generateAudio or a negative prompt). Pick the model by mode: happyhorse-1.1-t2v for text-to-video (no reference image), happyhorse-1.1-i2v for image-to-video from a single first frame, and happyhorse-1.1-r2v for reference-to-video with 1 to 9 reference images. For r2v, tag the images in the prompt as [Image 1]…[Image 9] and assign each a clear role. HappyHorse does not accept reference videos or reference audios. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
207
+ "description": "\"ltx25\" (default): LTX 2.5 with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Video model. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio. \"minimax-h3-t2v\": standard 20-step MiniMax H3 text-to-video; \"minimax-h3-t2v-turbo\": the existing 4-step LightX2V Turbo text-to-video; \"minimax-h3-fasth3-t2v-turbo\": the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple; \"minimax-h3-fasth3-t2v-turbo-2stage\": the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution picks the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 14 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 14 Spark per second, and 720 renders a 384px canvas (672x384 is delivered at 1344x768) for about the regular FastH3 price of 4 Spark per second. The estimate prices every request. It takes the same inputs, durations and LoRAs as \"minimax-h3-fasth3-t2v-turbo\". Choose it when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. All use native audio, fixed 24fps, 5.17-15.08s, and a 768p-class 32px-grid canvas; use animate_photo for H3 image-conditioned modes. Base and Turbo T2V/I2V/FLF2V prompts use the exact ordered fields integrated_multimodal_description, overall_soundscape, and non_diegetic_music; I2V/FLF2V prepend the official alignment line. \"minimax-h3-r2v\": standard 20-step MiniMax H3 reference-to-video; \"minimax-h3-r2v-turbo\": the dedicated LightX2V 4-step Ref2VA Turbo workflow using Euler/simple and a 960x544 default. FastH3 has no R2V mode. Both R2V selectors accept up to 9 images, 3 videos, and 3 audios (12 files total); at least one visual reference (image or video) is required and audio alone is invalid. Select references with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and address them with the official <Subject N>/<Picture N>/<Video N>/<Audio N> semantics. Seedance quality is selected only by model: use \"seedance2-mini\" for Seedance 2.0 Mini or faster/lower-cost 720p iteration, and use \"seedance2\" for the full Seedance 2.0 model, explicit full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft or Mini. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p and 720p ONLY (it cannot render 1080p or 4K), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 1080p/4K requests, which Seedance 2.5 cannot satisfy. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. Alibaba HappyHorse 1.1 video models (third-party vendor — requires Premium Spark). Select by mode: \"happyhorse-1.1-t2v\" for text-to-video, \"happyhorse-1.1-i2v\" for image-to-video from one first-frame image, and \"happyhorse-1.1-r2v\" for reference-to-video with up to 9 reference images. Resolutions 720P and 1080P; duration 3-15 seconds at 24 fps; native synchronized audio is always generated (do not set generateAudio or negativePrompt). Supported aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9. HappyHorse 1.1 takes image references only and renders a native synchronized audio track (always on; do not set generateAudio or a negative prompt). Pick the model by mode: happyhorse-1.1-t2v for text-to-video (no reference image), happyhorse-1.1-i2v for image-to-video from a single first frame, and happyhorse-1.1-r2v for reference-to-video with 1 to 9 reference images. For r2v, tag the images in the prompt as [Image 1]…[Image 9] and assign each a clear role. HappyHorse does not accept reference videos or reference audios. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
207
208
  },
208
209
  "generateAudio": {
209
210
  "type": "boolean",
@@ -240,15 +241,7 @@
240
241
  },
241
242
  "targetResolution": {
242
243
  "type": "number",
243
- "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini and Seedance 2.5 support 480p/720p only, so never set 1080p or 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. HappyHorse supports only 720p and 1080p. Never set 4K for Wan 3 or HappyHorse. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for H3 and never 1080p or 4K. Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, Wan 3, HappyHorse, or MiniMax H3, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead. For MiniMax H3 2K requests keep targetResolution at 768 (or omit it) and set outputScale to 2 instead."
244
- },
245
- "outputScale": {
246
- "type": "integer",
247
- "enum": [
248
- 1,
249
- 2
250
- ],
251
- "description": "MiniMax H3 only. 2 delivers 2K output: the clip renders on the normal H3 canvas and is delivered at twice its width and height (1344x768 becomes 2688x1536) with the same length and audio, for +10 Spark per second (+6 at 480p). Set 2 only when the user asks for 2K, 1440p-class or extra-sharp H3 output; leave unset otherwise. Ignored for other models."
244
+ "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini and Seedance 2.5 support 480p/720p only, so never set 1080p or 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. HappyHorse supports only 720p and 1080p. Never set 4K for Wan 3 or HappyHorse. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for the regular H3 selectors and never 1080p or 4K. The two-stage H3 selector \"minimax-h3-fasth3-t2v-turbo-2stage\" delivers twice the canvas, so there targetResolution names the delivered short-edge class: 1080 (544px canvas short edge: 960x544 delivered at 1920x1088), 1440 for 2K (the 1344x768 canvas delivered at 2688x1536), or 720 (384px canvas: 672x384 delivered at 1344x768); omit it for 2K. Never set 4K for H3. Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, Wan 3, HappyHorse, or MiniMax H3, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead."
252
245
  },
253
246
  "numberOfVariations": {
254
247
  "type": "number",
@@ -272,7 +265,7 @@
272
265
  "type": "string",
273
266
  "minLength": 1
274
267
  },
275
- "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-t2v\", \"minimax-h3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo\", \"minimax-h3-r2v\", \"minimax-h3-r2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
268
+ "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-t2v\", \"minimax-h3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo-2stage\", \"minimax-h3-r2v\", \"minimax-h3-r2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
276
269
  },
277
270
  "loraStrengths": {
278
271
  "type": "array",
@@ -773,13 +766,15 @@
773
766
  "minimax-h3-i2v",
774
767
  "minimax-h3-i2v-turbo",
775
768
  "minimax-h3-fasth3-i2v-turbo",
769
+ "minimax-h3-fasth3-i2v-turbo-2stage",
776
770
  "minimax-h3-flf2v",
777
771
  "minimax-h3-flf2v-turbo",
778
772
  "minimax-h3-fasth3-flf2v-turbo",
773
+ "minimax-h3-fasth3-flf2v-turbo-2stage",
779
774
  "wan3.0-video",
780
775
  "wan3.0-spicy-video"
781
776
  ],
782
- "description": "\"ltx25\" (default): LTX 2.5 I2V or first/last-frame video with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Which video model to use. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio, up to 10s. \"minimax-h3-i2v\" is standard MiniMax H3 from one first frame; \"minimax-h3-i2v-turbo\" is the existing 4-step LightX2V Turbo engine; \"minimax-h3-fasth3-i2v-turbo\" is the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. The matching FLF2V selectors provide standard, LightX2V Turbo, and FastH3 Turbo first/last-frame generation; FastH3 has no R2V mode; use frameRole=\"both\" and provide the end frame. H3 generates native audio at fixed 24fps for 5.17-15.08s and has no negative-prompt input. H3 Base and Turbo prompts use the exact three-field contract and the official mode-specific alignment line. Do not set Seedance here; use generate_video with Seedance references. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
777
+ "description": "\"ltx25\" (default): LTX 2.5 I2V or first/last-frame video with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Which video model to use. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio, up to 10s. \"minimax-h3-i2v\" is standard MiniMax H3 from one first frame; \"minimax-h3-i2v-turbo\" is the existing 4-step LightX2V Turbo engine; \"minimax-h3-fasth3-i2v-turbo\" is the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. \"minimax-h3-fasth3-i2v-turbo-2stage\" (first frame) and \"minimax-h3-fasth3-flf2v-turbo-2stage\" (first and last frame) are the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution picks the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 14 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 14 Spark per second, and 720 renders a 384px canvas (672x384 is delivered at 1344x768) for about the regular FastH3 price of 4 Spark per second. The estimate prices every request. They take the same inputs, durations and LoRAs as their FastH3 selectors. Choose them when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. The matching FLF2V selectors provide standard, LightX2V Turbo, FastH3 Turbo, and two-stage FastH3 first/last-frame generation; FastH3 has no R2V mode; use frameRole=\"both\" and provide the end frame. H3 generates native audio at fixed 24fps for 5.17-15.08s and has no negative-prompt input. H3 Base and Turbo prompts use the exact three-field contract and the official mode-specific alignment line. Do not set Seedance here; use generate_video with Seedance references. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
783
778
  },
784
779
  "negativePrompt": {
785
780
  "type": "string",
@@ -811,15 +806,7 @@
811
806
  },
812
807
  "targetResolution": {
813
808
  "type": "number",
814
- "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", or \"1080p\" without exact pixels or an output orientation. This preserves the source image aspect ratio. Wan 3 supports 480p, 720p, and 1080p; HappyHorse supports only 720p and 1080p. Never set 4K for either. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for H3 and never 1080p or 4K. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\" or \"720p landscape\", use exact-pixel aspectRatio instead. For MiniMax H3 2K requests keep targetResolution at 768 (or omit it) and set outputScale to 2 instead."
815
- },
816
- "outputScale": {
817
- "type": "integer",
818
- "enum": [
819
- 1,
820
- 2
821
- ],
822
- "description": "MiniMax H3 only. 2 delivers 2K output: the clip renders on the normal H3 canvas and is delivered at twice its width and height (1344x768 becomes 2688x1536) with the same length and audio, for +10 Spark per second (+6 at 480p). Set 2 only when the user asks for 2K, 1440p-class or extra-sharp H3 output; leave unset otherwise. Ignored for other models."
809
+ "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", or \"1080p\" without exact pixels or an output orientation. This preserves the source image aspect ratio. Wan 3 supports 480p, 720p, and 1080p; HappyHorse supports only 720p and 1080p. Never set 4K for either. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for the regular H3 selectors and never 1080p or 4K. The two-stage H3 selectors \"minimax-h3-fasth3-i2v-turbo-2stage\" and \"minimax-h3-fasth3-flf2v-turbo-2stage\" deliver twice the canvas, so there targetResolution names the delivered short-edge class: 1080 (544px canvas short edge: 960x544 delivered at 1920x1088), 1440 for 2K (the 1344x768 canvas delivered at 2688x1536), or 720 (384px canvas: 672x384 delivered at 1344x768); omit it for 2K. The source aspect is kept: a portrait source at 1080 renders 544x960 and is delivered at 1088x1920. Never set 4K for H3. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\" or \"720p landscape\", use exact-pixel aspectRatio instead."
823
810
  },
824
811
  "sourceImageIndex": {
825
812
  "type": "number",
@@ -887,7 +874,7 @@
887
874
  "type": "string",
888
875
  "minLength": 1
889
876
  },
890
- "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-i2v\", \"minimax-h3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo\", \"minimax-h3-flf2v\", \"minimax-h3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
877
+ "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-i2v\", \"minimax-h3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo-2stage\", \"minimax-h3-flf2v\", \"minimax-h3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo-2stage\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
891
878
  },
892
879
  "loraStrengths": {
893
880
  "type": "array",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@sogni-ai/sogni-protocol",
3
- "version": "1.0.0-alpha.39",
3
+ "version": "1.0.0-alpha.40",
4
4
  "description": "Language-neutral protocol artifacts for the Sogni ecosystem: tool schemas, prompts, OpenAI tool manifests, and enums. Consumed by every Sogni SDK (TypeScript, Swift, and future Python/Kotlin/Rust SDKs) so contracts stay in lockstep across languages.",
5
5
  "keywords": [
6
6
  "sogni",
@@ -1,14 +1,13 @@
1
1
  {
2
2
  "contractId": "animate_photo_v1",
3
- "version": "1.1.0",
3
+ "version": "1.2.0",
4
4
  "toolName": "animate_photo",
5
- "baseDescription": "animate_photo produces video from one or more source images using LTX 2.5 by default, with LTX 2.3 retained as a rollback path.\n\nVIDEO PROMPT QUOTING: In video prompts, ONLY use double quotes for spoken dialogue.\nSpeaker tags are allowed outside the quotes for screenplay-style dialogue, e.g.\nCHARACTER: \"We made it.\" Never put on-screen text, overlay text, titles, captions, signs,\nwatermarks, or any visual text in quotes — describe them without quotes (e.g. bold white text\nreading CONGRATULATIONS overlays the lower third). Quotes signal speech to the model;\nquoting non-speech text confuses audio generation.\n\nDIALOGUE DURATION: Spoken dialogue in video prompts must fit the clip duration. Estimate\nat 2.5 words per second for natural cinematic delivery, plus ~1 second per acting beat\n(pauses, gestures, glances between lines). If the user did NOT explicitly request a specific\nduration (using default 5s), extend the duration to fit the dialogue (max 20s). If the user\nexplicitly requested a specific duration, condense the dialogue to fit while preserving meaning.\nAlways check: total dialogue words ÷ 2.5 + beat count ≤ clip duration.\n\nLATEST GENERATED IMAGE FOLLOW-UP: When the newest user turn asks to animate, make a video,\nor make a clip from a generated image/result (for example \"the apple\", \"this one\",\n\"the latest image\"), use animate_photo with that latest generated image. Do not inherit an\nolder Seedance model, resolution, or duration from an unrelated prior turn unless the newest\nuser turn explicitly says Seedance or confirms an immediately suggested Seedance video stage.\nLTX supports exact 2-20s durations, so honor requests like 3s exactly.\n\nWORD BUDGET PER CLIP: The handler REJECTS clips whose spoken dialogue exceeds the budget\n— there is NO auto-trim, so plan dialogue lengths up-front. Hard maximum is 3.75 spoken\nwords per second. Ceilings: 5s = 18 words, 6s = 22 words, 8s = 30 words, 10s = 37 words,\n15s = 56 words, 20s = 75 words. Aim below these ceilings. If a scene's dialogue won't fit,\ntighten the lines, raise the per-clip duration, or split into two segments — do NOT submit\nand hope it works. Spoken words inside double quotes count toward the budget; speaker tags\nand visual/action prose are free.\n\nBATCH VIDEO PER-CLIP DURATION: For a multi-segment animate_photo batch\n(sourceImageIndices + prompts) when the user states a TOTAL video length but NO per-clip\nlength, target 15 seconds per clip when dialogue is involved, and pass that duration\nexplicitly. Example: 60s total → 4 segments × 15s, NOT 6×10s or 12×5s. There is NO 3-clip\nbatch cap: sourceImageIndices supports up to 16 clips, so never split one planned batch into\n\"first 3\" and \"remaining clips\" calls. Do NOT split a planned 15s dialogue scene into multiple\nshorter clips just because a retry complains about word budget; keep duration=15 and tighten\nthe line. Use 5s clips only for single short motion beats or one very short spoken phrase.\nIf the user explicitly specifies a per-clip duration, honor that instead.\n\nN-VERSIONS-OF-A-VIDEO PATTERN: NEVER call animate_photo N times sequentially — ALWAYS\nuse sourceImageIndices in ONE call so all N projects run in parallel. Two flavors:\n(A) SHARED CONTENT — one edit_image/generate_image call with numberOfVariations=N + {|}\nDynamic Prompts to make N distinct source images, then ONE animate_photo call with\nsourceImageIndices=[start..start+N-1] and a single shared prompt.\n(B) PER-CLIP CONTENT — when each clip has DIFFERENT dialogue, jokes, narration, or motion,\npass BOTH sourceImageIndices AND prompts (array of N strings, one per clip) in the SAME\nsingle animate_photo call. The top-level prompt is still required — pass a brief batch summary.\n\nCRITICAL: sourceImageIndices values MUST be read from the latest edit_image/generate_image\ntool result's startIndex field — if startIndex=3 and 4 images were generated, pass\nsourceImageIndices=[3,4,5,6], NOT [0,1,2,3]. Negative indices refer to uploaded images:\n-1 first upload, -2 second upload, -3 third upload. Use repeated -1 entries only when\nintentionally reusing the primary uploaded image. When prompts is supplied, prompts.length\nMUST equal sourceImageIndices.length.\n\nSEEDANCE UPLOADED STORYBOARD DEFAULT: If the user uploaded a storyboard, shot sheet,\nor visual trailer board and asks to make a trailer/video/movie/clip from it, do NOT use\nanimate_photo on the board image and do NOT split it into four LTX clips. Use generate_video\nwith Seedance referenceImageIndices for one continuous clip unless the user explicitly asks\nfor separate LTX clips or first-frame/last-frame animation.\n\nSCREENPLAY / STORYBOARD ANIMATE RULE: For full storyboard projects, use one\nanimate_photo batch with sourceImageIndices + prompts so each clip keeps its own exact\nscene text, stable cast anchors, and screenplay-style speaker-tagged dialogue, and all video\nclips render in parallel. Every speaking clip's video prompt must include that clip's actual\nquoted dialogue, not placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\",\nor \"final line lands\". If each generated scene keyframe should be both the first and last frame\nof its own stitched segment, call animate_photo with sourceImageIndices=[start..end],\nframeRole=\"both\", prompts=[...], and OMIT endImageIndex/endImageIndices so the handler\nuses each source as its own end frame.\n\nUPLOADED REFERENCE LOOPED SKITS: When the user supplies one uploaded reference image and\nasks for several scripted/storyboard/dialogue segments to reuse that same image as BOTH the\nfirst frame and last frame of each segment before stitching, do it in ONE animate_photo call:\nsourceImageIndices=[-1,-1,...], frameRole=\"both\", endImageIndex=-1 (or matching\nendImageIndices=[-1,-1,...]), duration equal to the requested per-segment duration, and\nprompts=[one full scene prompt per segment]. Each prompt must preserve the exact screenplay\nspeaker tags and quoted dialogue from that scene, e.g. HOST: \"...\" GUEST: \"...\". Do not\ndrop speaker tags, convert them to generic narration, omit the last-frame contract, analyze\nthe image first, generate new keyframes first, or split the batch into serial calls. After\nthe single animate_photo batch completes, call stitch_video with the returned video indices.\n\nFor adjacent transition chains: N images create N-1 clips — call animate_photo with\nframeRole=\"both\", sourceImageIndices=[start..end-1], endImageIndices=[start+1..end],\nprompts=[one transition prompt per adjacent pair], then stitch_video. If 5 uploaded images\nare the keyframe sequence, use sourceImageIndices=[-1,-2,-3,-4],\nendImageIndices=[-2,-3,-4,-5], frameRole=\"both\", prompts length 4, then stitch_video.\nDo NOT set endImageIndex=-1 in generated-keyframe patterns — that means every clip ends\non the primary uploaded image.\n\nUPLOADED FIRST-FRAME/LAST-FRAME TRANSITION CHAINS: If the user uploads multiple images\nand asks for a video that transitions from image to image, changes country/version every\nN seconds, or says to use first-frame/last-frame for each pair, call animate_photo directly.\nDo not call edit_image, generate_image, analyze_image, or map_assets_for_model first — the\nuploaded images are already the keyframes. For N uploaded images, create N-1 adjacent clips\nunless the user explicitly asks for a loop back to the first image. Use per-clip duration\nfrom \"every N seconds\" when present; otherwise divide the requested total by the number of\nadjacent clips. After animate_photo returns the batch videos, always call stitch_video with\nthose video indices before finalizing.\n\nMINIMAX H3 2K OUTPUT: outputScale=2 is available only on the MiniMax H3 selectors. It renders on the normal H3 canvas and delivers the clip at twice its width and height (1344x768 becomes 2688x1536) with the same length and audio, for +10 Spark per second (+6 at 480p). Set outputScale=2 only when the user asks for 2K, 1440p-class or extra-sharp H3 output; keep targetResolution at 768 or omit it, never set 1080p, 1440p or 4K for H3, and leave outputScale unset otherwise. Do not set it for any other model.",
5
+ "baseDescription": "animate_photo produces video from one or more source images using LTX 2.5 by default, with LTX 2.3 retained as a rollback path.\n\nVIDEO PROMPT QUOTING: In video prompts, ONLY use double quotes for spoken dialogue.\nSpeaker tags are allowed outside the quotes for screenplay-style dialogue, e.g.\nCHARACTER: \"We made it.\" Never put on-screen text, overlay text, titles, captions, signs,\nwatermarks, or any visual text in quotes — describe them without quotes (e.g. bold white text\nreading CONGRATULATIONS overlays the lower third). Quotes signal speech to the model;\nquoting non-speech text confuses audio generation.\n\nDIALOGUE DURATION: Spoken dialogue in video prompts must fit the clip duration. Estimate\nat 2.5 words per second for natural cinematic delivery, plus ~1 second per acting beat\n(pauses, gestures, glances between lines). If the user did NOT explicitly request a specific\nduration (using default 5s), extend the duration to fit the dialogue (max 20s). If the user\nexplicitly requested a specific duration, condense the dialogue to fit while preserving meaning.\nAlways check: total dialogue words ÷ 2.5 + beat count ≤ clip duration.\n\nLATEST GENERATED IMAGE FOLLOW-UP: When the newest user turn asks to animate, make a video,\nor make a clip from a generated image/result (for example \"the apple\", \"this one\",\n\"the latest image\"), use animate_photo with that latest generated image. Do not inherit an\nolder Seedance model, resolution, or duration from an unrelated prior turn unless the newest\nuser turn explicitly says Seedance or confirms an immediately suggested Seedance video stage.\nLTX supports exact 2-20s durations, so honor requests like 3s exactly.\n\nWORD BUDGET PER CLIP: The handler REJECTS clips whose spoken dialogue exceeds the budget\n— there is NO auto-trim, so plan dialogue lengths up-front. Hard maximum is 3.75 spoken\nwords per second. Ceilings: 5s = 18 words, 6s = 22 words, 8s = 30 words, 10s = 37 words,\n15s = 56 words, 20s = 75 words. Aim below these ceilings. If a scene's dialogue won't fit,\ntighten the lines, raise the per-clip duration, or split into two segments — do NOT submit\nand hope it works. Spoken words inside double quotes count toward the budget; speaker tags\nand visual/action prose are free.\n\nBATCH VIDEO PER-CLIP DURATION: For a multi-segment animate_photo batch\n(sourceImageIndices + prompts) when the user states a TOTAL video length but NO per-clip\nlength, target 15 seconds per clip when dialogue is involved, and pass that duration\nexplicitly. Example: 60s total → 4 segments × 15s, NOT 6×10s or 12×5s. There is NO 3-clip\nbatch cap: sourceImageIndices supports up to 16 clips, so never split one planned batch into\n\"first 3\" and \"remaining clips\" calls. Do NOT split a planned 15s dialogue scene into multiple\nshorter clips just because a retry complains about word budget; keep duration=15 and tighten\nthe line. Use 5s clips only for single short motion beats or one very short spoken phrase.\nIf the user explicitly specifies a per-clip duration, honor that instead.\n\nN-VERSIONS-OF-A-VIDEO PATTERN: NEVER call animate_photo N times sequentially — ALWAYS\nuse sourceImageIndices in ONE call so all N projects run in parallel. Two flavors:\n(A) SHARED CONTENT — one edit_image/generate_image call with numberOfVariations=N + {|}\nDynamic Prompts to make N distinct source images, then ONE animate_photo call with\nsourceImageIndices=[start..start+N-1] and a single shared prompt.\n(B) PER-CLIP CONTENT — when each clip has DIFFERENT dialogue, jokes, narration, or motion,\npass BOTH sourceImageIndices AND prompts (array of N strings, one per clip) in the SAME\nsingle animate_photo call. The top-level prompt is still required — pass a brief batch summary.\n\nCRITICAL: sourceImageIndices values MUST be read from the latest edit_image/generate_image\ntool result's startIndex field — if startIndex=3 and 4 images were generated, pass\nsourceImageIndices=[3,4,5,6], NOT [0,1,2,3]. Negative indices refer to uploaded images:\n-1 first upload, -2 second upload, -3 third upload. Use repeated -1 entries only when\nintentionally reusing the primary uploaded image. When prompts is supplied, prompts.length\nMUST equal sourceImageIndices.length.\n\nSEEDANCE UPLOADED STORYBOARD DEFAULT: If the user uploaded a storyboard, shot sheet,\nor visual trailer board and asks to make a trailer/video/movie/clip from it, do NOT use\nanimate_photo on the board image and do NOT split it into four LTX clips. Use generate_video\nwith Seedance referenceImageIndices for one continuous clip unless the user explicitly asks\nfor separate LTX clips or first-frame/last-frame animation.\n\nSCREENPLAY / STORYBOARD ANIMATE RULE: For full storyboard projects, use one\nanimate_photo batch with sourceImageIndices + prompts so each clip keeps its own exact\nscene text, stable cast anchors, and screenplay-style speaker-tagged dialogue, and all video\nclips render in parallel. Every speaking clip's video prompt must include that clip's actual\nquoted dialogue, not placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\",\nor \"final line lands\". If each generated scene keyframe should be both the first and last frame\nof its own stitched segment, call animate_photo with sourceImageIndices=[start..end],\nframeRole=\"both\", prompts=[...], and OMIT endImageIndex/endImageIndices so the handler\nuses each source as its own end frame.\n\nUPLOADED REFERENCE LOOPED SKITS: When the user supplies one uploaded reference image and\nasks for several scripted/storyboard/dialogue segments to reuse that same image as BOTH the\nfirst frame and last frame of each segment before stitching, do it in ONE animate_photo call:\nsourceImageIndices=[-1,-1,...], frameRole=\"both\", endImageIndex=-1 (or matching\nendImageIndices=[-1,-1,...]), duration equal to the requested per-segment duration, and\nprompts=[one full scene prompt per segment]. Each prompt must preserve the exact screenplay\nspeaker tags and quoted dialogue from that scene, e.g. HOST: \"...\" GUEST: \"...\". Do not\ndrop speaker tags, convert them to generic narration, omit the last-frame contract, analyze\nthe image first, generate new keyframes first, or split the batch into serial calls. After\nthe single animate_photo batch completes, call stitch_video with the returned video indices.\n\nFor adjacent transition chains: N images create N-1 clips — call animate_photo with\nframeRole=\"both\", sourceImageIndices=[start..end-1], endImageIndices=[start+1..end],\nprompts=[one transition prompt per adjacent pair], then stitch_video. If 5 uploaded images\nare the keyframe sequence, use sourceImageIndices=[-1,-2,-3,-4],\nendImageIndices=[-2,-3,-4,-5], frameRole=\"both\", prompts length 4, then stitch_video.\nDo NOT set endImageIndex=-1 in generated-keyframe patterns — that means every clip ends\non the primary uploaded image.\n\nUPLOADED FIRST-FRAME/LAST-FRAME TRANSITION CHAINS: If the user uploads multiple images\nand asks for a video that transitions from image to image, changes country/version every\nN seconds, or says to use first-frame/last-frame for each pair, call animate_photo directly.\nDo not call edit_image, generate_image, analyze_image, or map_assets_for_model first — the\nuploaded images are already the keyframes. For N uploaded images, create N-1 adjacent clips\nunless the user explicitly asks for a loop back to the first image. Use per-clip duration\nfrom \"every N seconds\" when present; otherwise divide the requested total by the number of\nadjacent clips. After animate_photo returns the batch videos, always call stitch_video with\nthose video indices before finalizing.\n\nMINIMAX H3 TWO-STAGE OUTPUT: 1080p and 2K MiniMax H3 delivery is its own selector, not an option. \"minimax-h3-fasth3-i2v-turbo-2stage\" (first frame) and \"minimax-h3-fasth3-flf2v-turbo-2stage\" (first and last frame, frameRole=\"both\" with an end frame) are the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution names the delivered class: 1080 renders a 544px short-edge canvas (960x544 becomes 1920x1088) for 14 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 becomes 2688x1536) for 14 Spark per second, and 720 renders a 384px canvas (672x384 becomes 1344x768) for about the regular FastH3 price of 4 Spark per second. The estimate prices every request. The source aspect is kept (a portrait source at 1080 renders 544x960 and is delivered at 1088x1920). They take the same prompt contract, durations and LoRAs as their FastH3 selectors. Choose them when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. Never set 4K for H3.",
6
6
  "parameterDocs": {
7
7
  "sourceImageIndices": "Batch source image indices. Read startIndex from prior generate_image/edit_image result. Negative = uploaded images (-1 = first upload).",
8
8
  "prompts": "Per-clip prompt array. Length MUST equal sourceImageIndices.length when both are set.",
9
9
  "duration": "Per-clip duration in seconds. Target 15s when dialogue is involved and total length is given without per-clip spec.",
10
10
  "frameRole": "Set to \"both\" for first+last frame transitions using sourceImageIndices + endImageIndices.",
11
- "endImageIndices": "End frames for adjacent-chain transitions. N images → N-1 clips.",
12
- "outputScale": "MiniMax H3 only: 2 = 2K delivery (twice the width and height, same length and audio, +10 Spark/s, +6 at 480p). Set only when the user asks for 2K, 1440p-class or extra-sharp H3 output; otherwise omit."
11
+ "endImageIndices": "End frames for adjacent-chain transitions. N images → N-1 clips."
13
12
  }
14
13
  }
@@ -1,11 +1,10 @@
1
1
  {
2
2
  "contractId": "generate_video_v1",
3
- "version": "1.2.0",
3
+ "version": "1.3.0",
4
4
  "toolName": "generate_video",
5
- "baseDescription": "generate_video produces text-to-video clips and Seedance multimodal reference videos.\nUse for text-only video generation with no source image input. For Seedance, also use this\ntool when uploaded/generated images, videos, or audio are loose references. Use animate_photo\nonly when a non-Seedance source image must become the first frame of an LTX/WAN animation.\n\nSEEDANCE UPLOADED STORYBOARD DEFAULT: When the user uploads a storyboard, shot sheet,\nmood board, or trailer concept image and asks to make a movie trailer/video/clip from it,\ndefault to one Seedance generate_video call with referenceImageIndices=[-1]. Do not first\nextract panels with edit_image, do not generate replacement keyframes, and do not make four\nseparate LTX animate_photo clips unless the user explicitly asks for separate clips or LTX.\nUse seedance2 when premium Spark access is available; if premium access is unavailable,\nexplain the limitation or use the best non-Seedance fallback the user accepts.\n\nSTORYTELLING / COMMERCIAL / TRAILER PROMPTS: For creative video requests, turn the brief\ninto timed, causally connected visual beats before writing the final prompt. Default social\nvideo is 15s 9:16 with a strong first 1-2s, visible escalation, payoff, and brand/CTA/final\nimage. Commercials should show audience desire/problem, transformation, proof/benefit, and\nCTA. Trailers should follow hook → world → disruption → escalation → reveal → title/CTA.\nEvery beat must be generatable: subject, setting, action, camera, lighting, audio, and text\nrole where relevant. Avoid vague \"cinematic\" filler, feature dumps, and beautiful images with\nno visible change.\n\nVIDEO PROMPT QUOTING: ONLY use double quotes for spoken dialogue in video prompts. Never\nquote on-screen text, titles, captions, or visual text elements — describe them without\nquotes. Quotes signal speech to the model and confuse audio generation.\n\nSTORYBOARD TEXT: Structural headings, section numbers, slide titles, panel titles, and\ncaptions in storyboard references may become short audio-only narration/VO or\nkey-message beats, but they are not subtitles, title cards, lower thirds, or visible\noverlays unless the user explicitly asks for visible text, on-screen text, a title\ncard, subtitle, lower third, signage, or CTA. Keep narration as separate brief phrases\nwith pauses; do not concatenate storyboard labels into run-on voiceover.\n\nDIALOGUE DURATION: Spoken dialogue must fit the clip. Estimate 2.5 words per second\nnatural delivery plus ~1s per acting beat. Hard maximum 3.75 words/second.\nCheck: dialogue words ÷ 2.5 + beats ≤ duration. Do not submit oversized dialogue.\n\nLATEST USER DURATION WINS: In follow-up turns, use the newest duration the user states,\neven if a previous assistant message mentioned a longer script/runtime. For example, if\nhistory says \"the full script is 66 seconds\" but the user now says \"do a 30 second version\",\ngenerate the 30 second version. Do not ask a clarification question just because history\ncontains another duration; treat the latest user request as the override.\n\nSEEDANCE DURATION LIMITS: Seedance 2.0 and Mini support 4-15s clips; Seedance 2.5 supports 4-30s clips. If the user explicitly asks\nfor Seedance below 4s, do not silently round up. Ask whether they prefer a 4s Seedance clip\nor an exact-duration LTX clip. If the user did not explicitly ask for Seedance, choose the\nmodel/tool that can satisfy the requested duration exactly.\n\nMINIMAX H3 2K OUTPUT: outputScale=2 is available only on the MiniMax H3 selectors. It renders on the normal H3 canvas and delivers the clip at twice its width and height (1344x768 becomes 2688x1536) with the same length and audio, for +10 Spark per second (+6 at 480p). Set outputScale=2 only when the user asks for 2K, 1440p-class or extra-sharp H3 output; keep targetResolution at 768 or omit it, never set 1080p, 1440p or 4K for H3, and leave outputScale unset otherwise. Do not set it for any other model.",
5
+ "baseDescription": "generate_video produces text-to-video clips and Seedance multimodal reference videos.\nUse for text-only video generation with no source image input. For Seedance, also use this\ntool when uploaded/generated images, videos, or audio are loose references. Use animate_photo\nonly when a non-Seedance source image must become the first frame of an LTX/WAN animation.\n\nSEEDANCE UPLOADED STORYBOARD DEFAULT: When the user uploads a storyboard, shot sheet,\nmood board, or trailer concept image and asks to make a movie trailer/video/clip from it,\ndefault to one Seedance generate_video call with referenceImageIndices=[-1]. Do not first\nextract panels with edit_image, do not generate replacement keyframes, and do not make four\nseparate LTX animate_photo clips unless the user explicitly asks for separate clips or LTX.\nUse seedance2 when premium Spark access is available; if premium access is unavailable,\nexplain the limitation or use the best non-Seedance fallback the user accepts.\n\nSTORYTELLING / COMMERCIAL / TRAILER PROMPTS: For creative video requests, turn the brief\ninto timed, causally connected visual beats before writing the final prompt. Default social\nvideo is 15s 9:16 with a strong first 1-2s, visible escalation, payoff, and brand/CTA/final\nimage. Commercials should show audience desire/problem, transformation, proof/benefit, and\nCTA. Trailers should follow hook → world → disruption → escalation → reveal → title/CTA.\nEvery beat must be generatable: subject, setting, action, camera, lighting, audio, and text\nrole where relevant. Avoid vague \"cinematic\" filler, feature dumps, and beautiful images with\nno visible change.\n\nVIDEO PROMPT QUOTING: ONLY use double quotes for spoken dialogue in video prompts. Never\nquote on-screen text, titles, captions, or visual text elements — describe them without\nquotes. Quotes signal speech to the model and confuse audio generation.\n\nSTORYBOARD TEXT: Structural headings, section numbers, slide titles, panel titles, and\ncaptions in storyboard references may become short audio-only narration/VO or\nkey-message beats, but they are not subtitles, title cards, lower thirds, or visible\noverlays unless the user explicitly asks for visible text, on-screen text, a title\ncard, subtitle, lower third, signage, or CTA. Keep narration as separate brief phrases\nwith pauses; do not concatenate storyboard labels into run-on voiceover.\n\nDIALOGUE DURATION: Spoken dialogue must fit the clip. Estimate 2.5 words per second\nnatural delivery plus ~1s per acting beat. Hard maximum 3.75 words/second.\nCheck: dialogue words ÷ 2.5 + beats ≤ duration. Do not submit oversized dialogue.\n\nLATEST USER DURATION WINS: In follow-up turns, use the newest duration the user states,\neven if a previous assistant message mentioned a longer script/runtime. For example, if\nhistory says \"the full script is 66 seconds\" but the user now says \"do a 30 second version\",\ngenerate the 30 second version. Do not ask a clarification question just because history\ncontains another duration; treat the latest user request as the override.\n\nSEEDANCE DURATION LIMITS: Seedance 2.0 and Mini support 4-15s clips; Seedance 2.5 supports 4-30s clips. If the user explicitly asks\nfor Seedance below 4s, do not silently round up. Ask whether they prefer a 4s Seedance clip\nor an exact-duration LTX clip. If the user did not explicitly ask for Seedance, choose the\nmodel/tool that can satisfy the requested duration exactly.\n\nMINIMAX H3 TWO-STAGE OUTPUT: 1080p and 2K MiniMax H3 delivery is its own selector, not an option. \"minimax-h3-fasth3-t2v-turbo-2stage\" is the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution names the delivered class: 1080 renders a 544px short-edge canvas (960x544 becomes 1920x1088) for 14 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 becomes 2688x1536) for 14 Spark per second, and 720 renders a 384px canvas (672x384 becomes 1344x768) for about the regular FastH3 price of 4 Spark per second. The estimate prices every request. It takes the same prompt contract, durations and LoRAs as \"minimax-h3-fasth3-t2v-turbo\". Choose it when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. Never set 4K for H3.",
6
6
  "parameterDocs": {
7
7
  "prompt": "Video prompt. Use double quotes ONLY for spoken dialogue. Describe visual text without quotes.",
8
- "duration": "Clip duration in seconds. Plan dialogue word count against the 3.75 words/second ceiling.",
9
- "outputScale": "MiniMax H3 only: 2 = 2K delivery (twice the width and height, same length and audio, +10 Spark/s, +6 at 480p). Set only when the user asks for 2K, 1440p-class or extra-sharp H3 output; otherwise omit."
8
+ "duration": "Clip duration in seconds. Plan dialogue word count against the 3.75 words/second ceiling."
10
9
  }
11
10
  }
@@ -30,13 +30,15 @@
30
30
  "minimax-h3-i2v",
31
31
  "minimax-h3-i2v-turbo",
32
32
  "minimax-h3-fasth3-i2v-turbo",
33
+ "minimax-h3-fasth3-i2v-turbo-2stage",
33
34
  "minimax-h3-flf2v",
34
35
  "minimax-h3-flf2v-turbo",
35
36
  "minimax-h3-fasth3-flf2v-turbo",
37
+ "minimax-h3-fasth3-flf2v-turbo-2stage",
36
38
  "wan3.0-video",
37
39
  "wan3.0-spicy-video"
38
40
  ],
39
- "description": "\"ltx25\" (default): LTX 2.5 I2V or first/last-frame video with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Which video model to use. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio, up to 10s. \"minimax-h3-i2v\" is standard MiniMax H3 from one first frame; \"minimax-h3-i2v-turbo\" is the existing 4-step LightX2V Turbo engine; \"minimax-h3-fasth3-i2v-turbo\" is the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. The matching FLF2V selectors provide standard, LightX2V Turbo, and FastH3 Turbo first/last-frame generation; FastH3 has no R2V mode; use frameRole=\"both\" and provide the end frame. H3 generates native audio at fixed 24fps for 5.17-15.08s and has no negative-prompt input. H3 Base and Turbo prompts use the exact three-field contract and the official mode-specific alignment line. Do not set Seedance here; use generate_video with Seedance references. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
41
+ "description": "\"ltx25\" (default): LTX 2.5 I2V or first/last-frame video with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Which video model to use. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio, up to 10s. \"minimax-h3-i2v\" is standard MiniMax H3 from one first frame; \"minimax-h3-i2v-turbo\" is the existing 4-step LightX2V Turbo engine; \"minimax-h3-fasth3-i2v-turbo\" is the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. \"minimax-h3-fasth3-i2v-turbo-2stage\" (first frame) and \"minimax-h3-fasth3-flf2v-turbo-2stage\" (first and last frame) are the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution picks the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 14 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 14 Spark per second, and 720 renders a 384px canvas (672x384 is delivered at 1344x768) for about the regular FastH3 price of 4 Spark per second. The estimate prices every request. They take the same inputs, durations and LoRAs as their FastH3 selectors. Choose them when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. The matching FLF2V selectors provide standard, LightX2V Turbo, FastH3 Turbo, and two-stage FastH3 first/last-frame generation; FastH3 has no R2V mode; use frameRole=\"both\" and provide the end frame. H3 generates native audio at fixed 24fps for 5.17-15.08s and has no negative-prompt input. H3 Base and Turbo prompts use the exact three-field contract and the official mode-specific alignment line. Do not set Seedance here; use generate_video with Seedance references. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
40
42
  },
41
43
  "negativePrompt": {
42
44
  "type": "string",
@@ -68,15 +70,7 @@
68
70
  },
69
71
  "targetResolution": {
70
72
  "type": "number",
71
- "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", or \"1080p\" without exact pixels or an output orientation. This preserves the source image aspect ratio. Wan 3 supports 480p, 720p, and 1080p; HappyHorse supports only 720p and 1080p. Never set 4K for either. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for H3 and never 1080p or 4K. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\" or \"720p landscape\", use exact-pixel aspectRatio instead. For MiniMax H3 2K requests keep targetResolution at 768 (or omit it) and set outputScale to 2 instead."
72
- },
73
- "outputScale": {
74
- "type": "integer",
75
- "enum": [
76
- 1,
77
- 2
78
- ],
79
- "description": "MiniMax H3 only. 2 delivers 2K output: the clip renders on the normal H3 canvas and is delivered at twice its width and height (1344x768 becomes 2688x1536) with the same length and audio, for +10 Spark per second (+6 at 480p). Set 2 only when the user asks for 2K, 1440p-class or extra-sharp H3 output; leave unset otherwise. Ignored for other models."
73
+ "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", or \"1080p\" without exact pixels or an output orientation. This preserves the source image aspect ratio. Wan 3 supports 480p, 720p, and 1080p; HappyHorse supports only 720p and 1080p. Never set 4K for either. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for the regular H3 selectors and never 1080p or 4K. The two-stage H3 selectors \"minimax-h3-fasth3-i2v-turbo-2stage\" and \"minimax-h3-fasth3-flf2v-turbo-2stage\" deliver twice the canvas, so there targetResolution names the delivered short-edge class: 1080 (544px canvas short edge: 960x544 delivered at 1920x1088), 1440 for 2K (the 1344x768 canvas delivered at 2688x1536), or 720 (384px canvas: 672x384 delivered at 1344x768); omit it for 2K. The source aspect is kept: a portrait source at 1080 renders 544x960 and is delivered at 1088x1920. Never set 4K for H3. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\" or \"720p landscape\", use exact-pixel aspectRatio instead."
80
74
  },
81
75
  "sourceImageIndex": {
82
76
  "type": "number",
@@ -144,7 +138,7 @@
144
138
  "type": "string",
145
139
  "minLength": 1
146
140
  },
147
- "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-i2v\", \"minimax-h3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo\", \"minimax-h3-flf2v\", \"minimax-h3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
141
+ "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-i2v\", \"minimax-h3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo-2stage\", \"minimax-h3-flf2v\", \"minimax-h3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo-2stage\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
148
142
  },
149
143
  "loraStrengths": {
150
144
  "type": "array",
@@ -65,6 +65,7 @@
65
65
  "minimax-h3-t2v",
66
66
  "minimax-h3-t2v-turbo",
67
67
  "minimax-h3-fasth3-t2v-turbo",
68
+ "minimax-h3-fasth3-t2v-turbo-2stage",
68
69
  "happyhorse-1.1-t2v",
69
70
  "happyhorse-1.1-i2v",
70
71
  "happyhorse-1.1-r2v",
@@ -73,7 +74,7 @@
73
74
  "wan3.0-video",
74
75
  "wan3.0-spicy-video"
75
76
  ],
76
- "description": "\"ltx25\" (default): LTX 2.5 with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Video model. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio. \"minimax-h3-t2v\": standard 20-step MiniMax H3 text-to-video; \"minimax-h3-t2v-turbo\": the existing 4-step LightX2V Turbo text-to-video; \"minimax-h3-fasth3-t2v-turbo\": the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. All use native audio, fixed 24fps, 5.17-15.08s, and a 768p-class 32px-grid canvas; use animate_photo for H3 image-conditioned modes. Base and Turbo T2V/I2V/FLF2V prompts use the exact ordered fields integrated_multimodal_description, overall_soundscape, and non_diegetic_music; I2V/FLF2V prepend the official alignment line. \"minimax-h3-r2v\": standard 20-step MiniMax H3 reference-to-video; \"minimax-h3-r2v-turbo\": the dedicated LightX2V 4-step Ref2VA Turbo workflow using Euler/simple and a 960x544 default. FastH3 has no R2V mode. Both R2V selectors accept up to 9 images, 3 videos, and 3 audios (12 files total); at least one visual reference (image or video) is required and audio alone is invalid. Select references with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and address them with the official <Subject N>/<Picture N>/<Video N>/<Audio N> semantics. Seedance quality is selected only by model: use \"seedance2-mini\" for Seedance 2.0 Mini or faster/lower-cost 720p iteration, and use \"seedance2\" for the full Seedance 2.0 model, explicit full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft or Mini. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p and 720p ONLY (it cannot render 1080p or 4K), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 1080p/4K requests, which Seedance 2.5 cannot satisfy. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. Alibaba HappyHorse 1.1 video models (third-party vendor — requires Premium Spark). Select by mode: \"happyhorse-1.1-t2v\" for text-to-video, \"happyhorse-1.1-i2v\" for image-to-video from one first-frame image, and \"happyhorse-1.1-r2v\" for reference-to-video with up to 9 reference images. Resolutions 720P and 1080P; duration 3-15 seconds at 24 fps; native synchronized audio is always generated (do not set generateAudio or negativePrompt). Supported aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9. HappyHorse 1.1 takes image references only and renders a native synchronized audio track (always on; do not set generateAudio or a negative prompt). Pick the model by mode: happyhorse-1.1-t2v for text-to-video (no reference image), happyhorse-1.1-i2v for image-to-video from a single first frame, and happyhorse-1.1-r2v for reference-to-video with 1 to 9 reference images. For r2v, tag the images in the prompt as [Image 1]…[Image 9] and assign each a clear role. HappyHorse does not accept reference videos or reference audios. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
77
+ "description": "\"ltx25\" (default): LTX 2.5 with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Video model. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio. \"minimax-h3-t2v\": standard 20-step MiniMax H3 text-to-video; \"minimax-h3-t2v-turbo\": the existing 4-step LightX2V Turbo text-to-video; \"minimax-h3-fasth3-t2v-turbo\": the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple; \"minimax-h3-fasth3-t2v-turbo-2stage\": the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution picks the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 14 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 14 Spark per second, and 720 renders a 384px canvas (672x384 is delivered at 1344x768) for about the regular FastH3 price of 4 Spark per second. The estimate prices every request. It takes the same inputs, durations and LoRAs as \"minimax-h3-fasth3-t2v-turbo\". Choose it when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. All use native audio, fixed 24fps, 5.17-15.08s, and a 768p-class 32px-grid canvas; use animate_photo for H3 image-conditioned modes. Base and Turbo T2V/I2V/FLF2V prompts use the exact ordered fields integrated_multimodal_description, overall_soundscape, and non_diegetic_music; I2V/FLF2V prepend the official alignment line. \"minimax-h3-r2v\": standard 20-step MiniMax H3 reference-to-video; \"minimax-h3-r2v-turbo\": the dedicated LightX2V 4-step Ref2VA Turbo workflow using Euler/simple and a 960x544 default. FastH3 has no R2V mode. Both R2V selectors accept up to 9 images, 3 videos, and 3 audios (12 files total); at least one visual reference (image or video) is required and audio alone is invalid. Select references with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and address them with the official <Subject N>/<Picture N>/<Video N>/<Audio N> semantics. Seedance quality is selected only by model: use \"seedance2-mini\" for Seedance 2.0 Mini or faster/lower-cost 720p iteration, and use \"seedance2\" for the full Seedance 2.0 model, explicit full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft or Mini. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p and 720p ONLY (it cannot render 1080p or 4K), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 1080p/4K requests, which Seedance 2.5 cannot satisfy. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. Alibaba HappyHorse 1.1 video models (third-party vendor — requires Premium Spark). Select by mode: \"happyhorse-1.1-t2v\" for text-to-video, \"happyhorse-1.1-i2v\" for image-to-video from one first-frame image, and \"happyhorse-1.1-r2v\" for reference-to-video with up to 9 reference images. Resolutions 720P and 1080P; duration 3-15 seconds at 24 fps; native synchronized audio is always generated (do not set generateAudio or negativePrompt). Supported aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9. HappyHorse 1.1 takes image references only and renders a native synchronized audio track (always on; do not set generateAudio or a negative prompt). Pick the model by mode: happyhorse-1.1-t2v for text-to-video (no reference image), happyhorse-1.1-i2v for image-to-video from a single first frame, and happyhorse-1.1-r2v for reference-to-video with 1 to 9 reference images. For r2v, tag the images in the prompt as [Image 1]…[Image 9] and assign each a clear role. HappyHorse does not accept reference videos or reference audios. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
77
78
  },
78
79
  "generateAudio": {
79
80
  "type": "boolean",
@@ -110,15 +111,7 @@
110
111
  },
111
112
  "targetResolution": {
112
113
  "type": "number",
113
- "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini and Seedance 2.5 support 480p/720p only, so never set 1080p or 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. HappyHorse supports only 720p and 1080p. Never set 4K for Wan 3 or HappyHorse. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for H3 and never 1080p or 4K. Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, Wan 3, HappyHorse, or MiniMax H3, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead. For MiniMax H3 2K requests keep targetResolution at 768 (or omit it) and set outputScale to 2 instead."
114
- },
115
- "outputScale": {
116
- "type": "integer",
117
- "enum": [
118
- 1,
119
- 2
120
- ],
121
- "description": "MiniMax H3 only. 2 delivers 2K output: the clip renders on the normal H3 canvas and is delivered at twice its width and height (1344x768 becomes 2688x1536) with the same length and audio, for +10 Spark per second (+6 at 480p). Set 2 only when the user asks for 2K, 1440p-class or extra-sharp H3 output; leave unset otherwise. Ignored for other models."
114
+ "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini and Seedance 2.5 support 480p/720p only, so never set 1080p or 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. HappyHorse supports only 720p and 1080p. Never set 4K for Wan 3 or HappyHorse. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for the regular H3 selectors and never 1080p or 4K. The two-stage H3 selector \"minimax-h3-fasth3-t2v-turbo-2stage\" delivers twice the canvas, so there targetResolution names the delivered short-edge class: 1080 (544px canvas short edge: 960x544 delivered at 1920x1088), 1440 for 2K (the 1344x768 canvas delivered at 2688x1536), or 720 (384px canvas: 672x384 delivered at 1344x768); omit it for 2K. Never set 4K for H3. Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, Wan 3, HappyHorse, or MiniMax H3, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead."
122
115
  },
123
116
  "numberOfVariations": {
124
117
  "type": "number",
@@ -142,7 +135,7 @@
142
135
  "type": "string",
143
136
  "minLength": 1
144
137
  },
145
- "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-t2v\", \"minimax-h3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo\", \"minimax-h3-r2v\", \"minimax-h3-r2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
138
+ "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-t2v\", \"minimax-h3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo-2stage\", \"minimax-h3-r2v\", \"minimax-h3-r2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Do not invent ids."
146
139
  },
147
140
  "loraStrengths": {
148
141
  "type": "array",
package/version.json CHANGED
@@ -1,4 +1,4 @@
1
1
  {
2
- "protocolVersion": "6.2.0",
2
+ "protocolVersion": "7.0.0",
3
3
  "description": "Sogni protocol artifact version. SDKs may refuse to operate against a protocolVersion they were not built for. Bump the major when removing or renaming any schema / enum / manifest field; bump the minor when adding new optional fields or new tools; bump the patch for description / prose changes only."
4
4
  }