@sogni-ai/sogni-protocol 1.0.0-alpha.13 → 1.0.0-alpha.14
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
|
@@ -171,12 +171,15 @@
|
|
|
171
171
|
"seedance2-mini",
|
|
172
172
|
"seedance2-fast",
|
|
173
173
|
"minimax-h3-t2v",
|
|
174
|
+
"minimax-h3-t2v-turbo",
|
|
175
|
+
"minimax-h3-i2v-turbo",
|
|
176
|
+
"minimax-h3-flf2v-turbo",
|
|
174
177
|
"happyhorse-1.1-t2v",
|
|
175
178
|
"happyhorse-1.1-i2v",
|
|
176
179
|
"happyhorse-1.1-r2v",
|
|
177
180
|
"minimax-h3-r2v"
|
|
178
181
|
],
|
|
179
|
-
"description": "Video model. \"ltx23\" (default): LTX 2.3 with native audio; Fast/HQ use the distilled 8-step variant and Default Media Quality Pro uses the non-distilled dev variant. \"wan22\": Fast 4-step, simple motion, no audio. Default: \"ltx23\". HappyHorse 1.1 can be used here for \"happyhorse-1.1-t2v\" text-to-video, \"happyhorse-1.1-i2v\" with one uploaded/generated first-frame image via referenceImageIndices, or \"happyhorse-1.1-r2v\" with 1-9 image references. For a locked still image/source-frame animation, animate_photo with videoModel=\"happyhorse-1.1-i2v\" is also valid. HappyHorse supports 720p/1080p, 3-15s clips, native synchronized audio that is always on, image-only references, and no negativePrompt or generateAudio input. MiniMax H3 text-to-video uses \"minimax-h3-t2v\". H3 renders 5.17-15.08s clips at a fixed 24 fps inside a 1344x768 pixel budget on a 32px grid, jointly generates its own stereo audio, takes no negativePrompt input, and supports generateAudio=false to return a video without an audio track. MiniMax H3 reference-to-video uses \"minimax-h3-r2v\", a separate ref2va checkpoint and the only H3 mode that takes loose references: up to 9 reference images, 3 reference videos (24 fps, 2-15s, each with an optional soundtrack) and 3 standalone audio tracks, no more than 12 reference files in total, passed with referenceImageIndices/referenceVideoIndices/referenceAudioIndices. At least one reference
|
|
182
|
+
"description": "Video model. \"ltx23\" (default): LTX 2.3 with native audio; Fast/HQ use the distilled 8-step variant and Default Media Quality Pro uses the non-distilled dev variant. \"wan22\": Fast 4-step, simple motion, no audio. Default: \"ltx23\". HappyHorse 1.1 can be used here for \"happyhorse-1.1-t2v\" text-to-video, \"happyhorse-1.1-i2v\" with one uploaded/generated first-frame image via referenceImageIndices, or \"happyhorse-1.1-r2v\" with 1-9 image references. For a locked still image/source-frame animation, animate_photo with videoModel=\"happyhorse-1.1-i2v\" is also valid. HappyHorse supports 720p/1080p, 3-15s clips, native synchronized audio that is always on, image-only references, and no negativePrompt or generateAudio input. MiniMax H3 standard text-to-video uses \"minimax-h3-t2v\"; use \"minimax-h3-t2v-turbo\" for the 4-step Turbo tier. The Turbo image-to-video and first-to-last-frame selectors are \"minimax-h3-i2v-turbo\" and \"minimax-h3-flf2v-turbo\"; they keep H3's standard geometry, frame grid, and native-audio contract. There is no H3 Turbo r2v selector. H3 renders 5.17-15.08s clips at a fixed 24 fps inside a 1344x768 pixel budget on a 32px grid, jointly generates its own stereo audio, takes no negativePrompt input, and supports generateAudio=false to return a video without an audio track. MiniMax H3 reference-to-video uses \"minimax-h3-r2v\", a separate ref2va checkpoint and the only H3 mode that takes loose references: up to 9 reference images, 3 reference videos (24 fps, 2-15s, each with an optional soundtrack) and 3 standalone audio tracks, no more than 12 reference files in total, passed with referenceImageIndices/referenceVideoIndices/referenceAudioIndices. At least one visual reference is required: one or more images and/or videos. A video can be the only visual input; audio cannot be the sole input. H3 r2v references are NOT locked frames — name them in the prompt with H3's own 1-based per-type labels <Picture 1>/<Video 1>/<Audio 1> and give every one an explicit job (identity, style, camera movement, voice character), stating which reference wins when two disagree. Use animate_photo with \"minimax-h3-i2v\" for a first-frame animation or \"minimax-h3-flf2v\" with frameRole=\"both\" for a first-to-last-frame transition. Seedance quality is selected only by model: use \"seedance2-mini\" for fast, lower-cost 720p Seedance draft iteration unless the user explicitly asks for legacy Fast, use \"seedance2-fast\" only when the user asks for Seedance Fast / seedance-fast, and use \"seedance2\" for the full Seedance 2.0 model, explicit non-fast/full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft, Mini, or the fast model. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices."
|
|
180
183
|
},
|
|
181
184
|
"generateAudio": {
|
|
182
185
|
"type": "boolean",
|
|
@@ -187,21 +190,21 @@
|
|
|
187
190
|
"items": {
|
|
188
191
|
"type": "number"
|
|
189
192
|
},
|
|
190
|
-
"description": "Image references for Seedance (@Image tags), HappyHorse 1.1 r2v, and MiniMax H3 r2v. Use negative indices for uploaded images (-1 first upload, -2 second upload) and non-negative indices for generated image results. For Seedance, omit by default: uploaded images are auto-forwarded as @Image references. Anchor frame intent in the prompt with @Image tags: \"Use @Image1 as the opening shot reference. Begin the video with a composition, subject placement, lighting, mood, and camera framing that closely match @Image1.\" (or @Image2 as the final shot reference). For seamless-loop or \"first frame and last frame identical\" requests with a single uploaded image, anchor it explicitly as both: \"Use @Image1 as both the first frame and last frame so the video loops cleanly back to the opening composition.\" Do not use animate_photo sourceImageIndex/frameRole/endImageIndex for Seedance. For HappyHorse 1.1 r2v, pass 1-9 image references. For MiniMax H3 r2v, at least one image reference is required
|
|
193
|
+
"description": "Image references for Seedance (@Image tags), HappyHorse 1.1 r2v, and MiniMax H3 r2v. Use negative indices for uploaded images (-1 first upload, -2 second upload) and non-negative indices for generated image results. For Seedance, omit by default: uploaded images are auto-forwarded as @Image references. Anchor frame intent in the prompt with @Image tags: \"Use @Image1 as the opening shot reference. Begin the video with a composition, subject placement, lighting, mood, and camera framing that closely match @Image1.\" (or @Image2 as the final shot reference). For seamless-loop or \"first frame and last frame identical\" requests with a single uploaded image, anchor it explicitly as both: \"Use @Image1 as both the first frame and last frame so the video loops cleanly back to the opening composition.\" Do not use animate_photo sourceImageIndex/frameRole/endImageIndex for Seedance. For HappyHorse 1.1 r2v, pass 1-9 image references. For MiniMax H3 r2v, up to 9 images are accepted; at least one image or video reference is required, and H3 references are loose references, not locked frames, and are addressed in the prompt as <Picture 1>, <Picture 2>, and so on in selection order."
|
|
191
194
|
},
|
|
192
195
|
"referenceVideoIndices": {
|
|
193
196
|
"type": "array",
|
|
194
197
|
"items": {
|
|
195
198
|
"type": "number"
|
|
196
199
|
},
|
|
197
|
-
"description": "Optional loose video references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded videos (-1 first uploaded video, -2 second uploaded video) and non-negative indices for generated video results. For Seedance, omit by default: uploaded videos are auto-forwarded as @Video references. Set to choose a subset or include previously generated video URLs. Do not use this for uploaded source-video transforms, upscales, enhancements, restyles, or remasters; use video_to_video with controlMode=\"seedance-v2v\" instead. For MiniMax H3 r2v, up to 3 reference videos (24 fps, 2-15s each, optional soundtrack), addressed as <Video 1>, <Video 2>, and so on in selection order; they
|
|
200
|
+
"description": "Optional loose video references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded videos (-1 first uploaded video, -2 second uploaded video) and non-negative indices for generated video results. For Seedance, omit by default: uploaded videos are auto-forwarded as @Video references. Set to choose a subset or include previously generated video URLs. Do not use this for uploaded source-video transforms, upscales, enhancements, restyles, or remasters; use video_to_video with controlMode=\"seedance-v2v\" instead. For MiniMax H3 r2v, up to 3 reference videos (24 fps, 2-15s each, optional soundtrack), addressed as <Video 1>, <Video 2>, and so on in selection order; they can satisfy the required visual reference without an image."
|
|
198
201
|
},
|
|
199
202
|
"referenceAudioIndices": {
|
|
200
203
|
"type": "array",
|
|
201
204
|
"items": {
|
|
202
205
|
"type": "number"
|
|
203
206
|
},
|
|
204
|
-
"description": "Optional loose audio references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded audio files (-1 first uploaded audio, -2 second uploaded audio) and non-negative indices for generated audio results. For Seedance, omit by default: uploaded audio is auto-forwarded as @Audio references when the Seedance request also has an image or video reference. Use this only for loose background, mood, timing, or style references under an image/video-anchored Seedance shot. If the uploaded audio is the primary sync target, lip-sync target, or requested as sound-to-video/audio-sync, use sound_to_video with videoModel=\"seedance2-mini\" instead unless the user asks for full Seedance. Audio-only Seedance requests are unsupported; use sound_to_video for uploaded-audio-only workflows. For MiniMax H3 r2v, up to 3 standalone audio tracks, addressed as <Audio 1>, <Audio 2>, and so on in selection order — a reference video's own soundtrack takes its Audio number before standalone tracks; they supplement
|
|
207
|
+
"description": "Optional loose audio references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded audio files (-1 first uploaded audio, -2 second uploaded audio) and non-negative indices for generated audio results. For Seedance, omit by default: uploaded audio is auto-forwarded as @Audio references when the Seedance request also has an image or video reference. Use this only for loose background, mood, timing, or style references under an image/video-anchored Seedance shot. If the uploaded audio is the primary sync target, lip-sync target, or requested as sound-to-video/audio-sync, use sound_to_video with videoModel=\"seedance2-mini\" instead unless the user asks for full Seedance. Audio-only Seedance requests are unsupported; use sound_to_video for uploaded-audio-only workflows. For MiniMax H3 r2v, up to 3 standalone audio tracks, addressed as <Audio 1>, <Audio 2>, and so on in selection order — a reference video's own soundtrack takes its Audio number before standalone tracks; they supplement a required image or video reference and cannot be the sole input."
|
|
205
208
|
},
|
|
206
209
|
"width": {
|
|
207
210
|
"type": "number",
|
|
@@ -152,12 +152,15 @@
|
|
|
152
152
|
"seedance2-mini",
|
|
153
153
|
"seedance2-fast",
|
|
154
154
|
"minimax-h3-t2v",
|
|
155
|
+
"minimax-h3-t2v-turbo",
|
|
156
|
+
"minimax-h3-i2v-turbo",
|
|
157
|
+
"minimax-h3-flf2v-turbo",
|
|
155
158
|
"happyhorse-1.1-t2v",
|
|
156
159
|
"happyhorse-1.1-i2v",
|
|
157
160
|
"happyhorse-1.1-r2v",
|
|
158
161
|
"minimax-h3-r2v"
|
|
159
162
|
],
|
|
160
|
-
"description": "Video model. \"ltx23\" (default): LTX 2.3 with native audio; Fast/HQ use the distilled 8-step variant and Default Media Quality Pro uses the non-distilled dev variant. \"wan22\": Fast 4-step, simple motion, no audio. Default: \"ltx23\". HappyHorse 1.1 can be used here for \"happyhorse-1.1-t2v\" text-to-video, \"happyhorse-1.1-i2v\" with one uploaded/generated first-frame image via referenceImageIndices, or \"happyhorse-1.1-r2v\" with 1-9 image references. For a locked still image/source-frame animation, animate_photo with videoModel=\"happyhorse-1.1-i2v\" is also valid. HappyHorse supports 720p/1080p, 3-15s clips, native synchronized audio that is always on, image-only references, and no negativePrompt or generateAudio input. MiniMax H3 text-to-video uses \"minimax-h3-t2v\". H3 renders 5.17-15.08s clips at a fixed 24 fps inside a 1344x768 pixel budget on a 32px grid, jointly generates its own stereo audio, takes no negativePrompt input, and supports generateAudio=false to return a video without an audio track. MiniMax H3 reference-to-video uses \"minimax-h3-r2v\", a separate ref2va checkpoint and the only H3 mode that takes loose references: up to 9 reference images, 3 reference videos (24 fps, 2-15s, each with an optional soundtrack) and 3 standalone audio tracks, no more than 12 reference files in total, passed with referenceImageIndices/referenceVideoIndices/referenceAudioIndices. At least one reference
|
|
163
|
+
"description": "Video model. \"ltx23\" (default): LTX 2.3 with native audio; Fast/HQ use the distilled 8-step variant and Default Media Quality Pro uses the non-distilled dev variant. \"wan22\": Fast 4-step, simple motion, no audio. Default: \"ltx23\". HappyHorse 1.1 can be used here for \"happyhorse-1.1-t2v\" text-to-video, \"happyhorse-1.1-i2v\" with one uploaded/generated first-frame image via referenceImageIndices, or \"happyhorse-1.1-r2v\" with 1-9 image references. For a locked still image/source-frame animation, animate_photo with videoModel=\"happyhorse-1.1-i2v\" is also valid. HappyHorse supports 720p/1080p, 3-15s clips, native synchronized audio that is always on, image-only references, and no negativePrompt or generateAudio input. MiniMax H3 standard text-to-video uses \"minimax-h3-t2v\"; use \"minimax-h3-t2v-turbo\" for the 4-step Turbo tier. The Turbo image-to-video and first-to-last-frame selectors are \"minimax-h3-i2v-turbo\" and \"minimax-h3-flf2v-turbo\"; they keep H3's standard geometry, frame grid, and native-audio contract. There is no H3 Turbo r2v selector. H3 renders 5.17-15.08s clips at a fixed 24 fps inside a 1344x768 pixel budget on a 32px grid, jointly generates its own stereo audio, takes no negativePrompt input, and supports generateAudio=false to return a video without an audio track. MiniMax H3 reference-to-video uses \"minimax-h3-r2v\", a separate ref2va checkpoint and the only H3 mode that takes loose references: up to 9 reference images, 3 reference videos (24 fps, 2-15s, each with an optional soundtrack) and 3 standalone audio tracks, no more than 12 reference files in total, passed with referenceImageIndices/referenceVideoIndices/referenceAudioIndices. At least one visual reference is required: one or more images and/or videos. A video can be the only visual input; audio cannot be the sole input. H3 r2v references are NOT locked frames — name them in the prompt with H3's own 1-based per-type labels <Picture 1>/<Video 1>/<Audio 1> and give every one an explicit job (identity, style, camera movement, voice character), stating which reference wins when two disagree. Use animate_photo with \"minimax-h3-i2v\" for a first-frame animation or \"minimax-h3-flf2v\" with frameRole=\"both\" for a first-to-last-frame transition. Seedance quality is selected only by model: use \"seedance2-mini\" for fast, lower-cost 720p Seedance draft iteration unless the user explicitly asks for legacy Fast, use \"seedance2-fast\" only when the user asks for Seedance Fast / seedance-fast, and use \"seedance2\" for the full Seedance 2.0 model, explicit non-fast/full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft, Mini, or the fast model. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices."
|
|
161
164
|
},
|
|
162
165
|
"generateAudio": {
|
|
163
166
|
"type": "boolean",
|
|
@@ -168,21 +171,21 @@
|
|
|
168
171
|
"items": {
|
|
169
172
|
"type": "number"
|
|
170
173
|
},
|
|
171
|
-
"description": "Image references for Seedance (@Image tags), HappyHorse 1.1 r2v, and MiniMax H3 r2v. Use negative indices for uploaded images (-1 first upload, -2 second upload) and non-negative indices for generated image results. For Seedance, omit by default: uploaded images are auto-forwarded as @Image references. Anchor frame intent in the prompt with @Image tags: \"Use @Image1 as the opening shot reference. Begin the video with a composition, subject placement, lighting, mood, and camera framing that closely match @Image1.\" (or @Image2 as the final shot reference). For seamless-loop or \"first frame and last frame identical\" requests with a single uploaded image, anchor it explicitly as both: \"Use @Image1 as both the first frame and last frame so the video loops cleanly back to the opening composition.\" Do not use animate_photo sourceImageIndex/frameRole/endImageIndex for Seedance. For HappyHorse 1.1 r2v, pass 1-9 image references. For MiniMax H3 r2v, at least one image reference is required
|
|
174
|
+
"description": "Image references for Seedance (@Image tags), HappyHorse 1.1 r2v, and MiniMax H3 r2v. Use negative indices for uploaded images (-1 first upload, -2 second upload) and non-negative indices for generated image results. For Seedance, omit by default: uploaded images are auto-forwarded as @Image references. Anchor frame intent in the prompt with @Image tags: \"Use @Image1 as the opening shot reference. Begin the video with a composition, subject placement, lighting, mood, and camera framing that closely match @Image1.\" (or @Image2 as the final shot reference). For seamless-loop or \"first frame and last frame identical\" requests with a single uploaded image, anchor it explicitly as both: \"Use @Image1 as both the first frame and last frame so the video loops cleanly back to the opening composition.\" Do not use animate_photo sourceImageIndex/frameRole/endImageIndex for Seedance. For HappyHorse 1.1 r2v, pass 1-9 image references. For MiniMax H3 r2v, up to 9 images are accepted; at least one image or video reference is required, and H3 references are loose references, not locked frames, and are addressed in the prompt as <Picture 1>, <Picture 2>, and so on in selection order."
|
|
172
175
|
},
|
|
173
176
|
"referenceVideoIndices": {
|
|
174
177
|
"type": "array",
|
|
175
178
|
"items": {
|
|
176
179
|
"type": "number"
|
|
177
180
|
},
|
|
178
|
-
"description": "Optional loose video references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded videos (-1 first uploaded video, -2 second uploaded video) and non-negative indices for generated video results. For Seedance, omit by default: uploaded videos are auto-forwarded as @Video references. Set to choose a subset or include previously generated video URLs. Do not use this for uploaded source-video transforms, upscales, enhancements, restyles, or remasters; use video_to_video with controlMode=\"seedance-v2v\" instead. For MiniMax H3 r2v, up to 3 reference videos (24 fps, 2-15s each, optional soundtrack), addressed as <Video 1>, <Video 2>, and so on in selection order; they
|
|
181
|
+
"description": "Optional loose video references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded videos (-1 first uploaded video, -2 second uploaded video) and non-negative indices for generated video results. For Seedance, omit by default: uploaded videos are auto-forwarded as @Video references. Set to choose a subset or include previously generated video URLs. Do not use this for uploaded source-video transforms, upscales, enhancements, restyles, or remasters; use video_to_video with controlMode=\"seedance-v2v\" instead. For MiniMax H3 r2v, up to 3 reference videos (24 fps, 2-15s each, optional soundtrack), addressed as <Video 1>, <Video 2>, and so on in selection order; they can satisfy the required visual reference without an image."
|
|
179
182
|
},
|
|
180
183
|
"referenceAudioIndices": {
|
|
181
184
|
"type": "array",
|
|
182
185
|
"items": {
|
|
183
186
|
"type": "number"
|
|
184
187
|
},
|
|
185
|
-
"description": "Optional loose audio references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded audio files (-1 first uploaded audio, -2 second uploaded audio) and non-negative indices for generated audio results. For Seedance, omit by default: uploaded audio is auto-forwarded as @Audio references when the Seedance request also has an image or video reference. Use this only for loose background, mood, timing, or style references under an image/video-anchored Seedance shot. If the uploaded audio is the primary sync target, lip-sync target, or requested as sound-to-video/audio-sync, use sound_to_video with videoModel=\"seedance2-mini\" instead unless the user asks for full Seedance. Audio-only Seedance requests are unsupported; use sound_to_video for uploaded-audio-only workflows. For MiniMax H3 r2v, up to 3 standalone audio tracks, addressed as <Audio 1>, <Audio 2>, and so on in selection order — a reference video's own soundtrack takes its Audio number before standalone tracks; they supplement
|
|
188
|
+
"description": "Optional loose audio references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded audio files (-1 first uploaded audio, -2 second uploaded audio) and non-negative indices for generated audio results. For Seedance, omit by default: uploaded audio is auto-forwarded as @Audio references when the Seedance request also has an image or video reference. Use this only for loose background, mood, timing, or style references under an image/video-anchored Seedance shot. If the uploaded audio is the primary sync target, lip-sync target, or requested as sound-to-video/audio-sync, use sound_to_video with videoModel=\"seedance2-mini\" instead unless the user asks for full Seedance. Audio-only Seedance requests are unsupported; use sound_to_video for uploaded-audio-only workflows. For MiniMax H3 r2v, up to 3 standalone audio tracks, addressed as <Audio 1>, <Audio 2>, and so on in selection order — a reference video's own soundtrack takes its Audio number before standalone tracks; they supplement a required image or video reference and cannot be the sole input."
|
|
186
189
|
},
|
|
187
190
|
"width": {
|
|
188
191
|
"type": "number",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@sogni-ai/sogni-protocol",
|
|
3
|
-
"version": "1.0.0-alpha.
|
|
3
|
+
"version": "1.0.0-alpha.14",
|
|
4
4
|
"description": "Language-neutral protocol artifacts for the Sogni ecosystem: tool schemas, prompts, OpenAI tool manifests, and enums. Consumed by every Sogni SDK (TypeScript, Swift, and future Python/Kotlin/Rust SDKs) so contracts stay in lockstep across languages.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"sogni",
|
|
@@ -38,12 +38,15 @@
|
|
|
38
38
|
"seedance2-mini",
|
|
39
39
|
"seedance2-fast",
|
|
40
40
|
"minimax-h3-t2v",
|
|
41
|
+
"minimax-h3-t2v-turbo",
|
|
42
|
+
"minimax-h3-i2v-turbo",
|
|
43
|
+
"minimax-h3-flf2v-turbo",
|
|
41
44
|
"happyhorse-1.1-t2v",
|
|
42
45
|
"happyhorse-1.1-i2v",
|
|
43
46
|
"happyhorse-1.1-r2v",
|
|
44
47
|
"minimax-h3-r2v"
|
|
45
48
|
],
|
|
46
|
-
"description": "Video model. \"ltx23\" (default): LTX 2.3 with native audio; Fast/HQ use the distilled 8-step variant and Default Media Quality Pro uses the non-distilled dev variant. \"wan22\": Fast 4-step, simple motion, no audio. Default: \"ltx23\". HappyHorse 1.1 can be used here for \"happyhorse-1.1-t2v\" text-to-video, \"happyhorse-1.1-i2v\" with one uploaded/generated first-frame image via referenceImageIndices, or \"happyhorse-1.1-r2v\" with 1-9 image references. For a locked still image/source-frame animation, animate_photo with videoModel=\"happyhorse-1.1-i2v\" is also valid. HappyHorse supports 720p/1080p, 3-15s clips, native synchronized audio that is always on, image-only references, and no negativePrompt or generateAudio input. MiniMax H3 text-to-video uses \"minimax-h3-t2v\". H3 renders 5.17-15.08s clips at a fixed 24 fps inside a 1344x768 pixel budget on a 32px grid, jointly generates its own stereo audio, takes no negativePrompt input, and supports generateAudio=false to return a video without an audio track. MiniMax H3 reference-to-video uses \"minimax-h3-r2v\", a separate ref2va checkpoint and the only H3 mode that takes loose references: up to 9 reference images, 3 reference videos (24 fps, 2-15s, each with an optional soundtrack) and 3 standalone audio tracks, no more than 12 reference files in total, passed with referenceImageIndices/referenceVideoIndices/referenceAudioIndices. At least one reference
|
|
49
|
+
"description": "Video model. \"ltx23\" (default): LTX 2.3 with native audio; Fast/HQ use the distilled 8-step variant and Default Media Quality Pro uses the non-distilled dev variant. \"wan22\": Fast 4-step, simple motion, no audio. Default: \"ltx23\". HappyHorse 1.1 can be used here for \"happyhorse-1.1-t2v\" text-to-video, \"happyhorse-1.1-i2v\" with one uploaded/generated first-frame image via referenceImageIndices, or \"happyhorse-1.1-r2v\" with 1-9 image references. For a locked still image/source-frame animation, animate_photo with videoModel=\"happyhorse-1.1-i2v\" is also valid. HappyHorse supports 720p/1080p, 3-15s clips, native synchronized audio that is always on, image-only references, and no negativePrompt or generateAudio input. MiniMax H3 standard text-to-video uses \"minimax-h3-t2v\"; use \"minimax-h3-t2v-turbo\" for the 4-step Turbo tier. The Turbo image-to-video and first-to-last-frame selectors are \"minimax-h3-i2v-turbo\" and \"minimax-h3-flf2v-turbo\"; they keep H3's standard geometry, frame grid, and native-audio contract. There is no H3 Turbo r2v selector. H3 renders 5.17-15.08s clips at a fixed 24 fps inside a 1344x768 pixel budget on a 32px grid, jointly generates its own stereo audio, takes no negativePrompt input, and supports generateAudio=false to return a video without an audio track. MiniMax H3 reference-to-video uses \"minimax-h3-r2v\", a separate ref2va checkpoint and the only H3 mode that takes loose references: up to 9 reference images, 3 reference videos (24 fps, 2-15s, each with an optional soundtrack) and 3 standalone audio tracks, no more than 12 reference files in total, passed with referenceImageIndices/referenceVideoIndices/referenceAudioIndices. At least one visual reference is required: one or more images and/or videos. A video can be the only visual input; audio cannot be the sole input. H3 r2v references are NOT locked frames — name them in the prompt with H3's own 1-based per-type labels <Picture 1>/<Video 1>/<Audio 1> and give every one an explicit job (identity, style, camera movement, voice character), stating which reference wins when two disagree. Use animate_photo with \"minimax-h3-i2v\" for a first-frame animation or \"minimax-h3-flf2v\" with frameRole=\"both\" for a first-to-last-frame transition. Seedance quality is selected only by model: use \"seedance2-mini\" for fast, lower-cost 720p Seedance draft iteration unless the user explicitly asks for legacy Fast, use \"seedance2-fast\" only when the user asks for Seedance Fast / seedance-fast, and use \"seedance2\" for the full Seedance 2.0 model, explicit non-fast/full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft, Mini, or the fast model. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. Seedance supports multimodal loose reference assets: images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices."
|
|
47
50
|
},
|
|
48
51
|
"generateAudio": {
|
|
49
52
|
"type": "boolean",
|
|
@@ -54,21 +57,21 @@
|
|
|
54
57
|
"items": {
|
|
55
58
|
"type": "number"
|
|
56
59
|
},
|
|
57
|
-
"description": "Image references for Seedance (@Image tags), HappyHorse 1.1 r2v, and MiniMax H3 r2v. Use negative indices for uploaded images (-1 first upload, -2 second upload) and non-negative indices for generated image results. For Seedance, omit by default: uploaded images are auto-forwarded as @Image references. Anchor frame intent in the prompt with @Image tags: \"Use @Image1 as the opening shot reference. Begin the video with a composition, subject placement, lighting, mood, and camera framing that closely match @Image1.\" (or @Image2 as the final shot reference). For seamless-loop or \"first frame and last frame identical\" requests with a single uploaded image, anchor it explicitly as both: \"Use @Image1 as both the first frame and last frame so the video loops cleanly back to the opening composition.\" Do not use animate_photo sourceImageIndex/frameRole/endImageIndex for Seedance. For HappyHorse 1.1 r2v, pass 1-9 image references. For MiniMax H3 r2v, at least one image reference is required
|
|
60
|
+
"description": "Image references for Seedance (@Image tags), HappyHorse 1.1 r2v, and MiniMax H3 r2v. Use negative indices for uploaded images (-1 first upload, -2 second upload) and non-negative indices for generated image results. For Seedance, omit by default: uploaded images are auto-forwarded as @Image references. Anchor frame intent in the prompt with @Image tags: \"Use @Image1 as the opening shot reference. Begin the video with a composition, subject placement, lighting, mood, and camera framing that closely match @Image1.\" (or @Image2 as the final shot reference). For seamless-loop or \"first frame and last frame identical\" requests with a single uploaded image, anchor it explicitly as both: \"Use @Image1 as both the first frame and last frame so the video loops cleanly back to the opening composition.\" Do not use animate_photo sourceImageIndex/frameRole/endImageIndex for Seedance. For HappyHorse 1.1 r2v, pass 1-9 image references. For MiniMax H3 r2v, up to 9 images are accepted; at least one image or video reference is required, and H3 references are loose references, not locked frames, and are addressed in the prompt as <Picture 1>, <Picture 2>, and so on in selection order."
|
|
58
61
|
},
|
|
59
62
|
"referenceVideoIndices": {
|
|
60
63
|
"type": "array",
|
|
61
64
|
"items": {
|
|
62
65
|
"type": "number"
|
|
63
66
|
},
|
|
64
|
-
"description": "Optional loose video references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded videos (-1 first uploaded video, -2 second uploaded video) and non-negative indices for generated video results. For Seedance, omit by default: uploaded videos are auto-forwarded as @Video references. Set to choose a subset or include previously generated video URLs. Do not use this for uploaded source-video transforms, upscales, enhancements, restyles, or remasters; use video_to_video with controlMode=\"seedance-v2v\" instead. For MiniMax H3 r2v, up to 3 reference videos (24 fps, 2-15s each, optional soundtrack), addressed as <Video 1>, <Video 2>, and so on in selection order; they
|
|
67
|
+
"description": "Optional loose video references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded videos (-1 first uploaded video, -2 second uploaded video) and non-negative indices for generated video results. For Seedance, omit by default: uploaded videos are auto-forwarded as @Video references. Set to choose a subset or include previously generated video URLs. Do not use this for uploaded source-video transforms, upscales, enhancements, restyles, or remasters; use video_to_video with controlMode=\"seedance-v2v\" instead. For MiniMax H3 r2v, up to 3 reference videos (24 fps, 2-15s each, optional soundtrack), addressed as <Video 1>, <Video 2>, and so on in selection order; they can satisfy the required visual reference without an image."
|
|
65
68
|
},
|
|
66
69
|
"referenceAudioIndices": {
|
|
67
70
|
"type": "array",
|
|
68
71
|
"items": {
|
|
69
72
|
"type": "number"
|
|
70
73
|
},
|
|
71
|
-
"description": "Optional loose audio references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded audio files (-1 first uploaded audio, -2 second uploaded audio) and non-negative indices for generated audio results. For Seedance, omit by default: uploaded audio is auto-forwarded as @Audio references when the Seedance request also has an image or video reference. Use this only for loose background, mood, timing, or style references under an image/video-anchored Seedance shot. If the uploaded audio is the primary sync target, lip-sync target, or requested as sound-to-video/audio-sync, use sound_to_video with videoModel=\"seedance2-mini\" instead unless the user asks for full Seedance. Audio-only Seedance requests are unsupported; use sound_to_video for uploaded-audio-only workflows. For MiniMax H3 r2v, up to 3 standalone audio tracks, addressed as <Audio 1>, <Audio 2>, and so on in selection order — a reference video's own soundtrack takes its Audio number before standalone tracks; they supplement
|
|
74
|
+
"description": "Optional loose audio references for Seedance and MiniMax H3 r2v. Use negative indices for uploaded audio files (-1 first uploaded audio, -2 second uploaded audio) and non-negative indices for generated audio results. For Seedance, omit by default: uploaded audio is auto-forwarded as @Audio references when the Seedance request also has an image or video reference. Use this only for loose background, mood, timing, or style references under an image/video-anchored Seedance shot. If the uploaded audio is the primary sync target, lip-sync target, or requested as sound-to-video/audio-sync, use sound_to_video with videoModel=\"seedance2-mini\" instead unless the user asks for full Seedance. Audio-only Seedance requests are unsupported; use sound_to_video for uploaded-audio-only workflows. For MiniMax H3 r2v, up to 3 standalone audio tracks, addressed as <Audio 1>, <Audio 2>, and so on in selection order — a reference video's own soundtrack takes its Audio number before standalone tracks; they supplement a required image or video reference and cannot be the sole input."
|
|
72
75
|
},
|
|
73
76
|
"width": {
|
|
74
77
|
"type": "number",
|
package/version.json
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
1
|
{
|
|
2
|
-
"protocolVersion": "1.
|
|
2
|
+
"protocolVersion": "1.9.0",
|
|
3
3
|
"description": "Sogni protocol artifact version. SDKs may refuse to operate against a protocolVersion they were not built for. Bump the major when removing or renaming any schema / enum / manifest field; bump the minor when adding new optional fields or new tools; bump the patch for description / prose changes only."
|
|
4
4
|
}
|