@sogni-ai/sogni-protocol 1.0.0-alpha.26 → 1.0.0-alpha.29
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/manifests/generation-tools.json +292 -73
- package/manifests/openai-tools.json +239 -56
- package/package.json +1 -1
- package/prompts/tools/enhance_prompt.json +1 -1
- package/schemas/tools/animate_photo.schema.json +59 -19
- package/schemas/tools/generate_video.schema.json +66 -18
- package/schemas/tools/sound_to_video.schema.json +41 -12
- package/schemas/tools/video_to_video.schema.json +79 -13
- package/version.json +1 -1
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"version": "2026-
|
|
2
|
+
"version": "2026-07-18.1",
|
|
3
3
|
"source": "sogni-creative-agent/src/tools/definitions/*/definition.ts",
|
|
4
4
|
"schemaRefs": {
|
|
5
5
|
"generate_image": "../schemas/tools/generate_image.schema.json",
|
|
@@ -33,7 +33,7 @@
|
|
|
33
33
|
"properties": {
|
|
34
34
|
"prompt": {
|
|
35
35
|
"type": "string",
|
|
36
|
-
"description": "Text description of the image (50-200 words). POSITIVE phrasing only. Be specific and vivid — reference real artists, franchises, and aesthetics by name.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT ORDER (follow this structure by default): [SUBJECT] → [ATTRIBUTES] → [ACTION/POSE] → [CAMERA/FRAMING] → [ENVIRONMENT] → [LIGHTING] → [STYLE/MEDIUM] → [MATERIALS/TEXTURES] → [SECONDARY DETAILS]. Lead with the main subject and its concrete, observable attributes unless the user explicitly asks for mood, atmosphere, or another prompt shape first. Put the most visually decisive details early.\n\nSPECIFICITY: Use concrete nouns and observable adjectives (\"weathered leather jacket\", not \"cool outfit\"). Specify framing (close-up, medium shot, full body, wide shot), angle (eye level, low angle, high angle, overhead), lighting type (\"soft overcast daylight\", \"warm golden-hour sunlight\", \"moody neon spill with deep shadows\"), and medium/style (\"photorealistic editorial photography\", \"cinematic still frame\", \"clean anime illustration\"). Include materials and textures when relevant (\"brushed aluminum\", \"wet asphalt reflections\", \"heavy wool texture\").\n\nDEFAULTS (fill in when user is underspecified): Framing: medium shot for portraits, wide shot for environments, full-body for fashion/outfits. Angle: eye level unless dramatic perspective requested. Lighting: soft natural light for realism, clean studio light for product shots. Style: photorealistic for realistic models, matching the model's native style for stylized models (e.g. anime illustration for pony/animagine). Reference real artists and franchises by name (\"in the style of Monet's Water Lilies\", \"Wes Anderson symmetrical pastel composition\", \"cyberpunk Blade Runner neon city\", \"shot on 85mm f/1.4 with shallow depth of field\").\n\nAVOID: Starting with abstract mood words alone. Burying the subject after a long style preamble. Stacking incompatible styles. Overloading with competing focal points. Vague phrases like \"very cool\" or \"epic vibes\".\n\nCHARACTER / MASCOT SHEETS: When the user asks for a character sheet, mascot sheet, model sheet, turnaround, expression sheet, or reusable character reference board, create ONE comprehensive professional reference-board image, not separate variations. Include a large hero pose, front / 3/4 / side / back turnaround views, an expression row, action/personality poses, accessories or props, color palette swatches, and compact notes such as personality, fun facts, or brand usage when appropriate. Preserve exact user-provided brand names, slogans, logo text, and requested copy verbatim; incidental tiny notes may be generated by the image model if the user did not provide exact wording. Keep the character consistent across every panel and use clean readable typography.\n\nBATCH VARIATIONS: When numberOfVariations > 1, the prompt describes one output image. Do not mention counts, \"versions\", \"different\", or \"multiple\" in the prompt text unless the user explicitly wants those words visible in the image. Do not describe multiple copies or duplicates of the subject in a single image unless the user asked for a collage, grid, or side-by-side composition. Use Dynamic Prompt syntax to vary one dimension across separate images. Example: user asks \"4 cats in different spots\" → numberOfVariations=4, prompt=\"a black cat {lounging in a sunlit window|prowling through autumn leaves|sitting on a vintage bookshelf|curled up by a fireplace}\" — each output is one cat in one spot. Vary setting, style, lighting, expression, or composition; preserve what the user specified. Preserve any requested orientation, aspect ratio, or exact pixel dimensions across every variation.\n\nSELECTION-GATED IMAGE STAGES: If the user asks for multiple image options/takes/versions and says they will pick one before a later dance, animation, or video, this tool call is still the first step. Generate the complete image batch now with the exact requested count, Dynamic Prompt options for each output, and the final video/image aspect ratio. Do not ask the user to choose before the images exist, and do not call video tools until after the user selects an image.\n\nLINKED VARIANTS: If multiple details must stay paired per output — visual style, outfit, label text, symbol, setting, character, prop, location, or before/after keyframe details — use ONE top-level Dynamic Prompt branch with one complete prompt per output. Do NOT use separate Dynamic Prompt groups for details that must stay together; unpaired groups can mix attributes. If the user asks for per-variant facial, identity, or appearance changes, repeat that guidance inside EVERY option. When the user names a subject or character, write that name or stable role inside every Dynamic Prompt option; a shared prefix outside the branch is not enough because each option must stand alone. Correct shape: \"{full prompt for variant 1 with all paired details|full prompt for variant 2 with all paired details|...}\".\n\nSCREENPLAY / STORYBOARD BATCHES: For multi-scene commercials, storyboards, or shot lists, numberOfVariations should equal the scene count and the prompt should be a single top-level dynamic branch containing one full scene prompt per option, e.g. \"{scene 1 full prompt|scene 2 full prompt|scene 3 full prompt}\". This is the required way to batch scenes with materially different content while still rendering one image per scene. If recurring characters appear, use stable character names and repeat the same visual anchors in every scene option where they appear (age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture). Do not rename, merge, redesign, or drift characters between scene keyframes unless the user asks. Include speaker-tagged dialogue details when dialogue affects the keyframe, e.g. CHARACTER: \"We made it.\" Do not set numberOfVariations=N with only scene 1's prompt; that creates N duplicate versions of scene 1, not N scenes. If the scene count is 16 or fewer, keep it in one call unless the user explicitly asks for separate projects or per-output settings require separate calls.\n\nCOMPOSITE GPT IMAGE 2 STORYBOARD SHEETS: When numberOfVariations=1 and the user asks for one composite video storyboard/keyframe sheet, the prompt must be a compiled storyboard prompt, not a concept summary. Include a SCENES: section with exactly the requested number of concrete entries named SCENE_01, SCENE_02, etc. Every scene entry must include Visual/Action, Camera/Motion, Dialogue/VO (or [no dialogue]), Audio/SFX, and any visible text or reference usage for that scene. Do not provide only the source brief or generic layout instructions; malformed compiled storyboard prompts are blocked by quality audit.\n\nVIDEO KEYFRAMES: When generating images intended as first+last frames for video (animate_photo with frameRole=\"both\"), use numberOfVariations=2 with Dynamic Prompts to create both frames in one call. Make each frame a distinct scene that creates a compelling transition. The video handler will inspect both generated frames and build a scene-aware transition prompt, so focus this image prompt on producing strong start/end visuals. Example: \"a serene lake {at dawn with mist rising and soft pink sky|at dusk with fireflies and deep blue twilight}\"."
|
|
36
|
+
"description": "Text description of the image (50-200 words). POSITIVE phrasing only. Be specific and vivid — reference real artists, franchises, and aesthetics by name.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT ORDER (follow this structure by default): [SUBJECT] → [ATTRIBUTES] → [ACTION/POSE] → [CAMERA/FRAMING] → [ENVIRONMENT] → [LIGHTING] → [STYLE/MEDIUM] → [MATERIALS/TEXTURES] → [SECONDARY DETAILS]. Lead with the main subject and its concrete, observable attributes unless the user explicitly asks for mood, atmosphere, or another prompt shape first. Put the most visually decisive details early.\n\nSPECIFICITY: Use concrete nouns and observable adjectives (\"weathered leather jacket\", not \"cool outfit\"). Specify framing (close-up, medium shot, full body, wide shot), angle (eye level, low angle, high angle, overhead), lighting type (\"soft overcast daylight\", \"warm golden-hour sunlight\", \"moody neon spill with deep shadows\"), and medium/style (\"photorealistic editorial photography\", \"cinematic still frame\", \"clean anime illustration\"). Include materials and textures when relevant (\"brushed aluminum\", \"wet asphalt reflections\", \"heavy wool texture\").\n\nDEFAULTS (fill in when user is underspecified): Framing: medium shot for portraits, wide shot for environments, full-body for fashion/outfits. Angle: eye level unless dramatic perspective requested. Lighting: soft natural light for realism, clean studio light for product shots. Style: photorealistic for realistic models, matching the model's native style for stylized models (e.g. anime illustration for pony/animagine). Reference real artists and franchises by name (\"in the style of Monet's Water Lilies\", \"Wes Anderson symmetrical pastel composition\", \"cyberpunk Blade Runner neon city\", \"shot on 85mm f/1.4 with shallow depth of field\").\n\nAVOID: Starting with abstract mood words alone. Burying the subject after a long style preamble. Stacking incompatible styles. Overloading with competing focal points. Vague phrases like \"very cool\" or \"epic vibes\".\n\nCHARACTER / MASCOT SHEETS: When the user asks for a character sheet, mascot sheet, model sheet, turnaround, expression sheet, or reusable character reference board, create ONE comprehensive professional reference-board image, not separate variations. Include a large hero pose, front / 3/4 / side / back turnaround views, an expression row, action/personality poses, accessories or props, color palette swatches, and compact notes such as personality, fun facts, or brand usage when appropriate. Preserve exact user-provided brand names, slogans, logo text, and requested copy verbatim; incidental tiny notes may be generated by the image model if the user did not provide exact wording. Keep the character consistent across every panel and use clean readable typography.\n\nBATCH VARIATIONS: When numberOfVariations > 1, the prompt describes one output image. Do not mention counts, \"versions\", \"different\", or \"multiple\" in the prompt text unless the user explicitly wants those words visible in the image. Do not describe multiple copies or duplicates of the subject in a single image unless the user asked for a collage, grid, or side-by-side composition. Use Dynamic Prompt syntax to vary one dimension across separate images. Example: user asks \"4 cats in different spots\" → numberOfVariations=4, prompt=\"a black cat {lounging in a sunlit window|prowling through autumn leaves|sitting on a vintage bookshelf|curled up by a fireplace}\" — each output is one cat in one spot. Vary setting, style, lighting, expression, or composition; preserve what the user specified. Preserve any requested orientation, aspect ratio, or exact pixel dimensions across every variation.\n\nSELECTION-GATED IMAGE STAGES: If the user asks for multiple image options/takes/versions and says they will pick one before a later dance, animation, or video, this tool call is still the first step. Generate the complete image batch now with the exact requested count, Dynamic Prompt options for each output, and the final video/image aspect ratio. Do not ask the user to choose before the images exist, and do not call video tools until after the user selects an image.\n\nLINKED VARIANTS: If multiple details must stay paired per output — visual style, outfit, label text, symbol, setting, character, prop, location, or before/after keyframe details — use ONE top-level Dynamic Prompt branch with one complete prompt per output. Do NOT use separate Dynamic Prompt groups for details that must stay together; unpaired groups can mix attributes. Treat every user requirement quantified across the batch (\"each\", \"every\", \"all\") as a hard per-option invariant and repeat it inside EVERY option. This includes identity/pose continuity, required clothing or styling, the actual setting, and literal visible names, labels, captions, flags, logos, or symbols. When the user names a subject or character, write that name or stable role inside every Dynamic Prompt option; a shared prefix outside the branch is not enough because each option must stand alone. Correct shape: \"{full prompt for variant 1 with all paired details|full prompt for variant 2 with all paired details|...}\".\n\nEach option must be a fully concrete standalone image description. Name the actual garment or styling, actual setting, actual accessories, and literal text or symbol shown on screen when requested. Never use meta-placeholder phrasing such as \"style-specific outfit\", \"variant-specific background\", \"include the requested symbol\", \"include a humorous alternate name\", or \"bake the name and symbol into the image\" — those describe the task instead of the image.\n\nORIGINAL + VARIANT BATCHES: When one option remakes or preserves an original and the other options are themed variants, the original option still needs a complete visual contract. Specify the original clothing and setting to preserve, plus every requested label, flag, logo, symbol, or prop. Do not leave that option as only \"the original\" or \"unchanged subject\" while the other options are concrete.\n\nSCREENPLAY / STORYBOARD BATCHES: For multi-scene commercials, storyboards, or shot lists, numberOfVariations should equal the scene count and the prompt should be a single top-level dynamic branch containing one full scene prompt per option, e.g. \"{scene 1 full prompt|scene 2 full prompt|scene 3 full prompt}\". This is the required way to batch scenes with materially different content while still rendering one image per scene. If recurring characters appear, use stable character names and repeat the same visual anchors in every scene option where they appear (age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture). Do not rename, merge, redesign, or drift characters between scene keyframes unless the user asks. Include speaker-tagged dialogue details when dialogue affects the keyframe, e.g. CHARACTER: \"We made it.\" Do not set numberOfVariations=N with only scene 1's prompt; that creates N duplicate versions of scene 1, not N scenes. If the scene count is 16 or fewer, keep it in one call unless the user explicitly asks for separate projects or per-output settings require separate calls.\n\nCOMPOSITE GPT IMAGE 2 STORYBOARD SHEETS: When numberOfVariations=1 and the user asks for one composite video storyboard/keyframe sheet, the prompt must be a compiled storyboard prompt, not a concept summary. Include a SCENES: section with exactly the requested number of concrete entries named SCENE_01, SCENE_02, etc. Every scene entry must include Visual/Action, Camera/Motion, Dialogue/VO (or [no dialogue]), Audio/SFX, and any visible text or reference usage for that scene. Do not provide only the source brief or generic layout instructions; malformed compiled storyboard prompts are blocked by quality audit.\n\nVIDEO KEYFRAMES: When generating images intended as first+last frames for video (animate_photo with frameRole=\"both\"), use numberOfVariations=2 with Dynamic Prompts to create both frames in one call. Make each frame a distinct scene that creates a compelling transition. The video handler will inspect both generated frames and build a scene-aware transition prompt, so focus this image prompt on producing strong start/end visuals. Example: \"a serene lake {at dawn with mist rising and soft pink sky|at dusk with fireflies and deep blue twilight}\".\n\nDISTINCT IMAGE SETS: When the user asks for a set, batch, or collection of distinct images/pages/designs/options in one project, do not write one composite prompt that lists all requested outputs as contents of every image. Set numberOfVariations to the requested output count and use exactly one Dynamic Prompt branch with the same number of options. Put shared style, medium, constraints, and dimensions outside the branch, and put one complete output concept in each branch option. Shape: \"shared constraints {complete prompt for output 1|complete prompt for output 2|...|complete prompt for output N}\".\n\nSEAMLESSLY REPEATING / TILING IMAGES: when the requested image is meant to repeat edge to edge without visible joins — a seamless pattern, repeating texture, wallpaper, tiling background, or an Escher-style tessellation of interlocking figures — it needs a specific configuration, because an ordinary render will not wrap. Set model=\"krea-2-turbo\" and width=1024 with height=1024. 1024x1024 is the only size that tiles reliably; 768, 1280, 1536 and non-square aspect ratios were measured at a 0% success rate, so do not honor a different size for a tiling request without telling the user it will not wrap. Build the prompt as: the subject, then \"a perfect crop from an infinite repeating pattern that continues beyond every edge\", then a motif-scale clause, then a lighting clause. MOTIF SCALE: add \"the motif repeats exactly once across and once down\" for large bold figures, or \"the motif repeats exactly two times across and two times down\" for a medium pattern. Use only 1 or 2 — both measured 63% while 3 and 4 were worse — and always keep the two counts equal. LIGHTING: the frame must not carry a global light gradient, because that is what makes opposite edges disagree — but individual figures may still be shaded. Use \"consistent even illumination from edge to edge, with natural shading and depth modeled within each object\" to keep three-dimensional depth, or \"uniform flat lighting with no shadows or vignette\" for a flatter graphic look and a slightly higher hit rate. Do not omit the lighting clause and do not soften it to something vague like \"evenly lit\": both drop the success rate to near zero. Keep the subject tonally close — one dominant colour family — because high-contrast palettes expose the seam. For an interlocking Escher tessellation rather than a flat pattern, phrase the subject as \"photorealistic Escher tessellation of <objects>\" and add \"every figure complete and recognizable, fitting its neighbors perfectly with no gaps, no overlaps\". Compliant rounded shapes tessellate (frogs, ducks, shells, leaves, feathers, lizards); rigid objects resist. Subject choice matters more than any clause. Tiling is probabilistic even with the right configuration — roughly half of renders wrap cleanly on a good subject and fewer on a hard one — so set numberOfVariations=4 and tell the user to pick whichever tiles, rather than promising every result will."
|
|
37
37
|
},
|
|
38
38
|
"model": {
|
|
39
39
|
"type": "string",
|
|
@@ -64,7 +64,7 @@
|
|
|
64
64
|
"pony-faetality",
|
|
65
65
|
"dreamshaper-xl"
|
|
66
66
|
],
|
|
67
|
-
"description": "DO NOT SET THIS PARAMETER unless the user names a specific model, asks for a very complex image render, asks for a video storyboard/storyboard sheet/contact sheet/panel layout image, asks for anime without naming a model, requests permitted NSFW/nudity content, or explicitly asks for Z-image/Z-image Turbo/Krea 2 Turbo image-to-image. The app auto-selects based on quality settings. Set \"gpt-image-2\" when the user asks for a ChatGPT, OpenAI, GPT, GPT-2, GPT Image, or gpt-image-2 image/model, when they explicitly request very strong text rendering, or by default for complex single-image renders that need dense labels, crisp typography, multi-panel composition, timing notes, foley notes, professional storyboard-sheet layout, or a comprehensive character/mascot/model sheet with turnarounds, expressions, accessories, palette swatches, and brand notes. Set \"one-obsession-v22\" when the user asks for an anime or anime-style image and has not named a specific image model. Set \"z-turbo\" when the user asks for Z-image Turbo; set \"z-image\" when they ask for Z-image without Turbo. Set \"krea-2-turbo\" when the user asks for Krea 2 Turbo. If the user names another image model, honor that requested model instead. A model preference usually does not change which tool to use; the Z-image and Krea 2 Turbo image-to-image exception uses sourceImageIndex plus starting_image_strength on this tool. NSFW rule:
|
|
67
|
+
"description": "DO NOT SET THIS PARAMETER unless the user names a specific model, asks for a very complex image render, asks for a video storyboard/storyboard sheet/contact sheet/panel layout image, asks for anime without naming a model, requests permitted NSFW/nudity content, or explicitly asks for Z-image/Z-image Turbo/Krea 2 Turbo image-to-image. The app auto-selects based on quality settings. Set \"gpt-image-2\" when the user asks for a ChatGPT, OpenAI, GPT, GPT-2, GPT Image, or gpt-image-2 image/model, when they explicitly request very strong text rendering, or by default for complex single-image renders that need dense labels, crisp typography, multi-panel composition, timing notes, foley notes, professional storyboard-sheet layout, or a comprehensive character/mascot/model sheet with turnarounds, expressions, accessories, palette swatches, and brand notes. Set \"one-obsession-v22\" when the user asks for an anime or anime-style image and has not named a specific image model. Set \"z-turbo\" when the user asks for Z-image Turbo; set \"z-image\" when they ask for Z-image without Turbo. Set \"krea-2-turbo\" when the user asks for Krea 2 Turbo. If the user names another image model, honor that requested model instead. A Krea request resolves by reference availability: when the user asks for a character sheet or storyboard with Krea and any uploaded, persona, or previously generated image is in context, call edit_image with model=\"krea-identity-edit\" instead of this tool; with no reference image in context render it here on \"krea-2-turbo\" (or \"dark-beast-krea2\" for Dark Beast), never substituting \"gpt-image-2\". A model preference usually does not change which tool to use; the Z-image and Krea 2 Turbo image-to-image exception uses sourceImageIndex plus starting_image_strength on this tool. NSFW rule: GPT Image 2 and Qwen image models CANNOT do nudity. For permitted NSFW/nudity content, prefer \"dark-beast-krea2\", then \"dark-beast-z-turbo\"; \"chroma1-hd\", \"pony-v7\", \"chroma-detail\", \"chroma-v46-flash\", and \"z-turbo\" are compatible fallbacks."
|
|
68
68
|
},
|
|
69
69
|
"width": {
|
|
70
70
|
"type": "number",
|
|
@@ -76,7 +76,7 @@
|
|
|
76
76
|
},
|
|
77
77
|
"numberOfVariations": {
|
|
78
78
|
"type": "number",
|
|
79
|
-
"description": "Number of variations (1-16). Set to the user's exact requested count in one call whenever they ask for multiple images and the outputs can share project settings. Trigger phrasings: \"draw N\", \"make N\", \"give me N\", \"show me N\", \"render N\", \"create N\", \"generate N\", \"N more\", \"another N\", \"N as separate\", \"N separate images\", \"N different images\", \"N options\", \"N takes\", \"N versions\", \"N variations\", \"N pictures of\", \"all at the same time\", \"in parallel\", \"side by side as separate\". This includes selection-gated image batches that will feed a later dance, animation, or video after the user picks one. Avoid multiple serial generate_image calls unless the user explicitly wants separate projects, isolated approvals, or per-output settings that cannot share one project. If the user previously got a composite \"N subjects in one image\" result and now says \"draw N more as separate images\" / \"as separate\" / \"separately\", set numberOfVariations=N for THIS call — the prior call's numberOfVariations does not carry forward when the user explicitly asks for separation. For screenplay/storyboard batches, this should equal the scene count and the prompt should contain one Dynamic Prompt branch with one full scene prompt per scene; do not set numberOfVariations=N with only one scene prompt. Default: 1 when the user clearly wants a single composite image (e.g. \"draw 2 goats in a meadow\" with no separation language, or explicit \"in one image\" / \"single image\" / \"composite\" / \"sheet\").",
|
|
79
|
+
"description": "Number of variations (1-16). Set to the user's exact requested count in one call whenever they ask for multiple images and the outputs can share project settings. Trigger phrasings: \"draw N\", \"make N\", \"give me N\", \"show me N\", \"render N\", \"create N\", \"generate N\", \"N more\", \"another N\", \"N as separate\", \"N separate images\", \"N different images\", \"N options\", \"N takes\", \"N versions\", \"N variations\", \"N pictures of\", \"all at the same time\", \"in parallel\", \"side by side as separate\". This includes selection-gated image batches that will feed a later dance, animation, or video after the user picks one. Avoid multiple serial generate_image calls unless the user explicitly wants separate projects, isolated approvals, or per-output settings that cannot share one project. If the user previously got a composite \"N subjects in one image\" result and now says \"draw N more as separate images\" / \"as separate\" / \"separately\", set numberOfVariations=N for THIS call — the prior call's numberOfVariations does not carry forward when the user explicitly asks for separation. For screenplay/storyboard batches, this should equal the scene count and the prompt should contain one Dynamic Prompt branch with one full scene prompt per scene; do not set numberOfVariations=N with only one scene prompt. Default: 1 when the user clearly wants a single composite image (e.g. \"draw 2 goats in a meadow\" with no separation language, or explicit \"in one image\" / \"single image\" / \"composite\" / \"sheet\").\n\nFor distinct image sets, numberOfVariations is the output count and the prompt must contain one Dynamic Prompt branch with the same option count. Do not satisfy a multi-output request by putting all requested items into one prompt; that makes every generated result contain the whole set.",
|
|
80
80
|
"minimum": 1,
|
|
81
81
|
"maximum": 16
|
|
82
82
|
},
|
|
@@ -100,15 +100,33 @@
|
|
|
100
100
|
"type": "number",
|
|
101
101
|
"description": "Guidance scale override. Higher values = more prompt adherence. Model-specific defaults are used if omitted. Only set when the user explicitly requests a guidance value."
|
|
102
102
|
},
|
|
103
|
+
"loras": {
|
|
104
|
+
"type": "array",
|
|
105
|
+
"minItems": 1,
|
|
106
|
+
"maxItems": 8,
|
|
107
|
+
"items": {
|
|
108
|
+
"type": "string",
|
|
109
|
+
"minLength": 1
|
|
110
|
+
},
|
|
111
|
+
"description": "Ordered LoRA IDs to apply to a compatible image model. Use only when the user explicitly requests LoRAs or asks for an effect one of these names directly. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only by the five Krea 2 based models: krea2_turbo_fp8_scaled (text-to-image), krea2_identity_edit_v1_2 and krea2_identity_edit_sogni_v0_3_alpha (identity edit), and the dark_beast_krea2_fp8 / dark_beast_krea2_identity_edit_v1_2 community variants.\n\nBipolar sliders — each id names its POSITIVE direction, a negative strength applies the opposite, and 0 disables it: krea2-detail-enhancer, krea2-scene-complexity, krea2-realism (+ = photoreal), krea2-amateur, krea2-candid, krea2-zoom (+ = zoomed in), krea2-skin-detail, krea2-wetness, krea2-age, krea2-height, krea2-weight, krea2-hourglass-figure, krea2-breast, krea2-chest-firmness, krea2-nipple-projection, krea2-warm-light, krea2-afterlight (+ = golden), krea2-skin-tone (+ = darker), krea2-purple-grainy (+ = grainy and muted). Positive-only fine-tunes: krea2-realism-engine (photographic realism), krea2-bloomgirls (polished influencer look), krea2-mystic-x (uncensored adult), krea2-aberrant (industrial body horror), krea2-filter-bypass-2 and krea2-filter-bypass-3 (restore expressions, anatomy and poses the base model flattens; try the 2-vector first). Exact per-LoRA ranges, maturity flags and the full contract: GET /v1/loras/comfy?modelId=<model>. Do not invent ids."
|
|
112
|
+
},
|
|
113
|
+
"loraStrengths": {
|
|
114
|
+
"type": "array",
|
|
115
|
+
"minItems": 1,
|
|
116
|
+
"maxItems": 8,
|
|
117
|
+
"items": {
|
|
118
|
+
"type": "number"
|
|
119
|
+
},
|
|
120
|
+
"description": "Strength for each LoRA in loras, in the same order. Omitting the array uses 1.0 for every LoRA, which is not each LoRA's catalog default — krea2-chest-firmness, krea2-nipple-projection and krea2-height default to 0 (no effect) — so prefer explicit values. Do NOT clamp to 0-1: most Krea 2 LoRAs are bipolar, so krea2-warm-light warms the grade at 2 and cools it at -2. Usable bands vary per LoRA — roughly -2..5 for krea2-detail-enhancer, -3..3 for krea2-warm-light, 3..9 for krea2-candid, 0.5..1 for krea2-realism-engine, 1..2 for the filter-bypass pair. Scale the magnitude to how strongly the user asked; the server clamps out-of-range values, and pushing past a LoRA's recommended band costs image quality rather than adding effect. Preserve explicit user values. Example: loras=[\"krea2-detail-enhancer\",\"krea2-amateur\"], loraStrengths=[3,-2]."
|
|
121
|
+
},
|
|
103
122
|
"gptImageQuality": {
|
|
104
123
|
"type": "string",
|
|
105
124
|
"enum": [
|
|
106
125
|
"low",
|
|
107
126
|
"medium",
|
|
108
|
-
"high"
|
|
109
|
-
"auto"
|
|
127
|
+
"high"
|
|
110
128
|
],
|
|
111
|
-
"description": "Optional GPT Image 2 rendering quality. Only set with model=\"gpt-image-2\" when the user explicitly asks for low/fast, medium/balanced, high/final
|
|
129
|
+
"description": "Optional GPT Image 2 rendering quality. Only set with model=\"gpt-image-2\" when the user explicitly asks for low/fast, medium/balanced, or high/final quality. Otherwise omit it and let the host app media quality setting map Fast to low, HQ to medium, and Pro to high."
|
|
112
130
|
},
|
|
113
131
|
"outputFormat": {
|
|
114
132
|
"type": "string",
|
|
@@ -135,31 +153,59 @@
|
|
|
135
153
|
"type": "function",
|
|
136
154
|
"function": {
|
|
137
155
|
"name": "generate_video",
|
|
138
|
-
"description": "Generate a video from text or Seedance multimodal references. LTX 2.5 is the default and generates audio natively; LTX 2.3 remains available as
|
|
156
|
+
"description": "Generate a video from text or Seedance multimodal references. LTX 2.5 is the default and generates audio natively; LTX 2.3 remains available as rollback and also generates audio natively (dialogue, sounds, ambient music) — describe audio in the prompt. If the user provides exact speech, include it in double quotes; if they only imply speech, describe the performance and voice without inventing quoted words. Never use placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". PERSONA VOICE: Only when the user explicitly asks to use/clone a registered persona voice clip, call resolve_personas first, then set voicePersonaName to select which persona's voice clip to use. Do not set voicePersonaName for ordinary character dialogue or inferred voices; describe those voices in the prompt for native LTX audio. For cross-persona narration (e.g. David narrates a video of Aleyna), resolve both personas and set voicePersonaName to the narrator only if that registered voice was requested. Persona voice requires ltx23 because LTX 2.5 has no compatible ID-LoRA and WAN 2.2 does not support voice identity. For non-Seedance syncing to a specific song or audio track, use sound_to_video instead. For non-Seedance animation from a locked source photo, use animate_photo. Do NOT use for My Personas unless generating a Seedance reference-based video — standard persona videos use resolve_personas → edit_image → animate_photo. SEEDANCE DEFAULT: For seedance2, seedance2-mini, or seedance2-5, default to exactly one video (4-15s on 2.0/Mini, 4-30s on seedance2-5) unless the user explicitly asks for multiple separate outputs. Multiple beats, shots, or scene descriptions in one up-to-15s Seedance prompt are still one video. If the user requests one continuous Seedance video longer than 15s, prefer seedance2-5, which renders up to 30s in a single call; beyond 30s (or on 2.0/Mini) preserve the requested total duration in the prompt/context and let chat orchestration split it into supported segment renders and stitch them instead of clamping it to a short excerpt. Uploaded/generated storyboard, shot-sheet, or trailer-concept images used as Seedance references should become one Seedance generate_video call by default; do not extract panels with edit_image and do not animate the storyboard sheet with LTX unless the user explicitly asks for separate non-Seedance clips. Seedance loose image, video, and audio references go through this tool; do not use animate_photo sourceImageIndex/frameRole/endImageIndex for Seedance. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. If the uploaded audio is the primary sync target, lip-sync target, or requested as sound-to-video/audio-sync, use sound_to_video with videoModel=\"seedance2-mini\" instead of this tool unless the user asks for full Seedance. Use referenceAudioIndices here only when audio is a loose reference under an image/video-anchored Seedance shot. For Seedance, every image — first frame, last frame, or loose reference — is passed through referenceImageIndices (auto-uploaded as referenceImageUrls). Anchor frame intent in the prompt with @Image tags such as \"Use @Image1 as the opening shot reference. Begin the video with a composition, subject placement, lighting, mood, and camera framing that closely match @Image1.\" (or @Image2 as the final shot reference). For seamless-loop or \"first frame and last frame identical\" requests with a single uploaded image, anchor it explicitly as both: \"Use @Image1 as both the first frame and last frame so the video loops cleanly back to the opening composition.\" Assign each useful @Image/@Video/@Audio tag a role. APPROVED STORYBOARD PRODUCTION: When the user asks for a production workflow from an approved storyboard, the chat orchestrator should use the durable CampaignStoryboard contract: render the composite board, audit it, generate per-scene GPT Image 2 keyframes, then render Seedance scene clips and stitch them. Do not replace that with a generic storyboard-reference video unless the user asks for a fast draft. PARTIAL VIDEO EDITS: Do NOT call generate_video to re-render an existing rendered/uploaded video just to change part of it (the bumper, the intro, the end card, a single scene, the last few seconds, etc.). Use replace_video_segment for that — it preserves the unchanged portion, keeps the original audio outside the replaced window, and costs far less. Likewise use extend_video to add new time to the end without rewriting the rest. If the request is vague, ask about vision/mood/style first. Only call once you have clear creative intent. WAN 3 uses the exact selector wan3.0-video: use this tool for text-to-video or loose Image 1/Video 1/Audio 1 references; use animate_photo for native first/last frames, sound_to_video when audio drives timing, and video_to_video for source-video edits.",
|
|
139
157
|
"parameters": {
|
|
140
158
|
"type": "object",
|
|
141
159
|
"properties": {
|
|
142
160
|
"prompt": {
|
|
143
161
|
"type": "string",
|
|
144
|
-
"description": "Write one flowing paragraph like a cinematographer describing a shot. Present tense, specific natural language. Longer clips need longer prompts; close-ups need more detail than wide shots.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. Set skipPromptProcessing=true; for Seedance also set expandPrompt=false.\n\nSTRUCTURE: shot/style → subject (age, clothing, hairstyle, distinguishing details) → environment, lighting, atmosphere → action beat by beat → camera movement → audio and dialogue.\n\nCAST CONTINUITY: For screenplay, script, storyboard, commercial, series, or other longer-form video tasks with recurring characters, use stable character names and repeat the same visual anchors every time they appear (age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture, voice). Do not rename, merge, redesign, or drift characters between scenes unless the user asks.\n\nMOTION PACING: Scale complexity to duration. <=6s: 1 main action beat + 1 simple camera move. Around 10s: 2-3 clear action beats + 1 camera move. >10s: up to 4 action beats in clear sequence. Prefer fewer readable beats over dense micro-actions, especially in short clips.\n\nBLOCKING: Direct the layout like scene blocking. State left/right placement, foreground/background, facing toward/away, and relative distance when multiple subjects or important objects are involved.\n\nACTION: Drive motion with concrete verbs. Specify who moves, what moves, how it moves, and what the camera does. Avoid generic phrases like \"comes alive.\"\n\nDIALOGUE: Put user-provided spoken lines in double quotes. For screenplay-style or longer-form tasks, prefix each spoken line with a stable speaker tag outside the quotes, e.g. CHARACTER: \"We made it.\" Break long speech into short quoted phrases with acting beats between them (gestures, pauses, glances). If the user asks for speech but provides no exact words, describe the visible delivery, voice quality, and emotion without inventing quoted dialogue; ask only when exact wording is the point of the request. Never write placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". Show emotion through visible behavior — not \"she is sad\", instead \"she looks down, pauses, and her voice cracks\". QUOTING RULE: ONLY use double quotes for spoken dialogue. Never quote on-screen text, overlay text, titles, captions, signs, or any visual text — describe them without quotes.\n\nSTORYBOARD TEXT: For storyboard references, structural headings, section numbers, slide titles, panel titles, and captions may become short audio-only narration/voiceover or key-message beats, but they are not subtitles, title cards, lower thirds, or visible overlays unless the user explicitly asks for visible text/on-screen text/title card/subtitle/lower third/signage/CTA. Do not concatenate storyboard labels into run-on voiceover; use separate brief phrases with pauses.\n\nAUDIO: Prompt sound intentionally — voice quality, volume, room tone, ambience, music, weather, footsteps. Include language or accent if relevant. Useful voice/volume anchors: whisper, mutter, shout, scream, energetic announcer, resonant voice with gravitas, distorted radio-style, robotic monotone, childlike curiosity.\n\nCAMERA: Cinematic terms — close-up, tracking shot, dolly in, handheld, slow arc, static frame. Describe movement relative to subject.\n\nFor specific characters (movies, TV): describe visual appearance — don't rely on names alone.\n\nFor complex/creative scenes (characters, dialogue, skits): capture the full creative intent. The system auto-expands into a detailed prompt.\n\nAVOID: Vague prompts, too many characters at once, conflicting lighting logic, readable text or logos, abstract emotions with no visible behavior, rigid numeric constraints (exact angles, counts, speeds).\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax. This is one Sogni project with multiple jobs, so prefer it when all outputs share the same references, model, duration, dimensions, and generation parameters and only prompt text varies. Lock in any camera/subject/style the user specified, vary the rest. Example: \"slow dolly in on a city street {at dawn with golden light|during a rainstorm|at night with neon reflections}\"."
|
|
162
|
+
"description": "Write one flowing paragraph like a cinematographer describing a shot. Present tense, specific natural language. Longer clips need longer prompts; close-ups need more detail than wide shots.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. Set skipPromptProcessing=true; for Seedance or Wan 3 also set expandPrompt=false.\n\nSTRUCTURE: shot/style → subject (age, clothing, hairstyle, distinguishing details) → environment, lighting, atmosphere → action beat by beat → camera movement → audio and dialogue.\n\nCAST CONTINUITY: For screenplay, script, storyboard, commercial, series, or other longer-form video tasks with recurring characters, use stable character names and repeat the same visual anchors every time they appear (age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture, voice). Do not rename, merge, redesign, or drift characters between scenes unless the user asks.\n\nMOTION PACING: Scale complexity to duration. <=6s: 1 main action beat + 1 simple camera move. Around 10s: 2-3 clear action beats + 1 camera move. >10s: up to 4 action beats in clear sequence. Prefer fewer readable beats over dense micro-actions, especially in short clips.\n\nBLOCKING: Direct the layout like scene blocking. State left/right placement, foreground/background, facing toward/away, and relative distance when multiple subjects or important objects are involved.\n\nACTION: Drive motion with concrete verbs. Specify who moves, what moves, how it moves, and what the camera does. Avoid generic phrases like \"comes alive.\"\n\nDIALOGUE: Put user-provided spoken lines in double quotes. For screenplay-style or longer-form tasks, prefix each spoken line with a stable speaker tag outside the quotes, e.g. CHARACTER: \"We made it.\" Break long speech into short quoted phrases with acting beats between them (gestures, pauses, glances). If the user asks for speech but provides no exact words, describe the visible delivery, voice quality, and emotion without inventing quoted dialogue; ask only when exact wording is the point of the request. Never write placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". Show emotion through visible behavior — not \"she is sad\", instead \"she looks down, pauses, and her voice cracks\". QUOTING RULE: ONLY use double quotes for spoken dialogue. Never quote on-screen text, overlay text, titles, captions, signs, or any visual text — describe them without quotes.\n\nSTORYBOARD TEXT: For storyboard references, structural headings, section numbers, slide titles, panel titles, and captions may become short audio-only narration/voiceover or key-message beats, but they are not subtitles, title cards, lower thirds, or visible overlays unless the user explicitly asks for visible text/on-screen text/title card/subtitle/lower third/signage/CTA. Do not concatenate storyboard labels into run-on voiceover; use separate brief phrases with pauses.\n\nAUDIO: Prompt sound intentionally — voice quality, volume, room tone, ambience, music, weather, footsteps. Include language or accent if relevant. Useful voice/volume anchors: whisper, mutter, shout, scream, energetic announcer, resonant voice with gravitas, distorted radio-style, robotic monotone, childlike curiosity.\n\nCAMERA: Cinematic terms — close-up, tracking shot, dolly in, handheld, slow arc, static frame. Describe movement relative to subject.\n\nFor specific characters (movies, TV): describe visual appearance — don't rely on names alone.\n\nFor complex/creative scenes (characters, dialogue, skits): capture the full creative intent. The system auto-expands into a detailed prompt.\n\nAVOID: Vague prompts, too many characters at once, conflicting lighting logic, readable text or logos, abstract emotions with no visible behavior, rigid numeric constraints (exact angles, counts, speeds).\n\nNON-SEEDANCE POSITIVE CONSTRAINTS: For videoModel=\"ltx25\", \"ltx23\", or \"wan22\", prompt is a positive prompt. Translate user avoid/no/don't constraints into affirmative production constraints instead of copying negative phrasing. Preserve exact quoted visible text or dialogue when the user explicitly requests it; keep surrounding surfaces blank.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax. This is one Sogni project with multiple jobs, so prefer it when all outputs share the same references, model, duration, dimensions, and generation parameters and only prompt text varies. Lock in any camera/subject/style the user specified, vary the rest. Example: \"slow dolly in on a city street {at dawn with golden light|during a rainstorm|at night with neon reflections}\"."
|
|
145
163
|
},
|
|
146
164
|
"expandPrompt": {
|
|
147
165
|
"type": "boolean",
|
|
148
|
-
"description": "Seedance only. Whether to run
|
|
166
|
+
"description": "Seedance and Wan 3 only. Whether to run Sogni's exact-model prompt shaper before dispatch. Defaults to true. Set false when the supplied prompt is already model-ready and must remain exact."
|
|
149
167
|
},
|
|
150
168
|
"skipPromptProcessing": {
|
|
151
169
|
"type": "boolean",
|
|
152
|
-
"description": "Bypass automatic prompt shaping/refinement and voice-identity prompt formatting so the prompt text is sent unchanged to the video model. Set true ONLY when the user explicitly says not to modify/rewrite/enhance/expand/change/improve the prompt, or to use/send it exactly, verbatim, or as-is, AND the provided prompt already satisfies the tool requirements. Continue to set non-prompt parameters such as model, duration, count, aspect ratio, and seed. For Seedance literal prompt requests, also set expandPrompt=false. Do not set for ordinary underspecified requests."
|
|
170
|
+
"description": "Bypass automatic prompt shaping/refinement and voice-identity prompt formatting so the prompt text is sent unchanged to the video model. Set true ONLY when the user explicitly says not to modify/rewrite/enhance/expand/change/improve the prompt, or to use/send it exactly, verbatim, or as-is, AND the provided prompt already satisfies the tool requirements. Continue to set non-prompt parameters such as model, duration, count, aspect ratio, and seed. For Seedance or Wan 3 literal prompt requests, also set expandPrompt=false. Do not set for ordinary underspecified requests."
|
|
153
171
|
},
|
|
154
172
|
"duration": {
|
|
155
173
|
"type": "number",
|
|
156
|
-
"description": "Video duration in seconds. Default: 5.
|
|
174
|
+
"description": "Video duration in seconds. Default: 5. Per-model range: LTX 2.3 and LTX 2.5 = 2-20s; Wan 3 = 2-30s; Seedance 2.0 and Mini = 4-15s; Seedance 2.5 = 4-30s; HappyHorse 1.1 = 3-15s. Use when the user explicitly requests a specific length. MiniMax H3 is quantized to a 17-frame grid at a fixed 24 fps and renders 124-362 frames, so an H3 clip runs 5.17-15.08 seconds and a requested length outside that window snaps to the nearest valid H3 length.",
|
|
157
175
|
"minimum": 2,
|
|
158
176
|
"maximum": 30
|
|
159
177
|
},
|
|
178
|
+
"smartDuration": {
|
|
179
|
+
"type": "boolean",
|
|
180
|
+
"description": "Wan 3 only. Set true to let Wan 3 choose an output length from 2-30 seconds. Do not also set duration. Sogni reserves the 30-second maximum before generation and settles the completed job down to Alibaba's reported output duration."
|
|
181
|
+
},
|
|
182
|
+
"ratio": {
|
|
183
|
+
"type": "string",
|
|
184
|
+
"enum": [
|
|
185
|
+
"adaptive",
|
|
186
|
+
"16:9",
|
|
187
|
+
"4:3",
|
|
188
|
+
"1:1",
|
|
189
|
+
"3:4",
|
|
190
|
+
"9:16"
|
|
191
|
+
],
|
|
192
|
+
"description": "Wan 3 only. Output ratio. Use \"adaptive\" to let the provider choose from the source or context; omit to use the provider default."
|
|
193
|
+
},
|
|
194
|
+
"watermark": {
|
|
195
|
+
"type": "boolean",
|
|
196
|
+
"description": "Wan 3 only. Add Alibaba's visible watermark. Defaults to false."
|
|
197
|
+
},
|
|
198
|
+
"referenceFileUrl": {
|
|
199
|
+
"type": "string",
|
|
200
|
+
"description": "Wan 3 only. One public HTTPS document URL for context (DOCX/DOC/XLSX/XLS/PPTX/PPT/PDF/TXT/KEY/PAGES/NUMBERS/Markdown, up to 100 MB; PDF/DOCX/DOC/PPTX/PPT/KEY/PAGES up to 50 pages). Mutually exclusive with referenceLinkUrl and first/last-frame inputs."
|
|
201
|
+
},
|
|
202
|
+
"referenceLinkUrl": {
|
|
203
|
+
"type": "string",
|
|
204
|
+
"description": "Wan 3 only. One public HTTPS webpage URL for context. Mutually exclusive with referenceFileUrl and first/last-frame inputs."
|
|
205
|
+
},
|
|
160
206
|
"negativePrompt": {
|
|
161
207
|
"type": "string",
|
|
162
|
-
"description": "Advanced LTX
|
|
208
|
+
"description": "Advanced LTX/WAN only. Use this field only when the user explicitly asks to set a separate negative prompt. MiniMax H3 has no negative-prompt input; put requested exclusions in prompt. Do not set for MiniMax H3, Seedance, or HappyHorse.\n\nWan 3 has no negativePrompt request field; do not set this for wan3.0-video."
|
|
163
209
|
},
|
|
164
210
|
"videoModel": {
|
|
165
211
|
"type": "string",
|
|
@@ -176,46 +222,47 @@
|
|
|
176
222
|
"happyhorse-1.1-i2v",
|
|
177
223
|
"happyhorse-1.1-r2v",
|
|
178
224
|
"minimax-h3-r2v",
|
|
179
|
-
"minimax-h3-r2v-turbo"
|
|
225
|
+
"minimax-h3-r2v-turbo",
|
|
226
|
+
"wan3.0-video"
|
|
180
227
|
],
|
|
181
|
-
"description": "Video model. \"ltx25\" (default)
|
|
228
|
+
"description": "Video model. \"ltx25\" (default) selects LTX 2.5 standard generation with native audio. Fast, HQ, and Pro currently use the release-validated official Distilled/Turbo workflow; Dev is withheld until upstream publishes and Sogni validates an official ComfyUI Dev recipe. \"ltx23\" remains available as an explicit rollback selector. \"wan22\" is the fast WAN path without native audio. LTX 2.5 V2V controls are exposed separately through video_to_video; voice ID-LoRA, transition LoRA, and 10Eros remain LTX 2.3-only. HappyHorse 1.1 can be used here for \"happyhorse-1.1-t2v\" text-to-video, \"happyhorse-1.1-i2v\" with one uploaded/generated first-frame image via referenceImageIndices, or \"happyhorse-1.1-r2v\" with 1-9 image references. For a locked still image/source-frame animation, animate_photo with videoModel=\"happyhorse-1.1-i2v\" is also valid. HappyHorse supports 720p/1080p, 3-15s clips, native synchronized audio that is always on, image-only references, and no negativePrompt or generateAudio input. MiniMax H3 text-to-video uses \"minimax-h3-t2v\". H3 renders 5.17-15.08s clips at a fixed 24 fps inside a 1344x768 pixel budget on a 32px grid, jointly generates its own stereo audio, takes no negativePrompt input, and supports generateAudio=false to return a video without an audio track. MiniMax H3 reference-to-video uses \"minimax-h3-r2v\", a separate ref2va checkpoint and the only H3 mode that takes loose references: up to 9 reference images, 3 reference videos (24 fps, 2-15s, each with an optional soundtrack) and 3 standalone audio tracks, no more than 12 reference files in total, passed with referenceImageIndices/referenceVideoIndices/referenceAudioIndices. At least one visual reference (image or video) is required; audio alone is invalid. H3 r2v references are NOT locked frames — name them in the prompt with H3's own 1-based per-type labels <Picture 1>/<Video 1>/<Audio 1> and give every one an explicit job (identity, style, camera movement, voice character), stating which reference wins when two disagree. Use animate_photo with \"minimax-h3-i2v\" or \"minimax-h3-i2v-turbo\" and frameRole=\"start\" for an opening-frame I2VA animation, or frameRole=\"end\" for a closing-frame L2VA animation. Use \"minimax-h3-flf2v\" or \"minimax-h3-flf2v-turbo\" with frameRole=\"both\" only for a first-to-last-frame transition. Seedance quality is selected only by model: use \"seedance2-mini\" for fast, lower-cost 720p Seedance draft iteration, and use \"seedance2\" for the full Seedance 2.0 model, explicit non-fast/full-quality requests, 1080p/4K requests, or generated/uploaded storyboard images unless the user explicitly asks for a draft, Mini, or the fast model. Do not use Default Media Quality Fast/HQ/Pro or targetResolution to represent Seedance quality. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p and 720p ONLY (it cannot render 1080p or 4K), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 1080p/4K requests, which Seedance 2.5 cannot satisfy. Seedance supports multimodal loose reference assets. Seedance 2.0 and Mini accept images (up to 9), videos (up to 3), and audios (up to 3), with no more than 12 asset files total; Seedance 2.5 accepts images (up to 30), videos (up to 10), and audios (up to 10), with up to 50 reference media files total, subject to those per-modality caps. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. MiniMax H3 Base and Turbo T2V, I2VA, L2VA, and FLF2VA prompts use exactly integrated_multimodal_description, overall_soundscape, then non_diegetic_music. I2VA prepends the official opening-frame alignment line, L2VA prepends the official duration-aware closing-frame alignment line, and FLF2VA prepends the official two-endpoint alignment line. For dialogue, use stable (S1) speaker IDs; keep identity, action, and delivery outside <d>, with only the language tag and exact spoken words inside <d>[Language] ...</d>. Use <scenetrans> at both connecting points when one line crosses a cut and explicitly state that its audio continues across the cut. Use the plain <cutoff> marker only when the video ending truncates speech; never emit tokenizer-internal <|...|> markers or plain caption/lyrics boundary tags. Do not merely delete pipe characters: caption markers become exact visible text in double quotes, lyrics markers become an ordinary <d>[Language] ...</d> singing block, and <|cutoff|> becomes plain <cutoff>. Standard uses 20 steps with res_multistep/simple. Turbo T2V, I2VA, L2VA, and FLF2VA use 4 steps with simple scheduling; er_sde is the default sampler, and direct CLI A/B overrides may select euler, er_sde, or sa_solver. Ref2VA Turbo is a separate 4-step Euler/simple workflow selected with minimax-h3-r2v-turbo; L2VA still uses the I2V selector with frameRole=\"end\". MiniMax H3 R2V requires at least one visual reference (image or video); audio alone is invalid. Its six ordered prompt sections are subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, and non_diegetic_music. Use <Subject N> for reusable visible content, <Picture N> only for concrete keyframes/composition anchors, <Video N> for whole-video relationships, and <Audio N> for copied or referenced audio. Use minimax-h3-r2v for the standard 20-step Ref2VA workflow and minimax-h3-r2v-turbo for the dedicated LightX2V 4-step Euler/simple Turbo workflow with its upstream-aligned 960x544 default. \"wan3.0-video\" is Alibaba Wan 3: one canonical premium-vendor model for text-to-video, first-frame and first+last-frame animation, loose multimodal references, audio-driven generation, and uploaded-video editing/extension. It renders 2-30s at fixed 30 fps with optional native audio, supports 480p/720p/1080p and 16:9/4:3/1:1/3:4/9:16, accepts up to 10 reference images, 5 reference videos, and 5 reference audios, and uses plain per-type prompt labels Image 1, Video 1, and Audio 1. Do not send negativePrompt. Use animate_photo for native first/last frames, generate_video for text or loose references, sound_to_video when audio is the primary driver, and video_to_video with controlMode=\"seedance-v2v\" for edits or extensions."
|
|
182
229
|
},
|
|
183
230
|
"generateAudio": {
|
|
184
231
|
"type": "boolean",
|
|
185
|
-
"description": "Whether
|
|
232
|
+
"description": "Whether the returned video should include generated/native audio. Omit to include audio by default; set false only when the user explicitly asks for silent output or no audio. Supported by LTX, MiniMax H3, and Seedance; not supported by WAN or HappyHorse.\n\nWan 3 supports this toggle; omit it for audio-on by default or set false only for an explicitly silent result."
|
|
186
233
|
},
|
|
187
234
|
"referenceImageIndices": {
|
|
188
235
|
"type": "array",
|
|
189
236
|
"items": {
|
|
190
237
|
"type": "number"
|
|
191
238
|
},
|
|
192
|
-
"description": "
|
|
239
|
+
"description": "Seedance or MiniMax H3 R2V image references. Use negative indices for uploaded images and non-negative indices for generated image results. Seedance uses @Image tags. H3 uses <Picture 1>, <Picture 2>, and so on in selection order; these are loose references, not locked frames. H3 requires at least one visual across referenceImageIndices and referenceVideoIndices, so this array may be empty when a reference video is supplied; audio alone is invalid. Wan 3 loose images use Image 1, Image 2, and so on, with up to 10 images."
|
|
193
240
|
},
|
|
194
241
|
"referenceVideoIndices": {
|
|
195
242
|
"type": "array",
|
|
196
243
|
"items": {
|
|
197
244
|
"type": "number"
|
|
198
245
|
},
|
|
199
|
-
"description": "
|
|
246
|
+
"description": "Seedance or MiniMax H3 R2V loose video references. Use negative indices for uploaded videos and non-negative indices for generated video results. Seedance uses @Video tags; H3 uses <Video 1>, <Video 2>, and so on in selection order. A reference video can be the only visual input for H3 Ref2VA. Do not use this for source-video transforms; use video_to_video instead. Wan 3 loose videos use Video 1, Video 2, and so on, with up to 5 videos."
|
|
200
247
|
},
|
|
201
248
|
"referenceAudioIndices": {
|
|
202
249
|
"type": "array",
|
|
203
250
|
"items": {
|
|
204
251
|
"type": "number"
|
|
205
252
|
},
|
|
206
|
-
"description": "
|
|
253
|
+
"description": "Seedance or MiniMax H3 R2V loose audio references. Use negative indices for uploaded audio files and non-negative indices for generated audio results. Seedance uses @Audio tags; H3 uses <Audio 1>, <Audio 2>, and so on in selection order. H3 audio may accompany an image or video but cannot be the sole reference input. Wan 3 loose audios use Audio 1, Audio 2, and so on, with up to 5 audios."
|
|
207
254
|
},
|
|
208
255
|
"width": {
|
|
209
256
|
"type": "number",
|
|
210
|
-
"description": "Video width in pixels. LTX 2.
|
|
257
|
+
"description": "Video width in pixels. LTX 2.3: 640-3840. WAN: 480-1536. Default resolution depends on model and quality tier: LTX Fast about 720p and High/Pro about 1080p; WAN Fast uses 480p short side and High/Pro uses 720p short side. Set width only when the user specifies an exact width or orientation-qualified exact pixels. A bare named resolution like \"720p resolution\" is a short-side target, not an instruction to make landscape 1280x720. If the user gives only one exact dimension, set only that dimension and preserve/infer the sensible aspect ratio. User-requested exact dimensions override the default media quality. Mappings when orientation is explicit: 480p landscape=854x480, 480p portrait=480x854, 720p landscape=1280x720, 720p portrait=720x1280, 1080p landscape=1920x1080, 1080p portrait=1080x1920, 4K landscape=3840x2160. Non-step values are accepted when in bounds; LTX snaps to the nearest 64px step and WAN snaps to the nearest 16px step internally, so do not ask the user to adjust by a few pixels."
|
|
211
258
|
},
|
|
212
259
|
"height": {
|
|
213
260
|
"type": "number",
|
|
214
|
-
"description": "Video height in pixels. LTX 2.
|
|
261
|
+
"description": "Video height in pixels. LTX 2.3: 640-3840. WAN: 480-1536. Set height only when the user specifies an exact height or orientation-qualified exact pixels. A bare named resolution like \"720p resolution\" is a short-side target; do not convert it to landscape dimensions unless the user says landscape/horizontal/widescreen. If the user gives only one exact dimension, set only that dimension and preserve/infer the sensible aspect ratio. User-requested exact dimensions override Default Media Quality, including Pro. Non-step values are accepted when in bounds; LTX snaps to the nearest 64px step and WAN snaps to the nearest 16px step internally, so do not ask the user to adjust by a few pixels."
|
|
215
262
|
},
|
|
216
263
|
"targetResolution": {
|
|
217
264
|
"type": "number",
|
|
218
|
-
"description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini and Seedance 2.5 support 480p/720p only, so never set 1080p or 4K for \"seedance2-5\". Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead."
|
|
265
|
+
"description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini and Seedance 2.5 support 480p/720p only, so never set 1080p or 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. HappyHorse supports only 720p and 1080p. Never set 4K for Wan 3 or HappyHorse. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for H3 and never 1080p or 4K. Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, Wan 3, HappyHorse, or MiniMax H3, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead."
|
|
219
266
|
},
|
|
220
267
|
"numberOfVariations": {
|
|
221
268
|
"type": "number",
|
|
@@ -230,6 +277,25 @@
|
|
|
230
277
|
"voicePersonaName": {
|
|
231
278
|
"type": "string",
|
|
232
279
|
"description": "ONLY when the user explicitly requests a registered/reference persona voice clip. Name of the persona whose voice clip to use as referenceAudioIdentity. Set this when the narrator/speaker is a different persona than the one described in the video (e.g. \"David\" narrates a scene featuring Aleyna), or to explicitly select a requested voice when multiple personas with voice clips are resolved. Do NOT set this for ordinary character dialogue, inferred voices, or personas without a voice clip — LTX 2.3 generates voice natively from the text prompt instead. Requires ltx23 because LTX 2.5 has no compatible ID-LoRA."
|
|
280
|
+
},
|
|
281
|
+
"loras": {
|
|
282
|
+
"type": "array",
|
|
283
|
+
"minItems": 1,
|
|
284
|
+
"maxItems": 8,
|
|
285
|
+
"items": {
|
|
286
|
+
"type": "string",
|
|
287
|
+
"minLength": 1
|
|
288
|
+
},
|
|
289
|
+
"description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-t2v\", \"minimax-h3-t2v-turbo\", \"minimax-h3-r2v\", \"minimax-h3-r2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nOne LoRA is published for MiniMax H3 today: h3-realism-people (fal), a realism pass trained on live-action footage of people. It restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It needs its trigger word: put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. Exact ranges and any LoRA published since: GET /v1/loras/comfy?modelId=<model>. Do not invent ids."
|
|
290
|
+
},
|
|
291
|
+
"loraStrengths": {
|
|
292
|
+
"type": "array",
|
|
293
|
+
"minItems": 1,
|
|
294
|
+
"maxItems": 8,
|
|
295
|
+
"items": {
|
|
296
|
+
"type": "number"
|
|
297
|
+
},
|
|
298
|
+
"description": "Strength for each LoRA in loras, in the same order. Omitting the array applies 1.0 to every LoRA, which is NOT the catalog default and for h3-realism-people is already at the top of its band, so send explicit values. Video LoRAs are positive-only — unlike the bipolar Krea 2 image sliders, a negative value is not an inverse effect and 0 is off. h3-realism-people takes 0-2 and its catalog default is 0.8; 0.6-1 is the usable band. It also pulls the camera in as it climbs: at 1.5 and above the shot reliably recomposes and the grade darkens, which on an image-conditioned mode can crop the subject out of the frame the user supplied. Raise it above 1 only when the user asks for more, and prefer the default when they supplied a first or last frame."
|
|
233
299
|
}
|
|
234
300
|
},
|
|
235
301
|
"required": [
|
|
@@ -248,7 +314,7 @@
|
|
|
248
314
|
"properties": {
|
|
249
315
|
"prompt": {
|
|
250
316
|
"type": "string",
|
|
251
|
-
"description": "Genre, mood, and style description for the music. Be specific about musical characteristics.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nExamples:\n- \"upbeat electronic dance music with driving bass and synth arpeggios\"\n- \"mellow jazz ballad with soft piano, brushed drums, and walking bass\"\n- \"epic orchestral soundtrack with soaring strings and powerful brass\"\n- \"lo-fi hip hop beat with vinyl crackle, muted keys, and chill vibes\"\n- \"acoustic folk song with fingerpicked guitar and warm harmonies\"\n\nInclude:\n- Genre (rock, jazz, electronic, classical, hip-hop, etc.)\n- Mood (happy, melancholic, energetic, relaxing, epic, etc.)\n- Instruments (piano, guitar, drums, synth, strings, etc.)\n- Style descriptors (driving, mellow, atmospheric, punchy, etc.)\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary ONE dimension across separate tracks. This is one Sogni project with multiple jobs, so prefer it when all tracks share the same duration, BPM, key, lyrics, model, and generation parameters and only prompt text varies. Lock in any genre/mood/instruments the user specified, vary the rest. Example: \"{lo-fi hip hop beat with muted keys|jazz piano trio with brushed drums|ambient electronic with soft pads} with warm reverb and vinyl texture\"."
|
|
317
|
+
"description": "Genre, mood, and style description for the music. Be specific about musical characteristics.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nExamples:\n- \"upbeat electronic dance music with driving bass and synth arpeggios\"\n- \"mellow jazz ballad with soft piano, brushed drums, and walking bass\"\n- \"epic orchestral soundtrack with soaring strings and powerful brass\"\n- \"lo-fi hip hop beat with vinyl crackle, muted keys, and chill vibes\"\n- \"acoustic folk song with fingerpicked guitar and warm harmonies\"\n\nInclude:\n- Genre (rock, jazz, electronic, classical, hip-hop, etc.)\n- Mood (happy, melancholic, energetic, relaxing, epic, etc.)\n- Instruments (piano, guitar, drums, synth, strings, etc.)\n- Style descriptors (driving, mellow, atmospheric, punchy, etc.)\n\nMODEL \"music3\": MiniMax Music 3 wants a structured caption instead of a tag list — write the prompt as one paragraph in three labeled parts: \"Global Metadata: genre, BPM, key, emotional progression across the song, production profile. Vocal Details: gender, timbre, delivery, harmonies (or: none, purely instrumental). Arrangement: primary and secondary instruments, groove, bass, percussion, textures, how sections evolve.\" The more specific, the closer the result. Fold tempo and key into this caption — music3 ignores the bpm/keyscale/timesig args.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary ONE dimension across separate tracks. This is one Sogni project with multiple jobs, so prefer it when all tracks share the same duration, BPM, key, lyrics, model, and generation parameters and only prompt text varies. Lock in any genre/mood/instruments the user specified, vary the rest. Example: \"{lo-fi hip hop beat with muted keys|jazz piano trio with brushed drums|ambient electronic with soft pads} with warm reverb and vinyl texture\"."
|
|
252
318
|
},
|
|
253
319
|
"duration": {
|
|
254
320
|
"type": "number",
|
|
@@ -268,7 +334,7 @@
|
|
|
268
334
|
},
|
|
269
335
|
"lyrics": {
|
|
270
336
|
"type": "string",
|
|
271
|
-
"description": "Song lyrics. Optional — omit for instrumental music. Format: write lyrics naturally with line breaks. The model will attempt to sing these lyrics with the generated music. Works best with clear, rhythmic phrasing that matches the BPM."
|
|
337
|
+
"description": "Song lyrics. Optional — omit for instrumental music. Format: write lyrics naturally with line breaks. The model will attempt to sing these lyrics with the generated music. Works best with clear, rhythmic phrasing that matches the BPM. For model \"music3\", structure lyrics with plain section tags on their own lines ([Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Solo], [Outro]) — no modifiers inside brackets, and write enough sections to fill the requested duration since the composer ends the song when the lyric sheet runs out. For instrumental music3 tracks, pass ONLY a skeleton of those tags one per line (e.g. [Intro] [Verse] [Chorus] [Verse] [Solo] [Chorus] [Outro]) — bare instrumental pieces end early without it."
|
|
272
338
|
},
|
|
273
339
|
"model": {
|
|
274
340
|
"type": "string",
|
|
@@ -277,7 +343,7 @@
|
|
|
277
343
|
"sft",
|
|
278
344
|
"music3"
|
|
279
345
|
],
|
|
280
|
-
"description": "Music model. \"
|
|
346
|
+
"description": "Music model. \"music3\" (default): MiniMax Music 3 — premium autoregressive composer with the best vocals, lyric adherence and song structure; 30 steps, up to 5 minutes, and it treats duration as a ceiling (may end the song early at a musical resolution). BPM/key/timesig args are ignored by music3 — fold tempo and key into the prompt instead. \"turbo\": ACE-Step 1.5 Turbo — fast 4-16 step drafts at roughly 1/20 the music3 cost; use only when the user asks for a quick, cheap, or draft track, or names ACE-Step. \"sft\": ACE-Step 1.5 SFT — experimental, strong lyric handling, 10-200 steps; use only when the user names it. Default to \"music3\" whenever the user does not ask for a draft or a specific model."
|
|
281
347
|
},
|
|
282
348
|
"timesig": {
|
|
283
349
|
"type": "number",
|
|
@@ -306,13 +372,13 @@
|
|
|
306
372
|
"type": "function",
|
|
307
373
|
"function": {
|
|
308
374
|
"name": "edit_image",
|
|
309
|
-
"description": "Generate images
|
|
375
|
+
"description": "Generate or edit images using reference photos. Supports GPT Image 2 up to 16 images, Qwen up to 3, and Krea 2 Identity Edit with 1-2. This is the required tool for identity-sensitive edits of a referenced person or character: wardrobe, makeover, pose/repositioning, face/head/body swap, background or lighting changes, character-consistent style transfer, persona scene creation, and non-Pro single-character sheets. Use model=\"krea-identity-edit\" for those by default unless the user explicitly names another model. Keep the primary base/scene image first and an optional person/detail reference second. Use this instead of generate_image whenever uploaded/persona assets must guide the result. Use restore_photo/refine_result/apply_style only for identity-neutral restoration or edits; a portrait follow-up that must preserve likeness stays on edit_image. Exception: explicit Z-image/Z-image Turbo/base Krea 2 Turbo img2img uses generate_image with sourceImageIndex and starting_image_strength.",
|
|
310
376
|
"parameters": {
|
|
311
377
|
"type": "object",
|
|
312
378
|
"properties": {
|
|
313
379
|
"prompt": {
|
|
314
380
|
"type": "string",
|
|
315
|
-
"description": "Edit instruction describing what to generate using the reference images as guidance. 50-200 words recommended.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT CONSTRUCTION ORDER — build the prompt in this sequence:\n1. IDENTITY LOCK — state which picture owns the person's identity (GOLDEN RULE: never leave identity ambiguous when editing a person)\n2. REQUESTED EDIT — describe only what CHANGES (the delta), not the whole image\n3. REFERENCE ROLE MAPPING — assign each picture ONE primary role: base_identity (face/person), pose_reference, outfit_reference, style_reference, background_reference, or color_reference\n4. POSE / COMPOSITION — pose, framing, camera angle (omit if unchanged)\n5. STYLE — artistic style, genre, era (omit if unchanged)\n6. LIGHTING / REALISM — \"maintain realistic anatomy, perspective, and lighting integration\"\n7. PRESERVE clause — always end with \"preserve all unmentioned details\"\n\nIDENTITY LOCK (required when a person is in any reference image):\n\"Preserve the exact facial likeness from picture N — face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline, apparent age, and overall recognizability.\"\nNever let a style, pose, or clothing reference silently override the face. If multiple images are provided, explicitly state \"identity comes only from picture N — do not borrow identity from other pictures.\"\n\nMINIMAL-CHANGE PRINCIPLE: The base image already contains the subject, composition, camera angle, expression, lighting, and background. Describe only the delta. Use positive constraints (\"preserve exact facial likeness\") not negative ones (\"don't change the face\").\n\nSINGLE-IMAGE PATTERN:\n\"Preserve the exact facial likeness and recognizability of the person from picture 1. [Describe only the requested change]. Keep the same pose, framing, camera angle, and expression unless the user specifically requests changes to these. Preserve all unmentioned details.\"\n\nMULTI-IMAGE PATTERN:\n\"Use the person from picture 1 as the final subject and preserve their exact facial likeness. [Requested edit]. Identity comes only from picture 1. Pose from picture 2. Outfit from picture 3. Do not borrow identity from pictures 2 or 3. Maintain realistic anatomy, perspective, and lighting integration. Preserve all unmentioned details.\"\n\nCREATIVE TRANSFORMATIONS — be vivid and reference-specific, name the artist, franchise, or era, but always anchor identity first:\n - \"Preserve the exact facial likeness from picture 1. Transform them into a Renaissance oil painting in the style of Vermeer — rich warm tones, dramatic chiaroscuro lighting, ornate period clothing. Maintain realistic anatomy. Preserve all unmentioned details.\"\n - \"Preserve the exact facial likeness from picture 1. Reimagine them as a Marvel superhero — cinematic dramatic lighting, heroic pose, detailed costume with cape, glowing energy effects. Preserve all unmentioned details.\"\n - \"Preserve the exact facial likeness from picture 1. Transform them into a Studio Ghibli anime character — soft watercolor backgrounds, gentle Ghibli-style rendering, whimsical atmosphere. Preserve all unmentioned details.\"\n - \"Preserve the exact facial likeness from picture 1. Place them into a Star Wars scene — Jedi robes, lightsaber glow, dramatic sci-fi backdrop. Preserve all unmentioned details.\"\n - \"Preserve the exact facial likeness from picture 1. Turn them into a GTA loading screen character — bold outlines, saturated colors, attitude-filled pose, urban backdrop. Preserve all unmentioned details.\"\n\nFAILURE MODES TO AVOID:\n- Face drift: identity source not specified, or style/pose reference overrides the face\n- Over-editing: for simple edits, prompt rewrites the entire image instead of describing the delta (creative transformations may intentionally change more)\n- Reference confusion: multiple images provided without explicit role mapping\n\nCHARACTER / MASCOT SHEETS: When the user asks for a character sheet, mascot sheet, model sheet, turnaround, expression sheet, or reusable character reference board using uploaded references, create ONE comprehensive professional reference-board image, not separate variations. Map reference roles clearly first (for example: picture 1 = character identity/style reference, picture 2 = logo/brand asset) and keep the character identity consistent across every panel. Include a large hero pose, front / 3/4 / side / back turnaround views, an expression row, action/personality poses, accessories or props, color palette swatches, and compact notes such as personality, fun facts, or brand usage when appropriate. Preserve exact user-provided brand names, slogans, logo text, and requested copy verbatim; incidental tiny notes may be generated by the image model if the user did not provide exact wording. Use clean readable typography.\n\nBATCH VARIATIONS: When numberOfVariations > 1, the prompt describes one output image. Do not mention counts, \"versions\", \"different\", or \"multiple\" in the prompt text unless the user explicitly wants those words visible in the image. Do not describe multiple copies or duplicates of the subject in a single image unless the user asked for a grid, collage, or side-by-side composition. Use Dynamic Prompt syntax to vary one dimension across separate images. For personas: vary scene, activity, expression, or environment; preserve identity. Example: user asks \"4 versions at the beach\" → numberOfVariations=4, prompt=\"[persona] at the beach {building a sandcastle|surfing a wave|reading under a palm tree|flying a kite}\" — each output is one person doing one activity. For direct edits: vary the approach, e.g., numberOfVariations=3, prompt=\"make the sky {a vibrant sunset|stormy and dramatic|clear blue}\". Preserve any requested orientation, aspect ratio, or exact pixel dimensions across every variation.\n\nSELECTION-GATED IMAGE STAGES: If the user asks for multiple reference-guided image options/takes/versions and says they will pick one before a later dance, animation, or video, this edit_image call is still the first step. Generate the complete image batch now with sourceImageIndex set to the relevant reference, the exact requested count, Dynamic Prompt options for each output, and the final video/image aspect ratio. Do not ask the user to choose before the images exist, and do not call video tools until after the user selects an image.\n\nLINKED VARIANTS: If multiple details must stay paired per output — visual style, identity cues, outfit, label text, symbols, setting, character, prop, location, or before/after keyframe details — use ONE top-level Dynamic Prompt branch with one complete prompt per output. Do NOT use separate Dynamic Prompt groups for details that must stay together; unpaired groups can mix attributes. If the user asks for per-variant facial, identity, or appearance changes, repeat that guidance inside EVERY option while also preserving recognizability. When the user names a subject or character, write that name or stable role inside every Dynamic Prompt option; a shared prefix outside the branch is not enough because each option must stand alone as a complete identity contract.\n\nEach option must be a fully concrete description — name the actual garment or styling, the actual setting, the actual accessories, and the literal text or symbol shown on screen when requested. Never use meta-placeholder phrasing such as \"style-specific outfit\", \"variant-specific background\", \"include the requested symbol\", \"include a humorous alternate name\", or \"bake the name and symbol into the image\" — those describe the task instead of the image.\n\nORIGINAL + VARIANT BATCHES: When one option is a remade/preserved original and the other options are themed variants, the original option still needs a concrete visual contract. Say to preserve the original clothing/wardrobe/outfit and original background/setting, then name any requested added text, label, flag, logo, symbol, or prop for that original option. Do not leave the original option as only \"unmodified original person\"; it must be as fully specified as every themed option.\n\nNEW SETTING PER OPTION: When the variant theme implies a new place, culture, era, or context, every option must name its own setting (location, props, lighting). Do NOT carry the source background forward, do NOT write \"in the same pose and placement as the original photo\" without also naming the new background, and do NOT rely on \"preserve all unmentioned details\" to handle the setting — the new setting IS a mentioned detail.\n\nRECOGNIZABILITY OVER FEATURE LOCK: For ethnic / age / character / art-style transformations, do NOT paste the strict IDENTITY LOCK feature list (\"face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline\") inside each option — that list contradicts the requested face change and the source face will pass through unchanged. Anchor recognizability per option through apparent age, signature hair silhouette, build, posture, and expression, and explicitly allow skin tone, facial features, and proportions to shift toward the target.\n\nCorrect shape (each option self-contained, concrete, with a fresh setting and a recognizability anchor instead of a strict feature lock):\n\"{The subject wearing [specific garment, color, cut, and material], standing in [specific NEW setting with props and lighting — never the source background], bold text at the bottom reads [literal requested text], [specific requested visual symbol] appears as a sign or prop, [requested per-variant facial or appearance shift, e.g. \"skin tone, eye shape, and bone structure shift toward <target> features\"], recognizable through apparent age, signature hair silhouette, build, posture, and expression|The subject wearing [second specific garment, color, cut, and material], standing in [second specific NEW setting with props and lighting], bold text at the bottom reads [second literal requested text], [second requested visual symbol] appears as a sign or prop, [second requested facial or appearance shift], recognizable through apparent age, signature hair silhouette, build, posture, and expression|...}\"\n\nWrong shape (placeholder labels masquerading as prompts):\n\"{First variant with variant-specific facial features, placeholder wardrobe, alternate name, and requested symbol baked in|Second variant with different variant-specific facial features, placeholder wardrobe, alternate name, and requested symbol baked in|...}\"\n\nAlso wrong (strict feature lock + no new setting — the source face and source background pass through unchanged):\n\"{Preserve the exact facial likeness — face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline. Reimagine as <variant>: [garment description], standing in the exact same pose and placement as the original photo. Preserve all unmentioned details.|Preserve the exact facial likeness — [same strict lock]. Reimagine as <other variant>: [other garment], standing in the exact same pose and placement as the original photo. Preserve all unmentioned details.|...}\"\n\nSCREENPLAY / STORYBOARD BATCHES: For multi-scene story, commercial, or longer-form video keyframes, use one Dynamic Prompt branch with one full scene prompt per option. Recurring characters must keep stable names and repeated visual anchors in every scene option where they appear: face/identity source if available, age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture, and role. Do not let style, scene changes, or pose references alter identity. Include screenplay-style speaker tags when dialogue matters, e.g. CHARACTER: \"We made it.\"\n\nCOMPOSITE GPT IMAGE 2 STORYBOARD SHEETS: When numberOfVariations=1 and the user asks for one composite video storyboard/keyframe sheet using uploaded or generated references, the prompt must be a compiled storyboard prompt, not a concept summary. Include a SCENES: section with exactly the requested number of concrete entries named SCENE_01, SCENE_02, etc. Every scene entry must include Visual/Action, Camera/Motion, Dialogue/VO (or [no dialogue]), Audio/SFX, and any visible text or reference usage for that scene. Do not provide only the source brief or generic layout instructions; malformed compiled storyboard prompts are blocked by quality audit."
|
|
381
|
+
"description": "Edit instruction describing what to generate using the reference images as guidance. 50-200 words recommended.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT CONSTRUCTION ORDER — build the prompt in this sequence:\n1. IDENTITY LOCK — state which picture owns the person's identity (GOLDEN RULE: never leave identity ambiguous when editing a person)\n2. REQUESTED EDIT — describe only what CHANGES (the delta), not the whole image\n3. REFERENCE ROLE MAPPING — assign each picture ONE primary role: base_identity (face/person), pose_reference, outfit_reference, style_reference, background_reference, or color_reference\n4. POSE / COMPOSITION — pose, framing, camera angle (omit if unchanged)\n5. STYLE — artistic style, genre, era (omit if unchanged)\n6. LIGHTING / REALISM — \"maintain realistic anatomy, perspective, and lighting integration\"\n7. PRESERVE clause — always end with \"preserve all unmentioned details\"\n\nIDENTITY LOCK (required when a person is in any reference image):\n\"Preserve the exact facial likeness from picture N — face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline, apparent age, and overall recognizability.\"\nNever let a style, pose, or clothing reference silently override the face. If multiple images are provided, explicitly state \"identity comes only from picture N — do not borrow identity from other pictures.\"\n\nMINIMAL-CHANGE PRINCIPLE: The base image already contains the subject, composition, camera angle, expression, lighting, and background. Describe only the delta. Use positive constraints (\"preserve exact facial likeness\") not negative ones (\"don't change the face\").\n\nSINGLE-IMAGE PATTERN:\n\"Preserve the exact facial likeness and recognizability of the person from picture 1. [Describe only the requested change]. Keep the same pose, framing, camera angle, and expression unless the user specifically requests changes to these. Preserve all unmentioned details.\"\n\nMULTI-IMAGE PATTERN:\n\"Use the person from picture 1 as the final subject and preserve their exact facial likeness. [Requested edit]. Identity comes only from picture 1. Pose from picture 2. Outfit from picture 3. Do not borrow identity from pictures 2 or 3. Maintain realistic anatomy, perspective, and lighting integration. Preserve all unmentioned details.\"\n\nKREA IDENTITY EDIT: Use model=\"krea-identity-edit\" whenever an edit of a referenced person or character must keep likeness or character identity while changing clothes, hair or makeup, pose or position, face/head/body, background, lighting, or visual style. Infer that semantic intent in any language; never route from keyword or regex matching. Also use it for a single-character sheet in non-Pro mode. This semantic default applies even when the user did not name Krea; an explicitly user-requested model always wins. Use model=\"dark-beast-krea2-identity-edit\" only when the user explicitly requests Dark Beast Krea 2 Identity Edit, its community/uncensored variant, or the dark_beast_krea2_identity_edit_v1_2 model id. These models require one reference image, accept up to two context images, work best at 512-2048 px, let the model tier and worker choose their current execution defaults, and do not use negative prompts. Put the primary scene/base image first and the person/detail reference second for scene-plus-person edits; reference them with context_image_0 and context_image_1 when model_ref tokens are needed.\n\nKREA 2 IDENTITY EDIT PROMPTING: Krea performs best with a concise, direct delta instruction rather than a generic 50-200 word expansion. For one reference, state the requested change in 1-4 concrete sentences and name only the details that must remain fixed; avoid restating the entire image or dumping a long facial-feature inventory. For two references, explicitly assign roles in a compact instruction: base scene/image first, person/detail/outfit/pose/style reference second. Use sourceImageIndex to make the base scene the first context image when uploads arrive in another order (-1 = first upload, -2 = second). End with a short preservation clause only when useful. Longer structured prompts remain appropriate for character sheets, grids, editorial layouts, or exact visible text.\n\nCREATIVE TRANSFORMATIONS — be vivid and reference-specific, name the artist, franchise, or era, but always anchor identity first:\n - \"Preserve the exact facial likeness from picture 1. Transform them into a Renaissance oil painting in the style of Vermeer — rich warm tones, dramatic chiaroscuro lighting, ornate period clothing. Maintain realistic anatomy. Preserve all unmentioned details.\"\n - \"Preserve the exact facial likeness from picture 1. Reimagine them as a Marvel superhero — cinematic dramatic lighting, heroic pose, detailed costume with cape, glowing energy effects. Preserve all unmentioned details.\"\n - \"Preserve the exact facial likeness from picture 1. Transform them into a Studio Ghibli anime character — soft watercolor backgrounds, gentle Ghibli-style rendering, whimsical atmosphere. Preserve all unmentioned details.\"\n - \"Preserve the exact facial likeness from picture 1. Place them into a Star Wars scene — Jedi robes, lightsaber glow, dramatic sci-fi backdrop. Preserve all unmentioned details.\"\n - \"Preserve the exact facial likeness from picture 1. Turn them into a GTA loading screen character — bold outlines, saturated colors, attitude-filled pose, urban backdrop. Preserve all unmentioned details.\"\n\nFAILURE MODES TO AVOID:\n- Face drift: identity source not specified, or style/pose reference overrides the face\n- Over-editing: for simple edits, prompt rewrites the entire image instead of describing the delta (creative transformations may intentionally change more)\n- Reference confusion: multiple images provided without explicit role mapping\n\nCHARACTER / MASCOT SHEETS: When the user asks for a character sheet, mascot sheet, model sheet, turnaround, expression sheet, or reusable character reference board using uploaded references, create ONE comprehensive professional reference-board image, not separate variations. Map reference roles clearly first (for example: picture 1 = character identity/style reference, picture 2 = logo/brand asset) and keep the character identity consistent across every panel. Include a large hero pose, front / 3/4 / side / back turnaround views, an expression row, action/personality poses, accessories or props, color palette swatches, and compact notes such as personality, fun facts, or brand usage when appropriate. Preserve exact user-provided brand names, slogans, logo text, and requested copy verbatim; incidental tiny notes may be generated by the image model if the user did not provide exact wording. Use clean readable typography.\n\nBATCH VARIATIONS: When numberOfVariations > 1, the prompt describes one output image. Do not mention counts, \"versions\", \"different\", or \"multiple\" in the prompt text unless the user explicitly wants those words visible in the image. Do not describe multiple copies or duplicates of the subject in a single image unless the user asked for a grid, collage, or side-by-side composition. Use Dynamic Prompt syntax to vary one dimension across separate images. For personas: vary scene, activity, expression, or environment; preserve identity. Example: user asks \"4 versions at the beach\" → numberOfVariations=4, prompt=\"[persona] at the beach {building a sandcastle|surfing a wave|reading under a palm tree|flying a kite}\" — each output is one person doing one activity. For direct edits: vary the approach, e.g., numberOfVariations=3, prompt=\"make the sky {a vibrant sunset|stormy and dramatic|clear blue}\". Preserve any requested orientation, aspect ratio, or exact pixel dimensions across every variation.\n\nSELECTION-GATED IMAGE STAGES: If the user asks for multiple reference-guided image options/takes/versions and says they will pick one before a later dance, animation, or video, this edit_image call is still the first step. Generate the complete image batch now with sourceImageIndex set to the relevant reference, the exact requested count, Dynamic Prompt options for each output, and the final video/image aspect ratio. Do not ask the user to choose before the images exist, and do not call video tools until after the user selects an image.\n\nLINKED VARIANTS: If multiple details must stay paired per output — visual style, identity cues, outfit, label text, symbols, setting, character, prop, location, or before/after keyframe details — use ONE top-level Dynamic Prompt branch with one complete prompt per output. Do NOT use separate Dynamic Prompt groups for details that must stay together; unpaired groups can mix attributes. If the user asks for per-variant facial, identity, or appearance changes, repeat that guidance inside EVERY option while also preserving recognizability. When the user names a subject or character, write that name or stable role inside every Dynamic Prompt option; a shared prefix outside the branch is not enough because each option must stand alone as a complete identity contract.\n\nEach option must be a fully concrete description — name the actual garment or styling, the actual setting, the actual accessories, and the literal text or symbol shown on screen when requested. Never use meta-placeholder phrasing such as \"style-specific outfit\", \"variant-specific background\", \"include the requested symbol\", \"include a humorous alternate name\", or \"bake the name and symbol into the image\" — those describe the task instead of the image.\n\nORIGINAL + VARIANT BATCHES: When one option is a remade/preserved original and the other options are themed variants, the original option still needs a concrete visual contract. Say to preserve the original clothing/wardrobe/outfit and original background/setting, then name any requested added text, label, flag, logo, symbol, or prop for that original option. Do not leave the original option as only \"unmodified original person\"; it must be as fully specified as every themed option.\n\nNEW SETTING PER OPTION: When the variant theme implies a new place, culture, era, or context, every option must name its own setting (location, props, lighting). Do NOT carry the source background forward, do NOT write \"in the same pose and placement as the original photo\" without also naming the new background, and do NOT rely on \"preserve all unmentioned details\" to handle the setting — the new setting IS a mentioned detail.\n\nRECOGNIZABILITY OVER FEATURE LOCK: For ethnic / age / character / art-style transformations, do NOT paste the strict IDENTITY LOCK feature list (\"face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline\") inside each option — that list contradicts the requested face change and the source face will pass through unchanged. Anchor recognizability per option through apparent age, signature hair silhouette, build, posture, and expression, and explicitly allow skin tone, facial features, and proportions to shift toward the target.\n\nCorrect shape (each option self-contained, concrete, with a fresh setting and a recognizability anchor instead of a strict feature lock):\n\"{The subject wearing [specific garment, color, cut, and material], standing in [specific NEW setting with props and lighting — never the source background], bold text at the bottom reads [literal requested text], [specific requested visual symbol] appears as a sign or prop, [requested per-variant facial or appearance shift, e.g. \"skin tone, eye shape, and bone structure shift toward <target> features\"], recognizable through apparent age, signature hair silhouette, build, posture, and expression|The subject wearing [second specific garment, color, cut, and material], standing in [second specific NEW setting with props and lighting], bold text at the bottom reads [second literal requested text], [second requested visual symbol] appears as a sign or prop, [second requested facial or appearance shift], recognizable through apparent age, signature hair silhouette, build, posture, and expression|...}\"\n\nWrong shape (placeholder labels masquerading as prompts):\n\"{First variant with variant-specific facial features, placeholder wardrobe, alternate name, and requested symbol baked in|Second variant with different variant-specific facial features, placeholder wardrobe, alternate name, and requested symbol baked in|...}\"\n\nAlso wrong (strict feature lock + no new setting — the source face and source background pass through unchanged):\n\"{Preserve the exact facial likeness — face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline. Reimagine as <variant>: [garment description], standing in the exact same pose and placement as the original photo. Preserve all unmentioned details.|Preserve the exact facial likeness — [same strict lock]. Reimagine as <other variant>: [other garment], standing in the exact same pose and placement as the original photo. Preserve all unmentioned details.|...}\"\n\nSCREENPLAY / STORYBOARD BATCHES: For multi-scene story, commercial, or longer-form video keyframes, use one Dynamic Prompt branch with one full scene prompt per option. Recurring characters must keep stable names and repeated visual anchors in every scene option where they appear: face/identity source if available, age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture, and role. Do not let style, scene changes, or pose references alter identity. Include screenplay-style speaker tags when dialogue matters, e.g. CHARACTER: \"We made it.\"\n\nCOMPOSITE GPT IMAGE 2 STORYBOARD SHEETS: When numberOfVariations=1 and the user asks for one composite video storyboard/keyframe sheet using uploaded or generated references, the prompt must be a compiled storyboard prompt, not a concept summary. Include a SCENES: section with exactly the requested number of concrete entries named SCENE_01, SCENE_02, etc. Every scene entry must include Visual/Action, Camera/Motion, Dialogue/VO (or [no dialogue]), Audio/SFX, and any visible text or reference usage for that scene. Do not provide only the source brief or generic layout instructions; malformed compiled storyboard prompts are blocked by quality audit.\n\nKREA 2 IDENTITY EDIT PROMPTING: Krea performs best with a concise, direct delta instruction rather than a generic 50-200 word expansion. For one reference, state the requested change in 1-4 concrete sentences and name only the details that must remain fixed; avoid restating the entire image or dumping a long facial-feature inventory. Good shapes include \"Put a retro red trench coat on her\", \"Move her backward so her feet are visible\", \"Restage the portrait in psychedelic rainbow light\", and \"Turn her head left while preserving her likeness.\" For two references, explicitly assign roles in a compact instruction: base scene/image first, person/detail/outfit/pose/style reference second. Use sourceImageIndex to make the base scene the first context image when uploads arrive in another order (-1 = first upload, -2 = second). End with a short preservation clause only when useful. Longer structured prompts remain appropriate for character sheets, grids, editorial layouts, or exact visible text."
|
|
316
382
|
},
|
|
317
383
|
"model": {
|
|
318
384
|
"type": "string",
|
|
@@ -323,12 +389,31 @@
|
|
|
323
389
|
"krea-identity-edit",
|
|
324
390
|
"dark-beast-krea2-identity-edit"
|
|
325
391
|
],
|
|
326
|
-
"description": "The app auto-selects Fast→Qwen Lightning and HQ/Pro→full Qwen
|
|
392
|
+
"description": "The app auto-selects Fast→Qwen Lightning and HQ/Pro→full Qwen for ordinary identity-neutral edits. REQUIRED IDENTITY DEFAULT: set \"krea-identity-edit\" (Krea 2 Identity Edit LoRA v1.2) whenever an edit of a referenced person or character must keep likeness or character identity while changing clothes, hair or makeup, pose or position, face/head/body, background, lighting, or visual style. Also use it for a single-character sheet in non-Pro mode. It renders character sheets and storyboard/panel sheets from a reference image too: when the user names Krea or Krea 2 for a sheet, character sheet, or storyboard and any reference image is in context, set \"krea-identity-edit\" and keep it — the GPT Image 2 layout default does not override a model the user asked for. With no reference image in context a Krea request is a generate_image render on \"krea-2-turbo\" (or \"dark-beast-krea2\"), not an edit_image call. This semantic default applies even when the user did not name Krea; an explicitly user-requested model always wins. Set \"dark-beast-krea2-identity-edit\" only when the user explicitly asks for Dark Beast Krea 2 Identity Edit, its uncensored/community variant, or dark_beast_krea2_identity_edit_v1_2. Set \"gpt-image-2\" when the user explicitly names GPT/OpenAI/ChatGPT Image, or when precise typography, dense labels, or a professional multi-panel layout is the primary requirement; Pro character sheets may retain GPT Image 2. If GPT Image 2 is unavailable for detail-critical layout work, fall back to full \"qwen\", never \"qwen-lightning\". Krea identity edit models require at least one reference image, accept up to two context images, and work best at 512-2048px. Let the model tier and worker choose their current steps, guidance, sampler, scheduler, grounding, and reference-boost defaults; do not send a negative prompt. Put the base scene/image first and a person/detail reference second. Z-image, Z-image Turbo, and base Krea 2 Turbo are generate_image img2img models, not edit_image selectors. If the user names another edit/image model, honor it. GPT Image 2 always processes input images at high fidelity; do not set input_fidelity."
|
|
327
393
|
},
|
|
328
394
|
"sourceImageIndex": {
|
|
329
395
|
"type": "number",
|
|
330
396
|
"description": "Index of the primary image to use as the main reference. For follow-up edits when generated image results already exist, use the 0-based generated image result index; for example, editing the latest generated storyboard/image should use that generated result index so the model modifies the existing image instead of redrawing from uploads. When no generated image results exist, use sourceImageIndex=-1 to use the uploaded image references. The primary image and any additional uploaded images are passed as context images to guide generation."
|
|
331
397
|
},
|
|
398
|
+
"loras": {
|
|
399
|
+
"type": "array",
|
|
400
|
+
"minItems": 1,
|
|
401
|
+
"maxItems": 8,
|
|
402
|
+
"items": {
|
|
403
|
+
"type": "string",
|
|
404
|
+
"minLength": 1
|
|
405
|
+
},
|
|
406
|
+
"description": "Ordered LoRA IDs for the edit. ONLY valid with model=\"krea-identity-edit\" or model=\"dark-beast-krea2-identity-edit\" — Qwen and GPT Image 2 accept no LoRAs and the IDs are dropped. Use when the user asks to shift a trait the identity edit itself does not change, such as age, build, skin, lighting or grain, while the identity LoRA holds the likeness. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nBipolar sliders — each id names its POSITIVE direction, a negative strength applies the opposite, and 0 disables it: krea2-detail-enhancer, krea2-scene-complexity, krea2-realism (+ = photoreal), krea2-amateur, krea2-candid, krea2-zoom (+ = zoomed in), krea2-skin-detail, krea2-wetness, krea2-age, krea2-height, krea2-weight, krea2-hourglass-figure, krea2-breast, krea2-chest-firmness, krea2-nipple-projection, krea2-warm-light, krea2-afterlight (+ = golden), krea2-skin-tone (+ = darker), krea2-purple-grainy (+ = grainy and muted). Positive-only fine-tunes: krea2-realism-engine (photographic realism), krea2-bloomgirls (polished influencer look), krea2-mystic-x (uncensored adult), krea2-aberrant (industrial body horror), krea2-filter-bypass-2 and krea2-filter-bypass-3 (restore expressions, anatomy and poses the base model flattens; try the 2-vector first). Exact per-LoRA ranges, maturity flags and the full contract: GET /v1/loras/comfy?modelId=<model>. Do not invent ids."
|
|
407
|
+
},
|
|
408
|
+
"loraStrengths": {
|
|
409
|
+
"type": "array",
|
|
410
|
+
"minItems": 1,
|
|
411
|
+
"maxItems": 8,
|
|
412
|
+
"items": {
|
|
413
|
+
"type": "number"
|
|
414
|
+
},
|
|
415
|
+
"description": "Strength for each LoRA in loras, in the same order. Omitting the array uses 1.0 for every LoRA, which is not each LoRA's catalog default — krea2-chest-firmness, krea2-nipple-projection and krea2-height default to 0 (no effect) — so prefer explicit values. Do NOT clamp to 0-1: most Krea 2 LoRAs are bipolar, so krea2-warm-light warms the grade at 2 and cools it at -2. Usable bands vary per LoRA — roughly -2..5 for krea2-detail-enhancer, -3..3 for krea2-warm-light, 3..9 for krea2-candid, 0.5..1 for krea2-realism-engine, 1..2 for the filter-bypass pair. Scale the magnitude to how strongly the user asked; the server clamps out-of-range values, and pushing past a LoRA's recommended band costs image quality rather than adding effect. Preserve explicit user values. Example: loras=[\"krea2-detail-enhancer\",\"krea2-amateur\"], loraStrengths=[3,-2]."
|
|
416
|
+
},
|
|
332
417
|
"numberOfVariations": {
|
|
333
418
|
"type": "number",
|
|
334
419
|
"description": "Number of variations (1-16). Pass the user's exact requested count in one call when the outputs can share project settings. \"4 variations\" → numberOfVariations=4 in a single call. Use the exact requested count for reference-guided images that will feed a later video after the user picks one. Use separate calls only when the user explicitly wants independent projects, isolated approvals, or per-output settings that cannot share one project. For screenplay/storyboard batches, the prompt should contain one Dynamic Prompt branch with one full scene prompt per scene; do not set numberOfVariations=N with only scene 1's prompt. Use 1 unless the user explicitly asks for multiple. Default: 1.",
|
|
@@ -352,10 +437,9 @@
|
|
|
352
437
|
"enum": [
|
|
353
438
|
"low",
|
|
354
439
|
"medium",
|
|
355
|
-
"high"
|
|
356
|
-
"auto"
|
|
440
|
+
"high"
|
|
357
441
|
],
|
|
358
|
-
"description": "Optional GPT Image 2 rendering quality. Only set with model=\"gpt-image-2\" when the user explicitly asks for low/fast, medium/balanced, high/final
|
|
442
|
+
"description": "Optional GPT Image 2 rendering quality. Only set with model=\"gpt-image-2\" when the user explicitly asks for low/fast, medium/balanced, or high/final quality. Otherwise omit it and let the host app media quality setting map Fast to low, HQ to medium, and Pro to high."
|
|
359
443
|
},
|
|
360
444
|
"outputFormat": {
|
|
361
445
|
"type": "string",
|
|
@@ -544,13 +628,13 @@
|
|
|
544
628
|
"type": "function",
|
|
545
629
|
"function": {
|
|
546
630
|
"name": "animate_photo",
|
|
547
|
-
"description": "Animate a photo into video with motion, audio, and dialogue using LTX 2.5 by default, LTX 2.3 as rollback, or WAN 2.2. Do NOT use this tool for seedance2, seedance2-mini, or seedance2-5. Seedance media references — including Seedance 2.5 first-and-last-frame requests — must go through generate_video with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and @Image/@Video/@Audio role text in the prompt; for seamless-loop Seedance requests with one uploaded image, the prompt should anchor it as both the first frame and last frame. LTX/WAN NOTE: uploaded audio files are not loose references for ltx23/wan22; use sound_to_video when uploaded audio is the primary sync target. DANCE REQUESTS (\"make them dance\", \"do the X dance\"): use dance_montage — NOT this tool. LTX 2.5
|
|
631
|
+
"description": "Animate a photo into video with motion, audio, and dialogue using LTX 2.5 by default, LTX 2.3 as rollback, or WAN 2.2. Do NOT use this tool for seedance2, seedance2-mini, or seedance2-5. Seedance media references — including Seedance 2.5 first-and-last-frame requests — must go through generate_video with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and @Image/@Video/@Audio role text in the prompt; for seamless-loop Seedance requests with one uploaded image, the prompt should anchor it as both the first frame and last frame. LTX/WAN NOTE: uploaded audio files are not loose references for ltx25/ltx23/wan22; use sound_to_video when uploaded audio is the primary sync target. DANCE REQUESTS (\"make them dance\", \"do the X dance\"): use dance_montage — NOT this tool. LTX 2.5 and LTX 2.3 generate audio natively — describe dialogue and ambient sounds directly in the prompt (do NOT pre-generate audio for this tool). If the user provides exact speech, include it in double quotes; if they only imply speech, describe the performance and voice without inventing quoted words. Avoid placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". PERSONA VOICE: Only when the user explicitly asks to use/clone a registered persona voice clip, call resolve_personas first, then set voicePersonaName to select which persona's voice clip to use. Do not set voicePersonaName for ordinary character dialogue or inferred voices; describe those voices in the prompt for native LTX audio. For cross-persona narration (e.g. David narrates a video of Aleyna), resolve both personas and set voicePersonaName to the narrator only if that registered voice was requested. Persona voice requires ltx23 because LTX 2.5 has no compatible ID-LoRA and WAN 2.2 does not support voice identity. PERSONA PIPELINE: For persona videos, ensure an image of the persona exists before calling animate_photo. The standard pipeline is: resolve_personas → edit_image → animate_photo. If a suitable persona image already exists (user uploaded one, a prior edit_image/generate_image result, OR the user explicitly says to use the Persona image/reference photo directly), skip edit_image and animate directly. After resolve_personas, this tool can animate the injected persona image directly when that explicit direct-use instruction is given. Auto-uses the latest result image (from any prior tool) unless sourceImageIndex is set. Supports start-frame (default), end-frame, and start+end interpolation modes for LTX/WAN and MiniMax H3. For H3, use an I2V selector with frameRole=\"start\" for I2VA or frameRole=\"end\" for L2VA, and use an FLF2V selector with frameRole=\"both\" only when both endpoints are supplied. This H3 requirement overrides the later generic self-loop omission rule: even when the opening and closing frame are the same image, explicitly repeat that image in endImageIndex or endImageIndices — ask the user which frame role their image should play if they mention \"end frame\", \"last frame\", or provide two images. FIRST+LAST FRAME WORKFLOW: When the user wants a non-Seedance video using two different scenes as start and end frames, prefer generating both images in a single generate_image/edit_image call with numberOfVariations=2 and Dynamic Prompts, then call animate_photo with frameRole=\"both\", sourceImageIndex=0, endImageIndex=1. If the user explicitly wants separately created frame assets, preserve that staged instruction while keeping indices correct. In frameRole=\"both\", the handler automatically inspects both images and upgrades the base prompt into a scene-aware smooth transition prompt, so your prompt should state the desired transition style, action, dialogue, and audio rather than trying to list every visible object. If the request is vague, analyze the image first and suggest 2-3 specific animation ideas tailored to what you see. Call once you have clear creative intent. N-VIDEOS PATTERN: Avoid sequential animate_photo calls for N outputs. For a single fixed source/end frame where only prompt text varies, use sourceImageIndex + numberOfVariations=N + one Dynamic Prompt branch in prompt so Sogni submits one project with multiple jobs. If the user explicitly asks for Dynamic Prompt or Dynamic Template syntax, prefer this one-project path whenever every output uses the same source/end frames and shared settings, even if they also ask to stitch the completed clips afterward. Use sourceImageIndices/prompts for multi-segment stitched non-Seedance video, different source/end assets, different audio windows, different durations/dimensions, isolated retry lifecycle, or other per-output parameters. sourceImageIndices supports up to 16 entries; there is NO 3-clip cap, so do not split one planned batch into \"first 3\" and \"remaining\" calls. For a dialogue-heavy total-duration request with no explicit per-clip duration, prefer 15-second clips on ltx25 by default (or ltx23 rollback) (30s total = 2 clips × 15s) and 10-second clips on wan22 (60s total on wan22 = 6 clips × 10s; do NOT pick 4 clips × 15s on wan22 — the wan22 worker rejects clips longer than 10s). Multi-source flavors: (A) SHARED CONTENT — when all N clips have the same dialogue/motion but different source visuals (different scenes, outfits, environments, persona looks), first generate N distinct images via ONE edit_image/generate_image call with numberOfVariations=N + Dynamic Prompts {|}, then call animate_photo with sourceImageIndices=[start..start+N-1] and a single shared `prompt`. If all segments intentionally reuse the primary uploaded image and only prompt text varies, use sourceImageIndex=-1, frameRole=\"both\" if requested, endImageIndex=-1 if requested, numberOfVariations=N, and one Dynamic Prompt branch in prompt. Each branch option must be a complete natural-language motion prompt; do not include \"clip N\", source-frame boilerplate, \"overall request context\", or instructions to follow the user request. For a long scripted/dialogue/storyboard video from a single supplied/uploaded image where each segment needs isolated exact dialogue or per-segment wiring, use sourceImageIndices=[-1,-1,...] and per-clip prompts. Only set frameRole=\"both\" and endImageIndex=-1 when the user explicitly says the same uploaded/source/original image should be both the first and last frame of every segment. If the user requests generated source images first, honor that image stage, then animate the generated result indices. When using generated scene keyframes and each clip should begin and end on its own scene image for stitching, call animate_photo with frameRole=\"both\" and sourceImageIndices=[start..end] but OMIT endImageIndex; do not set endImageIndex=-1 unless every source is the uploaded image. (B) PER-CLIP CONTENT — when source/end asset wiring or other per-output parameters differ, pass BOTH sourceImageIndices AND `prompts` (an array of N strings, one per clip) in the same single call. Each prompt must independently anchor the visible characters, scene action, camera, audio, exact screenplay-style speaker tags, and exact quoted dialogue for that segment. If you just wrote or displayed a script/table, copy the exact dialogue lines into the corresponding per-clip prompts; do not summarize them as speech activity. If using named speaker tags with any multi-person reference image or generated scene keyframe, include one explicit cast map in each prompt that binds each name to visible position, clothing, and props/actions, e.g. SPEAKER_A = left person holding a prop; SPEAKER_B = center person with tablet; SPEAKER_C = right person near table. Do not also describe the same people again as generic man/boy/girl/woman/character subjects. For screenplay, storyboard, commercial, series, or other longer-form tasks with recurring characters, preserve the same character names and repeated visual anchors in every per-clip prompt where each character appears. Use the standard single-source path (numberOfVariations only) when the user wants motion variety from a single fixed frame. Wan 3 first-frame and first+last-frame generation is supported with videoModel=\"wan3.0-video\".",
|
|
548
632
|
"parameters": {
|
|
549
633
|
"type": "object",
|
|
550
634
|
"properties": {
|
|
551
635
|
"prompt": {
|
|
552
636
|
"type": "string",
|
|
553
|
-
"description": "I2V RULE: Do NOT re-describe what is visible in the input image. Focus on the transition from stillness — motion, expression changes, what happens next, camera movement, and sound.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. Set skipPromptProcessing=true; for Seedance also set expandPrompt=false.\n\nSTRUCTURE: \"[How the subject begins to move]. [What changes next]. [Camera behavior]. [Audio].\"\n\nMOTION PACING: Scale complexity to duration. <=6s: 1 main action beat + 1 simple camera move. Around 10s: 2-3 clear action beats + 1 camera move. >10s: up to 4 action beats in clear sequence. Prefer fewer readable beats over dense micro-actions, especially in short clips.\n\nBLOCKING: Use the image as the anchor and direct only meaningful layout changes. If the prompt introduces multiple moving subjects, state left/right placement, foreground/background, facing toward/away, and relative distance.\n\nACTION: One flowing paragraph. Describe motion beat by beat with temporal connectors (\"as\", \"then\", \"while\"). Specify who moves, what moves, how it moves, and what the camera does. One main thread — avoid too many actions at once or generic phrases like \"comes alive.\"\n\nDIALOGUE: Put user-provided spoken lines in double quotes. For screenplay-style or longer-form tasks, prefix each spoken line with a stable speaker tag outside the quotes, e.g. CHARACTER: \"We made it.\" Break long speech into short quoted phrases with acting beats between them (gestures, pauses, glances). If the user asks for speech but provides no exact words, describe the visible delivery, voice quality, and emotion without inventing quoted dialogue; ask only when exact wording is the point of the request. Never write placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". Show emotion through visible behavior, not labels. LTX 2.5
|
|
637
|
+
"description": "I2V RULE: Do NOT re-describe what is visible in the input image. Focus on the transition from stillness — motion, expression changes, what happens next, camera movement, and sound.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. Set skipPromptProcessing=true; for Seedance or Wan 3 also set expandPrompt=false.\n\nSTRUCTURE: \"[How the subject begins to move]. [What changes next]. [Camera behavior]. [Audio].\"\n\nMOTION PACING: Scale complexity to duration. <=6s: 1 main action beat + 1 simple camera move. Around 10s: 2-3 clear action beats + 1 camera move. >10s: up to 4 action beats in clear sequence. Prefer fewer readable beats over dense micro-actions, especially in short clips.\n\nBLOCKING: Use the image as the anchor and direct only meaningful layout changes. If the prompt introduces multiple moving subjects, state left/right placement, foreground/background, facing toward/away, and relative distance.\n\nACTION: One flowing paragraph. Describe motion beat by beat with temporal connectors (\"as\", \"then\", \"while\"). Specify who moves, what moves, how it moves, and what the camera does. One main thread — avoid too many actions at once or generic phrases like \"comes alive.\"\n\nDIALOGUE: Put user-provided spoken lines in double quotes. For screenplay-style or longer-form tasks, prefix each spoken line with a stable speaker tag outside the quotes, e.g. CHARACTER: \"We made it.\" Break long speech into short quoted phrases with acting beats between them (gestures, pauses, glances). If the user asks for speech but provides no exact words, describe the visible delivery, voice quality, and emotion without inventing quoted dialogue; ask only when exact wording is the point of the request. Never write placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". Show emotion through visible behavior, not labels. LTX 2.5 and LTX 2.3 generate audio natively. QUOTING RULE: ONLY use double quotes for spoken dialogue. Never quote on-screen text, overlay text, titles, captions, signs, or any visual text — describe them without quotes (e.g. bold white text reading CONGRATULATIONS overlays the lower third).\n\nAUDIO: Prompt sound intentionally — voice quality, volume, room tone, ambience, music, weather, footsteps. Include language or accent if relevant. Useful voice/volume anchors: whisper, mutter, shout, scream, energetic announcer, resonant voice with gravitas, distorted radio-style, robotic monotone, childlike curiosity.\n\nCAMERA: Cinematic terms — slow push-in, static tripod, handheld, slow arc, dolly in. Describe movement relative to subject.\n\nFor first+last-frame transitions (frameRole=\"both\"), write a concise base request for the transition style, action, dialogue, and audio. The handler will inspect both frames and expand it into a scene-aware prompt that maps visible objects and subjects between frames.\n\nFor specific characters (movies, TV): describe visual appearance — don't rely on names alone.\n\nFor complex/creative scenes (characters talking, skits), capture full creative intent — system auto-expands into detailed prompt.\n\nAVOID: Re-describing the image, vague prompts, too many actions at once, abstract emotions without visible behavior, rigid numeric constraints, readable text or logos.\n\nPOSITIVE CONSTRAINT TRANSLATION: For LTX 2.3 and WAN 2.2, the prompt field is a positive prompt. Translate user avoid/no/don't constraints into affirmative production constraints instead of copying negative phrasing. Examples: \"no people in background\" -> single subject focus with an empty background; \"no text\" -> clean blank surfaces; \"don't make it blurry\" -> crisp sharp focus; \"no weird hands\" -> natural anatomically consistent hands; \"no mouth movement, no talking, no lip syncing\" -> silent expression-only physical performance with facial motion independent of speech timing; \"don't change the room\" -> the same room and layout remain consistent; \"keep flames consistent\" -> flame and ember movement remains consistent with the source scene. Preserve exact quoted visible text or dialogue when the user explicitly requests it, and keep surrounding surfaces blank. For Dynamic Prompt batches, put these translated shared constraints before the \"{...}\" branch so every variation inherits them.\n\nWAN 2.2 (\"wan22\"): 30-150 words, subtle natural movements. Motion-only visual prompt; omit soundtrack, ambience, room tone, music, hums, sighs, spoken words, voice, and SFX cues because WAN does not generate audio.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary motion, camera, or atmosphere while preserving the user's specified elements. This is one Sogni project with multiple jobs, so prefer it when all outputs share the same source/end frames and generation parameters and only prompt text varies. Example: \"{gentle sway with drifting embers|slow paw wave with a tiny head tilt|small hop with soft fur motion}\".\n\nMINIMAX H3 FRAME-ROLE PROMPTING: frameRole=\"start\" uses I2VA and describes motion forward from the supplied opening frame. frameRole=\"end\" uses L2VA: the supplied image is the closing frame, so infer a plausible earlier state and describe action, camera, objects, and scene gradually converging on it at the end; this overrides generic \"what happens next\" wording. frameRole=\"both\" uses FLF2VA and describes a coherent transition between the supplied opening and closing frames. The model-specific prompt shaper supplies and validates the exact official alignment line."
|
|
554
638
|
},
|
|
555
639
|
"expandPrompt": {
|
|
556
640
|
"type": "boolean",
|
|
@@ -558,7 +642,7 @@
|
|
|
558
642
|
},
|
|
559
643
|
"skipPromptProcessing": {
|
|
560
644
|
"type": "boolean",
|
|
561
|
-
"description": "Bypass automatic prompt shaping/refinement, image-description anchoring, transition-prompt rewriting, and voice-identity prompt formatting so the prompt text is sent unchanged to the video model. Set true ONLY when the user explicitly says not to modify/rewrite/enhance/expand/change/improve the prompt, or to use/send it exactly, verbatim, or as-is, AND the provided prompt already satisfies the tool requirements. Continue to set non-prompt parameters such as source indices, frameRole, model, duration, count, and aspect ratio. For Seedance literal prompt requests, also set expandPrompt=false. Do not set for ordinary underspecified requests."
|
|
645
|
+
"description": "Bypass automatic prompt shaping/refinement, image-description anchoring, transition-prompt rewriting, and voice-identity prompt formatting so the prompt text is sent unchanged to the video model. Set true ONLY when the user explicitly says not to modify/rewrite/enhance/expand/change/improve the prompt, or to use/send it exactly, verbatim, or as-is, AND the provided prompt already satisfies the tool requirements. Continue to set non-prompt parameters such as source indices, frameRole, model, duration, count, and aspect ratio. For Seedance or Wan 3 literal prompt requests, also set expandPrompt=false. Do not set for ordinary underspecified requests."
|
|
562
646
|
},
|
|
563
647
|
"videoModel": {
|
|
564
648
|
"type": "string",
|
|
@@ -571,29 +655,50 @@
|
|
|
571
655
|
"minimax-h3-i2v",
|
|
572
656
|
"minimax-h3-i2v-turbo",
|
|
573
657
|
"minimax-h3-flf2v",
|
|
574
|
-
"minimax-h3-flf2v-turbo"
|
|
658
|
+
"minimax-h3-flf2v-turbo",
|
|
659
|
+
"wan3.0-video"
|
|
575
660
|
],
|
|
576
|
-
"description": "Which video model to use. \"ltx25\" (default)
|
|
577
|
-
},
|
|
578
|
-
"generateAudio": {
|
|
579
|
-
"type": "boolean",
|
|
580
|
-
"description": "Whether to include generated/native audio for audio-capable models. Omit to include audio by default; set false only when the user explicitly asks for silent output or no audio. When false, the returned video has no audio track. Ignored by audio-less WAN."
|
|
661
|
+
"description": "Which video model to use. \"ltx25\" (default) selects LTX 2.5 I2V/FLF with native audio; frameRole=\"both\" uses the FLF template with the I2V public model ID. Fast, HQ, and Pro currently use the release-validated official Distilled/Turbo workflow; Dev is withheld until upstream publishes and Sogni validates an official ComfyUI Dev recipe. \"ltx23\" remains an explicit rollback selector and is the only LTX family with the legacy transition/identity LoRAs. \"wan22\" is the fast WAN path without native audio. \"happyhorse-1.1-i2v\": HappyHorse 1.1 image-to-video from one source frame; 3-15s, 720p/1080p, native synchronized audio always on. \"happyhorse-1.1-r2v\": HappyHorse 1.1 reference-to-video with image-only loose references; use only when referenceImageIndices provide additional image references for the same clip. \"minimax-h3-i2v\" and \"minimax-h3-i2v-turbo\": MiniMax H3 from exactly one endpoint image; use frameRole=\"start\" for an I2VA opening frame or frameRole=\"end\" for an L2VA closing frame. Both roles carry the sole endpoint in sourceImageIndex/sourceImageIndices; do not invent a minimax-h3-l2v selector and do not use endImageIndex/endImageIndices for last-frame-only L2VA. \"minimax-h3-flf2v\" and \"minimax-h3-flf2v-turbo\": MiniMax H3 first-and-last-frame interpolation; these require frameRole=\"both\" plus sourceImageIndex/sourceImageIndices for the opening frame and endImageIndex/endImageIndices for the closing frame. H3 renders 5.17-15.08s at a fixed 24 fps with jointly generated stereo audio inside a 1344x768 pixel budget on a 32px grid. Do not set seedance2, seedance2-mini, or seedance2-5 here; use generate_video with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and @Image/@Video/@Audio role text for Seedance. HappyHorse accepts neither negativePrompt nor generateAudio. MiniMax H3 accepts no negativePrompt; set generateAudio=false only when the user asks for silent output, and the returned video has no audio track. MiniMax H3 Base and Turbo T2V, I2VA, L2VA, and FLF2VA prompts use exactly integrated_multimodal_description, overall_soundscape, then non_diegetic_music. I2VA prepends the official opening-frame alignment line, L2VA prepends the official duration-aware closing-frame alignment line, and FLF2VA prepends the official two-endpoint alignment line. For dialogue, use stable (S1) speaker IDs; keep identity, action, and delivery outside <d>, with only the language tag and exact spoken words inside <d>[Language] ...</d>. Use <scenetrans> at both connecting points when one line crosses a cut and explicitly state that its audio continues across the cut. Use the plain <cutoff> marker only when the video ending truncates speech; never emit tokenizer-internal <|...|> markers or plain caption/lyrics boundary tags. Do not merely delete pipe characters: caption markers become exact visible text in double quotes, lyrics markers become an ordinary <d>[Language] ...</d> singing block, and <|cutoff|> becomes plain <cutoff>. Standard uses 20 steps with res_multistep/simple. Turbo T2V, I2VA, L2VA, and FLF2VA use 4 steps with simple scheduling; er_sde is the default sampler, and direct CLI A/B overrides may select euler, er_sde, or sa_solver. Ref2VA Turbo is a separate 4-step Euler/simple workflow selected with minimax-h3-r2v-turbo; L2VA still uses the I2V selector with frameRole=\"end\". \"wan3.0-video\" is Alibaba Wan 3: one canonical premium-vendor model for text-to-video, first-frame and first+last-frame animation, loose multimodal references, audio-driven generation, and uploaded-video editing/extension. It renders 2-30s at fixed 30 fps with optional native audio, supports 480p/720p/1080p and 16:9/4:3/1:1/3:4/9:16, accepts up to 10 reference images, 5 reference videos, and 5 reference audios, and uses plain per-type prompt labels Image 1, Video 1, and Audio 1. Do not send negativePrompt. Use animate_photo for native first/last frames, generate_video for text or loose references, sound_to_video when audio is the primary driver, and video_to_video with controlMode=\"seedance-v2v\" for edits or extensions."
|
|
581
662
|
},
|
|
582
663
|
"negativePrompt": {
|
|
583
664
|
"type": "string",
|
|
584
|
-
"description": "Advanced LTX
|
|
665
|
+
"description": "Advanced LTX/WAN only. Use this field only when the user explicitly asks to set a separate negative prompt. MiniMax H3 has no negative-prompt input; put requested exclusions in prompt.\n\nWan 3 has no negativePrompt request field; do not set this for wan3.0-video."
|
|
666
|
+
},
|
|
667
|
+
"generateAudio": {
|
|
668
|
+
"type": "boolean",
|
|
669
|
+
"description": "Whether the returned video should include generated/native audio. Omit to include audio by default; set false only when the user explicitly asks for silent output or no audio. Supported by LTX and MiniMax H3; ignored by audio-less WAN.\n\nWan 3 supports this toggle; omit it for audio-on by default or set false only for an explicitly silent result."
|
|
585
670
|
},
|
|
586
671
|
"duration": {
|
|
587
672
|
"type": "number",
|
|
588
|
-
"description": "Video duration in seconds. Default: 5. Use when the user explicitly requests a specific length (e.g., \"make a 10 second video\"). Per-model maximum: ltx25 and ltx23 = 20s, wan22 = 10s (clips longer than this are invalid), minimax-h3 = 15.08s with a 5.17s minimum because H3 renders 124-362 frames on a 17-frame grid at a fixed 24 fps. For totals beyond the per-model cap, batch multiple clips via sourceImageIndices instead of requesting a single oversized clip."
|
|
673
|
+
"description": "Video duration in seconds. Default: 5. Use when the user explicitly requests a specific length (e.g., \"make a 10 second video\"). Per-model maximum: ltx25 and ltx23 = 20s, wan22 = 10s (clips longer than this are invalid), wan3.0-video = 30s with a 2s minimum, minimax-h3 = 15.08s with a 5.17s minimum because H3 renders 124-362 frames on a 17-frame grid at a fixed 24 fps. For totals beyond the per-model cap, batch multiple clips via sourceImageIndices instead of requesting a single oversized clip."
|
|
674
|
+
},
|
|
675
|
+
"smartDuration": {
|
|
676
|
+
"type": "boolean",
|
|
677
|
+
"description": "Wan 3 only. Let the model choose 2-30 seconds. Do not also set duration. The 30-second maximum is reserved and the final charge settles down to reported duration."
|
|
678
|
+
},
|
|
679
|
+
"ratio": {
|
|
680
|
+
"type": "string",
|
|
681
|
+
"enum": [
|
|
682
|
+
"adaptive",
|
|
683
|
+
"16:9",
|
|
684
|
+
"4:3",
|
|
685
|
+
"1:1",
|
|
686
|
+
"3:4",
|
|
687
|
+
"9:16"
|
|
688
|
+
],
|
|
689
|
+
"description": "Wan 3 only. Use \"adaptive\" to preserve the source frame shape."
|
|
690
|
+
},
|
|
691
|
+
"watermark": {
|
|
692
|
+
"type": "boolean",
|
|
693
|
+
"description": "Wan 3 only. Add Alibaba's visible watermark. Defaults to false."
|
|
589
694
|
},
|
|
590
695
|
"targetResolution": {
|
|
591
696
|
"type": "number",
|
|
592
|
-
"description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\",
|
|
697
|
+
"description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", or \"1080p\" without exact pixels or an output orientation. This preserves the source image aspect ratio. Wan 3 supports 480p, 720p, and 1080p; HappyHorse supports only 720p and 1080p. Never set 4K for either. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for H3 and never 1080p or 4K. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\" or \"720p landscape\", use exact-pixel aspectRatio instead."
|
|
593
698
|
},
|
|
594
699
|
"sourceImageIndex": {
|
|
595
700
|
"type": "number",
|
|
596
|
-
"description": "Which image to use as the START frame. Use 0-based non-negative indices for generated result images. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. Omit to auto-select: uses the latest result for \"start\"/\"end\" modes, or the FIRST result for \"both\" mode. IMPORTANT: When frameRole is \"both\", set this to the start frame image index and endImageIndex to the end frame image index."
|
|
701
|
+
"description": "Which image to use as the START frame. Use 0-based non-negative indices for generated result images. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. Omit to auto-select: uses the latest result for \"start\"/\"end\" modes, or the FIRST result for \"both\" mode. IMPORTANT: When frameRole is \"both\", set this to the start frame image index and endImageIndex to the end frame image index.\n\nMiniMax H3: with an I2V selector this field is the sole endpoint image. frameRole=\"start\" treats it as the I2VA opening frame; frameRole=\"end\" treats it as the L2VA closing frame. With an FLF2V selector and frameRole=\"both\", this field is the required opening frame and endImageIndex is the required closing frame. For a same-image H3 loop, explicitly repeat this index in endImageIndex instead of omitting the closing field."
|
|
597
702
|
},
|
|
598
703
|
"sourceImageIndices": {
|
|
599
704
|
"type": "array",
|
|
@@ -602,7 +707,7 @@
|
|
|
602
707
|
},
|
|
603
708
|
"minItems": 1,
|
|
604
709
|
"maxItems": 16,
|
|
605
|
-
"description": "Array of source frame indices — one video is generated per entry as its own SDK project, all running in PARALLEL. Use this when outcomes need different source images, different end frames, isolated retry lifecycle, or other per-clip asset wiring/parameters. If every outcome uses the same source/end frames and only prompt text differs, prefer sourceImageIndex with numberOfVariations=N and one Dynamic Prompt branch in `prompt` so Sogni creates one project with multiple jobs. Use 0-based non-negative result indices for generated images. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. Repeating -1 is allowed for true multi-project workflows that intentionally reuse the same uploaded image while varying per-clip assets or parameters. By default all projects share the `prompt`/`voice`/`duration`, but you can pass `prompts` (array) to give each clip its own dialogue/motion when multi-project fan-out is required. Avoid sequential animate_photo calls for N outputs. Do NOT combine with `numberOfVariations` or `sourceImageIndex`. Use frameRole=\"end\" with sourceImageIndices only when the user explicitly says the repeated uploaded/generated image is the last/end frame for each clip and no first/start frame should be supplied; in that case omit endImageIndex/endImageIndices because each sourceImageIndices entry is the end frame. You MAY combine with frameRole=\"both\" when clips need start and end frames. For adjacent transition chains across generated images, use sourceImageIndices=[start..end-1] and endImageIndices=[start+1..end] so N images produce N-1 transition clips. If the uploaded/original image starts the chain and generated results are the remaining frames, use sourceImageIndices=[-1,start..end-1] and endImageIndices=[start..end]. If the user supplies multiple uploaded images as the actual keyframe sequence, use adjacent negative uploaded indices, e.g. 5 uploaded images become sourceImageIndices=[-1,-2,-3,-4], endImageIndices=[-2,-3,-4,-5], frameRole=\"both\", prompts length 4, then stitch_video. If the user specifies transition motion, camera behavior, actions, dialogue, or audio, copy those instructions into every corresponding per-clip prompt; only invent a generic smooth transition when the user does not specify one. If the user asks for a seamless loop or final transition from the last image back to the first, close the chain by including the last image as a source and the first image as the final end frame, e.g. 5 uploaded images become sourceImageIndices=[-1,-2,-3,-4,-5], endImageIndices=[-2,-3,-4,-5,-1]. For generated scene keyframes that should each loop to themselves, omit endImageIndex/endImageIndices so each source image is also its own end frame. Set endImageIndex=-1 only when every sourceImageIndices entry is also -1 and every segment reuses the first uploaded image. Range: 1–16 indices. For generated image batches, values MUST be read from the latest edit_image/generate_image tool result's `startIndex` field. If startIndex=3 and 4 images were generated in that batch, pass `[3,4,5,6]` (NOT `[0,1,2,3]`). Do NOT assume generated indices start at 0 — they don't if there are prior results in the conversation."
|
|
710
|
+
"description": "Array of source frame indices — one video is generated per entry as its own SDK project, all running in PARALLEL. Use this when outcomes need different source images, different end frames, isolated retry lifecycle, or other per-clip asset wiring/parameters. If every outcome uses the same source/end frames and only prompt text differs, prefer sourceImageIndex with numberOfVariations=N and one Dynamic Prompt branch in `prompt` so Sogni creates one project with multiple jobs. Use 0-based non-negative result indices for generated images. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. Repeating -1 is allowed for true multi-project workflows that intentionally reuse the same uploaded image while varying per-clip assets or parameters. By default all projects share the `prompt`/`voice`/`duration`, but you can pass `prompts` (array) to give each clip its own dialogue/motion when multi-project fan-out is required. Avoid sequential animate_photo calls for N outputs. Do NOT combine with `numberOfVariations` or `sourceImageIndex`. Use frameRole=\"end\" with sourceImageIndices only when the user explicitly says the repeated uploaded/generated image is the last/end frame for each clip and no first/start frame should be supplied; in that case omit endImageIndex/endImageIndices because each sourceImageIndices entry is the end frame. You MAY combine with frameRole=\"both\" when clips need start and end frames. For adjacent transition chains across generated images, use sourceImageIndices=[start..end-1] and endImageIndices=[start+1..end] so N images produce N-1 transition clips. If the uploaded/original image starts the chain and generated results are the remaining frames, use sourceImageIndices=[-1,start..end-1] and endImageIndices=[start..end]. If the user supplies multiple uploaded images as the actual keyframe sequence, use adjacent negative uploaded indices, e.g. 5 uploaded images become sourceImageIndices=[-1,-2,-3,-4], endImageIndices=[-2,-3,-4,-5], frameRole=\"both\", prompts length 4, then stitch_video. If the user specifies transition motion, camera behavior, actions, dialogue, or audio, copy those instructions into every corresponding per-clip prompt; only invent a generic smooth transition when the user does not specify one. If the user asks for a seamless loop or final transition from the last image back to the first, close the chain by including the last image as a source and the first image as the final end frame, e.g. 5 uploaded images become sourceImageIndices=[-1,-2,-3,-4,-5], endImageIndices=[-2,-3,-4,-5,-1]. For generated scene keyframes that should each loop to themselves, omit endImageIndex/endImageIndices so each source image is also its own end frame. Set endImageIndex=-1 only when every sourceImageIndices entry is also -1 and every segment reuses the first uploaded image. Range: 1–16 indices. For generated image batches, values MUST be read from the latest edit_image/generate_image tool result's `startIndex` field. If startIndex=3 and 4 images were generated in that batch, pass `[3,4,5,6]` (NOT `[0,1,2,3]`). Do NOT assume generated indices start at 0 — they don't if there are prior results in the conversation.\n\nMiniMax H3 fan-out: with an I2V selector, every entry is an opening frame when frameRole=\"start\" or a closing frame when frameRole=\"end\". With an FLF2V selector and frameRole=\"both\", every entry is an opening frame and uses an allowed shared endImageIndex or a corresponding endImageIndices entry under the existing fan-out rules. H3 overrides the generic self-loop omission rule: same-image loops must set endImageIndices equal to sourceImageIndices, or use an allowed shared endImageIndex where the existing fan-out rules permit it."
|
|
606
711
|
},
|
|
607
712
|
"prompts": {
|
|
608
713
|
"type": "array",
|
|
@@ -630,11 +735,11 @@
|
|
|
630
735
|
"end",
|
|
631
736
|
"both"
|
|
632
737
|
],
|
|
633
|
-
"description": "How to use the source image(s)
|
|
738
|
+
"description": "How to use the source image(s). \"start\" (default): first frame. \"end\": last frame. \"both\": interpolate between first and last frames. For MiniMax H3, use minimax-h3-i2v or minimax-h3-i2v-turbo with frameRole=\"start\" for I2VA or frameRole=\"end\" for L2VA; sourceImageIndex/sourceImageIndices carries the sole endpoint for either mode. Do not invent a minimax-h3-l2v selector and do not set endImageIndex/endImageIndices for last-frame-only L2VA. MiniMax H3 first-and-last-frame interpolation uses minimax-h3-flf2v or minimax-h3-flf2v-turbo and requires frameRole=\"both\" plus both source and end image fields, even when both fields repeat the same image for a loop. This explicit H3 closing-field requirement overrides generic self-loop guidance that permits omitting an end field. For single non-H3 clips using \"both\", set sourceImageIndex and endImageIndex; fan-out can use matching sourceImageIndices/endImageIndices."
|
|
634
739
|
},
|
|
635
740
|
"endImageIndex": {
|
|
636
741
|
"type": "number",
|
|
637
|
-
"description": "Which image to use as the END frame. Use 0-based non-negative indices for generated results. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. For a single frameRole=\"both\" transition between two different images, set this to the desired end frame. For sourceImageIndices fan-out where each generated keyframe should also be its own last frame, OMIT this field. Use a shared uploaded endImageIndex only when every sourceImageIndices entry is also an uploaded image; otherwise use endImageIndices for per-clip end frames."
|
|
742
|
+
"description": "Which image to use as the END frame. Use 0-based non-negative indices for generated results. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. For a single frameRole=\"both\" transition between two different images, set this to the desired end frame. For sourceImageIndices fan-out where each generated keyframe should also be its own last frame, OMIT this field. Use a shared uploaded endImageIndex only when every sourceImageIndices entry is also an uploaded image; otherwise use endImageIndices for per-clip end frames.\n\nMiniMax H3: use this only with an FLF2V selector and frameRole=\"both\" as the closing frame, including supported shared-end fan-out. Do not omit it for a single same-image H3 loop; repeat sourceImageIndex here. For last-frame-only L2VA, use the I2V selector with frameRole=\"end\" and carry the sole closing frame in sourceImageIndex instead."
|
|
638
743
|
},
|
|
639
744
|
"endImageIndices": {
|
|
640
745
|
"type": "array",
|
|
@@ -643,11 +748,30 @@
|
|
|
643
748
|
},
|
|
644
749
|
"minItems": 1,
|
|
645
750
|
"maxItems": 16,
|
|
646
|
-
"description": "Per-clip END frame indices for sourceImageIndices fan-out. Use ONLY with frameRole=\"both\". Length MUST exactly match sourceImageIndices. Use 0-based non-negative indices for generated results and negative indices for uploaded images (-1 first upload, -2 second upload, etc.). Use this for transition chains between generated images, e.g. 5 generated images at indices [0,1,2,3,4] should become 4 transition clips with sourceImageIndices=[0,1,2,3], endImageIndices=[1,2,3,4], prompts length 4, duration as requested, then stitch_video. If the chain starts on the uploaded image and continues through generated results [0,1,2,3], use sourceImageIndices=[-1,0,1,2] and endImageIndices=[0,1,2,3]. If the user supplies 5 uploaded images as the sequence, use sourceImageIndices=[-1,-2,-3,-4] and endImageIndices=[-2,-3,-4,-5]. If the user requests a seamless loop or final transition back to the first image, append that loop closure: sourceImageIndices=[-1,-2,-3,-4,-5], endImageIndices=[-2,-3,-4,-5,-1]. Do NOT also set endImageIndex when using this."
|
|
751
|
+
"description": "Per-clip END frame indices for sourceImageIndices fan-out. Use ONLY with frameRole=\"both\". Length MUST exactly match sourceImageIndices. Use 0-based non-negative indices for generated results and negative indices for uploaded images (-1 first upload, -2 second upload, etc.). Use this for transition chains between generated images, e.g. 5 generated images at indices [0,1,2,3,4] should become 4 transition clips with sourceImageIndices=[0,1,2,3], endImageIndices=[1,2,3,4], prompts length 4, duration as requested, then stitch_video. If the chain starts on the uploaded image and continues through generated results [0,1,2,3], use sourceImageIndices=[-1,0,1,2] and endImageIndices=[0,1,2,3]. If the user supplies 5 uploaded images as the sequence, use sourceImageIndices=[-1,-2,-3,-4] and endImageIndices=[-2,-3,-4,-5]. If the user requests a seamless loop or final transition back to the first image, append that loop closure: sourceImageIndices=[-1,-2,-3,-4,-5], endImageIndices=[-2,-3,-4,-5,-1]. Do NOT also set endImageIndex when using this.\n\nMiniMax H3 fan-out: use this only with an FLF2V selector and frameRole=\"both\", paired one-to-one with sourceImageIndices. For same-image H3 loops, set this array equal to sourceImageIndices instead of omitting it. For last-frame-only L2VA fan-out, use the I2V selector with frameRole=\"end\" and carry closing frames in sourceImageIndices instead."
|
|
647
752
|
},
|
|
648
753
|
"voicePersonaName": {
|
|
649
754
|
"type": "string",
|
|
650
755
|
"description": "ONLY when the user explicitly requests a registered/reference persona voice clip. Name of the persona whose voice clip to use as referenceAudioIdentity. Set this when the narrator/speaker is a different persona than the one shown in the video (e.g. \"David\" narrates a video of Aleyna), or to explicitly select a requested voice when multiple personas with voice clips are resolved. Do NOT set this for ordinary character dialogue, inferred voices, or personas without a voice clip — LTX 2.3 generates voice natively from the text prompt instead. Requires ltx23 because LTX 2.5 has no compatible ID-LoRA."
|
|
756
|
+
},
|
|
757
|
+
"loras": {
|
|
758
|
+
"type": "array",
|
|
759
|
+
"minItems": 1,
|
|
760
|
+
"maxItems": 8,
|
|
761
|
+
"items": {
|
|
762
|
+
"type": "string",
|
|
763
|
+
"minLength": 1
|
|
764
|
+
},
|
|
765
|
+
"description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-i2v\", \"minimax-h3-i2v-turbo\", \"minimax-h3-flf2v\", \"minimax-h3-flf2v-turbo\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nOne LoRA is published for MiniMax H3 today: h3-realism-people (fal), a realism pass trained on live-action footage of people. It restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It needs its trigger word: put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. Exact ranges and any LoRA published since: GET /v1/loras/comfy?modelId=<model>. Do not invent ids."
|
|
766
|
+
},
|
|
767
|
+
"loraStrengths": {
|
|
768
|
+
"type": "array",
|
|
769
|
+
"minItems": 1,
|
|
770
|
+
"maxItems": 8,
|
|
771
|
+
"items": {
|
|
772
|
+
"type": "number"
|
|
773
|
+
},
|
|
774
|
+
"description": "Strength for each LoRA in loras, in the same order. Omitting the array applies 1.0 to every LoRA, which is NOT the catalog default and for h3-realism-people is already at the top of its band, so send explicit values. Video LoRAs are positive-only — unlike the bipolar Krea 2 image sliders, a negative value is not an inverse effect and 0 is off. h3-realism-people takes 0-2 and its catalog default is 0.8; 0.6-1 is the usable band. It also pulls the camera in as it climbs: at 1.5 and above the shot reliably recomposes and the grade darkens, which on an image-conditioned mode can crop the subject out of the frame the user supplied. Raise it above 1 only when the user asks for more, and prefer the default when they supplied a first or last frame."
|
|
651
775
|
}
|
|
652
776
|
},
|
|
653
777
|
"required": [
|
|
@@ -691,17 +815,17 @@
|
|
|
691
815
|
"type": "function",
|
|
692
816
|
"function": {
|
|
693
817
|
"name": "video_to_video",
|
|
694
|
-
"description": "Transform an existing video using WAN 2.2 Animate, LTX 2.5 V2V controls by default, LTX 2.3 as rollback, or Seedance V2V when explicitly requested. LTX 2.5
|
|
818
|
+
"description": "Transform an existing video using AI. Uses WAN 2.2 Animate (move/replace) with a reference image, LTX 2.5 V2V controls by default (canny/pose/depth/detailer plus distilled inpaint/outpaint), LTX 2.3 as rollback, or Seedance V2V when explicitly requested. LTX 2.5 Fast, HQ, and Pro use the release-validated official Distilled workflow for canny/pose/depth/detailer/inpaint/outpaint; Dev is not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Requires an uploaded video file. Use when the user wants to animate a photo with video motion, replace subjects, restyle footage, extend its canvas, regenerate a region, or enhance quality. Wan 3 source-video editing and continuation uses videoModel=\"wan3.0-video\" with controlMode=\"seedance-v2v\". Wan 3 source-video editing and continuation uses videoModel=\"wan3.0-video\" with controlMode=\"seedance-v2v\".",
|
|
695
819
|
"parameters": {
|
|
696
820
|
"type": "object",
|
|
697
821
|
"properties": {
|
|
698
822
|
"prompt": {
|
|
699
823
|
"type": "string",
|
|
700
|
-
"description": "Describe the TARGET appearance
|
|
824
|
+
"description": "Describe the TARGET appearance (not the transformation process). 2-4 present-tense sentences.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. For Seedance or Wan 3, set expandPrompt=false.\n\nFor LTX 2.5 or 2.3 canny/depth/pose modes, the source video preserves composition, depth, or motion. Spend prompt detail on style, atmosphere, lighting, surface texture, color palette, scale, and pacing. LTX 2.5 Fast, HQ, and Pro use the release-validated official Distilled workflow for canny/pose/depth/detailer/inpaint/outpaint; Dev is not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe.\n\nExamples by mode:\n- animate-move (DEFAULT — WAN 2.2 Animate Move: applies camera/motion from source video to reference image): \"Smooth cinematic camera movement following the subject through the scene.\"\n- animate-replace (WAN 2.2 Animate Replace: replaces the subject in the source video with the reference image): \"The person from the reference photo performing the actions from the video.\"\n- canny (LTX 2.5 default; LTX 2.3 rollback — edge-detection restyle): \"Hand-drawn watercolor anime style with soft ink edges, muted teal and coral palette, rain mist, neon reflections, warm rim light, preserving original silhouettes and composition.\"\n- pose (LTX 2.5 or LTX 2.3 — tracks skeleton and transfers the reference-image subject): \"A glossy cartoon robot from the reference image performs the source video's motion, with brushed metal texture, glowing cyan joints, and energetic stage lighting.\" This mode requires a reference image as well as the source video.\n- depth (LTX 2.5 default; LTX 2.3 rollback — depth-map restyle): \"A misty alpine valley at golden hour, expansive scale, volumetric haze, cool blue shadows, warm rim light, cinematic depth, lingering continuous shot.\"\n- detailer (LTX 2.5 default; LTX 2.3 rollback — enhance quality): DESCRIBE THE SOURCE, do not request changes. Append quality qualifiers only. E.g. \"The same scene, ultra-sharp and clean, crisp high-resolution detail, preserving all original content, composition, and color.\" Avoid words like \"enhanced textures\", \"restyled\", or any new subjects/objects — they cause drift.\n- seedance-v2v (BytePlus Dreamina Seedance 2.0 V2V): \"Restyle the source clip in a watercolor look with soft ink edges, while preserving its motion and composition.\" Use natural prose; Seedance reads the reference video holistically rather than via control-net constraints, so describe target style/mood/dialogue rather than control strength.\n- outpaint (LTX 2.5 default; LTX 2.3 rollback — canvas extension): describe what fills the NEWLY REVEALED area around the original frame, consistent with the source scene. E.g. \"The same street scene continues seamlessly into the newly revealed space — more wet asphalt, parked cars, and glowing shopfronts, matching the original lighting and perspective.\" Set outpaintPosition (and optionally outpaintAspectRatio); no mask needed.\n- inpaint (LTX 2.5 default; LTX 2.3 rollback — masked region regeneration): describe ONLY what the inpainted region should become; the rest of the frame is preserved. E.g. \"A vintage red convertible parked at the curb, matching the scene's lighting and shadows.\" If the user supplied a mask, set maskImageIndex. If no mask was supplied, omit maskImageIndex so execution derives a mask from the source video and prompt.\n\nPresent tense. Positive phrasing. Concrete visual details.\n\nNON-SEEDANCE POSITIVE CONSTRAINTS: For LTX 2.5, LTX 2.3, and WAN 2.2 modes, prompt is a positive prompt. Translate user avoid/no/don't constraints into affirmative production constraints instead of copying negative phrasing. Preserve exact quoted visible text when the user explicitly requests it; keep surrounding surfaces blank.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary the artistic treatment while keeping control mode and structural intent consistent. Example: \"transform to {watercolor with soft edges|oil painting with bold strokes|anime with clean lines} style\"."
|
|
701
825
|
},
|
|
702
826
|
"expandPrompt": {
|
|
703
827
|
"type": "boolean",
|
|
704
|
-
"description": "Seedance only. Whether to run
|
|
828
|
+
"description": "Seedance and Wan 3 only. Whether to run Sogni's exact-model prompt shaper before dispatch. Defaults to true. Set false when the supplied prompt is already model-ready and must remain exact."
|
|
705
829
|
},
|
|
706
830
|
"videoSourceIndex": {
|
|
707
831
|
"type": "number",
|
|
@@ -716,13 +840,15 @@
|
|
|
716
840
|
"pose",
|
|
717
841
|
"depth",
|
|
718
842
|
"detailer",
|
|
843
|
+
"outpaint",
|
|
844
|
+
"inpaint",
|
|
719
845
|
"seedance-v2v"
|
|
720
846
|
],
|
|
721
|
-
"description": "How the source video and
|
|
847
|
+
"description": "How the source video and reference image interact. Pick by user intent:\n• \"animate-move\" (DEFAULT) — WAN 2.2 Animate Move. Applies camera movement and motion from the source video to the reference image, bringing a still photo to life. Requires sourceImageIndex.\n• \"animate-replace\" — WAN 2.2 Animate Replace. Replaces the subject in the source video with the person/character from the reference image, keeping the video's background and motion. Requires sourceImageIndex.\n• \"canny\" — LTX 2.5 (default) or 2.3 edge-detection control. Best for restyling while preserving exact composition and silhouettes. Video-only.\n• \"pose\" — LTX 2.5 (default) or 2.3 skeletal tracking. Best for replacing a person while keeping their motion. LTX 2.5 requires both the source video and sourceImageIndex for the subject appearance; LTX 2.3 keeps its existing optional-image rollback behavior.\n• \"depth\" — LTX 2.5 (default) or 2.3 depth-map control. Best for scenes with perspective, camera movement, or volumetric content. Video-only.\n• \"detailer\" — LTX 2.5 (default) or 2.3 quality enhancement. Describe the original scene with quality qualifiers and do not request content changes.\n• \"outpaint\" — distilled LTX 2.5 (default) or LTX 2.3 canvas extension. Set outpaintPosition and optionally outpaintAspectRatio; Pro/dev is not supported for this mode.\n• \"inpaint\" — LTX 2.5 by default (LTX 2.3 rollback) masked region regeneration. Regenerate or replace a specific region of the source video while preserving the rest (e.g. \"replace the billboard\", \"change what's on the table\"). If the user provides an uploaded mask image, set maskImageIndex to it. If no mask is provided, omit maskImageIndex; execution derives a mask from the source video and prompt before dispatch. The prompt describes only the target inpainted region. Video-only.\n• \"seedance-v2v\" — BytePlus Dreamina Seedance 2.0 video-to-video. Use only when the user explicitly asks for Seedance on the uploaded source video, such as a Seedance upscale, enhance, remaster, restyle, or transform. High-fidelity quality, native audio, time-coded scene control. Seedance V2V reads @Video1 holistically. Use it for restyling, motion transfer, extension, subject replacement, or scene transformation, and assign @Video1 a clear role such as source clip, camera movement, action timing, edit rhythm, or continuation anchor. Distinct from canny/depth/pose which use control-net constraints — Seedance treats the reference video holistically.\nCanny vs depth: canny preserves silhouettes and fine outlines — pick it for subject-led scenes and graphic restyles. Depth preserves 3D structure — pick it for scenes where the camera moves or spatial layout matters more than edge fidelity. Default: \"animate-move\".\n\nUse seedance-v2v with videoModel=\"wan3.0-video\" for Wan 3 source-video editing or continuation; describe Video 1 as the source in the prompt."
|
|
722
848
|
},
|
|
723
849
|
"negativePrompt": {
|
|
724
850
|
"type": "string",
|
|
725
|
-
"description": "
|
|
851
|
+
"description": "Advanced non-Seedance only. Use this field only when the user explicitly asks to set a separate negative prompt. For ordinary avoid/no/don't constraints on LTX 2.3 or WAN 2.2, translate them into affirmative production constraints inside prompt instead; do not move them here. Do not set when controlMode is seedance-v2v or videoModel is seedance2/seedance2-mini/seedance2-5.\n\nWan 3 has no negativePrompt request field; do not set this for wan3.0-video."
|
|
726
852
|
},
|
|
727
853
|
"videoModel": {
|
|
728
854
|
"type": "string",
|
|
@@ -732,28 +858,92 @@
|
|
|
732
858
|
"wan22-animate",
|
|
733
859
|
"seedance2",
|
|
734
860
|
"seedance2-mini",
|
|
735
|
-
"seedance2-5"
|
|
861
|
+
"seedance2-5",
|
|
862
|
+
"wan3.0-video"
|
|
736
863
|
],
|
|
737
|
-
"description": "Model selector for this video-to-video request. Usually omit
|
|
864
|
+
"description": "Model selector for this video-to-video request. Usually omit: LTX control modes choose \"ltx25-v2v\" by default, while \"ltx23-v2v\" remains an explicit rollback. LTX 2.5 Fast, HQ, and Pro currently use the release-validated official Distilled workflow for canny, depth, pose, detailer, inpaint, and outpaint. Dev is not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Voice ID-LoRA, transition LoRA, and 10Eros remain LTX 2.3-only. For controlMode=\"seedance-v2v\", Seedance quality is selected only by model: use \"seedance2-mini\" for fast, lower-cost 720p Seedance V2V, and use \"seedance2\" for full/non-fast Seedance or 1080p/4K. Do not infer the Seedance model from Default Media Quality Fast/HQ/Pro or from 480p/720p resolution requests alone. \"seedance2-5\" is 480p and 720p only, renders 4-30s at 24 fps, and supports native audio. \"wan3.0-video\" is Alibaba Wan 3: one canonical premium-vendor model for text-to-video, first-frame and first+last-frame animation, loose multimodal references, audio-driven generation, and uploaded-video editing/extension. It renders 2-30s at fixed 30 fps with optional native audio, supports 480p/720p/1080p and 16:9/4:3/1:1/3:4/9:16, accepts up to 10 reference images, 5 reference videos, and 5 reference audios, and uses plain per-type prompt labels Image 1, Video 1, and Audio 1. Do not send negativePrompt. Use animate_photo for native first/last frames, generate_video for text or loose references, sound_to_video when audio is the primary driver, and video_to_video with controlMode=\"seedance-v2v\" for edits or extensions."
|
|
738
865
|
},
|
|
739
866
|
"generateAudio": {
|
|
740
867
|
"type": "boolean",
|
|
741
|
-
"description": "Whether the
|
|
868
|
+
"description": "Whether the returned video should include generated or retained audio. Omit to include audio by default; set false when the user asks for silent output or no audio."
|
|
742
869
|
},
|
|
743
870
|
"targetResolution": {
|
|
744
871
|
"type": "number",
|
|
745
|
-
"description": "Seedance V2V only. Short-side output resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact dimensions. Seedance V2V full supports 4K; Seedance V2V Mini and Seedance 2.5 support 480p and 720p only, so never set 1080p or 4K for \"seedance2-5\". Preserve the source video shape instead of forcing landscape pixels."
|
|
872
|
+
"description": "Seedance or Wan 3 V2V only. Short-side output resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact dimensions. Seedance V2V full supports 4K; Seedance V2V Mini, Fast, and Seedance 2.5 support 480p and 720p only, so never set 1080p or 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. Preserve the source video shape instead of forcing landscape pixels."
|
|
746
873
|
},
|
|
747
874
|
"sourceImageIndex": {
|
|
748
875
|
"type": "number",
|
|
749
|
-
"description": "
|
|
876
|
+
"description": "Index of a reference image (0-based). Required for \"animate-move\" and \"animate-replace\". Required for LTX 2.5 \"pose\" because that workflow needs both the source video and a still image that defines the subject appearance; optional for LTX 2.3 \"pose\" rollback. Ignored by \"canny\", \"depth\", \"detailer\", \"outpaint\", and \"inpaint\"."
|
|
877
|
+
},
|
|
878
|
+
"outpaintPosition": {
|
|
879
|
+
"type": "string",
|
|
880
|
+
"enum": [
|
|
881
|
+
"center",
|
|
882
|
+
"top",
|
|
883
|
+
"bottom",
|
|
884
|
+
"left",
|
|
885
|
+
"right"
|
|
886
|
+
],
|
|
887
|
+
"description": "controlMode=\"outpaint\" only. Where the ORIGINAL frame is anchored inside the expanded canvas, which determines the direction the canvas grows: \"left\" anchors the original on the left and adds new space on the right; \"right\" adds space on the left; \"top\" adds space below; \"bottom\" adds space above; \"center\" expands all sides evenly. Default: \"center\". Pick by the user's direction (\"extend to the right\" → \"left\"; \"make it wider\"/\"widescreen\" → \"center\")."
|
|
888
|
+
},
|
|
889
|
+
"outpaintAspectRatio": {
|
|
890
|
+
"type": "string",
|
|
891
|
+
"enum": [
|
|
892
|
+
"16:9",
|
|
893
|
+
"9:16",
|
|
894
|
+
"1:1",
|
|
895
|
+
"4:3",
|
|
896
|
+
"3:4",
|
|
897
|
+
"21:9"
|
|
898
|
+
],
|
|
899
|
+
"description": "controlMode=\"outpaint\" only. OPTIONAL target aspect ratio for the expanded canvas (e.g. \"16:9\" to make a vertical clip widescreen). The canvas only grows to reach this ratio — the original content is never cropped. Omit to expand moderately in the direction implied by outpaintPosition. Only set when the user names a target shape or orientation."
|
|
900
|
+
},
|
|
901
|
+
"maskImageIndex": {
|
|
902
|
+
"type": "number",
|
|
903
|
+
"description": "controlMode=\"inpaint\" only. Optional 0-based index of an uploaded mask IMAGE that marks the region to regenerate (white pixels = regenerate, black = preserve). Omit when the user did not provide a mask; execution will derive one from the source video and prompt. Ignored by every other controlMode."
|
|
750
904
|
},
|
|
751
905
|
"duration": {
|
|
752
906
|
"type": "number",
|
|
753
|
-
"description": "Output video duration in seconds.
|
|
907
|
+
"description": "Output video duration in seconds. Per-model range: WAN 2.2/LTX modes = 2-20s; Wan 3 = 2-30s subject to input-video plus output duration staying at or below 30s; Seedance 2.0 and Mini = 4-15s; Seedance 2.5 = 4-30s. If omitted, the tool matches the uploaded source video duration when available (capped to the selected model range); otherwise it falls back to 10s for WAN Animate Move/Replace and 5s for LTX/Seedance/Wan 3 modes. For long stitched/bulk WAN Animate Move/Replace work with no explicit per-clip length, prefer about 10s clips rather than 5s chunks. Only pass this when the user explicitly requests a different length.",
|
|
754
908
|
"minimum": 2,
|
|
755
909
|
"maximum": 30
|
|
756
910
|
},
|
|
911
|
+
"smartDuration": {
|
|
912
|
+
"type": "boolean",
|
|
913
|
+
"description": "Wan 3 only. Let Wan 3 choose 2-30 output seconds. Do not also set duration; input plus output must stay within the provider's 30-second limit. Sogni reserves 30 seconds and settles down to actual duration."
|
|
914
|
+
},
|
|
915
|
+
"wan3TaskType": {
|
|
916
|
+
"type": "string",
|
|
917
|
+
"enum": [
|
|
918
|
+
"edit",
|
|
919
|
+
"extend"
|
|
920
|
+
],
|
|
921
|
+
"description": "Wan 3 only. Use \"edit\" to transform the source or \"extend\" to continue it. For extension, explicitly describe the intended continuation."
|
|
922
|
+
},
|
|
923
|
+
"ratio": {
|
|
924
|
+
"type": "string",
|
|
925
|
+
"enum": [
|
|
926
|
+
"adaptive",
|
|
927
|
+
"16:9",
|
|
928
|
+
"4:3",
|
|
929
|
+
"1:1",
|
|
930
|
+
"3:4",
|
|
931
|
+
"9:16"
|
|
932
|
+
],
|
|
933
|
+
"description": "Wan 3 only. Output ratio. Use \"adaptive\" to let the provider choose from the source; omit to use the provider default."
|
|
934
|
+
},
|
|
935
|
+
"referenceFileUrl": {
|
|
936
|
+
"type": "string",
|
|
937
|
+
"description": "Wan 3 only. One public HTTPS document URL for additional edit/extension context (DOCX/DOC/XLSX/XLS/PPTX/PPT/PDF/TXT/KEY/PAGES/NUMBERS/Markdown, up to 100 MB; PDF/DOCX/DOC/PPTX/PPT/KEY/PAGES up to 50 pages). Mutually exclusive with referenceLinkUrl."
|
|
938
|
+
},
|
|
939
|
+
"referenceLinkUrl": {
|
|
940
|
+
"type": "string",
|
|
941
|
+
"description": "Wan 3 only. One public HTTPS webpage URL for additional edit/extension context. Mutually exclusive with referenceFileUrl."
|
|
942
|
+
},
|
|
943
|
+
"watermark": {
|
|
944
|
+
"type": "boolean",
|
|
945
|
+
"description": "Wan 3 only. Add Alibaba's visible watermark. Defaults to false."
|
|
946
|
+
},
|
|
757
947
|
"numberOfVariations": {
|
|
758
948
|
"type": "number",
|
|
759
949
|
"description": "Number of video variations to generate (1-16). Default: 1.",
|
|
@@ -943,21 +1133,21 @@
|
|
|
943
1133
|
"type": "function",
|
|
944
1134
|
"function": {
|
|
945
1135
|
"name": "sound_to_video",
|
|
946
|
-
"description": "Generate video synchronized to audio. Use when the user has uploaded an audio file (mp3, wav, m4a, flac) and the audio is the primary sync target, especially uploaded-audio-only workflows. Also use after generate_music (\"turn that song into a video\", \"make a music video from that\"). Auto-detects generated audio from generate_music if no audio file is uploaded. Seedance animate_photo/generate_video can also attach uploaded audio as a loose @Audio reference when an image or video reference anchors the request; use this tool instead when the soundtrack itself should drive the video. If the user provides a reference image, use ltx25-ia2v by default (ltx23-ia2v is rollback); for lip-sync with a face image, use wan-s2v; if no image, use ltx25-a2v by default (ltx23-a2v is rollback). If the user wants dialogue/audio WITHOUT pre-existing audio, use animate_photo instead (LTX 2.5
|
|
1136
|
+
"description": "Generate video synchronized to audio. Use when the user has uploaded an audio file (mp3, wav, m4a, flac) and the audio is the primary sync target, especially uploaded-audio-only workflows. Also use after generate_music (\"turn that song into a video\", \"make a music video from that\"). Auto-detects generated audio from generate_music if no audio file is uploaded. Seedance animate_photo/generate_video can also attach uploaded audio as a loose @Audio reference when an image or video reference anchors the request; use this tool instead when the soundtrack itself should drive the video. If the user provides a reference image, use ltx25-ia2v by default (ltx23-ia2v is rollback); for lip-sync with a face image, use wan-s2v; if no image, use ltx25-a2v by default (ltx23-a2v is rollback). If the user wants dialogue/audio WITHOUT pre-existing audio, use animate_photo instead (LTX 2.5 and LTX 2.3 generate audio natively). Note: Persona voice clips from resolve_personas are NOT used by this tool — for persona voice identity in video, use animate_photo or generate_video with videoModel=\"ltx23\" because LTX 2.5 has no compatible ID-LoRA. LONG AUDIO ON SEEDANCE: Seedance 2.0 and Mini cap each clip at 15s; Seedance 2.5 renders up to 30s in one call, so prefer seedance2-5 for 16-30s audio instead of splitting. When the user uploads audio longer than the per-clip cap of the selected model and Seedance is selected (seedance2, seedance2-mini, or seedance2-5), do NOT clamp to 15s and drop the rest — split the run into multiple sound_to_video calls in the same turn (one per 15s segment, so a 20s audio becomes two clips: audioStart=0 duration=15, then audioStart=15 duration=5) and finish with a single stitch_video call referencing the resulting clip indices in order with audioIndex pointing at the same uploaded audio so the stitched output carries the full original soundtrack. LTX/WAN models accept up to 20s per clip, so single-call is fine for them. Use videoModel=\"wan3.0-video\" when the user explicitly requests Wan 3 audio-driven video.",
|
|
947
1137
|
"parameters": {
|
|
948
1138
|
"type": "object",
|
|
949
1139
|
"properties": {
|
|
950
1140
|
"prompt": {
|
|
951
1141
|
"type": "string",
|
|
952
|
-
"description": "Describe the video like a cinematographer. Let the audio define timing — use the prompt for visual interpretation. One flowing paragraph, present tense, specific natural language.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. For Seedance, set expandPrompt=false.\n\nSTRUCTURE: shot/style and scale → subject → environment, lighting, color, texture, atmosphere → visual action synced to audio → camera movement. For LTX 2.
|
|
1142
|
+
"description": "Describe the video like a cinematographer. Let the audio define timing — use the prompt for visual interpretation. One flowing paragraph, present tense, specific natural language.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. For Seedance or Wan 3, set expandPrompt=false.\n\nSTRUCTURE: shot/style and scale → subject → environment, lighting, color, texture, atmosphere → visual action synced to audio → camera movement. For LTX 2.3 image+audio mode, do not re-describe static details already visible in the reference image; focus on motion, action, camera, and how the image responds to the audio.\n\nMOTION PACING: Scale complexity to duration. <=6s: 1 main visual beat + 1 simple camera move. Around 10s: 2-3 clear beats + 1 camera move. >10s: up to 4 beats in clear sequence. Let the audio define timing, but avoid stacking subject, camera, and environment motion in short clips.\n\nBLOCKING: Direct layout when it affects the shot: left/right placement, foreground/background, facing direction, and relative distance between subjects.\n\nLIP-SYNC: Shot framing, speaker's appearance and setting, physical performance synced to audio — gestures, expressions, jaw movement between phrases. Include acting beats.\n\nMUSIC VISUALIZATION: Visual style, environment, and how elements react to rhythm and energy.\n\nAUDIO-REACTIVE: Motion and visual changes that correspond to sounds in the track.\n\nLTX VOCABULARY: camera (tracking, dolly, pan, tilt, handheld, static frame), lighting/atmosphere (golden hour, neon glow, dramatic shadows, fog, rain, smoke, reflections), scale/pacing (expansive, epic, intimate, claustrophobic, slow motion, time-lapse, lingering shot, continuous shot), style/genre (film noir, painterly, cyberpunk, stop-motion, claymation, 2D/3D animation, hand-drawn, fantasy, thriller, experimental film).\n\nAVOID: Vague prompts, too many competing visual elements, abstract descriptions without visible behavior, rigid numeric constraints, readable text or logos. QUOTING RULE: ONLY use double quotes for spoken dialogue. Never quote on-screen text, overlay text, titles, captions, signs, or any visual text — describe them without quotes.\n\nNON-SEEDANCE POSITIVE CONSTRAINTS: For ltx25-ia2v, ltx25-a2v, ltx23-ia2v, ltx23-a2v, and wan-s2v, prompt is a positive prompt. Translate user avoid/no/don't constraints into affirmative production constraints instead of copying negative phrasing. Preserve exact quoted visible text or dialogue when the user explicitly requests it; keep surrounding surfaces blank.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary the visual interpretation while keeping audio sync intent consistent. This is one Sogni project with multiple jobs, so prefer it when all outputs share the same audio source/window, image source, model, duration, dimensions, and parameters and only prompt text varies. Example: \"{abstract neon visualization|nature scene with swaying trees|urban street with rain} synced to the beat\"."
|
|
953
1143
|
},
|
|
954
1144
|
"expandPrompt": {
|
|
955
1145
|
"type": "boolean",
|
|
956
|
-
"description": "Seedance only. Whether to run
|
|
1146
|
+
"description": "Seedance and Wan 3 only. Whether to run Sogni's exact-model prompt shaper before dispatch. Defaults to true. Set false when the supplied prompt is already model-ready and must remain exact."
|
|
957
1147
|
},
|
|
958
1148
|
"negativePrompt": {
|
|
959
1149
|
"type": "string",
|
|
960
|
-
"description": "Advanced LTX 2.5/LTX 2.3/WAN only. The LTX A2V and IA2V workflows accept this separate negative prompt. Use it only when the user explicitly asks to set one. Do not set for Seedance."
|
|
1150
|
+
"description": "Advanced LTX 2.5/LTX 2.3/WAN only. The LTX A2V and IA2V workflows accept this separate negative prompt. Use it only when the user explicitly asks to set one. Do not set for Seedance.\n\nWan 3 has no negativePrompt request field; do not set this for wan3.0-video."
|
|
961
1151
|
},
|
|
962
1152
|
"audioSourceIndex": {
|
|
963
1153
|
"type": "number",
|
|
@@ -965,7 +1155,7 @@
|
|
|
965
1155
|
},
|
|
966
1156
|
"sourceImageIndex": {
|
|
967
1157
|
"type": "number",
|
|
968
|
-
"description": "Optional index of an uploaded image to use as the starting frame (0-based). Required for lip-sync models (WAN S2V). For audio-only-to-video models (LTX 2.
|
|
1158
|
+
"description": "Optional index of an uploaded image to use as the starting frame (0-based). Required for lip-sync models (WAN S2V). For audio-only-to-video models (LTX 2.3 A2V), this is optional — omit it to generate video purely from text + audio."
|
|
969
1159
|
},
|
|
970
1160
|
"audioStart": {
|
|
971
1161
|
"type": "number",
|
|
@@ -974,10 +1164,38 @@
|
|
|
974
1164
|
},
|
|
975
1165
|
"duration": {
|
|
976
1166
|
"type": "number",
|
|
977
|
-
"description": "Video duration in seconds. Default: 5.
|
|
1167
|
+
"description": "Video duration in seconds. Default: 5. Per-model range: LTX/WAN 2.2 = 2-20s; Wan 3 = 2-30s; Seedance 2.0 and Mini = 4-15s; Seedance 2.5 = 4-30s. For music videos, use the maximum duration the selected model allows because the audio is usually longer than the video limit. Use when the user explicitly requests a specific length.",
|
|
978
1168
|
"minimum": 2,
|
|
979
1169
|
"maximum": 30
|
|
980
1170
|
},
|
|
1171
|
+
"smartDuration": {
|
|
1172
|
+
"type": "boolean",
|
|
1173
|
+
"description": "Wan 3 only. Let the model choose 2-30 seconds. Do not also set duration. Sogni reserves 30 seconds and settles down to the provider-reported duration."
|
|
1174
|
+
},
|
|
1175
|
+
"ratio": {
|
|
1176
|
+
"type": "string",
|
|
1177
|
+
"enum": [
|
|
1178
|
+
"adaptive",
|
|
1179
|
+
"16:9",
|
|
1180
|
+
"4:3",
|
|
1181
|
+
"1:1",
|
|
1182
|
+
"3:4",
|
|
1183
|
+
"9:16"
|
|
1184
|
+
],
|
|
1185
|
+
"description": "Wan 3 only. Output ratio; \"adaptive\" derives it from the input."
|
|
1186
|
+
},
|
|
1187
|
+
"watermark": {
|
|
1188
|
+
"type": "boolean",
|
|
1189
|
+
"description": "Wan 3 only. Add Alibaba's visible watermark. Defaults to false."
|
|
1190
|
+
},
|
|
1191
|
+
"referenceFileUrl": {
|
|
1192
|
+
"type": "string",
|
|
1193
|
+
"description": "Wan 3 only. One public HTTPS document URL for additional audio-driven context (DOCX/DOC/XLSX/XLS/PPTX/PPT/PDF/TXT/KEY/PAGES/NUMBERS/Markdown, up to 100 MB; PDF/DOCX/DOC/PPTX/PPT/KEY/PAGES up to 50 pages). Mutually exclusive with referenceLinkUrl."
|
|
1194
|
+
},
|
|
1195
|
+
"referenceLinkUrl": {
|
|
1196
|
+
"type": "string",
|
|
1197
|
+
"description": "Wan 3 only. One public HTTPS webpage URL for additional audio-driven context. Mutually exclusive with referenceFileUrl."
|
|
1198
|
+
},
|
|
981
1199
|
"videoModel": {
|
|
982
1200
|
"type": "string",
|
|
983
1201
|
"enum": [
|
|
@@ -988,13 +1206,14 @@
|
|
|
988
1206
|
"ltx25-ia2v",
|
|
989
1207
|
"ltx25-a2v",
|
|
990
1208
|
"ltx23-ia2v",
|
|
991
|
-
"ltx23-a2v"
|
|
1209
|
+
"ltx23-a2v",
|
|
1210
|
+
"wan3.0-video"
|
|
992
1211
|
],
|
|
993
|
-
"description": "
|
|
1212
|
+
"description": "\"ltx25-ia2v\" (default with image) and \"ltx25-a2v\" (default without image) use the release-validated official LTX 2.5 Distilled/Turbo workflows for Fast, HQ, and Pro. Dev is not publicly routed until upstream publishes and Sogni validates official ComfyUI Dev recipes. \"ltx23-ia2v\" and \"ltx23-a2v\" remain rollback selectors with their existing quality-tier routing. \"wan-s2v\" is WAN 2.2 sound-to-video for lip-sync with a face image. Seedance quality is selected only by model: use \"seedance2-mini\" for faster/lower-cost 720p drafts, \"seedance2\" for full/non-fast Seedance or 1080p/4K, and \"seedance2-5\" for 480p/720p clips up to 30 seconds. Do not infer the Seedance model from Default Media Quality Fast/HQ/Pro. Omit to auto-select based on whether an image is present. \"wan3.0-video\" is Alibaba Wan 3: one canonical premium-vendor model for text-to-video, first-frame and first+last-frame animation, loose multimodal references, audio-driven generation, and uploaded-video editing/extension. It renders 2-30s at fixed 30 fps with optional native audio, supports 480p/720p/1080p and 16:9/4:3/1:1/3:4/9:16, accepts up to 10 reference images, 5 reference videos, and 5 reference audios, and uses plain per-type prompt labels Image 1, Video 1, and Audio 1. Do not send negativePrompt. Use animate_photo for native first/last frames, generate_video for text or loose references, sound_to_video when audio is the primary driver, and video_to_video with controlMode=\"seedance-v2v\" for edits or extensions."
|
|
994
1213
|
},
|
|
995
1214
|
"generateAudio": {
|
|
996
1215
|
"type": "boolean",
|
|
997
|
-
"description": "Whether the
|
|
1216
|
+
"description": "Whether the returned video should include audio. Omit to include audio by default; set false when the user asks for silent output or no audio. The reference audio is still required and still drives generation even when the returned video has no audio track."
|
|
998
1217
|
},
|
|
999
1218
|
"numberOfVariations": {
|
|
1000
1219
|
"type": "number",
|
|
@@ -1021,7 +1240,7 @@
|
|
|
1021
1240
|
"type": "function",
|
|
1022
1241
|
"function": {
|
|
1023
1242
|
"name": "extend_video",
|
|
1024
|
-
"description": "Extend a video by adding new time to the end. Works on BOTH videos previously rendered in this session AND user-uploaded videos — set videoIndex to a negative number (e.g. -1) to target an uploaded video when no prior render exists. The base video is auto-selected from the most recent video in this session unless videoIndex is set. For LTX
|
|
1243
|
+
"description": "Extend a video by adding new time to the end. Works on BOTH videos previously rendered in this session AND user-uploaded videos — set videoIndex to a negative number (e.g. -1) to target an uploaded video when no prior render exists. The base video is auto-selected from the most recent video in this session unless videoIndex is set. For LTX 2.5/2.3 base clips, the tool extracts the last frame and renders an image-to-video continuation; new non-Seedance continuations default to LTX 2.5. For Seedance base clips, the tool extracts a trailing reference segment and renders a video-to-video continuation. Returns both the standalone new segment and a spliced composite (base + new segment). Use when the user asks to \"make it longer\", \"extend the video\", \"add another N seconds\", \"continue the scene\", \"add an outro/bumper to the end\", etc. Prefer this over generate_image+animate_photo+stitch_video for \"add a bumper/outro to this video\" — extend_video preserves the original base bytes, audio, and timing instead of re-encoding them. Do not use this tool to render fresh videos from scratch — call generate_video or animate_photo for that. Output durations follow each model's native limits (LTX 2-20s; Seedance 2.0 and Mini 4-15s; Seedance 2.5 4-30s) for the new segment alone.",
|
|
1025
1244
|
"parameters": {
|
|
1026
1245
|
"type": "object",
|
|
1027
1246
|
"properties": {
|
|
@@ -1031,7 +1250,7 @@
|
|
|
1031
1250
|
},
|
|
1032
1251
|
"duration": {
|
|
1033
1252
|
"type": "number",
|
|
1034
|
-
"description": "Length in seconds of the new appended segment (NOT total final length). LTX 2-
|
|
1253
|
+
"description": "Length in seconds of the new appended segment (NOT total final length). LTX 2-20s; Seedance 2.0 and Mini 4-15s; Seedance 2.5 4-30s. Default: 5.",
|
|
1035
1254
|
"minimum": 2,
|
|
1036
1255
|
"maximum": 30
|
|
1037
1256
|
},
|
|
@@ -1049,7 +1268,7 @@
|
|
|
1049
1268
|
"seedance2-mini",
|
|
1050
1269
|
"seedance2-5"
|
|
1051
1270
|
],
|
|
1052
|
-
"description": "Which model to use for the new segment. Default: \"auto\" — preserve Seedance for a Seedance base and otherwise use LTX 2.5. Use ltx23 only for explicit rollback.
|
|
1271
|
+
"description": "Which model to use for the new segment. Default: \"auto\" — preserve Seedance for a Seedance base and otherwise use LTX 2.5. Use ltx23 only for explicit rollback. Override only when the user explicitly requests a different model."
|
|
1053
1272
|
},
|
|
1054
1273
|
"keepOriginalAudio": {
|
|
1055
1274
|
"type": "boolean",
|
|
@@ -1066,7 +1285,7 @@
|
|
|
1066
1285
|
"type": "function",
|
|
1067
1286
|
"function": {
|
|
1068
1287
|
"name": "replace_video_segment",
|
|
1069
|
-
"description": "Modify a portion of an existing video while keeping the rest intact — either by regenerating that slice fresh or by splicing in another existing clip. Operates on a [startSeconds, endSeconds] window inside a base video; everything outside the window stays exactly as it was. Works on BOTH videos previously rendered in this session AND user-uploaded videos (set videoIndex to a negative number to target an uploaded video when no prior render exists). WHEN TO USE: any request to change part of one video while keeping the rest, to put another clip inside another video at a specific position, or to interleave time slices of multiple videos. Plain language: \"regenerate from 5s to 10s\", \"redo the last 3 seconds\", \"swap out the middle\", \"replace the bumper at the end\", \"swap the end card\", \"change the outro / intro / ending / last clip\", \"replace 2s-4s with a stronger expression\", \"splice video 2 into video 1\", \"stitch video 2 into the middle of video 1\", \"insert the second clip at 5s\", \"alternate 1 second from each video\". The word \"stitch\" in the user request does not by itself mean stitch_video — when the user clearly wants insertion or in-place replacement, this tool is the right one. WHEN NOT TO USE — prefer stitch_video instead: the user wants to concatenate whole clips end-to-end without modifying their interiors (\"stitch these together\", \"play A then B\", \"add a bumper before / after\"). SPLICING EXISTING CLIPS: pass replacementVideoIndex when the replacement already exists as an uploaded or generated video — do not call generate_video / animate_photo / video_to_video in that case. Set endSeconds=startSeconds when the user asks for an insertion that should not remove time from the base video. TIME-SLICED INTERLEAVING (\"alternate 1 second from each video\"): pass replacementStartSeconds and replacementEndSeconds to cut the next source slice out of the replacement video before splicing it into the base. Repeat this call for each alternating window. By default use replacement windows (endSeconds = startSeconds + sliceDuration); use insertion windows (endSeconds = startSeconds) only when the user explicitly asks to lengthen the output by inserting extra slices. replacementStartSeconds and replacementEndSeconds must be concrete non-negative seconds; never use -1 as an end-of-source sentinel. PREFER this over re-running generate_video / animate_photo on the original prompt when the user only wants part of the video changed — re-rendering wastes credits, loses the unchanged sections, and breaks the original timing. If the user does not specify the exact start/end seconds (e.g. \"replace the bumper at the end\"), call analyze_video first to identify the correct window, OR derive it from the storyboard timing already in the conversation (e.g. last beat's time range). Do not guess wildly — pick a sensible bumper/end-card window such as the final 1-3 seconds when the storyboard says scene_07 is 14-15s. Returns both the standalone replacement clip and the spliced composite. For LTX
|
|
1288
|
+
"description": "Modify a portion of an existing video while keeping the rest intact — either by regenerating that slice fresh or by splicing in another existing clip. Operates on a [startSeconds, endSeconds] window inside a base video; everything outside the window stays exactly as it was. Works on BOTH videos previously rendered in this session AND user-uploaded videos (set videoIndex to a negative number to target an uploaded video when no prior render exists). WHEN TO USE: any request to change part of one video while keeping the rest, to put another clip inside another video at a specific position, or to interleave time slices of multiple videos. Plain language: \"regenerate from 5s to 10s\", \"redo the last 3 seconds\", \"swap out the middle\", \"replace the bumper at the end\", \"swap the end card\", \"change the outro / intro / ending / last clip\", \"replace 2s-4s with a stronger expression\", \"splice video 2 into video 1\", \"stitch video 2 into the middle of video 1\", \"insert the second clip at 5s\", \"alternate 1 second from each video\". The word \"stitch\" in the user request does not by itself mean stitch_video — when the user clearly wants insertion or in-place replacement, this tool is the right one. WHEN NOT TO USE — prefer stitch_video instead: the user wants to concatenate whole clips end-to-end without modifying their interiors (\"stitch these together\", \"play A then B\", \"add a bumper before / after\"). SPLICING EXISTING CLIPS: pass replacementVideoIndex when the replacement already exists as an uploaded or generated video — do not call generate_video / animate_photo / video_to_video in that case. Set endSeconds=startSeconds when the user asks for an insertion that should not remove time from the base video. TIME-SLICED INTERLEAVING (\"alternate 1 second from each video\"): pass replacementStartSeconds and replacementEndSeconds to cut the next source slice out of the replacement video before splicing it into the base. Repeat this call for each alternating window. By default use replacement windows (endSeconds = startSeconds + sliceDuration); use insertion windows (endSeconds = startSeconds) only when the user explicitly asks to lengthen the output by inserting extra slices. replacementStartSeconds and replacementEndSeconds must be concrete non-negative seconds; never use -1 as an end-of-source sentinel. PREFER this over re-running generate_video / animate_photo on the original prompt when the user only wants part of the video changed — re-rendering wastes credits, loses the unchanged sections, and breaks the original timing. If the user does not specify the exact start/end seconds (e.g. \"replace the bumper at the end\"), call analyze_video first to identify the correct window, OR derive it from the storyboard timing already in the conversation (e.g. last beat's time range). Do not guess wildly — pick a sensible bumper/end-card window such as the final 1-3 seconds when the storyboard says scene_07 is 14-15s. Returns both the standalone replacement clip and the spliced composite. For LTX 2.5/2.3 and Wan 2.2 base videos the tool locks both ends with first/last-frame keyframes for seamless edges; new non-Seedance LTX segments default to 2.5. For Seedance base videos the tool uses the original window as a reference for video-to-video transformation. If a requested window is shorter than the selected model's native render minimum, the handler renders a slightly larger handled clip, trims the result back to the requested seconds, then splices exactly that requested range. By default the regenerated segment's audio replaces the original audio in the [startSeconds, endSeconds] window, so new motion stays in sync with new sound. Pass keepOriginalAudio=true only when the user explicitly asks to keep the existing audio — phrasings like \"keep the audio\", \"leave the original audio\", \"preserve the music/score/dialogue\", \"don't change the audio\". If the user uses an ambiguous phrasing such as \"with the audio\" (which could mean either \"with the original audio kept\" or \"with new audio\"), DO NOT call this tool yet — first ask the user whether to preserve or replace the original audio in the replaced window. When replacementVideoIndex is set, the existing replacement clip's own audio is used; pass keepOriginalAudio=true only when the user explicitly wants the base video audio to stay over the replacement window.",
|
|
1070
1289
|
"parameters": {
|
|
1071
1290
|
"type": "object",
|
|
1072
1291
|
"properties": {
|
|
@@ -1111,7 +1330,7 @@
|
|
|
1111
1330
|
"seedance2-mini",
|
|
1112
1331
|
"seedance2-5"
|
|
1113
1332
|
],
|
|
1114
|
-
"description": "Which model to use for the new segment. Default: \"auto\" — preserve Seedance or WAN for matching base clips and otherwise use LTX 2.5. Use ltx23 only for explicit rollback.
|
|
1333
|
+
"description": "Which model to use for the new segment. Default: \"auto\" — preserve Seedance or WAN for matching base clips and otherwise use LTX 2.5. Use ltx23 only for explicit rollback. Override only when the user explicitly requests a different model."
|
|
1115
1334
|
},
|
|
1116
1335
|
"keepOriginalAudio": {
|
|
1117
1336
|
"type": "boolean",
|