@nodaro/prompts 1.6.0 → 1.7.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.cjs +628 -15
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +499 -2
- package/dist/index.d.ts +499 -2
- package/dist/index.js +622 -17
- package/dist/index.js.map +1 -1
- package/package.json +1 -1
- package/src/__tests__/doctrine-roster-completeness.test.ts +59 -0
- package/src/__tests__/gemini-omni-inputs.test.ts +68 -0
- package/src/__tests__/picker-analyzer-registry.test.ts +11 -1
- package/src/__tests__/provider-prompt-doctrine.test.ts +5 -3
- package/src/__tests__/seedance-2-inputs.test.ts +48 -0
- package/src/__tests__/veo-i2v-inputs.test.ts +58 -0
- package/src/gemini-omni-inputs.ts +73 -0
- package/src/index.ts +3 -0
- package/src/picker-analyzer-registry.ts +190 -0
- package/src/picker-wiring.ts +336 -0
- package/src/prompt-wizard-categories.ts +3 -2
- package/src/provider-prompt-doctrine.ts +228 -5
- package/src/resolve-prompt.ts +12 -1
- package/src/seedance-2-inputs.ts +10 -3
- package/src/veo-i2v-inputs.ts +68 -0
- package/src/video-reference-resolver.ts +12 -0
|
@@ -47,6 +47,7 @@ const SEEDANCE_2_DOCTRINE: ProviderPromptDoctrine = {
|
|
|
47
47
|
"References go by ordinal (@Image 1, Video 2) in attachment order; earlier = higher priority. Identity = ONE headshot + ONE full-body (multi-view sheets cause ID drift). 4-5 assets total beats maxing the 9/3/3 caps.",
|
|
48
48
|
"No negative-prompt parameter — put constraints in the prompt: 'keep it subtitle-free, do not generate a watermark, do not generate a logo'.",
|
|
49
49
|
"seedance-2-5 only: one shot runs to 30s (the 2.0 SKUs stop at 15s), so storyboard a whole beat instead of planning a stitch. Ref caps are wider (30/10/10), but 4-5 assets still gives the best identity fidelity.",
|
|
50
|
+
"Auto-path formula: Subject → Action → Environment → Camera → Style → Constraints in 60-100 words; ONE camera instruction (chain with 'then'); separate camera motion from subject motion; always add one lighting phrase.",
|
|
50
51
|
],
|
|
51
52
|
doctrine: `Prompt structure (front-load what matters most):
|
|
52
53
|
precise subject → action details → scene/environment → lighting & color tone → camera movement → visual style → image quality → constraints.
|
|
@@ -61,7 +62,7 @@ precise subject → action details → scene/environment → lighting & color to
|
|
|
61
62
|
**Generation differences (seedance-2-5 vs the 2.0 SKUs)**
|
|
62
63
|
- A single 2.5 shot runs to 30s, where every 2.0 SKU stops at 15s. Plan a complete 4-6 shot beat inside ONE generation instead of splitting it into two clips and stitching — no seam to hide, and continuity holds because it never leaves the model.
|
|
63
64
|
- 2.5 also takes far more reference material (30 images / 10 videos / 10 audio vs 9/3/3). Treat that as room for COVERAGE — more distinct characters, locations and props in one shot — not as licence to pile refs onto one identity. The "ONE headshot + ONE full-body, 4-5 assets total" rule above still produces the best likeness on 2.5.
|
|
64
|
-
- 2.5 renders at 480p/720p
|
|
65
|
+
- 2.5 renders at 480p/720p/1080p (1080p since 2026-08-17): there is no 4K tier, so route a job that needs 4K to seedance-2 (which has it) or upscale afterwards.
|
|
65
66
|
- With a start frame, 2.5 always derives the output aspect from that frame — an explicit aspect ratio is rejected outright, so compose the frame at the ratio you want.
|
|
66
67
|
|
|
67
68
|
**References (when reference media is attached)**
|
|
@@ -84,12 +85,35 @@ precise subject → action details → scene/environment → lighting & color to
|
|
|
84
85
|
**Known weaknesses → workarounds**
|
|
85
86
|
- Text rendering is weak: keep on-screen text to short common words; for exact text or logos, attach the artwork as a reference image and instruct "the logo from Image N stays in the corner unchanged".
|
|
86
87
|
- More than 4 referenced people gets unstable: group people into composite images of ≤4 first (image generation), then reference those composites.
|
|
87
|
-
- Repeated extension degrades quality: prefer high-definition reference assets and avoid stacking many continuations
|
|
88
|
+
- Repeated extension degrades quality: prefer high-definition reference assets and avoid stacking many continuations.
|
|
89
|
+
|
|
90
|
+
**Auto-path formula (community-sourced enrichment — apiyi.com Seedance 2.0 prompt guide,
|
|
91
|
+
higgsfield.ai 4K breakdown; captured 2026-08-09)**
|
|
92
|
+
- Six steps IN ORDER, 60-100 words total (longer measurably degrades): Subject → Action → Environment → Camera → Style → Constraints.
|
|
93
|
+
- ONE primary camera instruction per shot. Compound moves chain with "then": "camera slow tracking then subtle rise" — never two competing verbs. The 8 reliable camera types: push-in, pull-out, pan, tracking, orbit/arc, aerial, handheld, locked-off.
|
|
94
|
+
- SEPARATE camera movement from subject movement — the single biggest quality lever: "The dancer spins slowly. Camera holds fixed framing." — never "spinning camera around a dancing person".
|
|
95
|
+
- Pace with human words (slow / gentle / gradual / smooth / controlled) — never fps numbers or f-stops in the basic path.
|
|
96
|
+
- ALWAYS add one lighting phrase (highest-impact single addition): golden hour / rim light / neon glow / backlit / overcast.
|
|
97
|
+
- Bake stability constraints in: "avoid jitter and bent limbs", "avoid temporal flicker", "avoid identity drift".
|
|
98
|
+
- Ban vague adjectives standing alone ("epic", "amazing", "beautiful", bare "cinematic") — every adjective needs a concrete noun.
|
|
99
|
+
- Mode notes: i2v — skip subject description (the frame has it), focus on motion, append "preserve composition and colors". v2v — describe the style TRANSFORM, keep motion + identity.
|
|
100
|
+
- Advanced (pro path): focal angles in degrees ("47° normal", "29° telephoto", "107° wide"); "180° shutter" for filmic motion blur; handheld texture as "organic shake, micro-drift, subtle dutch"; "white balance locked 5200K"; explicit POSITIVE LOCKS section + "100% matches the reference" for identity-critical shots.
|
|
101
|
+
|
|
102
|
+
**Camera-path control — the magenta-line method (STORYBOARD community technique; the
|
|
103
|
+
manual pro path for precise trajectories, NOT the auto path)**
|
|
104
|
+
1. Duplicate the start frame; on the COPY draw a thick magenta line + arrowhead — the line is the camera's flight path, the arrow its end point. Keep the clean original.
|
|
105
|
+
2. Attach BOTH frames and declare the guide: "Image N contains a magenta line and arrow — a hidden camera trajectory guide, NOT part of the scene. Completely remove it: no line, no arrow, no paint, no trail, no reflection." Skipping the removal order RENDERS the line.
|
|
106
|
+
3. Command the path: "one continuous FPV drone glide following the S-shaped curve as closely as possible — do not shortcut. Camera motion is the priority." Lock the clean frame as first frame + scene reference; lock the destination frame if wired.
|
|
107
|
+
4. Pace with timing blocks ("[00:00-00:02] rise over the rooftop … [00:07-00:09] settle on the doorway") and keep any dialogue SHORT — long lines fight the move.
|
|
108
|
+
5. Assign image-input roles explicitly: first-frame/scene-ref · destination frame · path-guide · 3-6 character-identity refs — and bind identities with @-mentions exactly like the platform's reference pills.`,
|
|
88
109
|
}
|
|
89
110
|
|
|
90
111
|
const KLING_AUDIO_DOCTRINE: ProviderPromptDoctrine = {
|
|
91
|
-
|
|
92
|
-
|
|
112
|
+
// kling-turbo (2.5 Turbo Pro) + kling-master (2.1 Master) are SILENT tiers of
|
|
113
|
+
// the same engine: the structure/motion guidance applies, the Audio block
|
|
114
|
+
// does not (variant note in the doctrine body).
|
|
115
|
+
providers: ["kling", "kling-3.0", "kling-3-omni", "kling-turbo", "kling-master"],
|
|
116
|
+
heading: "Kling 2.1 / 2.5 / 2.6 / 3.0 / 3 Omni (kling, kling-3.0, kling-3-omni, kling-turbo, kling-master)",
|
|
93
117
|
tips: [
|
|
94
118
|
"Kling speaks scripted dialogue natively with lip sync — quote the line and enable sound: [Anna: warm calm voice]: \"We made it.\" On kling/kling-3.0 audio raises the credit cost; kling-3-omni includes it.",
|
|
95
119
|
"Structure prompts as Scene → character/element → Motion → Audio → style. Put ALL sound in one 'Audio:' block: dialogue in quotes, then SFX and ambience described plainly ('rain tapping on glass, no music').",
|
|
@@ -121,7 +145,11 @@ const KLING_AUDIO_DOCTRINE: ProviderPromptDoctrine = {
|
|
|
121
145
|
|
|
122
146
|
**Limits**
|
|
123
147
|
- Kling 2.6 prompts cap at 1000 characters — front-load scene + dialogue and trim style tails first. kling-3.0 accepts long prompts.
|
|
124
|
-
- Durations: 2.6 = 5/10s; 3.0/omni = 3-15s. A spoken line needs roughly 1s per 2-3 words — don't script more dialogue than the clip can hold
|
|
148
|
+
- Durations: 2.6 = 5/10s; 3.0/omni = 3-15s. A spoken line needs roughly 1s per 2-3 words — don't script more dialogue than the clip can hold.
|
|
149
|
+
|
|
150
|
+
**Variant note — kling-turbo (2.5 Turbo Pro) & kling-master (2.1 Master)**
|
|
151
|
+
- SILENT tiers: no audio parameter, so the entire Audio block above does not apply — skip dialogue/SFX cues; the Scene → Character → Motion → Style structure and motion guidance carry over unchanged.
|
|
152
|
+
- Durations 5/10s; kling-turbo takes an end frame (tail_image_url); kling-master is single-image i2v.`,
|
|
125
153
|
}
|
|
126
154
|
|
|
127
155
|
const MINIMAX_H3_DOCTRINE: ProviderPromptDoctrine = {
|
|
@@ -160,10 +188,205 @@ precise subject → action details → scene/environment → lighting & color to
|
|
|
160
188
|
- There is NO negative-prompt parameter — all constraints belong in the prompt text itself: "keep it subtitle-free, do not generate a watermark, do not generate a logo, stable picture".`,
|
|
161
189
|
}
|
|
162
190
|
|
|
191
|
+
const VEO_31_DOCTRINE: ProviderPromptDoctrine = {
|
|
192
|
+
providers: ["veo3", "veo3.1", "veo3_lite", "veo-1080p", "veo-4k", "veo-extend"],
|
|
193
|
+
heading: "VEO 3.1 — Quality / Fast / Lite (veo3, veo3.1, veo3_lite)",
|
|
194
|
+
tips: [
|
|
195
|
+
"Structure prompts as [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance] — lead with the camera, not the subject (Google's official formula).",
|
|
196
|
+
"Dialogue: quote the exact line with attribution — A woman says, \"We have to leave now.\" (no subtitles). Cue sound as separate lines: SFX: thunder cracks; Ambient noise: quiet hum of a starship bridge.",
|
|
197
|
+
"Multi-shot pacing via timestamp blocks: [00:00-00:02] medium shot… [00:02-00:04] reverse shot… — VEO honors per-window actions inside one 8s generation.",
|
|
198
|
+
"Negative prompting is positive phrasing: not 'no buildings' but 'a desolate landscape with no buildings or roads'. Keep prompts under ~175 words — longer overloads the generation.",
|
|
199
|
+
"Start+end frame: pass both and describe the transition move ('smooth 180-degree arc ending on the POV behind her'). References (ingredients) keep characters/objects consistent and DO generate audio.",
|
|
200
|
+
],
|
|
201
|
+
doctrine: `Prompt structure (Google's official Veo 3.1 formula — lead with the camera):
|
|
202
|
+
[Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance].
|
|
203
|
+
Example: "Medium shot, a tired corporate worker, rubbing his temples in exhaustion, in front of a bulky 1980s computer in a cluttered office late at night, lit by harsh fluorescents and the green monitor glow. Retro aesthetic, 1980s color film, slightly grainy."
|
|
204
|
+
|
|
205
|
+
**Camera vocabulary (use the exact terms)**
|
|
206
|
+
- Movement: dolly shot, tracking shot, crane shot, aerial view, slow pan, POV shot, 180-degree arc shot.
|
|
207
|
+
- Composition: wide shot, medium shot, close-up, extreme close-up, two-shot, low angle, high angle.
|
|
208
|
+
- Lens/focus: shallow depth of field, deep focus, wide-angle lens, macro lens, soft focus.
|
|
209
|
+
|
|
210
|
+
**Audio (native, multi-track — dialogue / SFX / ambience)**
|
|
211
|
+
- Dialogue: quote the exact line with attribution: The detective says in a weary voice, "Of all the offices in this town, you had to walk into mine." Append "(no subtitles)" — VEO otherwise tends to burn captions in.
|
|
212
|
+
- Sound effects on their own line: "SFX: a crystal wine glass shatters on the marble floor". Ambient bed: "Ambient noise: rain against the window, distant traffic".
|
|
213
|
+
- Sound can drive the visual ("the sound reverberating through the empty ballroom") — VEO syncs audio-visual timing.
|
|
214
|
+
|
|
215
|
+
**Multi-shot timestamp prompting (inside one generation)**
|
|
216
|
+
- Split the clip into [mm:ss-mm:ss] windows, one action per window:
|
|
217
|
+
[00:00-00:02] Medium shot from behind a young explorer walking toward a clearing.
|
|
218
|
+
[00:02-00:04] Reverse shot of her freckled face, eyes widening.
|
|
219
|
+
[00:04-00:08] Wide, high-angle crane shot revealing the ruins below.
|
|
220
|
+
- 4 / 6 / 8 second clips; budget ~2s per window.
|
|
221
|
+
|
|
222
|
+
**Frames & references**
|
|
223
|
+
- Start + end frame: wire both (imageUrls [start, end]) and describe the camera path between them — "a smooth 180-degree arc shot, starting front-facing and circling to end on the POV from behind her".
|
|
224
|
+
- Reference images (ingredients): attach character/object/scene refs and name them in the prompt ("using the provided images for the detective and the office, …"). Reference runs DO generate audio.
|
|
225
|
+
|
|
226
|
+
**Constraints**
|
|
227
|
+
- Negative prompting works by positive description: write "a desolate landscape with no buildings or roads", not "no buildings".
|
|
228
|
+
- Keep prompts ≤ ~175 words — beyond that instructions conflict and adherence drops. Resolution 720p/1080p; aspect 16:9 / 9:16.
|
|
229
|
+
|
|
230
|
+
Sources: Google Cloud "Ultimate prompting guide for Veo 3.1"
|
|
231
|
+
(cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1),
|
|
232
|
+
KIE VEO API docs (docs.kie.ai/veo3-api/generate-veo-3-video). Captured 2026-08-09.`,
|
|
233
|
+
}
|
|
234
|
+
|
|
235
|
+
const GEMINI_OMNI_DOCTRINE: ProviderPromptDoctrine = {
|
|
236
|
+
providers: ["gemini-omni-video"],
|
|
237
|
+
heading: "Gemini Omni Video (gemini-omni-video)",
|
|
238
|
+
tips: [
|
|
239
|
+
"Multimodal Google video with native audio: text-to-video, image-to-video, and video-edit through the same prompt surface. 4/6/8/10s; 720p/1080p or 4K tier.",
|
|
240
|
+
"Structure like the platform default: subject → action → scene → lighting → camera → style. Quote dialogue lines to have them spoken; describe SFX/ambience plainly in the prompt.",
|
|
241
|
+
"Text-to-video REQUIRES a concrete aspect ratio (the API hard-rejects a missing one); image runs infer aspect from the input.",
|
|
242
|
+
"Reference images ride along as additional imageUrls — bind them in the prompt ('the woman from the first image'). Video-edit: wire a source clip and describe the change, not the whole scene.",
|
|
243
|
+
],
|
|
244
|
+
doctrine: `Prompt structure (no public Google prompt guide exists for the Omni video endpoint —
|
|
245
|
+
the API contract is the doctrine source, like MiniMax H3; structure guidance mirrors the
|
|
246
|
+
platform's ordinal-reference conventions):
|
|
247
|
+
subject → action → scene/environment → lighting → camera movement → style → constraints.
|
|
248
|
+
|
|
249
|
+
**Modes (picked from the wired inputs)**
|
|
250
|
+
- Nothing visual → text-to-video. A concrete aspect ratio is REQUIRED — the API hard-rejects a missing one (Nodaro sends the node's ratio; there is no adaptive).
|
|
251
|
+
- Image(s) wired → image-to-video: the first image anchors the scene; extra images are references — bind each in the prompt ("the woman from the first image", "the interior from the second image").
|
|
252
|
+
- Source video wired → video-edit (served through the same handle): describe the CHANGE ("replace the daylight with dusk, keep the motion and framing"), not a full re-description.
|
|
253
|
+
|
|
254
|
+
**Audio (native)**
|
|
255
|
+
- Audio is generated with the clip. Quote dialogue to have it spoken; describe SFX and ambience plainly ("rain on glass, low synth bed"). State exclusions ("no music") or a bed may be invented.
|
|
256
|
+
|
|
257
|
+
**Duration & tiers**
|
|
258
|
+
- 4 / 6 / 8 / 10 seconds. 720p/1080p tier or the pricier 4K tier — pick 4K only when the deliverable needs it (nearly 2× the credits).
|
|
259
|
+
|
|
260
|
+
Source: KIE gemini-omni-video market contract (parameters + live behavior probed for the
|
|
261
|
+
aspect-ratio hard-reject, see providers/kie/video.ts). Captured 2026-08-09.`,
|
|
262
|
+
}
|
|
263
|
+
|
|
264
|
+
const GROK_IMAGINE_DOCTRINE: ProviderPromptDoctrine = {
|
|
265
|
+
providers: ["grok-i2v", "grok-imagine-video-1.5"],
|
|
266
|
+
heading: "Grok Imagine (grok-i2v, grok-imagine-video-1.5)",
|
|
267
|
+
tips: [
|
|
268
|
+
"Keep prompts simple and direct — Subject + Action + Setting + Camera + Mood. Grok expands the prompt itself; over-specification fights the expander.",
|
|
269
|
+
"Image-to-video: the input image IS the first frame (composition, identity, and style are preserved) — describe the MOTION, don't re-describe the still.",
|
|
270
|
+
"Video 1.5: 1-15s (default 8), 480p/720p/1080p, up to 7 input images (1080p allows only one). Native audio incl. music, SFX, and lip-synced dialogue — quote the line to have it spoken.",
|
|
271
|
+
"Aspect ratio applies to text runs (1:1/16:9/9:16/3:2/2:3/auto); a single input image locks the output to the image's own aspect.",
|
|
272
|
+
],
|
|
273
|
+
doctrine: `Prompt structure (xAI's guidance is minimal by design — the model auto-expands prompts):
|
|
274
|
+
Subject + Action + Setting + Camera + Lighting/Mood, written simply and directly. Reduce
|
|
275
|
+
descriptions of static/unchanged parts — spend the words on what MOVES.
|
|
276
|
+
|
|
277
|
+
**Image-to-video (the primary mode)**
|
|
278
|
+
- The input image is the FIRST FRAME, not a loose reference: composition, subject identity, and visual style carry over. Describe motion and camera only ("she turns toward the window as the camera slowly pushes in"); re-describing the still wastes adherence.
|
|
279
|
+
- grok-imagine-video-1.5 accepts up to 7 images (identity/scene references beyond the first frame); at 1080p only ONE image is allowed.
|
|
280
|
+
|
|
281
|
+
**Audio (video-1.5)**
|
|
282
|
+
- Native audio generates with the clip — background music, SFX, and lip-synced dialogue. Quote the spoken line; describe the music/SFX plainly. There is no audio toggle on the KIE contract — cue (or exclude) sound in the prompt text.
|
|
283
|
+
|
|
284
|
+
**Durations / tiers**
|
|
285
|
+
- grok-i2v: 6 or 10 seconds. grok-imagine-video-1.5: 1-15 seconds in 1s steps (default 8), 480p (default) / 720p / 1080p. Prompt cap 4096 chars — but shorter is better here.
|
|
286
|
+
|
|
287
|
+
Sources: KIE Grok Imagine contracts (docs.kie.ai/market/grok-imagine/image-to-video,
|
|
288
|
+
docs.kie.ai/market/grok-imagine/1-5-preview), xAI Grok Imagine 1.5 release notes
|
|
289
|
+
(x.ai/news/grok-imagine-1-5). Captured 2026-08-09.`,
|
|
290
|
+
}
|
|
291
|
+
|
|
292
|
+
const WAN_DOCTRINE: ProviderPromptDoctrine = {
|
|
293
|
+
providers: ["wan", "wan-i2v", "wan-turbo", "wan-flash", "wan-2.7", "wan-2.7-i2v", "wan-2.7-t2v", "wan-2.7-pro", "wan-videoedit"],
|
|
294
|
+
heading: "Wan 2.x (wan, wan-i2v, wan-turbo, wan-2.7 family)",
|
|
295
|
+
tips: [
|
|
296
|
+
"Alibaba's official formula: Entity + Scene + Motion (basic) → add Aesthetic control + Stylization (advanced). Image-to-video: Motion + Camera only — the image already defines entity and scene.",
|
|
297
|
+
"Sound (2.5+): append a sound description block — voice / sound effects / background music. Avoid scripting EXACT lip-synced lines (official anti-pattern); describe the voice and intent instead.",
|
|
298
|
+
"Multi-shot (2.6/2.7): Overall description + shot number + timestamp + per-shot content. For ONE continuous take write 'Generate single shot' (the shot_type parameter is gone in 2.7).",
|
|
299
|
+
"References go by 'Image 1' / 'Video 1' (capitalized, with a space). Anti-patterns: naming real people, demanding exact legible text, rapid scene changes in one clip, very long choreography.",
|
|
300
|
+
"Style words are strong levers: cyberpunk, claymation, pixel style, felt style, tilt-shift, time-lapse. wan-videoedit: describe the transform, keep motion + identity.",
|
|
301
|
+
],
|
|
302
|
+
doctrine: `Prompt structure (Alibaba Model Studio's official formulas):
|
|
303
|
+
- Basic: Entity + Scene + Motion.
|
|
304
|
+
- Advanced: Entity (description) + Scene (description) + Motion (description) + Aesthetic control + Stylization.
|
|
305
|
+
- Image-to-video: Motion + Camera movement ONLY — the wired image already defines entity and scene; re-describing it fights the frame.
|
|
306
|
+
- Sound (2.5/2.6/2.7): … + Sound description (voice / sound effects / background music).
|
|
307
|
+
- Multi-shot (2.6/2.7): Overall description + Shot number + Timestamp + Shot content.
|
|
308
|
+
- Reference-to-video (2.6/2.7): Reference identifier + Action + Scene + optional Lines + optional BGM.
|
|
309
|
+
|
|
310
|
+
**Camera vocabulary**
|
|
311
|
+
push-in (intimacy/tension), pull-out (scale/isolation), tracking shot, orbit, fixed camera, and compound movements chained sequentially for epic scale.
|
|
312
|
+
|
|
313
|
+
**Single-shot control (2.7)**
|
|
314
|
+
- The shot_type parameter no longer exists — write "Generate single shot" in the prompt to force one continuous take; otherwise 2.7's planner may cut.
|
|
315
|
+
|
|
316
|
+
**References**
|
|
317
|
+
- English format is "Image 1" / "Video 1" (capitalized, space-separated) — bind every wired asset by that name or it may be ignored.
|
|
318
|
+
|
|
319
|
+
**Official anti-patterns (from Alibaba's guide)**
|
|
320
|
+
- Do NOT name specific real people.
|
|
321
|
+
- Do NOT script exact lip-synced dialogue — describe the voice and intent ("she murmurs a reassurance, warm and low") instead of demanding word-perfect lips.
|
|
322
|
+
- Avoid rapid scene changes inside a single clip, very long choreographed sequences, and demands for exactly legible on-screen text.
|
|
323
|
+
|
|
324
|
+
**Stylization**
|
|
325
|
+
- Style words are strong levers: cyberpunk, line-art illustration, felt style, 3D cartoon, pixel style, puppet animation, claymation, black-and-white animation, tilt-shift, time-lapse.
|
|
326
|
+
|
|
327
|
+
Source: Alibaba Cloud Model Studio — "Text-to-video / image-to-video prompt guide"
|
|
328
|
+
(alibabacloud.com/help/en/model-studio/text-to-video-prompt). Captured 2026-08-09.`,
|
|
329
|
+
}
|
|
330
|
+
|
|
331
|
+
const HAPPYHORSE_DOCTRINE: ProviderPromptDoctrine = {
|
|
332
|
+
providers: ["happyhorse", "happyhorse-i2v", "happyhorse-ref2v", "happyhorse-edit"],
|
|
333
|
+
heading: "HappyHorse 1.1 (happyhorse, happyhorse-i2v, happyhorse-ref2v)",
|
|
334
|
+
tips: [
|
|
335
|
+
"Any-language prompts up to 5000 chars (2500 Chinese) — excess is silently truncated, so front-load subject → action → scene → camera → style.",
|
|
336
|
+
"3-15 seconds per second of billing; 720p or 1080p; ratios 16:9 / 9:16 / 1:1 / 4:3 / 3:4. Pick the shortest duration that serves the shot.",
|
|
337
|
+
"ref2v is one of the few true REFERENCE modes on the roster: wired refs keep identity across the clip — bind each reference explicitly in the prompt.",
|
|
338
|
+
"No published vendor style guide — the platform's standard structure applies; keep one camera move per shot and quantify motion physically.",
|
|
339
|
+
],
|
|
340
|
+
doctrine: `Prompt structure (no public HappyHorse prompt guide exists — the KIE API contract is the
|
|
341
|
+
doctrine source; platform-standard structure applies):
|
|
342
|
+
subject → action → scene/environment → lighting → camera movement → style → constraints.
|
|
343
|
+
|
|
344
|
+
**Contract facts (KIE, per-mode pages)**
|
|
345
|
+
- Prompts: any language, up to 5000 non-Chinese / 2500 Chinese characters — excess is TRUNCATED silently, so put the load-bearing content first.
|
|
346
|
+
- Duration 3-15s (default 5), billed per second. Resolution 720p / 1080p (default). Aspect 16:9 (default) / 9:16 / 1:1 / 4:3 / 3:4.
|
|
347
|
+
- Modes: text-to-video (happyhorse), image-to-video (happyhorse-i2v), reference-to-video (happyhorse-ref2v) — ref2v preserves wired identities; name each reference in the prompt so the binding is explicit.
|
|
348
|
+
|
|
349
|
+
**Style guidance (platform-standard, honestly generic)**
|
|
350
|
+
- One camera movement per shot; physical, quantified action ("slowly raises a hand") over abstract emotion words; state exclusions ("no on-screen text, no watermark") in the prompt.
|
|
351
|
+
|
|
352
|
+
Source: KIE HappyHorse 1.1 contracts (docs.kie.ai/market/happyhorse/text-to-video,
|
|
353
|
+
…/happyhorse-1-1/image-to-video, …/happyhorse-1-1/reference-to-video). Captured 2026-08-09.`,
|
|
354
|
+
}
|
|
355
|
+
|
|
356
|
+
const RUNWAY_KIE_DOCTRINE: ProviderPromptDoctrine = {
|
|
357
|
+
providers: ["runway-kie", "runway-extend", "runway-aleph"],
|
|
358
|
+
heading: "Runway via KIE (runway-kie)",
|
|
359
|
+
tips: [
|
|
360
|
+
"Prompt cap is 1800 chars; KIE's own guidance: be specific about subject, action, style, and setting. No native audio — plan sound as a separate pass.",
|
|
361
|
+
"Durations 5 or 10s with a hard trade-off: 10s cannot be 1080p, 1080p cannot exceed 5s — pick per deliverable.",
|
|
362
|
+
"Text runs REQUIRE an aspect ratio (16:9/4:3/1:1/3:4/9:16); image runs IGNORE it — the input image dictates output dimensions.",
|
|
363
|
+
"Image-to-video treats the image as the anchor frame: describe motion and camera, not the still.",
|
|
364
|
+
],
|
|
365
|
+
doctrine: `Prompt structure (KIE contract guidance): "be specific about subject, action, style, and
|
|
366
|
+
setting" — subject → action → scene → camera → style, within the 1800-character cap.
|
|
367
|
+
|
|
368
|
+
**Contract facts (KIE Runway endpoint)**
|
|
369
|
+
- Duration 5 or 10 seconds; quality 720p or 1080p — 10s@1080p does NOT exist (10s forces 720p; 1080p forces 5s). Choose by deliverable: crisp hero shot → 5s/1080p; longer beat → 10s/720p.
|
|
370
|
+
- Text-to-video REQUIRES aspectRatio (16:9 / 4:3 / 1:1 / 3:4 / 9:16). Image-to-video IGNORES aspectRatio — the input image dictates output dimensions.
|
|
371
|
+
- No audio is generated — score/SFX are a separate pass (merge-video-audio / video-sfx downstream).
|
|
372
|
+
|
|
373
|
+
**Style guidance**
|
|
374
|
+
- The image input anchors composition and identity — describe the motion ("she pushes the door open as the camera tracks left"), not the still.
|
|
375
|
+
- Keep one continuous camera idea per clip; front-load the subject and action.
|
|
376
|
+
|
|
377
|
+
Source: KIE Runway contract (docs.kie.ai/runway-api/generate-ai-video). Captured 2026-08-09.`,
|
|
378
|
+
}
|
|
379
|
+
|
|
163
380
|
export const PROVIDER_PROMPT_DOCTRINES: readonly ProviderPromptDoctrine[] = [
|
|
164
381
|
SEEDANCE_2_DOCTRINE,
|
|
165
382
|
KLING_AUDIO_DOCTRINE,
|
|
166
383
|
MINIMAX_H3_DOCTRINE,
|
|
384
|
+
VEO_31_DOCTRINE,
|
|
385
|
+
GEMINI_OMNI_DOCTRINE,
|
|
386
|
+
GROK_IMAGINE_DOCTRINE,
|
|
387
|
+
WAN_DOCTRINE,
|
|
388
|
+
HAPPYHORSE_DOCTRINE,
|
|
389
|
+
RUNWAY_KIE_DOCTRINE,
|
|
167
390
|
]
|
|
168
391
|
|
|
169
392
|
const DOCTRINE_BY_PROVIDER: ReadonlyMap<string, ProviderPromptDoctrine> = new Map(
|
package/src/resolve-prompt.ts
CHANGED
|
@@ -94,7 +94,18 @@ export function computeNodePrompt(
|
|
|
94
94
|
let typed: ReadonlyArray<string | undefined>
|
|
95
95
|
if (nodeType === "text-to-speech") {
|
|
96
96
|
// data.text is a phantom field on TTS; only directText (gated) is real.
|
|
97
|
-
|
|
97
|
+
//
|
|
98
|
+
// The gate is a PREFERENCE, not a lock: when textSource is "connected" we
|
|
99
|
+
// still fall back to typed text if nothing is wired. Writers flip the gate
|
|
100
|
+
// (PromptFieldSpec.promptGate), but data reaches nodes from places no
|
|
101
|
+
// writer touches — workflows saved before that fix, JSON imports, MCP
|
|
102
|
+
// writes, templates — and there the text sat visible in the node while the
|
|
103
|
+
// run failed with "no text found" (founder, 2026-08-14). Coerce rather
|
|
104
|
+
// than reject, same principle as normalizeModelInput.
|
|
105
|
+
typed =
|
|
106
|
+
data.textSource === "direct" || !present(wired)
|
|
107
|
+
? [data.directText as string | undefined]
|
|
108
|
+
: []
|
|
98
109
|
} else {
|
|
99
110
|
const fields = NODE_PROMPT_CANDIDATE_FIELDS[nodeType] ?? ["prompt"]
|
|
100
111
|
typed = fields.map((f) => data[f] as string | undefined)
|
package/src/seedance-2-inputs.ts
CHANGED
|
@@ -12,6 +12,12 @@ export interface Seedance2InputsArgs {
|
|
|
12
12
|
refImageUrls?: readonly string[]
|
|
13
13
|
refVideoUrls?: readonly string[]
|
|
14
14
|
refAudioUrls?: readonly string[]
|
|
15
|
+
/** Per-provider input caps (2026-08-15). The Seedance 2.x GENERATIONS share
|
|
16
|
+
* this resolver's whole mode logic but not their caps — 2.5 takes the same
|
|
17
|
+
* three kinds at 30/10/10 where 2.0 stops at 9/3/3. Defaults to the 2.0
|
|
18
|
+
* caps so every existing caller is byte-identical; the adapter passes the
|
|
19
|
+
* provider's own entry from VIDEO_REF_LIMITS_BY_PROVIDER. */
|
|
20
|
+
limits?: { images: number; videos: number; audio: number }
|
|
15
21
|
}
|
|
16
22
|
|
|
17
23
|
export interface Seedance2InputsResult {
|
|
@@ -54,11 +60,12 @@ export function promptBindsFirstFrame(prompt: string | undefined): boolean {
|
|
|
54
60
|
}
|
|
55
61
|
|
|
56
62
|
export function resolveSeedance2Inputs(args: Seedance2InputsArgs): Seedance2InputsResult {
|
|
63
|
+
const limits = args.limits ?? SEEDANCE_2_REF_LIMITS
|
|
57
64
|
const firstFrameUrl = clean(args.firstFrameUrl)
|
|
58
65
|
const lastFrameUrl = clean(args.lastFrameUrl)
|
|
59
66
|
const refImages = cleanList(args.refImageUrls)
|
|
60
|
-
const refVideos = cleanList(args.refVideoUrls).slice(0,
|
|
61
|
-
const refAudios = cleanList(args.refAudioUrls).slice(0,
|
|
67
|
+
const refVideos = cleanList(args.refVideoUrls).slice(0, limits.videos)
|
|
68
|
+
const refAudios = cleanList(args.refAudioUrls).slice(0, limits.audio)
|
|
62
69
|
|
|
63
70
|
const hasAnyReference = refImages.length > 0 || refVideos.length > 0 || refAudios.length > 0
|
|
64
71
|
|
|
@@ -81,7 +88,7 @@ export function resolveSeedance2Inputs(args: Seedance2InputsArgs): Seedance2Inpu
|
|
|
81
88
|
// if the 9-image cap is exceeded. Frames are appended AFTER the kept user
|
|
82
89
|
// images so existing user @Image ordinals are preserved.
|
|
83
90
|
const frameCount = (firstFrameUrl ? 1 : 0) + (lastFrameUrl ? 1 : 0)
|
|
84
|
-
const userImageSlots = Math.max(0,
|
|
91
|
+
const userImageSlots = Math.max(0, limits.images - frameCount)
|
|
85
92
|
const keptUserImages = refImages.slice(0, userImageSlots)
|
|
86
93
|
const droppedRefImages = refImages.length - keptUserImages.length
|
|
87
94
|
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
import { promptBindsFirstFrame } from "./seedance-2-inputs.js"
|
|
2
|
+
import { identityRefsSentence, REF_BINDING } from "./video-reference-resolver.js"
|
|
3
|
+
|
|
4
|
+
/**
|
|
5
|
+
* VEO 3.x i2v input resolution — the mutually-exclusive sibling of
|
|
6
|
+
* `resolveGeminiOmniI2vInputs`. VEO's API carries ONE `imageUrls` array
|
|
7
|
+
* (≤3) whose meaning flips with `generationType`: plain i2v reads it as
|
|
8
|
+
* [first(, last)] frames; REFERENCE_2_VIDEO reads every entry as a
|
|
9
|
+
* reference ingredient. Frames and identities cannot ride separate
|
|
10
|
+
* channels, so an anchored call that must carry identity references moves
|
|
11
|
+
* to REFERENCE_2_VIDEO with the anchor in seat 1, bound in prose as the
|
|
12
|
+
* opening frame (requested, not pixel-guaranteed — the accepted trade,
|
|
13
|
+
* same as seedance-2's reference mode).
|
|
14
|
+
*
|
|
15
|
+
* REFERENCES WIN THE SEATS (the 2026-08-14 standing rule: refs are a must,
|
|
16
|
+
* frames additional): the end anchor is dropped in reference mode rather
|
|
17
|
+
* than spending one of three seats on a closing guess. The caller logs it.
|
|
18
|
+
*
|
|
19
|
+
* BYTE-IDENTICAL with no references: plain frame mode, frames kept, no
|
|
20
|
+
* generationType, no suffix — exactly what every veo i2v call has always
|
|
21
|
+
* sent.
|
|
22
|
+
*/
|
|
23
|
+
|
|
24
|
+
export interface VeoI2vInputsArgs {
|
|
25
|
+
/** Used only to detect an existing first-frame binding (seedance rule). */
|
|
26
|
+
prompt?: string
|
|
27
|
+
firstFrameUrl: string
|
|
28
|
+
endFrameUrl?: string
|
|
29
|
+
refImageUrls?: Array<string | undefined>
|
|
30
|
+
}
|
|
31
|
+
|
|
32
|
+
export interface VeoI2vInputsResult {
|
|
33
|
+
/** The `imageUrls` payload: frames in plain mode, [anchor, ...refs] in
|
|
34
|
+
* reference mode — never more than VEO's 3-ingredient cap. */
|
|
35
|
+
imageUrls: string[]
|
|
36
|
+
/** Present (REFERENCE_2_VIDEO) exactly when references ride. */
|
|
37
|
+
generationType?: "REFERENCE_2_VIDEO"
|
|
38
|
+
promptSuffix: string
|
|
39
|
+
droppedRefImages: number
|
|
40
|
+
/** True when an end anchor was surrendered to reference mode. */
|
|
41
|
+
droppedEndFrame: boolean
|
|
42
|
+
}
|
|
43
|
+
|
|
44
|
+
/** VEO's ingredient cap — the adapter's REFERENCE_2_VIDEO path has always
|
|
45
|
+
* sliced to 3 (kie/video.ts), mirrored in VIDEO_REF_LIMITS_BY_PROVIDER. */
|
|
46
|
+
const VEO_INGREDIENT_SLOTS = 3
|
|
47
|
+
|
|
48
|
+
export function resolveVeoI2vInputs(args: VeoI2vInputsArgs): VeoI2vInputsResult {
|
|
49
|
+
const refs = (args.refImageUrls ?? []).filter((u): u is string => typeof u === "string" && u.length > 0)
|
|
50
|
+
if (refs.length === 0) {
|
|
51
|
+
return {
|
|
52
|
+
imageUrls: args.endFrameUrl ? [args.firstFrameUrl, args.endFrameUrl] : [args.firstFrameUrl],
|
|
53
|
+
promptSuffix: "",
|
|
54
|
+
droppedRefImages: 0,
|
|
55
|
+
droppedEndFrame: false,
|
|
56
|
+
}
|
|
57
|
+
}
|
|
58
|
+
const kept = refs.slice(0, VEO_INGREDIENT_SLOTS - 1)
|
|
59
|
+
const droppedRefImages = refs.length - kept.length
|
|
60
|
+
const frameSentence = promptBindsFirstFrame(args.prompt) ? "" : REF_BINDING.frame(1, "opening")
|
|
61
|
+
return {
|
|
62
|
+
imageUrls: [args.firstFrameUrl, ...kept],
|
|
63
|
+
generationType: "REFERENCE_2_VIDEO",
|
|
64
|
+
promptSuffix: [frameSentence, identityRefsSentence(2, kept.length + 1)].filter(Boolean).join(" "),
|
|
65
|
+
droppedRefImages,
|
|
66
|
+
droppedEndFrame: Boolean(args.endFrameUrl),
|
|
67
|
+
}
|
|
68
|
+
}
|
|
@@ -50,6 +50,18 @@ import type { ConnectedReference } from "@nodaro/shared"
|
|
|
50
50
|
* the body `{image:N}` tokens through `REF_BINDING[kind]` — so the five arrows
|
|
51
51
|
* are the ONLY emission sites for the binding surface string.
|
|
52
52
|
*/
|
|
53
|
+
/**
|
|
54
|
+
* The identity-reference binding sentence shared by the flat-image-list
|
|
55
|
+
* resolvers (gemini-omni, veo i2v): names the ordinal span as identities and
|
|
56
|
+
* says the two things a multimodal model needs to hear — match exactly, and
|
|
57
|
+
* these are not frames. One spelling; both resolvers ride it.
|
|
58
|
+
*/
|
|
59
|
+
export function identityRefsSentence(firstOrdinal: number, lastOrdinal: number): string {
|
|
60
|
+
return firstOrdinal === lastOrdinal
|
|
61
|
+
? `${REF_BINDING.ordinal(firstOrdinal)} is an identity reference for this shot's subjects — match its subject's exact appearance; it is not a frame.`
|
|
62
|
+
: `${REF_BINDING.ordinal(firstOrdinal)} through ${REF_BINDING.ordinal(lastOrdinal)} are identity references for this shot's subjects — match each subject's exact appearance; they are not frames.`
|
|
63
|
+
}
|
|
64
|
+
|
|
53
65
|
export const REF_BINDING = {
|
|
54
66
|
image: (label: string, n: number) => `the ${label} from @image_${n}`,
|
|
55
67
|
video: (label: string, n: number) => `the ${label} from @video_${n}`,
|