venice-video-harness 2.15.0 → 2.18.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/sd25-pe/SKILL.md +1020 -0
- package/.agents/skills/venice-video-model-routing/SKILL.md +13 -8
- package/AGENTS.md +41 -10
- package/CHANGELOG.md +332 -0
- package/README.md +17 -12
- package/capabilities.json +187 -8
- package/dist/agent/guide.d.ts.map +1 -1
- package/dist/agent/guide.js +12 -2
- package/dist/agent/guide.js.map +1 -1
- package/dist/agent/pipeline.d.ts.map +1 -1
- package/dist/agent/pipeline.js +8 -0
- package/dist/agent/pipeline.js.map +1 -1
- package/dist/mini-drama/assembler.d.ts.map +1 -1
- package/dist/mini-drama/assembler.js +39 -12
- package/dist/mini-drama/assembler.js.map +1 -1
- package/dist/mini-drama/character-reference-generator.d.ts +27 -0
- package/dist/mini-drama/character-reference-generator.d.ts.map +1 -0
- package/dist/mini-drama/character-reference-generator.js +119 -0
- package/dist/mini-drama/character-reference-generator.js.map +1 -0
- package/dist/mini-drama/choices.d.ts +25 -4
- package/dist/mini-drama/choices.d.ts.map +1 -1
- package/dist/mini-drama/choices.js +14 -2
- package/dist/mini-drama/choices.js.map +1 -1
- package/dist/mini-drama/cli.d.ts.map +1 -1
- package/dist/mini-drama/cli.js +976 -144
- package/dist/mini-drama/cli.js.map +1 -1
- package/dist/mini-drama/generation-planner.d.ts.map +1 -1
- package/dist/mini-drama/generation-planner.js +12 -2
- package/dist/mini-drama/generation-planner.js.map +1 -1
- package/dist/mini-drama/location-generator.d.ts +45 -8
- package/dist/mini-drama/location-generator.d.ts.map +1 -1
- package/dist/mini-drama/location-generator.js +261 -87
- package/dist/mini-drama/location-generator.js.map +1 -1
- package/dist/mini-drama/montage.d.ts +64 -0
- package/dist/mini-drama/montage.d.ts.map +1 -0
- package/dist/mini-drama/montage.js +330 -0
- package/dist/mini-drama/montage.js.map +1 -0
- package/dist/mini-drama/panel-fixer.d.ts.map +1 -1
- package/dist/mini-drama/panel-fixer.js +37 -109
- package/dist/mini-drama/panel-fixer.js.map +1 -1
- package/dist/mini-drama/prompt-builder.d.ts +23 -0
- package/dist/mini-drama/prompt-builder.d.ts.map +1 -1
- package/dist/mini-drama/prompt-builder.js +267 -26
- package/dist/mini-drama/prompt-builder.js.map +1 -1
- package/dist/mini-drama/reference-slots.d.ts.map +1 -1
- package/dist/mini-drama/reference-slots.js +44 -7
- package/dist/mini-drama/reference-slots.js.map +1 -1
- package/dist/mini-drama/storyboard-reference-generator.d.ts.map +1 -1
- package/dist/mini-drama/storyboard-reference-generator.js +88 -51
- package/dist/mini-drama/storyboard-reference-generator.js.map +1 -1
- package/dist/mini-drama/timeline-export/davinci-fcpxml.d.ts.map +1 -1
- package/dist/mini-drama/timeline-export/davinci-fcpxml.js +28 -19
- package/dist/mini-drama/timeline-export/davinci-fcpxml.js.map +1 -1
- package/dist/mini-drama/timeline-export/fcpxml.d.ts.map +1 -1
- package/dist/mini-drama/timeline-export/fcpxml.js +32 -10
- package/dist/mini-drama/timeline-export/fcpxml.js.map +1 -1
- package/dist/mini-drama/timeline-export/premiere-xmeml.d.ts.map +1 -1
- package/dist/mini-drama/timeline-export/premiere-xmeml.js +23 -8
- package/dist/mini-drama/timeline-export/premiere-xmeml.js.map +1 -1
- package/dist/mini-drama/timeline-export/probe.d.ts +21 -0
- package/dist/mini-drama/timeline-export/probe.d.ts.map +1 -1
- package/dist/mini-drama/timeline-export/probe.js +30 -0
- package/dist/mini-drama/timeline-export/probe.js.map +1 -1
- package/dist/mini-drama/timeline-export/types.d.ts +8 -0
- package/dist/mini-drama/timeline-export/types.d.ts.map +1 -1
- package/dist/mini-drama/treatment.d.ts.map +1 -1
- package/dist/mini-drama/treatment.js +8 -2
- package/dist/mini-drama/treatment.js.map +1 -1
- package/dist/mini-drama/video-generator.d.ts.map +1 -1
- package/dist/mini-drama/video-generator.js +136 -7
- package/dist/mini-drama/video-generator.js.map +1 -1
- package/dist/mini-drama/video-qa.d.ts +102 -0
- package/dist/mini-drama/video-qa.d.ts.map +1 -0
- package/dist/mini-drama/video-qa.js +230 -0
- package/dist/mini-drama/video-qa.js.map +1 -0
- package/dist/mini-drama/voice-reference.d.ts +25 -0
- package/dist/mini-drama/voice-reference.d.ts.map +1 -1
- package/dist/mini-drama/voice-reference.js +56 -0
- package/dist/mini-drama/voice-reference.js.map +1 -1
- package/dist/mini-drama/workshop.d.ts +15 -1
- package/dist/mini-drama/workshop.d.ts.map +1 -1
- package/dist/mini-drama/workshop.js +61 -4
- package/dist/mini-drama/workshop.js.map +1 -1
- package/dist/series/manager.d.ts +9 -0
- package/dist/series/manager.d.ts.map +1 -1
- package/dist/series/manager.js +18 -1
- package/dist/series/manager.js.map +1 -1
- package/dist/series/types.d.ts +120 -19
- package/dist/series/types.d.ts.map +1 -1
- package/dist/series/types.js +93 -22
- package/dist/series/types.js.map +1 -1
- package/dist/session/status.d.ts +2 -0
- package/dist/session/status.d.ts.map +1 -1
- package/dist/session/status.js +6 -1
- package/dist/session/status.js.map +1 -1
- package/dist/venice/client.d.ts +4 -0
- package/dist/venice/client.d.ts.map +1 -1
- package/dist/venice/client.js +9 -2
- package/dist/venice/client.js.map +1 -1
- package/dist/venice/edit-post.d.ts +21 -0
- package/dist/venice/edit-post.d.ts.map +1 -0
- package/dist/venice/edit-post.js +129 -0
- package/dist/venice/edit-post.js.map +1 -0
- package/dist/venice/generate.d.ts +17 -24
- package/dist/venice/generate.d.ts.map +1 -1
- package/dist/venice/generate.js +20 -22
- package/dist/venice/generate.js.map +1 -1
- package/dist/venice/image-bytes.d.ts +10 -4
- package/dist/venice/image-bytes.d.ts.map +1 -1
- package/dist/venice/image-bytes.js +19 -5
- package/dist/venice/image-bytes.js.map +1 -1
- package/dist/venice/models.d.ts +25 -0
- package/dist/venice/models.d.ts.map +1 -1
- package/dist/venice/models.js +73 -0
- package/dist/venice/models.js.map +1 -1
- package/dist/venice/reference-draft.d.ts +59 -0
- package/dist/venice/reference-draft.d.ts.map +1 -0
- package/dist/venice/reference-draft.js +125 -0
- package/dist/venice/reference-draft.js.map +1 -0
- package/dist/venice/types.d.ts +7 -0
- package/dist/venice/types.d.ts.map +1 -1
- package/dist/venice/video.d.ts +7 -0
- package/dist/venice/video.d.ts.map +1 -1
- package/dist/venice/video.js +6 -1
- package/dist/venice/video.js.map +1 -1
- package/dist/web/events.d.ts +15 -0
- package/dist/web/events.d.ts.map +1 -0
- package/dist/web/events.js +59 -0
- package/dist/web/events.js.map +1 -0
- package/dist/web/jobs.d.ts +74 -0
- package/dist/web/jobs.d.ts.map +1 -0
- package/dist/web/jobs.js +212 -0
- package/dist/web/jobs.js.map +1 -0
- package/dist/web/server.d.ts +13 -0
- package/dist/web/server.d.ts.map +1 -0
- package/dist/web/server.js +400 -0
- package/dist/web/server.js.map +1 -0
- package/dist/web/settings.d.ts +45 -0
- package/dist/web/settings.d.ts.map +1 -0
- package/dist/web/settings.js +151 -0
- package/dist/web/settings.js.map +1 -0
- package/dist/web/state.d.ts +97 -0
- package/dist/web/state.d.ts.map +1 -0
- package/dist/web/state.js +298 -0
- package/dist/web/state.js.map +1 -0
- package/dist/web/ui/dist/assets/index-3mXnkSfB.js +41 -0
- package/dist/web/ui/dist/assets/index-B1szWLlq.js +41 -0
- package/dist/web/ui/dist/assets/index-B8GV_ORG.css +1 -0
- package/dist/web/ui/dist/assets/index-BhVDUhlN.js +41 -0
- package/dist/web/ui/dist/assets/index-BlWGeQRh.js +41 -0
- package/dist/web/ui/dist/assets/index-BnjJa6vV.js +41 -0
- package/dist/web/ui/dist/assets/index-BqCWjC77.css +1 -0
- package/dist/web/ui/dist/assets/index-BzZTKlyH.css +1 -0
- package/dist/web/ui/dist/assets/index-C5VCsEeF.js +41 -0
- package/dist/web/ui/dist/assets/index-CDliCUHR.js +40 -0
- package/dist/web/ui/dist/assets/index-CJl8zd9Y.js +41 -0
- package/dist/web/ui/dist/assets/index-CP2etMzp.js +41 -0
- package/dist/web/ui/dist/assets/index-C_z8KiSR.js +41 -0
- package/dist/web/ui/dist/assets/index-D5jahF05.js +41 -0
- package/dist/web/ui/dist/assets/index-D6I39LuH.css +1 -0
- package/dist/web/ui/dist/assets/index-D6recjFU.css +1 -0
- package/dist/web/ui/dist/assets/index-DjrrcmCm.js +41 -0
- package/dist/web/ui/dist/assets/index-Dy7Nb9jE.js +41 -0
- package/dist/web/ui/dist/assets/index-F8cUuofv.js +41 -0
- package/dist/web/ui/dist/assets/index-NJEUerGY.js +41 -0
- package/dist/web/ui/dist/assets/index-Tt_Xu3go.js +41 -0
- package/dist/web/ui/dist/index.html +14 -0
- package/dist/web/watcher.d.ts +14 -0
- package/dist/web/watcher.d.ts.map +1 -0
- package/dist/web/watcher.js +77 -0
- package/dist/web/watcher.js.map +1 -0
- package/package.json +6 -2
|
@@ -0,0 +1,1020 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: sd25-pe
|
|
3
|
+
description: Use when a user asks an Agent to optimize text, stories, or optional multimodal references for Seedance 2.5 text-to-video, multi-reference generation, keyframes, storyboards, blockouts, video editing, audio editing, or extension.
|
|
4
|
+
metadata:
|
|
5
|
+
skill_version: 0.3.3
|
|
6
|
+
owner: seedance
|
|
7
|
+
tags:
|
|
8
|
+
- seedance
|
|
9
|
+
- prompt-optimization
|
|
10
|
+
- multimodal-video
|
|
11
|
+
supported_runtimes: []
|
|
12
|
+
required_capabilities:
|
|
13
|
+
filesystem_read: false
|
|
14
|
+
filesystem_write: false
|
|
15
|
+
tool_use: false
|
|
16
|
+
network: false
|
|
17
|
+
binary_outputs: false
|
|
18
|
+
io_contract:
|
|
19
|
+
output_kind: text
|
|
20
|
+
primary_outputs:
|
|
21
|
+
- optimized_prompt
|
|
22
|
+
exports: []
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
# Seedance 2.5 Prompt Optimizer
|
|
26
|
+
|
|
27
|
+
## Purpose
|
|
28
|
+
|
|
29
|
+
Compile the user's raw text, novel excerpts, and optional image, video, and audio references into one clean Prompt that can be submitted directly to Seedance 2.5. Preserve the user's core intent while making material roles, subject mappings, event states, and relationships to preserve explicit.
|
|
30
|
+
|
|
31
|
+
This Skill's responsibility ends with understanding the input and producing the Prompt. It must never invoke a generation API by itself. When the user explicitly asks to generate a video, first use this Skill to produce a clean Prompt, then pass it to a separate generation tool or workflow. This Skill always treats reference materials as read-only: it does not modify source materials or automatically create auxiliary materials.
|
|
32
|
+
|
|
33
|
+
## When to Use
|
|
34
|
+
|
|
35
|
+
Load this Skill when:
|
|
36
|
+
|
|
37
|
+
- The user asks to optimize, rewrite, or complete a Seedance 2.5 Prompt.
|
|
38
|
+
- The user provides a rough idea, long-form plot, novel excerpt, or unstructured complex Prompt.
|
|
39
|
+
- The user provides images, videos, audio, file paths, or another multimodal request that requires material-role mapping.
|
|
40
|
+
- The user wants multi-reference generation, a long video, keyframe control, a storyboard grid, blockout rendering, video editing, audio editing, or video extension.
|
|
41
|
+
- The user wants to improve emotional performance, camera movement, audio, dialogue language, or a product/process demonstration.
|
|
42
|
+
|
|
43
|
+
Do not load this Skill automatically when:
|
|
44
|
+
|
|
45
|
+
- The user only asks about API parameters, pricing, quota, errors, or model capabilities and does not need a Prompt.
|
|
46
|
+
- The user only asks to evaluate a generated video and does not ask for a rewritten Prompt.
|
|
47
|
+
|
|
48
|
+
When the user explicitly asks to generate a video, still compile the raw input into a clean Prompt with this Skill before passing it to the downstream generation workflow. This Skill does not submit the generation task itself.
|
|
49
|
+
|
|
50
|
+
## Non-Negotiable Principles
|
|
51
|
+
|
|
52
|
+
1. **Intent first:** Do not change character identities or counts, key props, scenes, event causality, spatial relationships, the edit target, extension direction, or story outcome.
|
|
53
|
+
2. **Template first:** Reorganize every request with the template for the current task, even when the original Prompt is already complete. Do not merely paraphrase it.
|
|
54
|
+
3. **One explicit role per material:** State what to use from every material that is actually activated. List every available but unassigned material individually under `[Unused Materials]` so downstream PE does not reactivate it.
|
|
55
|
+
4. **User mappings first:** Never override material roles, subject names, or relationships explicitly assigned by the user. Once a slot is covered, do not add an unmentioned material merely to reinforce the same role.
|
|
56
|
+
5. **Ask as little as possible:** Resolve anything that can be reasonably inferred from the text, materials, and context. Ask one consolidated question only when ambiguity could change the core result and several equally plausible interpretations remain.
|
|
57
|
+
6. **Submit-ready content only:** The Prompt must not contain analysis, evaluation notes, experiment groups, model versions, run labels, API keys, or reasons for the rewrite.
|
|
58
|
+
7. **Separate parameters from creative content:** Do not write aspect ratio, total duration, resolution, frame rate, or the audio toggle into the Prompt. For ordinary generation, set them on the generation page or through the API. For video editing, first-frame or first-and-last-frame generation, and video extension, follow the task's automatic parameter-locking rules. Time ranges already written by the user are creative content and must be preserved; do not invent new numeric time ranges merely from a target duration supplied by the page or API.
|
|
59
|
+
8. **Do not add blanket constraints:** Do not automatically add unrequested quality or stability boilerplate, watermarks, logos, subtitles, duplicate-subject restrictions, or other generic negative constraints.
|
|
60
|
+
9. **Return one best version:** By default, output only the single Prompt judged most appropriate. Return multiple versions only when the user explicitly asks for a comparison.
|
|
61
|
+
10. **Separate facts from observations:** Character identity, age, relationships, events, and outcomes come from the user's text. Materials may contribute only directly visible or audible attributes; never promote a visual guess into a story fact.
|
|
62
|
+
11. **Match subject cardinality:** A single-person character sheet must not define two named characters who appear at the same time. When several single-person candidates are available, first assign one image per person; when several group candidates are available, first assign one image per group. Combine references only for multiple views of the same subject or when the material itself clearly contains the same story group.
|
|
63
|
+
|
|
64
|
+
## Input States
|
|
65
|
+
|
|
66
|
+
Determine which input state applies before processing the content.
|
|
67
|
+
|
|
68
|
+
### Text-Only Video Generation
|
|
69
|
+
|
|
70
|
+
The user provides no materials, and the text does not express a need to reference a particular person, product, scene, action, camera movement, or sound. Extract the subject, event, scene, visual treatment, camera treatment, and audio, then apply the video-generation template. Do not invent material numbers or suggest that the user add materials.
|
|
71
|
+
|
|
72
|
+
### A Reference Is Required but Missing
|
|
73
|
+
|
|
74
|
+
The user explicitly asks to "reference this character," "preserve the source video," "replace it with this product," or makes an equivalent request, but the corresponding material or location is not available. Still produce the best Prompt supported by the text; do not block on the missing material. Outside the Prompt, add at most one suggestion explaining that the mapping can be made more explicit after the material is provided.
|
|
75
|
+
|
|
76
|
+
Compare every material reference in the Prompt against the actual available-material list. If even one explicit reference is missing, treat that reference as "required but missing" even when other materials are available. Complete the Prompt with the remaining valid materials, but do not claim to have seen the missing material, guess its appearance, voice, or content, or continue referencing its nonexistent number in the Prompt. Preserve the character, event, or dialogue role that the missing material was meant to support using the user's text; omit only appearance, voice, action, or scene details that could be established solely by that material.
|
|
77
|
+
|
|
78
|
+
### Materials Are Readable
|
|
79
|
+
|
|
80
|
+
Actively inspect the images, videos, and audio supplied by the user. Inventory all materials first, then establish mappings in combination with the Prompt. Do not infer content from filenames alone, and do not force every material into the Prompt.
|
|
81
|
+
|
|
82
|
+
### The Current Agent Cannot Read the Materials
|
|
83
|
+
|
|
84
|
+
The user provides an attachment, path, or URL that the current runtime cannot access. Do not pretend to have inspected it. Continue optimizing everything that can be confirmed from the text, and identify which materials could not be read. Request access or confirmation only when the inaccessible item is the sole core subject reference, sole editing master, or sole source video for extension.
|
|
85
|
+
|
|
86
|
+
Skip and list a damaged or unreadable non-core material by number; continue processing the rest of the request.
|
|
87
|
+
|
|
88
|
+
### Material-Specification Preflight
|
|
89
|
+
|
|
90
|
+
Distinguish hard input limits from stability recommendations:
|
|
91
|
+
|
|
92
|
+
- Up to 30 images, each no larger than 4K.
|
|
93
|
+
- Up to 10 videos, with a combined duration of no more than 30 seconds.
|
|
94
|
+
- Up to 10 audio clips, with a combined duration of no more than 30 seconds.
|
|
95
|
+
- Up to 50 reference materials across images, videos, and audio combined.
|
|
96
|
+
|
|
97
|
+
Recommended ranges are not hard limits. Prefer 1-8 distinct subjects across subject-reference images; 1-5 distinct subjects across subject audio/video, with 5-10 seconds per clip; and, for video editing, a source video under 20 seconds with 1-5 reference images. When the input exceeds a recommendation but remains within the hard limits, continue optimizing rather than blocking because there are many materials. Reduce cross-contamination through one-role-per-material declarations, subject mapping, and scene-based activation.
|
|
98
|
+
|
|
99
|
+
When the input exceeds a hard limit, prioritize materials explicitly selected by the user, materials that cover required entities, the sole editing master or extension source video, and critical action, audio, or boundary-frame materials. Omit the rest from the Prompt, then append one `Material Note:` after the Prompt explaining which material type or numbers must be reduced before submission. Never claim that Prompt wording can bypass an input limit.
|
|
100
|
+
|
|
101
|
+
## Core Workflow
|
|
102
|
+
|
|
103
|
+
### 1. Parse the User's Goal
|
|
104
|
+
|
|
105
|
+
Build a "story contract" from the user's text before inspecting the materials. Extract and lock:
|
|
106
|
+
|
|
107
|
+
- Subjects and subject count.
|
|
108
|
+
- Actions, events, and causal order.
|
|
109
|
+
- Scene, time, weather, and spatial relationships.
|
|
110
|
+
- Prop ownership, handoffs, and final states.
|
|
111
|
+
- Visual style, camera, audio, dialogue, and subtitle requirements.
|
|
112
|
+
- Content the user explicitly requires or excludes.
|
|
113
|
+
- Whether the primary task is generation, editing, or extension; whether generation uses keyframes, a storyboard grid, or a blockout; and whether an editing task modifies audio only.
|
|
114
|
+
|
|
115
|
+
The story contract is the factual boundary for every later rewrite. Do not replace, merge, or delete a character or event from the story contract merely because a material contains a more visually prominent person, outfit, prop, or scene.
|
|
116
|
+
|
|
117
|
+
Preserve the meaning of the user's notation. `Shot 45`, `shot 45`, and equivalent forms refer to a shot number by default, not a `45-degree camera angle`. Interpret the number as a camera angle only when the user explicitly writes "45 degrees" or an equivalent photographic angle. Never rewrite a shot number, material number, chapter number, or step number as a camera parameter.
|
|
118
|
+
|
|
119
|
+
Whenever the input contains speech, dialogue, narration, or an audio role, build an internal "dialogue ledger" for every stage. Record the speaker, whether the speaker is vocalizing, the exact dialogue, the audio role, the language, and whether the voice is on-screen or off-screen. Do not turn a segment with a specified speaker or audio role into silence, and do not swap speakers or audio references. If the user explicitly requires a character to speak with a bound audio reference but supplies no dialogue text, or the current environment cannot transcribe it reliably, preserve that character's responsibility to speak using the reference audio without inventing dialogue.
|
|
120
|
+
|
|
121
|
+
When the user provides only a quoted fragment, keyword, or short phrase, treat only that fragment as speakable text. You may add the speaker's expression, action, and delivery, but never invent a complete sentence before or after it. When the user describes only a speaking intention such as "assert authority" or "continue arguing" without providing any exact words, do not invent dialogue. Express the intention through observable mouth movement, pauses, posture, and the other character's reaction.
|
|
122
|
+
|
|
123
|
+
For example, if the input says only that she says "You're an outsider" with an intimidating sense of status, the final dialogue may contain only `{You're an outsider.}` Do not splice narrative explanations such as "family relationship" or "status pressure" into the line, and do not expand it into "You're an outsider. What right do you have..." or any other full sentence.
|
|
124
|
+
|
|
125
|
+
Create a "required-entity checklist" for the story contract, with one slot for every explicitly present character, group, key prop, and scene. This checklist is for internal verification only and must not be shown to the user. Record each slot's story role, activated materials, and observable attributes, then confirm before output that every slot appears in the Prompt.
|
|
126
|
+
|
|
127
|
+
Translate internal thoughts into visible actions, expressions, dialogue, or narration without adding a story-changing event.
|
|
128
|
+
|
|
129
|
+
For a causal-reveal shot in which a character changes expression or behavior only after a cause appears, show the trigger clearly before showing the reaction. When trigger and reaction are spatially separated, connect them with one explicit shift of gaze or camera. When both can fit in one frame, preserve the trigger and reaction in the same composition. Do not push into a face-only close-up before the audience can see the trigger, and do not show only the reaction while omitting its cause.
|
|
130
|
+
|
|
131
|
+
When the plot explicitly depicts pretending to be hurt, simulated injury, or a near miss, state an observable uninjured condition in any ambiguous shot: for example, intact skin, clothing, and props. Do not rewrite acting or an event that did not occur as a real wound, bleeding, or damage.
|
|
132
|
+
|
|
133
|
+
### 2. Compile Novels and Long-Form Text
|
|
134
|
+
|
|
135
|
+
First convert the text into filmable events:
|
|
136
|
+
|
|
137
|
+
- When the user asks for a trailer, overview, or ensemble piece, choose a montage structure that summarizes the theme.
|
|
138
|
+
- When the user asks for one scene, preserve the core events that occur continuously within that scene.
|
|
139
|
+
- When the user does not specify a range, select one causally complete main event that can fit in the target video, and briefly disclose the selection outside the Prompt.
|
|
140
|
+
- Ask one consolidated question only when several mutually exclusive main threads are equally important and the choice would change the core story.
|
|
141
|
+
|
|
142
|
+
Compress repeated descriptions and information that cannot be represented on screen. Preserve character relationships, key dialogue, trigger events, and the ending state.
|
|
143
|
+
|
|
144
|
+
### 3. Inventory and Understand the Materials
|
|
145
|
+
|
|
146
|
+
For a large material set, use a two-pass review:
|
|
147
|
+
|
|
148
|
+
1. In the first pass, inventory every material lightly and identify candidates for characters, products, props, scenes, actions, camera movement, pacing, audio, and style.
|
|
149
|
+
2. In the second pass, inspect in depth only materials that match the story, conflict with another candidate, act as keyframes, or will be activated in the current scene.
|
|
150
|
+
|
|
151
|
+
When inspecting a video, confirm at minimum its subjects, main action, camera changes, opening state, and ending state. When inspecting audio, confirm at minimum its sound type, voice characteristics, language, dialogue content, or ambience role.
|
|
152
|
+
|
|
153
|
+
Separate two kinds of information:
|
|
154
|
+
|
|
155
|
+
- **Story assignment:** A character's name, age, relationships, event role, prop ownership, and outcome. These come from the user's text.
|
|
156
|
+
- **Material observation:** Directly visible or audible attributes such as facial features, hairstyle, clothing, material, color, spatial layout, action, camera movement, and voice characteristics.
|
|
157
|
+
|
|
158
|
+
Materials may supply only directly observable attributes. Do not infer or rewrite identities, ages, relationships, or plot from appearance. Do not invent a brand, color, profession, personality, or prop function that the material does not establish.
|
|
159
|
+
|
|
160
|
+
When the user calls someone only a "person" or "subject," retain that neutral term. Do not rename the subject a dancer, actor, worker, or another identity based on action, posture, or clothing. Reuse a character name or role only when the user has supplied it.
|
|
161
|
+
|
|
162
|
+
Forms of address, rank, profession, and relationships define story slots; they do not automatically imply a young, old, vulnerable, dominant, or otherwise characterized appearance. Include such attributes only when the user's text states them explicitly or when the material shows a directly observable manifestation.
|
|
163
|
+
|
|
164
|
+
By default, describe only facial features, hairstyle, clothing, accessories, and directly observable posture from a character reference. Do not replace observable detail with summary labels such as "girlish," "mature," "fragile," or "dominant." When the user explicitly requests those performance directions, translate them into expressions, posture, and actions.
|
|
165
|
+
|
|
166
|
+
Use this fixed mapping priority:
|
|
167
|
+
|
|
168
|
+
```text
|
|
169
|
+
User's explicit assignment > Prompt description > Material content > Filename and metadata > Upload order
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
When the user's explicit mappings cover all required entities, leave every other available but unmentioned material inactive by default. Do not infer a second character, second scene, or auxiliary reference merely because its content is similar, and do not add an unmentioned material to reinforce an already covered role. Generic background guests, furniture, decoration, and environmental elements do not create a material gap; the selected scene material and text description cover them. Evaluate an unmentioned material only when the user explicitly asks the Agent to select from all materials, asks to combine multiple references, or leaves a required core character, prop, scene, action, or sound without a source. Still list all default-inactive numbers under `[Unused Materials]` in the Prompt.
|
|
173
|
+
|
|
174
|
+
Upload order is not semantic evidence of identity or role; it normally serves only to create stable numbering. Use upload order as a final stable tie-breaker only when the user requests a direct answer, all candidates are equally plausible, and the number of story slots equals the number of candidates. Pair them one to one according to the first appearance of story slots and the order of candidate materials. This ordering resolves assignment stability only; it does not prove identity and must not be used to add an unconfirmed age, relationship, or personality.
|
|
175
|
+
|
|
176
|
+
If the input is JSON or long-form text containing Asset IDs, assign `@Image N`, `@Video N`, and `@Audio N` in each media type's order of appearance, then replace the Asset IDs in the text with those references. The final Prompt must not expose raw Asset IDs.
|
|
177
|
+
|
|
178
|
+
If only local paths are available and no reference tags exist, assign stable aliases by media type and in the order provided by the user. Under "Material Understanding," list each path and alias. Use those aliases in the final Prompt and remind the downstream uploader to preserve the same order.
|
|
179
|
+
|
|
180
|
+
### 4. Establish Material Roles and Subject Mappings
|
|
181
|
+
|
|
182
|
+
Every activated material must have one explicit role:
|
|
183
|
+
|
|
184
|
+
- Images: character appearance and clothing, product structure and material, props, scene layout, lighting, or keyframes.
|
|
185
|
+
- Videos: actions, camera movement, pacing, timeline, or the sole editing master or source video for extension.
|
|
186
|
+
- Audio: a specified speaker's voice and dialogue, ambience, sound effects, or music.
|
|
187
|
+
|
|
188
|
+
Before mapping, check every character, prop, and scene required by the story. Then follow this order:
|
|
189
|
+
|
|
190
|
+
1. Take the next unassigned character, group, prop, or scene from the required-entity checklist.
|
|
191
|
+
2. Review all material candidates and compare subject count, clothing layers, silhouette, structure, props, and scene role. Do not stop after finding the first usable material.
|
|
192
|
+
3. Select the best match for the current slot, then move to the next slot. Different named characters or groups should use different best-matching candidates by default.
|
|
193
|
+
4. Perform an omission audit: every required entity must have exactly one explicit role in the Prompt, and every activated material must perform only its declared role.
|
|
194
|
+
5. Subtract assigned materials from the complete available-material list. If the remainder is nonempty, add `[Unused Materials]` immediately after the material-role section. List every unused number by media type and state explicitly that these materials do not define people, scenes, props, actions, camera treatment, or audio. Do not explain unused materials only outside the Prompt, and do not replace explicit numbers with "other materials."
|
|
195
|
+
|
|
196
|
+
Do not merge two named characters into one subject or omit a character because its material is less visually prominent. When distinguishable candidates are available, do not reuse one material for several named characters or groups. Reuse is allowed only when the material itself clearly contains those characters together and no better independent candidates exist. When several materials jointly define one entity, say so explicitly. Materials not invoked by the story may remain unused.
|
|
197
|
+
|
|
198
|
+
In the final material-role section, each line must define exactly one subject, group, prop, or scene and its primary reference. Do not compress multiple subjects and materials into a range mapping such as "Characters A and B reference @Images 1 and 2." Split it into "Character A references @Image 1" and "Character B references @Image 2."
|
|
199
|
+
|
|
200
|
+
If two core identities have no appearance clues in the text and their candidate materials are equally plausible, do not pretend that the hidden answer can be inferred from the images. Ask one consolidated question under "Handle Mapping Confidence." If the user requests a direct output based on a reasonable assumption, apply the stable tie-breaker above to create a conservative one-to-one mapping, and avoid adding unconfirmed differences such as age or personality.
|
|
201
|
+
|
|
202
|
+
When one subject has several references, state the view or attribute contributed by each material and make clear that all of them define one entity rather than multiple copies.
|
|
203
|
+
|
|
204
|
+
Bind distinct subjects one by one, for example:
|
|
205
|
+
|
|
206
|
+
```text
|
|
207
|
+
<Character A> corresponds to @Image 1. Use only the facial features, hairstyle, and clothing.
|
|
208
|
+
<Character B> corresponds to @Image 2. Use only the facial features, hairstyle, and clothing.
|
|
209
|
+
<Prop A> corresponds to @Image 3. Use only the structure, material, and color.
|
|
210
|
+
<Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.
|
|
211
|
+
|
|
212
|
+
[Unused Materials]
|
|
213
|
+
@Images 5 and 6 are not used in this task and must not define people, scenes, props, actions, or camera treatment.
|
|
214
|
+
@Audio 2 is not used in this task and must not define dialogue, voice characteristics, ambience, sound effects, or music.
|
|
215
|
+
```
|
|
216
|
+
|
|
217
|
+
Do not replace one-to-one character mappings with a range sentence.
|
|
218
|
+
|
|
219
|
+
When a reference video's role is unspecified, select only the task-relevant dimensions from action, camera movement, pacing, scene, and audio. A generation task must not inherit the video's character identity, clothing, or entire scene by default. When the reference video already defines the action, camera movement, and sequence accurately, state only which dimensions to inherit; there is no need to restate every action. Repeating the action may conflict with the reference itself.
|
|
220
|
+
|
|
221
|
+
When written dialogue conflicts with the content of reference audio, the user's text controls the words. By default, the audio supplies only voice characteristics, accent, speed, and emotion. The exception is when the user explicitly asks to reuse the audio's dialogue.
|
|
222
|
+
|
|
223
|
+
When several references for the same subject conflict, first follow the user's assignment. If the user has not specified one, assign appearance, clothing, structure, or material to the best reference based on clarity and story fit. Ask for clarification only when the core identity still cannot be determined.
|
|
224
|
+
|
|
225
|
+
### 5. Handle Mapping Confidence
|
|
226
|
+
|
|
227
|
+
- **High confidence:** Map automatically and continue.
|
|
228
|
+
- **Medium confidence:** Use the most reasonable mapping and disclose the key assumption under "Material Understanding" outside the Prompt.
|
|
229
|
+
- **Low confidence without core impact:** Do not use the material and do not ask the user.
|
|
230
|
+
- **Low confidence affecting core identity, count, prop ownership, front/back, left/right, facing direction, edit target, sole master video, extension direction, or keyframe role:** Consolidate the related ambiguity into one short question.
|
|
231
|
+
|
|
232
|
+
If a user-defined mapping visibly conflicts with material content, still follow the user's mapping. Confirm once only when the conflict would almost certainly produce the wrong result.
|
|
233
|
+
|
|
234
|
+
Do not interrupt the user over missing style, lighting, ordinary camera movement, image quality, or other details that can be decided conservatively from context.
|
|
235
|
+
|
|
236
|
+
When clarification is required, ask one consolidated question and stop. Do not also output a temporary Prompt that could mislead downstream submission. After the user answers, resume from mapping and task routing. Adopt an assumption and disclose it outside the Prompt only when the user explicitly asks for a version based on a reasonable assumption.
|
|
237
|
+
|
|
238
|
+
If the user provides materials without any generation, editing, or extension goal, ask once for the minimum creative goal. Do not invent a story from the materials.
|
|
239
|
+
|
|
240
|
+
### 6. Select One Primary Task
|
|
241
|
+
|
|
242
|
+
Select exactly one of the three primary tasks. First determine whether the request modifies an existing video. If not, determine whether it generates content before or after an existing video. All other requests are generation:
|
|
243
|
+
|
|
244
|
+
1. **Video editing:** Use one source video as the sole editing master and modify only specified objects, regions, or audio.
|
|
245
|
+
2. **Video extension:** Generate a new continuous segment before or after the source video without rewriting the original segment.
|
|
246
|
+
3. **Video generation:** Generate a new video from text and optional reference materials.
|
|
247
|
+
|
|
248
|
+
Multi-reference organization, long videos, time ranges, keyframes, storyboard grids, blockouts, audio, emotion, and cinematography are composable modules, not new primary tasks. Re-rendering a blockout as a reference for action, space, or complete structure is video generation; the presence of a video input does not automatically route it to video editing.
|
|
249
|
+
|
|
250
|
+
When the user requests both editing and extension, do not discard either operation or force both into one Prompt:
|
|
251
|
+
|
|
252
|
+
- If replaced content must continue from the original video into the new segment, edit the source video first to create a new master, then extend the edited master.
|
|
253
|
+
- If a new subject appears only in the extended segment and does not modify the source video, perform only extension and define the new material's role for the extended segment.
|
|
254
|
+
- If several equally plausible operation orders remain, ask one consolidated question.
|
|
255
|
+
|
|
256
|
+
When two sequential operations are clearly required, output a two-step execution Prompt. They are consecutive steps of one request, not alternative versions.
|
|
257
|
+
|
|
258
|
+
### 7. Apply Task-Parameter Rules
|
|
259
|
+
|
|
260
|
+
These rules are for planning and notes only. Do not write them into the final Prompt:
|
|
261
|
+
|
|
262
|
+
- **Ordinary video generation:** Aspect ratio and total duration are set on the generation page or through the API. They may inform composition and event density.
|
|
263
|
+
- **Video editing:** Editing automatically locks the input video's aspect ratio and approximate duration; neither can be set separately. Input-frame processing may cause the output duration to differ from the source by up to approximately 0.3 seconds.
|
|
264
|
+
- **First-frame or first-and-last-frame generation:** The output aspect ratio is locked to the first image, while duration can be set. The first and last images should use the same aspect ratio. When the current Agent can read the images, verify the ratios proactively; when it cannot, do not pretend they were checked.
|
|
265
|
+
- **Video extension:** Extension automatically locks the input video's aspect ratio, while extension duration can be set.
|
|
266
|
+
|
|
267
|
+
When the user asks to set a parameter automatically locked by the task, do not ask a question, write the request into the Prompt, or use it for incorrect planning. Still produce the best Prompt first, then append at most one `Parameter Note:` after it. Use the same note when the first and last images have mismatched aspect ratios, explaining that they should be adjusted to the same ratio before submission to avoid stretching the last frame. Output no parameter note when there is no conflict.
|
|
268
|
+
|
|
269
|
+
### 8. Apply the Template and Clean the Prompt
|
|
270
|
+
|
|
271
|
+
Select the corresponding template below and retain only the sections required by the current task. Replace every `<placeholder>` with concrete content; never leave template instructions in the final Prompt.
|
|
272
|
+
|
|
273
|
+
Afterward, run the "Final Checklist" and deliver according to the "Output Contract."
|
|
274
|
+
|
|
275
|
+
## Video Generation Templates
|
|
276
|
+
|
|
277
|
+
### Basic Generation
|
|
278
|
+
|
|
279
|
+
Use this for text-only requests or simple events with few references:
|
|
280
|
+
|
|
281
|
+
```text
|
|
282
|
+
<Subject> performs <primary action or event> in <scene and environment>.
|
|
283
|
+
The visuals feature <visual style or emotion>.
|
|
284
|
+
Use <shot size, camera angle, camera movement, or cuts>.
|
|
285
|
+
Audio includes <dialogue, ambience, sound effects, or music>.
|
|
286
|
+
```
|
|
287
|
+
|
|
288
|
+
Delete any visual, camera, or audio line that the task does not need, but always make the subject and primary action or event explicit.
|
|
289
|
+
|
|
290
|
+
Example:
|
|
291
|
+
|
|
292
|
+
```text
|
|
293
|
+
A young ceramic artist throws clay in a studio at dawn, steadily supporting the spinning clay with both hands until it becomes a narrow-necked vase.
|
|
294
|
+
Soft morning light enters through the window on the left. The wooden table and clay retain natural warm tones.
|
|
295
|
+
Begin with a medium shot of the artist's hands, then slowly push in toward the mouth of the vase.
|
|
296
|
+
Retain the low hum of the wheel, the sound of palms rubbing wet clay, and distant birds outside the window.
|
|
297
|
+
```
|
|
298
|
+
|
|
299
|
+
### Generation with Reference Materials
|
|
300
|
+
|
|
301
|
+
```text
|
|
302
|
+
[Generation Goal]
|
|
303
|
+
Generate <video type or core event>. The central subject is <subject>, and the primary event is <summary>.
|
|
304
|
+
|
|
305
|
+
[Reference Material Roles]
|
|
306
|
+
@Image 1 defines <subject>'s <appearance, clothing, structure, or material>.
|
|
307
|
+
@Video 1 defines <action, camera movement, or pacing>. Do not use <identity, clothing, or scene likely to carry over unintentionally>.
|
|
308
|
+
@Audio 1 defines <character or sound type>'s <voice characteristics, dialogue, ambience, or music>.
|
|
309
|
+
|
|
310
|
+
[Unused Materials]
|
|
311
|
+
@Image 2, @Video 2, and @Audio 2 are not used in this task and must not define people, scenes, props, actions, camera treatment, or audio.
|
|
312
|
+
|
|
313
|
+
[Subjects and Relationships]
|
|
314
|
+
<Subject A> corresponds to @Image 1 and always retains <fixed attributes>.
|
|
315
|
+
The spatial, prop, or identity relationship between <Subject A> and <Subject B> is <relationship>.
|
|
316
|
+
|
|
317
|
+
[Event Script]
|
|
318
|
+
Opening state: <state of characters, props, and scene>.
|
|
319
|
+
Primary event: <continuous action or event>.
|
|
320
|
+
Ending state: <character positions, prop ownership, or final visible state>.
|
|
321
|
+
|
|
322
|
+
[Maintain Consistency]
|
|
323
|
+
Keep <character identities and count, clothing, prop ownership, spatial direction, and audio relationships> consistent.
|
|
324
|
+
```
|
|
325
|
+
|
|
326
|
+
Do not invent visual style, camera movement, or audio merely to fill the template. For a simple task, adjacent sections may be combined, but the role of every activated material must remain explicit.
|
|
327
|
+
|
|
328
|
+
Example:
|
|
329
|
+
|
|
330
|
+
```text
|
|
331
|
+
[Generation Goal]
|
|
332
|
+
Generate a video of a carpenter repairing an old wooden chair. The carpenter first inspects the loose backrest, then applies wood glue and secures the joint. At the end, the chair is stable again.
|
|
333
|
+
|
|
334
|
+
[Reference Material Roles]
|
|
335
|
+
@Image 1 defines the carpenter's facial features, short hair, and dark blue work apron. Do not use the image background.
|
|
336
|
+
@Image 2 defines the old chair's curved backrest, dark wood grain, and worn areas. Do not use the person in the image.
|
|
337
|
+
@Video 1 defines the hand movements for applying glue and pressing the joint closed. Do not use the person's identity, clothing, or workbench from the video.
|
|
338
|
+
|
|
339
|
+
[Subjects and Relationships]
|
|
340
|
+
The carpenter always wears the dark blue apron defined by @Image 1. The entire video contains only one old wooden chair as defined by @Image 2. The tools remain on the right side of the wooden worktable.
|
|
341
|
+
|
|
342
|
+
[Event Script]
|
|
343
|
+
At the start, the chair is centered on the worktable and the backrest joint is loose. After inspecting the joint, the carpenter applies wood glue and presses the backrest into place with both hands. At the end, the carpenter releases both hands, the backrest remains secure, and the chair's count and appearance remain unchanged.
|
|
344
|
+
|
|
345
|
+
[Maintain Consistency]
|
|
346
|
+
Keep the carpenter's identity and clothing, the chair's structure and count, the tool positions, and the workshop orientation consistent.
|
|
347
|
+
```
|
|
348
|
+
|
|
349
|
+
## Multi-Reference Organization
|
|
350
|
+
|
|
351
|
+
Organize multiple reference materials in this order:
|
|
352
|
+
|
|
353
|
+
```text
|
|
354
|
+
Define Each Material's Role -> Map Subjects -> Group by Type -> Create Subject Profiles -> Select References by Scene
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
A large material set does not mean every material belongs in the Prompt. Activate only materials relevant to the current story and scene. Put every actually available but unassigned material under `[Unused Materials]` inside the Prompt, list each number explicitly, and prohibit its activation. Unused images, videos, and audio may each be combined into one compact line, but no number may be omitted and the line must not become a range-based subject mapping.
|
|
358
|
+
|
|
359
|
+
Seedance 2.5 accepts up to 50 reference materials. Even near the limit, classify every item individually by character, prop, scene, action, and audio, then activate materials by scene. Do not use one umbrella sentence that asks the model to allocate roles by itself, and do not make every material appear at once merely to demonstrate quantity.
|
|
360
|
+
|
|
361
|
+
### Group by Type
|
|
362
|
+
|
|
363
|
+
```text
|
|
364
|
+
[Characters]
|
|
365
|
+
<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
|
|
366
|
+
<Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
|
|
367
|
+
Do not interchange the two characters' appearances, clothing, actions, positions, or dialogue.
|
|
368
|
+
|
|
369
|
+
[Props]
|
|
370
|
+
<Prop A> corresponds to @Image 3 and belongs only to <Character A>.
|
|
371
|
+
<Prop B> corresponds to @Image 4 and belongs only to <Character B>.
|
|
372
|
+
|
|
373
|
+
[Scenes]
|
|
374
|
+
<Scene A> references @Image 5. Use only the space, materials, and lighting.
|
|
375
|
+
<Scene B> references @Image 6. Use only the space, materials, and lighting.
|
|
376
|
+
|
|
377
|
+
[Motion and Audio]
|
|
378
|
+
@Video 1 defines <Character A>'s <action or camera movement>. Do not use the people or scene from the video.
|
|
379
|
+
@Audio 1 defines <Character B>'s <voice characteristics and specified dialogue>.
|
|
380
|
+
```
|
|
381
|
+
|
|
382
|
+
### Subject Profiles and Scene-Based Activation
|
|
383
|
+
|
|
384
|
+
```text
|
|
385
|
+
[Subject Profile: Character A]
|
|
386
|
+
Appearance and clothing: @Image 1.
|
|
387
|
+
Fixed prop: <Prop A> from @Image 3.
|
|
388
|
+
Allowed locations: <Scene A> and <Scene B>.
|
|
389
|
+
Motion reference: <action> from @Video 1.
|
|
390
|
+
Do not use: <other subjects' clothing, props, or audio>.
|
|
391
|
+
|
|
392
|
+
Scene 1 | <Scene Name>
|
|
393
|
+
Use: <subjects, props, scene, motion, and audio activated in this scene>.
|
|
394
|
+
Event: <one primary event>.
|
|
395
|
+
End state: <observable state>.
|
|
396
|
+
|
|
397
|
+
Scene 2 | <Scene Name>
|
|
398
|
+
Use: <subjects, props, scene, motion, and audio activated in this scene>.
|
|
399
|
+
Event: <one primary event>.
|
|
400
|
+
End state: <observable state>.
|
|
401
|
+
```
|
|
402
|
+
|
|
403
|
+
For multiple views of one subject, define the role of each image separately, such as front, left, right, and rear views, and state the number of entities that must appear in the output.
|
|
404
|
+
|
|
405
|
+
### Space and Blocking
|
|
406
|
+
|
|
407
|
+
Describe inside/outside, front/back, facing direction, distance, and separating structures relative to stable objects such as doors, tables, vehicles, counters, and roads. Do not rely only on screen-left or screen-right. For example:
|
|
408
|
+
|
|
409
|
+
```text
|
|
410
|
+
<Store Clerk> always stands on the inner side of the glass counter, facing outward. <Customer A> and <Customer B> stand side by side outside the counter and speak to <Store Clerk> across it.
|
|
411
|
+
```
|
|
412
|
+
|
|
413
|
+
When the user provides a clean blocking diagram, use it for composition, character positions, facing directions, and spatial relationships. Do not reproduce arrows, annotation boxes, or explanatory text from the diagram in the output video.
|
|
414
|
+
|
|
415
|
+
Multi-reference example:
|
|
416
|
+
|
|
417
|
+
```text
|
|
418
|
+
[Characters]
|
|
419
|
+
<Inspector> corresponds to @Image 1. Use only the facial features, short hair, and orange windbreaker.
|
|
420
|
+
<Archivist> corresponds to @Image 2. Use only the facial features, glasses, and gray knitted cardigan.
|
|
421
|
+
Do not interchange the two characters' appearances, clothing, actions, or dialogue.
|
|
422
|
+
|
|
423
|
+
[Props]
|
|
424
|
+
<Portable Data Recorder> corresponds to @Image 3 and belongs only to <Inspector>. There is only one recorder throughout.
|
|
425
|
+
|
|
426
|
+
[Scenes]
|
|
427
|
+
<Mountain Observatory> references @Image 4. Use only the building structure, metal platform, and overcast lighting.
|
|
428
|
+
<Records Room> references @Image 5. Use only the shelf layout, wooden table, and warm interior light.
|
|
429
|
+
|
|
430
|
+
[Motion and Audio]
|
|
431
|
+
@Video 1 defines the action of <Inspector> opening the equipment bay and removing the memory card. Do not use the person or scene from the video.
|
|
432
|
+
@Audio 1 defines <Inspector>'s voice characteristics and dialogue.
|
|
433
|
+
|
|
434
|
+
[Event Script]
|
|
435
|
+
Stage 1: <Inspector> stands alone in front of the equipment bay at <Mountain Observatory>, with <Portable Data Recorder> attached at the waist. The inspector opens the equipment bay and removes the only memory card. End state: the memory card is only in <Inspector>'s right hand.
|
|
436
|
+
Stage 2: Inside the observatory, <Inspector> inserts the memory card into <Portable Data Recorder>. End state: the recorder screen shows that data reading is complete, and the memory card remains inside the recorder.
|
|
437
|
+
Stage 3: Cut to <Records Room>. <Archivist> stands on the inner side of the wooden table, while <Inspector> stands on the outer side. <Inspector> places the recorder in the center of the table and says in natural conversational English: {This week's observation data has been exported.}
|
|
438
|
+
Stage 4: <Archivist> picks up the recorder and checks the data while <Inspector> listens with the mouth naturally closed. End state: the recorder is only in <Archivist>'s hands.
|
|
439
|
+
Stage 5: <Archivist> places the recorder back on the table. Both characters look at the completion status on the screen. End on a medium shot with clear identities, prop count, and blocking.
|
|
440
|
+
|
|
441
|
+
[Maintain Consistency]
|
|
442
|
+
Keep the two characters' identities and clothing, the recorder's count and ownership, the spatial orientation of both scenes, and the speaker relationship consistent.
|
|
443
|
+
```
|
|
444
|
+
|
|
445
|
+
## Long Videos and Time Ranges
|
|
446
|
+
|
|
447
|
+
Seedance 2.5 supports videos up to 30 seconds long. Total duration is set as a generation parameter; the Prompt is responsible only for organizing events within that duration.
|
|
448
|
+
|
|
449
|
+
Numeric event ranges explicitly written by the user are creative content and must be preserved. Parameter separation removes only interface settings such as "generate a 20-second video" or "use a 16:9 aspect ratio." If the user's existing numeric ranges conflict with the target total duration on the page or API, preserve event order, relative pacing, and story outcome, then reorganize the content into nonnumeric stages. Retain the numeric ranges only when the user explicitly states that they are hard constraints, and ask the user to resolve the parameter conflict.
|
|
450
|
+
|
|
451
|
+
For a long video, prefer stages. Give each stage only one primary state change and state its ending condition:
|
|
452
|
+
|
|
453
|
+
```text
|
|
454
|
+
[Stage 1]
|
|
455
|
+
Opening state: <initial state>.
|
|
456
|
+
Primary event: <one primary action or event>.
|
|
457
|
+
End state: <observable state>.
|
|
458
|
+
|
|
459
|
+
[Stage 2]
|
|
460
|
+
Continue from the previous stage: <state that must remain unchanged>.
|
|
461
|
+
Primary event: <one primary action or event>.
|
|
462
|
+
End state: <observable state>.
|
|
463
|
+
|
|
464
|
+
[Stage 3]
|
|
465
|
+
Primary event: <closing event>.
|
|
466
|
+
End state: <final visible state>.
|
|
467
|
+
```
|
|
468
|
+
|
|
469
|
+
The number of stages follows the number of events; it may increase or decrease and is not fixed at three. For a long Prompt, preserve subject mappings, material roles, events, and ending states first. Compress repeated style words, repeated constraints, and inactive-material descriptions. Do not impose a fixed word limit.
|
|
470
|
+
|
|
471
|
+
### Target-Duration Override
|
|
472
|
+
|
|
473
|
+
When a target duration supplied by the generation page or API is longer than the user's existing event timeline, first preserve the characters, event order, causality, and story outcome, then redistribute the pacing of existing events. When the duration comes only from the page or API, use nonnumeric stages so they collectively fill the narrative capacity of the target duration. Do not invent ranges such as `0-8 seconds` or `8-18 seconds` in the Prompt merely to match that parameter.
|
|
474
|
+
|
|
475
|
+
You may lengthen only the progression of an action, a reaction, a pause, or a scene transition: for example, allow a shift of gaze, change in breathing, pickup motion, or reaction after entering to unfold naturally. Do not add a character, main plot event, or new story outcome, and do not mechanically fill time with repeated motion or empty establishing shots. Do not write the target duration itself as an interface parameter sentence in the Prompt.
|
|
476
|
+
|
|
477
|
+
For handoffs, pickups, and placements, state the ownership change of the single object. After a handoff, the original holder no longer possesses it; at the end, it is only in the recipient's hands.
|
|
478
|
+
|
|
479
|
+
Use consecutive integer time ranges only when the user already supplied numeric ranges or explicitly asks to control a handoff, entrance/exit, or beat with time ranges. Do not split sparse events merely to create more ranges; when there are too many events, merge secondary events first:
|
|
480
|
+
|
|
481
|
+
```text
|
|
482
|
+
0-5 seconds: <opening state>; <primary event>; end state: <observable state>.
|
|
483
|
+
5-10 seconds: continue from <previous state>; <primary event>; end state: <observable state>.
|
|
484
|
+
10-15 seconds: <closing event>; end state: <final state>.
|
|
485
|
+
```
|
|
486
|
+
|
|
487
|
+
When the user requests numeric ranges, make them consecutive and non-overlapping. They are event budgets, not frame-accurate edit points. Do not claim 0.5-second precision or state an unverified maximum number of ranges. If the sequence is too dense, reduce the number of stages instead of subdividing further.
|
|
488
|
+
|
|
489
|
+
Overall output duration is still set as a generation parameter. Do not also write "generate an N-second video" in the Prompt.
|
|
490
|
+
|
|
491
|
+
## Video Editing Template
|
|
492
|
+
|
|
493
|
+
Every editing task must define one source video as the sole editing master. A Prompt that describes only the target appearance degrades into regeneration and is not a valid editing Prompt.
|
|
494
|
+
|
|
495
|
+
For visual editing, first inventory every category of visible subject in the source video, including named and unnamed characters, live-action people, models or mannequins, animals, props, foreground objects, and background subjects. For every category, state explicitly whether to replace it, remove it, or keep it unchanged. Do not omit an unnamed person, model, prop, or background subject. Objects the user did not ask to modify remain unchanged by default. Include an entire category in the replacement or removal scope only when the user explicitly requests a group-level change or asks the target image to retain only specified objects.
|
|
496
|
+
|
|
497
|
+
If the current Agent cannot inspect the full source video or has only sparse previews, it must not claim to have exhaustively inventoried every object. After explicitly describing all user-specified replacements, removals, and preserved objects, add this fallback sentence: `Except for the objects explicitly modified above, all other visible people, props, and background elements in @Video 1 remain unchanged and must not be replaced or removed.` When the user explicitly asks to retain only the target objects, replace the fallback with a sentence that removes every other object as requested.
|
|
498
|
+
|
|
499
|
+
Every video-editing Prompt must include one of these two closed-scope statements:
|
|
500
|
+
|
|
501
|
+
- Local modification: `Except for the objects explicitly modified above, all other visible people, props, and background elements in @Video 1 remain unchanged and must not be replaced or removed.`
|
|
502
|
+
- The user explicitly requests only the target objects: `Except for the objects explicitly retained above, remove all other visible subjects from @Video 1. Do not add unspecified objects.`
|
|
503
|
+
|
|
504
|
+
The source-video inventory determines only which objects are replaced, removed, or preserved. It must not change the target material set already specified by the user. If the user explicitly assigns @Image 3 to an edit target, do not add an unmentioned image as an appearance, clothing, group, or auxiliary reference merely because it is clearer or similar. Continue listing those materials under `[Unused Materials]`.
|
|
505
|
+
|
|
506
|
+
```text
|
|
507
|
+
[Edit Goal]
|
|
508
|
+
Edit @Video 1. Change only <original object or region> to <target content>.
|
|
509
|
+
|
|
510
|
+
[Source Video Role]
|
|
511
|
+
@Video 1 is the sole editing master. It defines the original scene, camera position, camera movement, motion paths, occlusion relationships, and event order.
|
|
512
|
+
|
|
513
|
+
[Target Material Role]
|
|
514
|
+
@Image 1 defines <target subject, background, or product>'s <appearance, structure, or material>. Do not use <irrelevant background, people, or composition>.
|
|
515
|
+
|
|
516
|
+
[Edit Objects and Scope]
|
|
517
|
+
Modify only <explicit objects and regions>. The entire video contains <number> target object(s). Do not modify <content to preserve>.
|
|
518
|
+
Except for the objects explicitly modified above, all other visible people, props, and background elements in @Video 1 remain unchanged and must not be replaced or removed.
|
|
519
|
+
|
|
520
|
+
[Timeline Inheritance]
|
|
521
|
+
<Target object> inherits every appearance, motion, occlusion, and exit of <original object>, including timing, duration, path, and speed changes.
|
|
522
|
+
Keep the other character actions, camera movements, cuts, and event order from @Video 1.
|
|
523
|
+
```
|
|
524
|
+
|
|
525
|
+
For subject replacement, state the original and target objects. For background replacement, modify only the background outside the subject silhouette. For a local edit, state the region, attribute, and content to preserve. When adding or removing an object, state its count, position, appearance timing, and affected scope.
|
|
526
|
+
|
|
527
|
+
Dynamic subject-replacement example:
|
|
528
|
+
|
|
529
|
+
```text
|
|
530
|
+
[Edit Goal]
|
|
531
|
+
Edit @Video 1. Replace only the red bicycle and rider passing in front of the bench with the dark gray electric patrol vehicle in @Image 1.
|
|
532
|
+
|
|
533
|
+
[Source Video Role]
|
|
534
|
+
@Video 1 is the sole editing master. It defines the park road, the two people on the bench, camera position, camera movement, the original rider's motion slot, occlusion relationships, and event order.
|
|
535
|
+
|
|
536
|
+
[Target Material Role]
|
|
537
|
+
@Image 1 defines only the dark gray electric patrol vehicle's body structure, color, and clear windshield. Do not use the image background or driver.
|
|
538
|
+
|
|
539
|
+
[Edit Objects and Scope]
|
|
540
|
+
Remove the red bicycle and rider from the source video. The entire video contains only one electric patrol vehicle. Keep the two people on the bench, the trees, road, and background from @Video 1.
|
|
541
|
+
|
|
542
|
+
[Timeline Inheritance]
|
|
543
|
+
The electric patrol vehicle fills the original rider's motion slot with exactly the same appearance timing, motion path, speed, and occlusion positions. The finished video no longer contains the red bicycle or rider. Keep all other character actions, camera movements, cuts, and event order from @Video 1.
|
|
544
|
+
```
|
|
545
|
+
|
|
546
|
+
For a cross-category dynamic replacement, prefer "motion-slot replacement": explicitly remove the original subject, make the target subject inherit exactly the same appearance timing, motion path, speed, and occlusion positions, and state that the original subject no longer appears. Keep the preservation list focused on adjacent regions that genuinely must not change so that it does not dilute the edit target.
|
|
547
|
+
|
|
548
|
+
```text
|
|
549
|
+
Remove <original moving subject> that passes in front of <foreground subject> in @Video 1. Replace it at exactly the same appearance timing, motion path, speed, and occlusion positions with <target moving subject> defined by @Image 1. The finished video no longer contains <original moving subject>.
|
|
550
|
+
```
|
|
551
|
+
|
|
552
|
+
Use global master-video inheritance by default. Add only a few observable event conditions, and only when a pickup, handoff, placement, entrance, or exit in the source video can be observed accurately:
|
|
553
|
+
|
|
554
|
+
```text
|
|
555
|
+
Only after <observable completion state of Event A> appears may <Event B> occur.
|
|
556
|
+
Only after <observable completion state of Event B> appears may <Event C> occur.
|
|
557
|
+
```
|
|
558
|
+
|
|
559
|
+
Do not reconstruct time ranges from memory or add event conditions from incomplete extracted frames. Prompt wording can improve the probability that critical events follow the source timeline, but it cannot guarantee frame-by-frame overlap after editing.
|
|
560
|
+
|
|
561
|
+
### Audio Editing
|
|
562
|
+
|
|
563
|
+
When modifying only dialogue, language, voice characteristics, background music, ambience, or action sound effects, still define the source video as the sole editing master. State separately which speaker or sound category changes, the target change, the time range, and which other sounds and visuals remain unchanged. Do not redesign character actions, lip-sync timing, camera treatment, or editing rhythm merely because audio is being edited.
|
|
564
|
+
|
|
565
|
+
```text
|
|
566
|
+
[Edit Goal]
|
|
567
|
+
Edit @Video 1. Within <the entire video or an explicit time range>, <remove, replace, or adjust> <speaker or sound category> only.
|
|
568
|
+
|
|
569
|
+
[Source Video Role]
|
|
570
|
+
@Video 1 is the sole editing master. It defines the original visuals, character actions, lip-sync timing, camera treatment, editing rhythm, other sounds, and event order.
|
|
571
|
+
|
|
572
|
+
[Target Audio Role]
|
|
573
|
+
@Audio 1 defines <target speaker or sound type>'s <voice characteristics, dialogue, ambience, sound effect, or music>. Do not use <irrelevant audio>.
|
|
574
|
+
|
|
575
|
+
[Audio Edit Scope]
|
|
576
|
+
Modify only <explicit speaker, sound category, or time range>.
|
|
577
|
+
|
|
578
|
+
[Content to Preserve]
|
|
579
|
+
Keep <other dialogue, lip-sync timing, ambience, action sound effects, visuals, camera treatment, and editing rhythm> from @Video 1 unchanged.
|
|
580
|
+
```
|
|
581
|
+
|
|
582
|
+
When the user only asks to remove the original background music, do not invent a target audio material. State directly that the background music is removed while character dialogue, lip-sync, ambience, action sound effects, and all visuals remain unchanged. When changing dialogue language or voice characteristics, preserve the dialogue content and speaking times from the source video by default unless the user explicitly asks to rewrite them as well.
|
|
583
|
+
|
|
584
|
+
## Video Extension Templates
|
|
585
|
+
|
|
586
|
+
If the user says only "extend" and the direction cannot be determined from context, extension direction is high-impact information; ask one consolidated question.
|
|
587
|
+
|
|
588
|
+
For subjects present at the extension boundary, use only names confirmed by the user. If the user says only "person" or "subject," retain that neutral term. Do not infer a profession, performance type, or story identity from the subject's action, posture, or clothing in the source video.
|
|
589
|
+
|
|
590
|
+
Throughout the extension, each subject remains one continuous instance. Do not duplicate, split, or generate a second copy of the same subject. Keep body structure, component count, and topology consistent with the boundary frame. When a subject turns, becomes occluded, leaves the frame, or re-enters, it is still the same continuous object and must not be replaced by a new instance.
|
|
591
|
+
|
|
592
|
+
The final Prompt must state the single-instance and topology requirements explicitly. Checking them only in internal analysis is insufficient.
|
|
593
|
+
|
|
594
|
+
### Forward Extension (After the Original Video)
|
|
595
|
+
|
|
596
|
+
A forward extension generates content after the source video ends. The first frame of the new segment continues from the source video's last frame.
|
|
597
|
+
|
|
598
|
+
```text
|
|
599
|
+
@Video 1 is the source video to extend forward.
|
|
600
|
+
|
|
601
|
+
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain continuity in <subject pose and orientation>, <prop position>, <background and spatial relationships>, <camera position and composition>, <lighting>, <audio state>, and <motion direction>.
|
|
602
|
+
|
|
603
|
+
Then, <new action, event, camera treatment, or audio to generate beyond the boundary>.
|
|
604
|
+
|
|
605
|
+
Throughout the extension, maintain continuity in <character identity and clothing>, <key props>, <background layout>, <camera axis>, and <existing audio environment>.
|
|
606
|
+
Each subject remains the same continuous object without duplication or splitting. Keep character anatomy and object component counts stable.
|
|
607
|
+
```
|
|
608
|
+
|
|
609
|
+
Example with additional reference materials:
|
|
610
|
+
|
|
611
|
+
```text
|
|
612
|
+
@Video 1 is the source video to extend forward.
|
|
613
|
+
@Image 1 defines <Gardener>'s facial features, short hair, and light green work apron. Do not use the image background.
|
|
614
|
+
@Image 2 defines <Wicker Flower Basket>'s structure and material. Do not use the garden or people in the image.
|
|
615
|
+
|
|
616
|
+
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain the greenhouse workbench, <Gardener>'s position and facing direction, the wooden rack's position, the locked-off medium shot, and the afternoon side light. Other reference materials must not replace this boundary image.
|
|
617
|
+
|
|
618
|
+
Then, <Gardener> picks up <Wicker Flower Basket> defined by @Image 2 from beneath the workbench and places it with both hands on the middle shelf of the wooden rack behind them. End state: the basket is only on the middle shelf, and <Gardener> has released both hands and taken half a step back.
|
|
619
|
+
|
|
620
|
+
Throughout the extension, keep <Gardener>'s face and apron, the greenhouse layout, wooden-rack position, camera direction, and greenhouse ambience continuous.
|
|
621
|
+
<Gardener> and <Wicker Flower Basket> each remain one continuous object without duplication or splitting. Keep body structure and basket component count stable.
|
|
622
|
+
```
|
|
623
|
+
|
|
624
|
+
### Backward Extension (Before the Original Video)
|
|
625
|
+
|
|
626
|
+
A backward extension generates content before the source video begins. The last frame of the new segment connects to the source video's first frame.
|
|
627
|
+
|
|
628
|
+
First describe what happens before the source video begins, then define the source video's first frame as the explicit end state of the extended segment. Writing only "then connect to the source video" may introduce later characters, props, or effects too early, or cause the image to change again after reaching the target state.
|
|
629
|
+
|
|
630
|
+
```text
|
|
631
|
+
@Video 1 is the source video to extend backward.
|
|
632
|
+
|
|
633
|
+
Extend @Video 1 backward. Before the source video begins, <preceding action, event, camera treatment, or audio>.
|
|
634
|
+
|
|
635
|
+
The last frame of the extended segment naturally connects to the first frame of @Video 1. Match <subject pose and orientation>, <prop position>, <background and spatial relationships>, <camera position and composition>, <lighting>, <audio state>, and <motion direction>.
|
|
636
|
+
|
|
637
|
+
Throughout the extension, maintain continuity in <character identity and clothing>, <key props>, <background layout>, <camera axis>, and <existing audio environment>.
|
|
638
|
+
Each subject remains the same continuous object without duplication or splitting. Keep character anatomy and object component counts stable.
|
|
639
|
+
Characters, props, or effects that belong only to later events in the source video must not appear early.
|
|
640
|
+
```
|
|
641
|
+
|
|
642
|
+
When additional references are available, define their character, clothing, prop, or audio roles one by one, then state that the source video controls the extension boundary. New materials must not override the source video's last-frame or first-frame control of the boundary image.
|
|
643
|
+
|
|
644
|
+
Extension creates only the new segment outside the boundary; it must not edit the original video at the same time. The goal is natural visual and audio continuity, not pixel-identical boundary frames. The extended segment's volume may differ slightly from the source video.
|
|
645
|
+
|
|
646
|
+
## Keyframe Anchors
|
|
647
|
+
|
|
648
|
+
Keyframe images are still uploaded as ordinary reference images; assign each one a role in the Prompt. Never combine first-frame and last-frame roles into a range sentence.
|
|
649
|
+
|
|
650
|
+
To fix a first frame, write the exact standalone sentence `@Image N is the first frame.` To fix a last frame, write the exact standalone sentence `@Image N is the last frame.` Do not weaken these into "used only as a first-frame reference," "reference the opening composition," or "first-frame composition reference," and do not compress the frame-role sentence and subsequent action into one sentence. Preserve the exact role sentence in the final Prompt, then use a new sentence to define that frame's composition, subject position, pose, prop state, scene, and camera direction.
|
|
651
|
+
|
|
652
|
+
The first frame locks the output aspect ratio. The first and last images should use the same aspect ratio to avoid stretching the last frame. Duration is still set on the generation page or through the API. This rule is for input preflight only; do not write it as an output parameter in the Prompt.
|
|
653
|
+
|
|
654
|
+
### First Frame with Additional References
|
|
655
|
+
|
|
656
|
+
```text
|
|
657
|
+
@Image 1 is the first frame.
|
|
658
|
+
This first frame defines the opening composition, subject position, pose, prop state, scene, and camera direction.
|
|
659
|
+
@Image 2 defines <subject>'s <appearance, clothing, structure, or material> without changing the first-frame composition defined by @Image 1.
|
|
660
|
+
@Image 3 defines <scene, prop, or lighting> without changing the first-frame composition defined by @Image 1.
|
|
661
|
+
|
|
662
|
+
The video begins naturally from the first frame defined by @Image 1. Then, <continuous action or event>.
|
|
663
|
+
Maintain continuity in <character identity, prop ownership, spatial relationships, and visual style>.
|
|
664
|
+
```
|
|
665
|
+
|
|
666
|
+
### First and Last Frames with Additional References
|
|
667
|
+
|
|
668
|
+
```text
|
|
669
|
+
@Image 1 is the first frame.
|
|
670
|
+
This first frame defines the opening composition, subject position, pose, prop state, scene, and camera direction.
|
|
671
|
+
@Image 2 is the last frame.
|
|
672
|
+
This last frame defines the ending composition, subject position, pose, prop state, scene, and camera direction.
|
|
673
|
+
@Image 3 defines <Subject A>'s <appearance, clothing, structure, or material> without changing the first-frame composition from @Image 1 or the last-frame composition from @Image 2.
|
|
674
|
+
@Image 4 defines <specified attributes> of <Subject B, prop, or scene> without changing the first-frame composition from @Image 1 or the last-frame composition from @Image 2.
|
|
675
|
+
|
|
676
|
+
<One continuous action or event>.
|
|
677
|
+
The video begins naturally from the first frame defined by @Image 1 and reaches the last frame defined by @Image 2 after the continuous action.
|
|
678
|
+
Between the first and last frames, maintain continuity in <character identity, prop structure and ownership, scene layout, and camera direction>.
|
|
679
|
+
```
|
|
680
|
+
|
|
681
|
+
Example:
|
|
682
|
+
|
|
683
|
+
```text
|
|
684
|
+
@Image 1 is the first frame.
|
|
685
|
+
This first frame defines the baking counter, <Pastry Chef>'s position, the undecorated cake, tool placement, and a frontal medium shot.
|
|
686
|
+
@Image 2 is the last frame.
|
|
687
|
+
This last frame defines the completed cake centered on the turntable, <Pastry Chef>'s hands released from the cake, and the same frontal medium shot.
|
|
688
|
+
@Image 3 defines <Pastry Chef>'s facial features, pinned-up hair, and white uniform without changing the first-frame composition from @Image 1 or the last-frame composition from @Image 2.
|
|
689
|
+
@Image 4 defines the cake's two-tier structure, white frosting material, and blueberry decoration without changing the first-frame composition from @Image 1 or the last-frame composition from @Image 2.
|
|
690
|
+
|
|
691
|
+
The video begins naturally from the first frame defined by @Image 1. <Pastry Chef> turns the cake stand, pipes an even cream border around both tiers, places the blueberries one by one, releases both hands from the cake, and naturally reaches the last frame defined by @Image 2.
|
|
692
|
+
Between the first and last frames, maintain continuity in <Pastry Chef>'s identity and clothing, the cake's count and two-tier structure, tool positions, baking-counter layout, and camera direction.
|
|
693
|
+
```
|
|
694
|
+
|
|
695
|
+
When the user explicitly requests an intermediate key state, an additional image may define the characters, action, props, and spatial relationships that must be visible near that moment. Describe the video as naturally reaching that state around the specified time. The image is a semantic anchor, not a static hold or pixel-locked frame. When the first and last boundaries are the priority, reduce other materials unrelated to the current event.
|
|
696
|
+
|
|
697
|
+
### Multi-Keyframe Sequence Control
|
|
698
|
+
|
|
699
|
+
When separate images define process stages, declare the keyframe order in the first sentence, then define every key state individually. Keyframes control stage order and visible states; they do not promise frame-by-frame reproduction or require a static hold at any key state.
|
|
700
|
+
|
|
701
|
+
```text
|
|
702
|
+
Use @Image 1 through @Image N in order as keyframes.
|
|
703
|
+
|
|
704
|
+
@Image 1 is the first frame.
|
|
705
|
+
This first frame defines <opening composition, subject position, pose, prop state, and camera direction>.
|
|
706
|
+
@Image 2 defines the second keyframe: <visible state at the end of the first stage>.
|
|
707
|
+
@Image 3 defines the third keyframe: <visible state at the end of the second stage>.
|
|
708
|
+
@Image N is the last frame.
|
|
709
|
+
This last frame defines <ending composition, subject position, pose, prop state, and camera direction>.
|
|
710
|
+
|
|
711
|
+
The video passes in order through the states defined by @Image 1, @Image 2, @Image 3, through @Image N, using continuous action to transition naturally between stages.
|
|
712
|
+
Throughout the sequence, maintain continuity in <subject identity, prop structure and ownership, scene layout, lighting, and camera axis>.
|
|
713
|
+
```
|
|
714
|
+
|
|
715
|
+
### Storyboard Grids
|
|
716
|
+
|
|
717
|
+
A storyboard grid defines the overall story, shot order, and approximate composition. It is not intended to reproduce every panel's details strictly. Prefer a clean storyboard with no more than 15 panels and minimal text labels. State the reading order, the shot structure to use, and the line-art style, annotations, or placeholder characters not to use.
|
|
718
|
+
|
|
719
|
+
```text
|
|
720
|
+
@Image 1 provides an <N-panel storyboard grid> for shot order and approximate composition. Read it <left to right, top to bottom>. Do not use the grid's <line-art style, text labels, or placeholder characters>.
|
|
721
|
+
@Image 2 defines <Subject A>'s <appearance and clothing>.
|
|
722
|
+
@Image 3 defines <key prop or scene>'s <structure, material, or lighting>.
|
|
723
|
+
|
|
724
|
+
Shot 1: <shot size, subject action, and scene state>.
|
|
725
|
+
Shot 2: <shot size, subject action, and camera movement or transition>.
|
|
726
|
+
...
|
|
727
|
+
Shot N: <closing action and final visible state>.
|
|
728
|
+
|
|
729
|
+
The final video uses <visual style>. Audio includes <dialogue, ambience, action sound effects, or music>.
|
|
730
|
+
```
|
|
731
|
+
|
|
732
|
+
### Blockout References and Rendering
|
|
733
|
+
|
|
734
|
+
First determine whether the blockout video provides a motion skeleton or complete structure:
|
|
735
|
+
|
|
736
|
+
- **Coarse blockout:** Simple geometry mainly provides action paths, motion direction, subject blocking, entrances and exits, camera position, camera movement, cuts, lighting changes, sound rhythm, or spatial relationships. Map every geometric object separately to its final subject or prop.
|
|
737
|
+
- **Fine blockout:** Character, prop, or scene structure is already complete and is mainly used to change character appearance, materials, colors, scene, or visual style. Preserve the original structure, action, space, and camera treatment.
|
|
738
|
+
|
|
739
|
+
A blockout video is a generation reference; it does not automatically become an editing master merely because it is a video. When the blockout contains path lines, coordinate axes, controllers, camera frustums, or text markers, state explicitly that those production markers must not be used.
|
|
740
|
+
|
|
741
|
+
#### Coarse Blockouts
|
|
742
|
+
|
|
743
|
+
```text
|
|
744
|
+
@Video 1 is a coarse blockout reference. It provides only <motion paths, subject blocking, camera position, camera movement, cuts, lighting changes, sound rhythm, or spatial relationships>. Do not use its blockout appearance, materials, or scene.
|
|
745
|
+
<Blockout Subject A> in @Video 1 corresponds to <Subject A>.
|
|
746
|
+
<Blockout Subject B or geometric prop> in @Video 1 corresponds to <Subject B or key prop>.
|
|
747
|
+
@Image 1 defines <Subject A>'s <appearance, clothing, or structure>.
|
|
748
|
+
@Image 2 defines <specified attributes> of <Subject B, key prop, or scene>.
|
|
749
|
+
|
|
750
|
+
<Subject> completes <primary action or event> in <scene>.
|
|
751
|
+
Keep <motion path, blocking, camera movement, cuts, lighting, or sound rhythm> from @Video 1.
|
|
752
|
+
The final video uses <characters, scene, materials, and visual style>. Audio includes <dialogue, ambience, or action sound effects>.
|
|
753
|
+
```
|
|
754
|
+
|
|
755
|
+
#### Fine Blockouts
|
|
756
|
+
|
|
757
|
+
```text
|
|
758
|
+
@Video 1 is a fine blockout reference. Preserve <subject structure, action, spatial layout, camera position, camera movement, and cuts>. Do not use its original gray materials, empty background, or production markers.
|
|
759
|
+
@Image 1 defines <subject>'s <character appearance, material, color, or surface details>.
|
|
760
|
+
@Image 2 defines <scene>'s <space, materials, lighting, or visual style>.
|
|
761
|
+
|
|
762
|
+
Re-render <subject> from @Video 1 as <final subject>, and re-render the scene as <final scene>.
|
|
763
|
+
Keep <structure, action, camera treatment, and spatial relationships> from @Video 1. Use <materials, colors, and style>. Audio includes <ambience, sound effects, or music>.
|
|
764
|
+
```
|
|
765
|
+
|
|
766
|
+
## Emotion, Cinematography, and Audio
|
|
767
|
+
|
|
768
|
+
### Emotion and Observable Performance
|
|
769
|
+
|
|
770
|
+
Writing only a direction such as tense, warm, or oppressive gives the model room to interpret the specific performance. When the user needs tighter acting control, use:
|
|
771
|
+
|
|
772
|
+
```text
|
|
773
|
+
Emotional or atmospheric direction + Triggering event + Observable performance + Observable camera, lighting, or audio change
|
|
774
|
+
```
|
|
775
|
+
|
|
776
|
+
Select a small number of the clearest cues from eye movement, brow tension, mouth movement, breathing, gaze direction, hand movement, and posture. Do not pile on every possible micro-expression. Organize performance by triggering events only when the emotion changes several times.
|
|
777
|
+
|
|
778
|
+
```text
|
|
779
|
+
The overall emotion shifts from <starting emotion> to <ending emotion>.
|
|
780
|
+
After <triggering event>, <subject> first shows <immediate observable reaction>.
|
|
781
|
+
Then, <eyes, brows, mouth, breathing, gaze, or hand movement> gradually <changes>.
|
|
782
|
+
Finally, <subject> expresses <target emotion> through <outward behavior>.
|
|
783
|
+
```
|
|
784
|
+
|
|
785
|
+
Example:
|
|
786
|
+
|
|
787
|
+
```text
|
|
788
|
+
The overall emotion shifts from restrained anticipation to an effort to remain composed after disappointment.
|
|
789
|
+
When a server places a returned letter on the table, the woman's fingers suddenly stop tracing the rim of the cup, and her gaze settles on the return mark on the envelope.
|
|
790
|
+
Her brows draw together slightly, the faint smile at the corners of her mouth gradually disappears, and after a slow breath she turns the envelope face down on the table.
|
|
791
|
+
Finally, she looks up at the empty chair opposite her, keeps her shoulders straight, and says in a calm but slightly tightened voice: {I understand.}
|
|
792
|
+
```
|
|
793
|
+
|
|
794
|
+
### Professional Cinematography
|
|
795
|
+
|
|
796
|
+
Basic camera language and popular camera techniques can be written directly into the Prompt. When the frame contains several subjects, still state which subject the camera follows or revolves around, where the movement begins, and where it ends. Do not write a detached camera term without a target.
|
|
797
|
+
|
|
798
|
+
```text
|
|
799
|
+
Popular camera technique + Target subject + Starting position or state + Movement direction + Destination position or state
|
|
800
|
+
```
|
|
801
|
+
|
|
802
|
+
Popular techniques include a one-take shot, dolly zoom, aerial view, FPV, bullet time, handheld camera, and bounce speed ramp. For a one-take shot, state the subjects, spaces, and events the continuous camera passes through in order. For handheld camera, state the subject being followed and the amount of shake. For a bounce speed ramp, state where the action accelerates, decelerates, or rebounds and its final resting state.
|
|
803
|
+
|
|
804
|
+
For a niche cinematography term, a term with inconsistent industry usage, or a term that requires precise control of the image change, keep the term itself and expand it into an observable result:
|
|
805
|
+
|
|
806
|
+
```text
|
|
807
|
+
Cinematography term + Target subject + Visual change + Foreground/background relationship + Direction or speed
|
|
808
|
+
```
|
|
809
|
+
|
|
810
|
+
For example, shallow depth of field should state which subject remains sharp and how the background blurs. A tracking shot should state that the camera matches the subject's speed and specify the direction of background motion blur. Rack focus should state which object loses focus, which object gains focus, and how their sharpness changes. A vignette should state that the corners darken gradually while center brightness remains natural. Focal length, aperture, and shutter values may supplement these visible results but must not replace them.
|
|
811
|
+
|
|
812
|
+
### Audio, Dialogue, and Text
|
|
813
|
+
|
|
814
|
+
Use the following syntax when content types must be distinguished explicitly:
|
|
815
|
+
|
|
816
|
+
| Content | Syntax | Example |
|
|
817
|
+
|-|-|-|
|
|
818
|
+
| Music | `()` | `(Soft piano music plays in the background)` |
|
|
819
|
+
| Sound Effects | `<>` | `<A bell rings in the distance>` |
|
|
820
|
+
| Dialogue | `{}` | `{Hello, welcome back.}` |
|
|
821
|
+
| Subtitles | `【】` | `【Chapter One: Departure】` |
|
|
822
|
+
|
|
823
|
+
For dialogue language, use: `Language + Optional regional variety or accent + Delivery style + Speaker + {Dialogue}`. Label each speaker separately rather than making one blanket declaration at the beginning. When English dialogue is likely to be spoken in Chinese, state at minimum that it must be spoken in English. Add a full direction such as "natural, conversational American English" or "authentic Los Angeles English" only when the user explicitly specifies that regional variety or accent, or explicitly asks for stronger reinforcement. If the user supplies only English dialogue, do not invent an American, British, or regional accent. If the user does not specify a language, do not infer Mandarin, a dialect, or a regional accent from the script's written form. When subtitles are not required, also state that no subtitles appear on screen.
|
|
824
|
+
|
|
825
|
+
For multi-speaker dialogue, bind the speaker and audio reference within each stage and state that the other characters listen with their mouths naturally closed. Describe the source of ambience, sound effects, and music separately so irrelevant audio is not treated as background music.
|
|
826
|
+
|
|
827
|
+
If the final Prompt uses `<>` to mark sound effects, do not also place subject names in angle brackets. Use plain character names so one symbol does not perform two roles.
|
|
828
|
+
|
|
829
|
+
For a no-dialogue task, constrain speech, mouth movement, sound sources, and written text together. For example: characters keep their mouths naturally closed, there is no narration, only the specified ambience remains, and no subtitles or signs appear on screen.
|
|
830
|
+
|
|
831
|
+
### Products and Real-World Processes
|
|
832
|
+
|
|
833
|
+
Convert abstract claims such as efficient, intelligent, or reliable into `Initial State -> Concrete Operation -> Observable Result`. Each stage should demonstrate only one operation or functional result while keeping product appearance, component positions, operator identity, and scene relationships consistent.
|
|
834
|
+
|
|
835
|
+
```text
|
|
836
|
+
Stage 1: Opening state: <initial state of the product and components>. <Operator completes one concrete action>. End state: <directly visible state>.
|
|
837
|
+
Stage 2: Continue from <previous state>. <Product performs one function>. End state: <directly visible result>.
|
|
838
|
+
Stage 3: <Operator completes the closing action>. End state: <final state of the product, components, and finished output>.
|
|
839
|
+
```
|
|
840
|
+
|
|
841
|
+
Example:
|
|
842
|
+
|
|
843
|
+
```text
|
|
844
|
+
Stage 1: At the start, the desktop humidifier is off, the water tank is empty, and the top cover lies to the right of the body. The operator removes the tank and fills it with clean water. End state: the water level remains below the maximum line.
|
|
845
|
+
Stage 2: The operator reinstalls the tank, closes the top cover, and presses the power button once. End state: the operator's hand has left the button, and the humidifier remains fixed in place.
|
|
846
|
+
Stage 3: Fine white mist flows continuously from the outlet and rises vertically without leaving water on the table. End state: the body, tank, and top cover remain intact.
|
|
847
|
+
```
|
|
848
|
+
|
|
849
|
+
Do not write only "show the complete process of product installation, operation, and completion." Exact screen text, formulas, and product parameters still follow the capability boundaries in the Output Contract.
|
|
850
|
+
|
|
851
|
+
## Output Contract
|
|
852
|
+
|
|
853
|
+
Match the user's language. Preserve the user's existing reference format, such as `@Image 1` or an equivalent runtime label. Do not translate or renumber an explicit reference on your own.
|
|
854
|
+
|
|
855
|
+
### Default Delivery
|
|
856
|
+
|
|
857
|
+
By default, output only the optimized Prompt body. Do not add Markdown headings, code fences, prefatory or closing explanations, or outer wrappers such as "Optimized Prompt," "Material Understanding," or "Optimization Notes." Do not instruct the model to wrap the final result in a code fence.
|
|
858
|
+
|
|
859
|
+
When the Prompt itself requires structure, retain task-internal labels such as `[Generation Goal]`, `[Reference Material Roles]`, `[Event Script]`, and `[Maintain Consistency]`. These labels are submit-ready Prompt content, not response wrappers.
|
|
860
|
+
|
|
861
|
+
When the Agent infers material mappings automatically, do not output a separate reasoning table. Write the selected mappings directly into the Prompt's material-role section, with one line stating the adopted scope of every activated material. When the Agent can access the complete material list, it must append `[Unused Materials]` inside the Prompt and list every available but unassigned material number individually. Do not explain unused materials only outside the Prompt. When the complete material list is unavailable, do not invent unused numbers.
|
|
862
|
+
|
|
863
|
+
### Partially Missing Materials
|
|
864
|
+
|
|
865
|
+
When some materials are missing, first output a complete Prompt body that no longer references the missing numbers. Then append at most one single-line suggestion beginning with `Supplementary Suggestion:` that identifies the missing number and intended role. This suggestion is not Prompt content: put it on a separate line rather than attaching it to the Prompt's final sentence. Do not split it into several suggestions or claim to have inspected the missing material.
|
|
866
|
+
|
|
867
|
+
Example: `Supplementary Suggestion: @Audio 3 was not provided. The current Prompt preserves the corresponding character and dialogue without specifying a reference voice; provide the audio to bind the voice characteristics more precisely.`
|
|
868
|
+
|
|
869
|
+
### Input and Parameter Notes
|
|
870
|
+
|
|
871
|
+
When input materials exceed a hard limit, first output a complete Prompt based on the selected materials, then append one `Material Note:` listing the material types or numbers that must be reduced before submission. Recommended ranges are not hard limits and do not trigger a material note.
|
|
872
|
+
|
|
873
|
+
When the user requests a parameter automatically locked for editing, first-frame or first-and-last-frame generation, or extension, or when the first and last images use different aspect ratios, first output the complete Prompt and then append one `Parameter Note:`. The note should state only the conflicting rule and the required pre-submission action. Do not repeat the Prompt, expand into API documentation, or write the parameter back into the Prompt.
|
|
874
|
+
|
|
875
|
+
Example: `Parameter Note: Video editing automatically preserves the input video's aspect ratio and approximate duration, so it cannot also be set to 16:9 and 20 seconds. The Prompt above follows the source video's timeline.`
|
|
876
|
+
|
|
877
|
+
### Required Capability Disclosures
|
|
878
|
+
|
|
879
|
+
Exact subtitles, formulas, signs, product specifications, or frame-level timing cannot be fully guaranteed by Prompt wording alone. Still output the best Prompt first. When genuinely necessary, append at most one `Additional Note:` explaining that prepared materials or post-production should also be used. Output no note when it is unnecessary.
|
|
880
|
+
|
|
881
|
+
For a mixed request that requires two sequential operations, output submit-ready Prompts under `Step 1:` and `Step 2:`, and state that Step 2 uses the output of Step 1 as its new master. In every other case, output one Prompt by default.
|
|
882
|
+
|
|
883
|
+
For ordinary generation, duration, aspect ratio, and other generation parameters supplied by the user may inform event density and composition, but do not write them into the Prompt. For video editing, first-frame or first-and-last-frame generation, and video extension, apply the task's automatic parameter-locking rules first and do not plan from conflicting parameters. List configurable parameters outside the Prompt only when the user explicitly asks for an invocation example.
|
|
884
|
+
|
|
885
|
+
### Complete Input-to-Output Example
|
|
886
|
+
|
|
887
|
+
User input: `Have the barista in the reference image perform a pour-over in the reference cafe, follow the action in the uploaded video, and hand the coffee to the customer at the end.`
|
|
888
|
+
|
|
889
|
+
After inspecting the three materials, the Agent outputs only this Prompt body:
|
|
890
|
+
|
|
891
|
+
[Generation Goal]
|
|
892
|
+
Generate a continuous video in which a barista completes a pour-over and hands the coffee to a customer.
|
|
893
|
+
|
|
894
|
+
[Reference Material Roles]
|
|
895
|
+
@Image 1 defines the barista's facial features, short hair, and brown apron. Do not use the image background.
|
|
896
|
+
@Image 2 defines the cafe's wooden counter, glass window, and warm afternoon light. Do not use the people in the image.
|
|
897
|
+
@Video 1 defines the pacing and wrist movement of pouring water. Do not use the person's identity, clothing, or scene from the video.
|
|
898
|
+
|
|
899
|
+
[Event Script]
|
|
900
|
+
At the start, the barista stands on the inner side of the counter, the customer stands on the outer side, and the dripper and serving carafe are between them. The barista pours slowly using the movement and pacing from @Video 1. After finishing, the barista returns the dripper to its holder, then uses both hands to pass the only cup of coffee to the customer. End state: the coffee is only in the customer's hands, and the barista has released the cup with both hands.
|
|
901
|
+
|
|
902
|
+
[Maintain Consistency]
|
|
903
|
+
Keep the barista's identity and apron, the cafe layout, inside/outside counter blocking, and coffee-cup count consistent.
|
|
904
|
+
|
|
905
|
+
## Final Checklist
|
|
906
|
+
|
|
907
|
+
Before output, confirm every item:
|
|
908
|
+
|
|
909
|
+
- One Prompt has exactly one primary task: generation, editing, or extension. A clearly mixed request has been split into two sequential Prompts, each with exactly one primary task.
|
|
910
|
+
- Subject identity and count, scene, prop ownership, spatial relationships, and story outcome have not changed.
|
|
911
|
+
- Every required character, prop, and scene in the story contract is covered. No two named characters were merged, and no material-based guess was written as identity, age, relationship, or plot fact.
|
|
912
|
+
- Every activated material has one unique, explicit role.
|
|
913
|
+
- The Prompt body contains no unavailable material number or raw Asset ID. For a partially missing material, only the single-line `Supplementary Suggestion:` after the body may name that missing number.
|
|
914
|
+
- Every material number in the Prompt has been compared against the actually available material list. No partially missing material is presented as inspected, and there is at most one single-line supplementary suggestion.
|
|
915
|
+
- Irrelevant materials are not forced into the video. When the complete material list is available, every available but unassigned material is listed individually under `[Unused Materials]`.
|
|
916
|
+
- No inaccessible material is presented as understood.
|
|
917
|
+
- Distinct characters, products, and props are mapped one by one rather than with a range sentence.
|
|
918
|
+
- A single-person character sheet does not define two appearing characters. When several individual or group candidates exist, one image per person and one image per group were assigned first.
|
|
919
|
+
- Each long-video stage contains only one primary state change and a clear ending state.
|
|
920
|
+
- Event time ranges explicitly requested by the user have not been removed without permission. When only an external target duration exists, existing events are redistributed with nonnumeric stages, and the parameter has not been converted into new time ranges in the Prompt.
|
|
921
|
+
- An editing task defines the sole editing master, edit scope, target count, content to preserve, and timeline inheritance. An audio edit also defines changed audio, preserved audio, lip-sync timing, and all visuals.
|
|
922
|
+
- The current task's automatically locked aspect-ratio and duration rules are followed. A conflict produces only one parameter note outside the Prompt.
|
|
923
|
+
- Extension direction and boundary-frame roles are correct, and the original video is not also rewritten. Boundary image, audio state, motion direction, and continuous subjects are all covered.
|
|
924
|
+
- The first frame, last frame, and other reference images each have their own roles without overriding one another. A fixed frame uses the standalone exact sentence `@Image N is the first frame.` or `@Image N is the last frame.` rather than a weakened composition-reference phrase. First/last aspect ratios were checked or handled according to the inaccessible-material rule.
|
|
925
|
+
- Multiple keyframes define each key state in sequence without promising frame-by-frame reproduction or a static hold.
|
|
926
|
+
- A storyboard grid states the reading order, each panel's shot role, and which line art, annotations, or placeholders not to use.
|
|
927
|
+
- A blockout is identified as coarse or fine, with subject mappings, inherited information, and excluded content defined accordingly.
|
|
928
|
+
- Abstract emotions and cinematography terms include observable results when control is needed.
|
|
929
|
+
- Within every speaking stage, the speaker, vocal state, dialogue, audio role, language, and on-screen/off-screen position match the input. Speech was not turned into silence, and no Mandarin, dialect, or regional accent was invented.
|
|
930
|
+
- Shot numbers, material numbers, chapter numbers, and step numbers remain identifiers and were not rewritten as camera angles or other photographic parameters.
|
|
931
|
+
- The Prompt contains no output parameters, internal analysis, evaluation metadata, keys, endpoints, or reasons for the rewrite.
|
|
932
|
+
- No negative constraint unrelated to the user's request was added automatically.
|
|
933
|
+
- Every placeholder has been replaced. By default, the output is one complete Prompt body without Markdown headings, code fences, or surrounding explanation.
|
|
934
|
+
|
|
935
|
+
## Compatibility and Runtime Notes
|
|
936
|
+
|
|
937
|
+
- **Text-only Agent:** Process text and explicit reference labels. Do not claim to have inspected attachments or path contents.
|
|
938
|
+
- **Multimodal Agent:** Within runtime capabilities, inspect images, videos, and audio, then perform the two-pass material review and automatic mapping.
|
|
939
|
+
- **No filesystem access:** Preserve the user's existing labels and apply the "Current Agent Cannot Read the Materials" fallback to inaccessible paths.
|
|
940
|
+
- **No network access:** Do not attempt to resolve remote URL content or infer material content from the URL name.
|
|
941
|
+
- **Output only:** This Skill outputs text and requires no file writes, network access, tool calls, or binary-output capability.
|
|
942
|
+
|
|
943
|
+
---
|
|
944
|
+
|
|
945
|
+
# Venice Video Harness Bridge
|
|
946
|
+
|
|
947
|
+
This section is specific to **this repo** (the Venice Video Harness). The
|
|
948
|
+
optimizer above is provider-neutral; here is how its output maps onto the
|
|
949
|
+
harness's fields and where the harness — not the optimizer — is authoritative.
|
|
950
|
+
|
|
951
|
+
Seedance 2.5 (`seedance-2-5-*`) is the harness's **default video model** across
|
|
952
|
+
every lane (montage, singles, multi-shot, action, atmosphere, character, and
|
|
953
|
+
in-family lip-sync). So this skill's compiled Prompt is what those lanes render.
|
|
954
|
+
Use it to write the creative content **before** the harness builds the API body.
|
|
955
|
+
|
|
956
|
+
## Where the compiled Prompt goes
|
|
957
|
+
|
|
958
|
+
| When you are… | Compile with this skill into… |
|
|
959
|
+
|---|---|
|
|
960
|
+
| Turning a vague idea into a brief | the `episode.workshop` **concept** string |
|
|
961
|
+
| Writing one shot's look | that shot's **`description`** in `script.json` (or `insert-shot --description`) |
|
|
962
|
+
| Writing a montage scene (the default lane) | the per-beat `[m:ss-m:ss] …` blocks the harness lays into `buildMontagePrompt`'s SEQUENCE section |
|
|
963
|
+
| Directing a voice/delivery | `character.voiceDesc` + the shot's `delivery` cue |
|
|
964
|
+
|
|
965
|
+
## Reconciliations — the harness owns these, do NOT hand-write them
|
|
966
|
+
|
|
967
|
+
The optimizer's "Non-Negotiable Principles" #7 (separate parameters from
|
|
968
|
+
creative content) is exactly right, and the harness enforces it. Concretely:
|
|
969
|
+
|
|
970
|
+
1. **Identity & references.** The harness builds the `@Image` slot plan
|
|
971
|
+
(`src/mini-drama/reference-slots.ts`) and pushes `reference_image_urls` in
|
|
972
|
+
that order. Name characters (`VIVIENNE`) and direct what they *do* — do NOT
|
|
973
|
+
write full physical descriptions or invent `@ImageN` numbers into a
|
|
974
|
+
`description`; the harness assigns them (AGENTS.md rules 19, 37, 42, 49).
|
|
975
|
+
Seedance 2.5 R2V takes up to **30** references, so the "up to 30 images"
|
|
976
|
+
budget in the optimizer's preflight is real here.
|
|
977
|
+
2. **Duration / resolution / aspect / audio toggle are parameters, never
|
|
978
|
+
Prompt text** (optimizer principle #7). The harness sets duration from the
|
|
979
|
+
plan (2.5 ladder: every integer 4-30s), pins `720p` for Seedance, and passes
|
|
980
|
+
`aspect_ratio` explicitly. Do not write "generate a 20-second 16:9 video"
|
|
981
|
+
into the Prompt. Timestamp ranges inside a montage SEQUENCE are creative
|
|
982
|
+
content and ARE kept — they double as the cutter's beat boundaries
|
|
983
|
+
(`GenerationUnit.montageBeats`).
|
|
984
|
+
3. **No music / diegetic sound only.** The optimizer's audio syntax (`()` music,
|
|
985
|
+
`<>` sfx, `{}` dialogue, `【】` subtitles) is compatible, but the harness owns
|
|
986
|
+
the music lane: keep the mandatory close "Diegetic sound only, no music, no
|
|
987
|
+
on-screen text." (montage) / the `negative_prompt` music suffix (singles),
|
|
988
|
+
and add music/ambient in post (`assemble.mix_audio`). Don't let the optimizer
|
|
989
|
+
talk you into baking a music bed into the Prompt.
|
|
990
|
+
|
|
991
|
+
## Montage SEQUENCE grammar = the optimizer's Long-Video stages
|
|
992
|
+
|
|
993
|
+
The harness's default lane (`buildMontagePrompt`, AGENTS.md rule 50) is the
|
|
994
|
+
"Make a full trailer with Seedance 2.5" grammar, which is the same shape as the
|
|
995
|
+
optimizer's **Long Videos and Time Ranges** template with explicit
|
|
996
|
+
`[0:03-0:05]` beats. When compiling a montage scene, follow the optimizer's
|
|
997
|
+
stage discipline (one primary state change per beat, an observable end state per
|
|
998
|
+
beat, consecutive non-overlapping ranges) and the harness will slice the render
|
|
999
|
+
at those exact timestamps.
|
|
1000
|
+
|
|
1001
|
+
## Do NOT import from the optimizer
|
|
1002
|
+
|
|
1003
|
+
- Any interface parameters into the Prompt body (durations, resolution, aspect,
|
|
1004
|
+
audio toggle) — those are harness-side (reconciliation #2).
|
|
1005
|
+
- Exhaustive identity descriptions or `@ImageN` tags you assigned by hand —
|
|
1006
|
+
the slot planner owns them (reconciliation #1).
|
|
1007
|
+
- A music/soundtrack instruction — the assembler owns music (reconciliation #3).
|
|
1008
|
+
|
|
1009
|
+
Everything else — intent-first compilation, one-role-per-material mapping,
|
|
1010
|
+
event scripts, keyframe/first-last-frame discipline, emotion/cinematography
|
|
1011
|
+
translation into observable results, and the final checklist — applies verbatim.
|
|
1012
|
+
|
|
1013
|
+
## Cross-references
|
|
1014
|
+
|
|
1015
|
+
- **AGENTS.md** rule 50 (montage-first / Seedance 2.5 default), rule 38 (direct,
|
|
1016
|
+
don't decorate), rules 19/37/42/49 (identity & spatial consistency).
|
|
1017
|
+
- `.agents/skills/venice-video-model-routing/SKILL.md` — model routing + the
|
|
1018
|
+
reference/audio capability matrix.
|
|
1019
|
+
- `src/mini-drama/prompt-builder.ts` — `buildMontagePrompt` / `buildMultiShotPrompt`
|
|
1020
|
+
(where the compiled creative content is assembled into the API body).
|