@kolbo/mcp 1.91.3 → 1.91.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@kolbo/mcp",
3
- "version": "1.91.3",
3
+ "version": "1.91.4",
4
4
  "description": "Kolbo AI MCP Server - Generate images, videos, music, speech, and sound effects from Claude Code",
5
5
  "main": "src/index.js",
6
6
  "bin": {
@@ -1,6 +1,6 @@
1
1
  # AUTO-GENERATED — do not edit
2
2
 
3
- This tree is mirrored from kolbo-code@c1a6a60, the single source of truth.
3
+ This tree is mirrored from kolbo-code@972d2b4, the single source of truth.
4
4
  Canonical source: packages/opencode/skills/kolbo/
5
5
  Distribution: .github/workflows/sync-skill-to-plugin.yml
6
6
 
package/skill/SKILL.md CHANGED
@@ -1,5 +1,5 @@
1
1
  ---
2
- version: 0.9.16
2
+ version: 0.9.17
3
3
  name: kolbo
4
4
  description: |
5
5
  Generate, edit, analyze, and direct creative media through Kolbo AI: images,
package/skill/VERSION CHANGED
@@ -1 +1 @@
1
- 0.9.16
1
+ 0.9.17
@@ -9,6 +9,9 @@
9
9
  "generation_mode": "reference-to-video",
10
10
  "control_density": "anchored",
11
11
  "duration_seconds": 8,
12
+ "shot_count": 1,
13
+ "continuous_take": true,
14
+ "post_voiceover": [],
12
15
  "aspect_ratio": "16:9",
13
16
  "active_assets": [
14
17
  {
@@ -2,6 +2,20 @@
2
2
 
3
3
  ## Pre-generation audit
4
4
 
5
+ Read the current creative brief, including rejected concepts and scene-specific
6
+ exceptions. Compare it to the final prompt and actual tool arguments, not the
7
+ assistant's explanation. Run `scripts/filmmaking/lint_prompt.py` against the shot
8
+ card for saved prompts. It checks Total duration/aspect/count, SHOT numbering,
9
+ continuous-take conflicts and exact multilingual tags. Optional card fields:
10
+ `shot_count`, `continuous_take`, dialogue items' `prompt_text` (exact text sent
11
+ to the model, including requested transliteration), and `post_voiceover` (a list
12
+ of exact narration strings excluded from generation). These checks do not prove
13
+ camera semantics, humor, visual quality or model adherence; review those separately.
14
+
15
+ Elements also rejects explicit Total declarations contradicting tool duration,
16
+ aspect or shot flags before submission. Resolve the mismatch against the brief;
17
+ do not strip declarations just to bypass validation. Free-form prompts remain supported.
18
+
5
19
  ### Story and edit
6
20
 
7
21
  - Does the shot have a necessary dramatic/editorial job?
@@ -11,6 +11,10 @@ Load this file when the user wants a **Seedance 2 / Seedance 2.0** (ByteDance) v
11
11
 
12
12
  **Elements uses this same file.** `generate_elements` is not a second prompt language. Do not write `SCENE CONTEXT` / `OPTICS` / `ACTION` department packs for Elements or Seedance — those are filmmaking audit contracts, not the generation compile shape.
13
13
 
14
+ ## Creative direction takes precedence
15
+
16
+ The current user brief overrides template defaults and illustrative examples. Keep the two-layer organization, but include only relevant locks. State concrete camera trajectory and visible action prominently; optics numbers, equipment names, repetition and word counts are not guarantees of fidelity. Preserve a continuous-shot exception even when other scenes are multishot. For one shot use `Single continuous shot`, `Total: Xs / 1 shot / AR`, one SHOT heading and `multi_shots: false`; for multiple shots use `Multishot ON` and matching counts. AR comes from the brief, never a copied example. Keep dialogue in the user's requested language or phonetic spelling; test pronunciation rather than claiming guaranteed support or impossibility. Narration reserved for post does not belong in the generation prompt.
17
+
14
18
  ## Universal Rules (apply to EVERY Seedance / Elements prompt)
15
19
 
16
20
  - **NO MUSIC BY DEFAULT (HARD):** Unless the user explicitly asks for music, every final Seedance prompt—including every Elements/reference-driven prompt—must explicitly say `No music. No musical score.` Keep requested dialogue, synchronized production sound, ambience, and SFX; "no music" does not mean "no audio." If the user explicitly requests music, describe that music instead and omit the no-music lock. Never invent background music from cinematic tone alone.
@@ -26,13 +30,13 @@ Load this file when the user wants a **Seedance 2 / Seedance 2.0** (ByteDance) v
26
30
  - A prompt with only shot body and no Total / Multishot header is a **failed turn** — rewrite before calling `generate_*`.
27
31
  - **MCP `duration` must match the Total line.** Pass `duration: X` (whole seconds) on `generate_video` / `generate_elements` / `generate_video_from_image` equal to the `Xs` in `Total: Xs / …`. Mismatch = wrong-length clip.
28
32
  - **Then the Locked Intro** — `[GLOBAL LOOK]` / `[CAST]` / `[LOCATION]` (+ LOCATION MAP / CONTINUITY / PHYSICS for multi-shot) — before any shot. A one-liner `same character throughout` is not a character lock.
29
- - **Order inside each shot**: Subject → Action → Camera → Constraints → (Audio/SFX if relevant). Do NOT restack GLOBAL LOOK style inside the shot.
30
- - **Prompt length**: simple single-idea pieces ~120–280 words. Locked-intro cinematic typically 400–900 words. Shorter than ~120 words = random output. The 10,000-char cap below always wins.
33
+ - **Inside each shot:** make the camera trajectory, subject action and timing easy to find. Put a requested signature camera move in the heading. Do not restack GLOBAL LOOK style inside the shot.
34
+ - **Prompt length**: simple single-idea pieces ~120–280 words. Locked-intro cinematic typically 400–900 words. These are examples, not minimum lengths. Do not pad. The selected model catalog cap wins; Seedance 2.5 uses its own adapter.
31
35
  - **Shot count is user-directed.** If the user asks for N shots, deliver exactly N in one prompt unless they ask to split.
32
- - **Always describe at least one camera movement per shot.**
36
+ - **Describe the camera behavior requested for each shot.** Preserve intentional static shots; do not replace requested dynamic moves with static dialogue coverage.
33
37
  - **Tell Seedance what the camera is NOT doing** (e.g. `no cuts, no zoom, natural head movement`) — this is what locks POV.
34
- - **Final prompt is always English**, wrapped in a copy-ready code block. Detect intent in any language and reply in the user's language, but the prompt itself is English.
35
- - **HARD CAP: 10,000 characters TOTAL for the ENTIRE prompt** — measured as one single string including all shots, boilerplate, SFX lines, and the Total lines. It is per PROMPT, not per shot. **Never** split into multiple prompts, code blocks, or "part 1 / part 2" to evade the cap. Count the final prompt before output; if over, trim (cut adjectives, collapse boilerplate, shorten SFX lists, merge or drop shots) and re-count until it fits.
38
+ - **Control prose defaults to English; exact dialogue and asset tags retain the user-requested language**, wrapped in a copy-ready code block. Reply in the user's language; preserve any explicit request for the prompt language too.
39
+ - **HARD CAP: 10,000 characters TOTAL for the ENTIRE prompt** — measured as one single string including all shots, boilerplate, SFX lines, and the Total lines. It is per PROMPT, not per shot. **Never** split into multiple prompts, code blocks, or "part 1 / part 2" to evade the cap. Count the final prompt before output; if over, trim (cut adjectives, collapse boilerplate, shorten SFX lists, tighten redundant shot prose without dropping user-requested shots) and re-count until it fits.
36
40
 
37
41
  ## Locked Intro (DEFAULT for any multi-shot cinematic — including Elements)
38
42
 
@@ -138,7 +142,7 @@ Appearance locks WHO. Persona locks HOW THEY BEHAVE — without it Seedance rend
138
142
  ## Dialogue & expression
139
143
 
140
144
  - **Dialogue is PERFORMED by the model, never by a TTS tool.** Quoted lines in the prompt come back as synced speech with lip movement and room tone, together with the SFX you name in AUDIO. Scene dialogue therefore never routes through `generate_speech` or `generate_lipsync` — write the line in quotes inside its shot beat and let Seedance act it.
141
- - **Write dialogue in ENGLISH.** Seedance does not reliably perform other languages, and Hebrew in particular does not work — it comes back as accented gibberish or English-shaped mouth movement. Never offer a user "Hebrew dialogue directly". If the delivered film has to be Hebrew, the honest routes are: (a) keep the spoken lines English, or (b) stage the beat as expression + on-screen text, or (c) generate the scene clean and dub it afterwards as an explicit, separately-priced pass. Say which one you are doing.
145
+ - **Preserve requested dialogue and its language.** If the user requests Hebrew in Latin letters, preserve that phonetic text as dialogue, not an English translation. Native pronunciation and lip-sync require actual output inspection. Do not promise success or claim the language is impossible without current evidence. Offer a separately authorized dubbing pass only when needed; keep narration reserved for post out of the prompt.
142
146
  - `list_models` reports `sound_generation_type: "none"` for Seedance 2 / 2.5 because there is no in-app sound toggle (`sound_baked_in: true`). That field does NOT mean the model is silent. Do not read it as a reason to add TTS.
143
147
  - For silent tension, deliver it as expression, not speech: `He does not speak. His expression clearly says: "…"`.
144
148
 
@@ -231,7 +235,7 @@ Use only the discrete steps. Not "23°" — use 18° or 29°.
231
235
 
232
236
  ### Camera placement
233
237
 
234
- Place CAMERA in the **3rd position** of each shot's core layers (Subject → Action → Camera → Style → Constraints). FOV gets ignored at the end, conflicts with identity at the front.
238
+ Place a requested signature CAMERA trajectory prominently in the shot heading, then describe its timing with the subject action. GLOBAL LOOK owns shared lens / stock / grade. No fixed word order guarantees adherence.
235
239
 
236
240
  ### Pre-flight checklist (before output)
237
241
 
@@ -11,7 +11,7 @@ Load this file when the user wants a **Seedance 2.5** video (they said "2.5" / "
11
11
 
12
12
  **Audio:** Seedance 2.5 emits real synced audio. `list_models` shows `sound_generation_type: none` only because there is no in-app toggle (`sound_baked_in: true`) — it does NOT mean the model is silent, and it is never a reason to reach for TTS. Quoted dialogue is PERFORMED (synced voices, lip movement, room tone) alongside the SFX named in AUDIO, so scene dialogue never goes through `generate_speech` or `generate_lipsync`; write the lines in quotes inside their shot beats.
13
13
 
14
- **Dialogue language: English.** Other languages are not reliably performed, and Hebrew does not work — it returns accented gibberish or English-shaped mouth movement. Never offer a user "Hebrew dialogue directly". This restriction applies to rendered dialogue/prose only, never to binding identifiers: preserve an exact stored Hebrew Visual DNA or moodboard tag such as `@אביב` / `#ישראל` literally. See `models/seedance.md` for the three honest alternatives.
14
+ **Dialogue language follows the user.** Preserve requested Hebrew or Hebrew-in-Latin transliteration; do not translate it into English or change the spoken content. Pronunciation and lip-sync must be inspected in the generated output, not guaranteed from the prompt. Keep post-production VO out of the generation prompt. Asset tags always retain their exact stored spelling, including `@אביב` / `#ישראל` literally.
15
15
 
16
16
  **Use the cheapest supported tier unless the user selected an output resolution.** Resolution is a credit MULTIPLIER, not a flat rate. Relative to 720p: 480p ×0.44, 1080p ×2.25. A 30s pass costs ~540cr at 480p against ~1230cr at 720p and ~2770cr at 1080p. When no output resolution was selected and 480p is the cheapest supported tier, block the film at 480p, get the user's sign-off on staging, performance and timing, then re-run only the approved cut at a higher delivery resolution if the user explicitly authorizes that resolution increase. Approval of the creative cut alone does not authorize a more expensive resolution. If no output resolution was selected, use the cheapest supported tier from the live catalog even for final work; pass it explicitly.
17
17
 
@@ -27,6 +27,8 @@ Load this file when the user wants a **Seedance 2.5** video (they said "2.5" / "
27
27
 
28
28
  ## Universal Rules (HARD — same as help widget OUTPUT CONTRACT)
29
29
 
30
+ User-selected shot structure wins over examples. One continuous take uses `Single continuous shot`, `Total: Xs / 1 shot / AR`, one SHOT heading and `multi_shots: false`. Use Multishot ON only for multiple shots. Copy aspect, duration, dialogue and camera direction from the current scene brief. Expand craft blocks only when they resolve a real staging need; do not pad or introduce contradictory locks.
31
+
30
32
  - **First lines ALWAYS declare shot structure** (text-to-video / Elements / reference gen — NOT video-edit):
31
33
  1. `N connected cinematic shots, Xs total, AR, Multishot ON`
32
34
  2. `Total: Xs / N shots / AR`
@@ -6,6 +6,10 @@ Operate as a filmmaking system, not merely a prompt writer. Preserve project tru
6
6
 
7
7
  ## Start here
8
8
 
9
+ For a continuing production, read `.kolbo/creative-brief.md` before compiling or revising scenes. Keep it as a compact working table: stable scene ID, current world/action/tone, first-frame and camera trajectory/framing, shot count, duration/aspect, asset IDs and exact tags, spoken lines versus VO reserved for post, approval state, rejected concepts, and pending job IDs. Update only the dimensions changed by the user. Preserve the brief across compaction; do not put unapproved generations into `.kolbo/production.md`.
10
+
11
+ User direction wins over template defaults and examples. A custom skill supplements this workflow; do not require its missing supporting files without explaining the gap, and never claim to have read them. Check the final prompt and tool arguments against the brief before dispatch. A request for active movement does not require running; bright lighting does not imply pastel colors or restrained action. Replace rejected worlds substantively. Preserve approved scenes and local single-shot exceptions.
12
+
9
13
  1. Identify the requested production stage and deliverable.
10
14
  2. Read only the reference files required by the routing table below.
11
15
  3. Preserve or establish the relevant production truth before writing a shot.
@@ -115,6 +119,8 @@ When a generation fails:
115
119
  5. Log the change and verdict.
116
120
  6. After repeated failures, redesign the shot: bake the state into an asset, add a staging/layout reference, reduce actions, split the shot, change the angle, or switch model/mode.
117
121
 
122
+ Validate the revised approach on one representative shot before another batch, within existing authorization and budget; an explicit request for the whole batch wins. Do not silently split a user-requested continuous take, switch their selected model, or add paid tests. Distinguish defects observed in video/audio from hypotheses inferred from the prompt. Completion alone never verifies camera movement, cuts or performance.
123
+
118
124
  Do not keep polishing adjectives when the shot is physically or structurally overconstrained.
119
125
 
120
126
  ## Validate
@@ -8,6 +8,10 @@ Visual DNA profiles capture the visual "identity" of a character, style, product
8
8
 
9
9
  ## Workflow
10
10
 
11
+ ### Creation failure recovery
12
+
13
+ Create DNA through the available tool when the user requested it; do not default to asking for manual wizard work. A reference-preparation error **before submission** means this attempt did not create a profile. For an uncertain submission, reconcile with a personal/project-scoped list and inspect the matching profile before retrying. Existence alone does not establish that an errored call created it. Stop repeating an identical runtime error; report the exact failure and the verified scope, without inventing an auth outage or claiming reconnect will fix it. After success, check the returned ID, exact stored name, type and references; correct confirmed metadata errors in place rather than recreating. A headless wardrobe sheet for a recurring person remains character DNA.
14
+
11
15
  1. **Sheet first, then DNA.** For any production asset (character / location / prop), resolve the sheet **preset** (`list_presets` with `search`) and `generate_image` with that `preset_id` — custom instructions live on the preset. Then `create_visual_dna` with the sheet as `character_sheet_url` (max 4 extra images — if the user gives more, pick the 4 most representative **that share the same identity and vibe**; never pass 5+). Optionally video and audio. See **Purity** below before you generate those stills.
12
16
  2. **Types**: `character` (default), `style`, `product`, `scene`, `environment`.
13
17
  3. **Use** the profile by passing its `id` in `visual_dna_ids` in: `generate_image`, `generate_creative_director`, `generate_elements`, `generate_video_from_image`, `generate_video_from_video`, `generate_first_last_frame`.
@@ -10,8 +10,9 @@ from pathlib import Path
10
10
  from typing import Any
11
11
 
12
12
 
13
- TAG_RE = re.compile(r"@[A-Za-z][A-Za-z0-9_.:-]*(?:\s+\d+)?")
13
+ TAG_RE = re.compile(r"@[\w][\w.:-]*(?:\s+\d+)?", re.UNICODE)
14
14
  SHOT_RE = re.compile(r"(?im)^\s*(?:SHOT|SEGMENT)\s+(\d+)\b")
15
+ TOTAL_RE = re.compile(r"(?im)^\s*Total:\s*(\d+(?:\.\d+)?)s\s*/\s*(\d+)\s*shots?\s*/\s*(\d+:\d+)\s*$")
15
16
  RANGE_RE = re.compile(
16
17
  r"(?i)(\d+(?:\.\d+)?)\s*s\s*(?:-|–|—|to)\s*(\d+(?:\.\d+)?)\s*s"
17
18
  )
@@ -98,7 +99,7 @@ def check_adapter(
98
99
 
99
100
  def check_contradictions(prompt: str, report: Report) -> None:
100
101
  continuous = re.search(
101
- r"(?i)\b(one|single)\s+(?:unbroken\s+)?continuous\s+take\b|\bno\s+cuts?\b",
102
+ r"(?i)\b(one|single)\s+(?:unbroken\s+)?continuous\s+(?:take|shot)\b|\bno\s+cuts?\b",
102
103
  prompt,
103
104
  )
104
105
  cuts = re.search(r"(?i)\b(hard|match|smash|jump|whip)\s+cut\b|\bhard\s+cuts\b", prompt)
@@ -144,8 +145,18 @@ def check_stale_language(prompt: str, report: Report) -> None:
144
145
 
145
146
 
146
147
  def check_tags(prompt: str, card: dict[str, Any] | None, report: Report) -> tuple[set[str], set[str]]:
147
- found = {value.rstrip(".,;:!?") for value in TAG_RE.findall(prompt)}
148
148
  expected = expected_tags(card)
149
+ # Exact stored names may contain spaces, punctuation and non-Latin letters.
150
+ # Match those first, then scan the remainder for unregistered tags. Never
151
+ # normalize a canonical name, or accept @maya2 as a match for @maya.
152
+ remaining = prompt
153
+ found: set[str] = set()
154
+ for tag in sorted(expected, key=len, reverse=True):
155
+ pattern = re.escape(tag) + r"(?![\w:-])"
156
+ if re.search(pattern, remaining):
157
+ found.add(tag)
158
+ remaining = re.sub(pattern, "", remaining)
159
+ found.update(value.rstrip(".,;:!?") for value in TAG_RE.findall(remaining))
149
160
  for tag in sorted(expected - found):
150
161
  report.error("missing_active_asset", f"Shot-card asset is absent from prompt: {tag}")
151
162
  for tag in sorted(found - expected):
@@ -162,6 +173,40 @@ def check_tags(prompt: str, card: dict[str, Any] | None, report: Report) -> tupl
162
173
  return found, expected
163
174
 
164
175
 
176
+ def check_contract(prompt: str, card: dict[str, Any] | None, report: Report) -> None:
177
+ """Check explicit structure, not subjective direction or rendered quality."""
178
+ totals = [(float(seconds), int(count), aspect) for seconds, count, aspect in TOTAL_RE.findall(prompt)]
179
+ if totals and any(total != totals[0] for total in totals):
180
+ report.error("conflicting_totals", "Opening and closing Total declarations disagree.")
181
+ shots = [int(value) for value in re.findall(r"(?im)^\s*SHOT\s+(\d+)\b", prompt)]
182
+ if shots and shots != list(range(1, len(shots) + 1)):
183
+ report.error("shot_sequence", "SHOT headings must be unique and sequential from 1.")
184
+ if totals and shots and totals[0][1] != len(shots):
185
+ report.error("shot_count_mismatch", "Total shot count differs from the SHOT headings.")
186
+ if not card:
187
+ return
188
+ for duration, count, aspect in totals:
189
+ if card.get("duration_seconds") is not None and duration != card["duration_seconds"]:
190
+ report.error("duration_mismatch", "Prompt Total duration differs from the shot card.")
191
+ if card.get("aspect_ratio") and aspect != card["aspect_ratio"]:
192
+ report.error("aspect_mismatch", "Prompt Total aspect differs from the shot card.")
193
+ if card.get("shot_count") is not None and count != card["shot_count"]:
194
+ report.error("brief_shot_count_mismatch", "Prompt Total shot count differs from the shot card.")
195
+ if card.get("shot_count") is not None and shots and len(shots) != card["shot_count"]:
196
+ report.error("brief_shot_count_mismatch", "SHOT headings differ from the shot card count.")
197
+ if card.get("continuous_take") is True:
198
+ if len(shots) > 1 or any(count != 1 for _, count, _ in totals) or re.search(r"(?i)\bMultishot\s+ON\b", prompt):
199
+ report.error("brief_continuous_take", "The brief requires one continuous take; prompt declares multiple shots.")
200
+ for item in card.get("dialogue", []):
201
+ if isinstance(item, dict):
202
+ text = item.get("prompt_text")
203
+ if isinstance(text, str) and text not in prompt:
204
+ report.error("missing_dialogue", "Exact prompt_text from the dialogue card is missing or rewritten.")
205
+ for text in card.get("post_voiceover", []):
206
+ if isinstance(text, str) and text and text in prompt:
207
+ report.error("post_voiceover_in_prompt", "Narration reserved for post-production appears in the generation prompt.")
208
+
209
+
165
210
  def check_timing(prompt: str, card: dict[str, Any] | None, report: Report) -> None:
166
211
  duration = card.get("duration_seconds") if card else None
167
212
  for start_text, end_text in RANGE_RE.findall(prompt):
@@ -220,6 +265,7 @@ def main() -> int:
220
265
  found, expected = check_tags(prompt, card, report)
221
266
  check_timing(prompt, card, report)
222
267
  check_dialogue_card(card, report)
268
+ check_contract(prompt, card, report)
223
269
 
224
270
  result = {
225
271
  "prompt": str(args.prompt.resolve()),
@@ -24,7 +24,10 @@ const fs = require('fs');
24
24
  const path = require('path');
25
25
  const net = require('net');
26
26
  const dns = require('dns').promises;
27
- const { Agent, fetch: undiciFetch } = require('undici');
27
+ // Never load Bun's incomplete built-in undici shim. Node uses the installed
28
+ // dispatcher below; Bun uses native fetch pinned to vetted IPs (its node:tls
29
+ // compatibility transport can return empty peer certificates intermittently).
30
+ const { Agent, fetch: undiciFetch } = require('undici/index.js');
28
31
 
29
32
  const MAX_FILE_BYTES = 500 * 1024 * 1024; // 500 MB — larger than visual_dna because
30
33
  // lipsync/v2v/transcription accept full
@@ -201,42 +204,73 @@ async function resolvePublicAddresses(hostname) {
201
204
  return rows;
202
205
  }
203
206
 
204
- function pinnedDispatcher(addresses) {
207
+ function pinnedDispatcher(addresses) {
205
208
  let cursor = 0;
206
- return new Agent({
207
- connect: {
208
- lookup(_hostname, options, callback) {
209
+ return new Agent({
210
+ connect: {
211
+ lookup(_hostname, options, callback) {
209
212
  if (options?.all) return callback(null, addresses);
210
213
  const row = addresses[cursor++ % addresses.length];
211
214
  return callback(null, row.address, row.family);
212
215
  },
213
216
  },
214
217
  });
215
- }
218
+ }
219
+
220
+ async function release(dispatcher) {
221
+ try { await dispatcher.close(); } catch (_) { /* Cleanup must not mask fetch errors. */ }
222
+ }
223
+
224
+ async function bunFetch(url, addresses, signal) {
225
+ let error;
226
+ for (const row of addresses) {
227
+ const pinned = new URL(url);
228
+ pinned.hostname = row.family === 6 ? `[${row.address}]` : row.address;
229
+ try {
230
+ return await fetch(pinned, {
231
+ redirect: 'manual', signal,
232
+ // Disable environment proxy routing: the connection must use the
233
+ // checked address, while HTTP routing and TLS verification use the
234
+ // original hostname. Never disable certificate verification.
235
+ proxy: '',
236
+ headers: { Host: url.host },
237
+ tls: url.protocol === 'https:' ? {
238
+ serverName: url.hostname.replace(/^\[|\]$/g, ''), rejectUnauthorized: true,
239
+ } : undefined,
240
+ });
241
+ } catch (err) {
242
+ error = err;
243
+ if (signal?.aborted) throw err;
244
+ }
245
+ }
246
+ throw error;
247
+ }
216
248
 
217
249
  async function safeFetch(rawUrl, opts = {}) {
218
250
  let current = rawUrl;
219
251
  for (let i = 0; i <= MAX_REDIRECTS; i++) {
220
252
  const url = assertSafeUrl(current);
221
253
  const addresses = await resolvePublicAddresses(url.hostname);
222
- const dispatcher = pinnedDispatcher(addresses);
254
+ const dispatcher = process.versions.bun ? null : pinnedDispatcher(addresses);
223
255
  let res;
224
256
  try {
225
- res = await undiciFetch(current, { redirect: 'manual', signal: opts.signal, dispatcher });
257
+ res = process.versions.bun
258
+ ? await bunFetch(url, addresses, opts.signal)
259
+ : await undiciFetch(current, { redirect: 'manual', signal: opts.signal, dispatcher });
226
260
  } catch (err) {
227
- await dispatcher.close().catch(() => {});
261
+ if (dispatcher) await release(dispatcher);
228
262
  throw err;
229
263
  }
230
264
  if (res.status >= 300 && res.status < 400 && res.headers.get('location')) {
231
265
  const next = new URL(res.headers.get('location'), current).toString();
232
266
  await res.body?.cancel().catch(() => {});
233
- await dispatcher.close().catch(() => {});
267
+ if (dispatcher) await release(dispatcher);
234
268
  current = next;
235
269
  continue;
236
270
  }
237
271
  // close() waits for this response body to be consumed, so schedule it but
238
272
  // do not await it before returning the Response to the caller.
239
- dispatcher.close().catch(() => {});
273
+ if (dispatcher) void release(dispatcher);
240
274
  return res;
241
275
  }
242
276
  throw new Error(`Too many redirects fetching ${rawUrl}`);
@@ -11,6 +11,7 @@ const { ownedUrl } = require('./owned-url');
11
11
  const { UI, uiResult, canonicalModelId, assertModelSupportsType, modelInfo, voiceInfo, resolveCatalogAspectRatio } = require('../apps');
12
12
  const { modelTypeForEditOperation, assertExecutableEditModel } = require('./editModelCatalog');
13
13
  const { withLocalRehost } = require('./local-rehost');
14
+ const { validateVideoPrompt } = require('./video-prompt');
14
15
 
15
16
  // ─── Cinematic Dimensions schema (shared by generate_image + generate_image_edit) ───
16
17
  // Kolbo's "Cinema mode": eight independent photographic dimensions, each an OPTIONAL
@@ -1454,8 +1455,10 @@ function registerGenerateTools(server, client, options = {}) {
1454
1455
  session_id: sessionIdField
1455
1456
  },
1456
1457
  async ({ prompt, model, reference_images, reference_videos, reference_audio_urls, audio_url, files, duration, aspect_ratio, motion, preset_id, enhance_prompt = false, visual_dna_ids, resolution, sound_enabled, keyframes, multi_shots, multi_shot_count, session_name, project_id, session_id }) => {
1458
+ validateVideoPrompt({ prompt, duration, aspect_ratio, multi_shots, multi_shot_count });
1457
1459
  model = await canonicalModelId(client, model, 'elements'); // lenient id resolution ("z-image" → "z-image/turbo")
1458
1460
  aspect_ratio = await resolveCatalogAspectRatio(client, model, aspect_ratio, 'elements');
1461
+ validateVideoPrompt({ prompt, duration, aspect_ratio, multi_shots, multi_shot_count });
1459
1462
  if (!prompt) throw new Error('prompt is required');
1460
1463
 
1461
1464
  // Elements is the one tool that takes all three modalities, and either a
@@ -0,0 +1,20 @@
1
+ // Validate only explicit, machine-readable declarations. Never infer creative
2
+ // intent from adjectives, require a template, or rewrite a user's prompt.
3
+ function validateVideoPrompt({ prompt, duration, aspect_ratio, multi_shots, multi_shot_count }) {
4
+ const totals = [...prompt.matchAll(/^\s*Total:\s*(\d+(?:\.\d+)?)s\s*\/\s*(\d+)\s*shots?\s*\/\s*(\d+:\d+)\s*$/gim)]
5
+ .map((m) => ({ duration: Number(m[1]), count: Number(m[2]), aspect: m[3] }));
6
+ if (!totals.length) return; // Existing free-form clients remain supported.
7
+ const total = totals[0];
8
+ const errors = [];
9
+ if (totals.some((t) => t.duration !== total.duration || t.count !== total.count || t.aspect !== total.aspect)) errors.push('Total declarations disagree');
10
+ if (duration != null && duration !== total.duration) errors.push('duration differs from prompt Total');
11
+ if (/^\d+:\d+$/.test(aspect_ratio || '') && aspect_ratio !== total.aspect) errors.push('aspect_ratio differs from prompt Total');
12
+ if (multi_shots === false && total.count > 1) errors.push('multi_shots=false conflicts with multiple declared shots');
13
+ if (multi_shots === true && total.count === 1) errors.push('multi_shots=true conflicts with one declared shot');
14
+ if (multi_shot_count != null && multi_shot_count !== total.count) errors.push('multi_shot_count differs from prompt Total');
15
+ const shots = [...prompt.matchAll(/^\s*SHOT\s+(\d+)\b/gim)].map((m) => Number(m[1]));
16
+ if (shots.length && (shots.length !== total.count || shots.some((n, i) => n !== i + 1))) errors.push('SHOT headings disagree with declared count or order');
17
+ if (errors.length) throw new Error(`Video prompt conflict before submission: ${errors.join('; ')}. Reconcile with the user's approved brief; no generation was submitted.`);
18
+ }
19
+
20
+ module.exports = { validateVideoPrompt };
@@ -54,11 +54,13 @@ function registerVisualDnaTools(server, client, options = {}) {
54
54
  }
55
55
 
56
56
  // Resolve all sources to buffers in parallel.
57
- const [imageFiles, videoFile, audioFile] = await Promise.all([
57
+ const [imageFiles, videoFile, audioFile] = await Promise.all([
58
58
  Promise.all(imageList.map(src => resolveToBuffer(src, 'image', options))),
59
59
  video ? resolveToBuffer(video, 'video', options) : Promise.resolve(null),
60
60
  audio ? resolveToBuffer(audio, 'audio', options) : Promise.resolve(null)
61
- ]);
61
+ ]).catch((cause) => {
62
+ throw new Error(`Visual DNA reference preparation failed before submission; this attempt did not create a profile. ${cause.message}`, { cause });
63
+ });
62
64
 
63
65
  const form = new FormData();
64
66
  form.append('name', name);