comfyui-mcp 0.52.13 → 0.52.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/README.md +3 -2
  2. package/dist/boot.js +6 -0
  3. package/dist/boot.js.map +1 -1
  4. package/dist/handshake-instructions.js +68 -0
  5. package/dist/handshake-instructions.js.map +1 -0
  6. package/dist/orchestrator/agent-backend.js +17 -0
  7. package/dist/orchestrator/agent-backend.js.map +1 -1
  8. package/dist/orchestrator/backend-readiness.js +43 -0
  9. package/dist/orchestrator/backend-readiness.js.map +1 -1
  10. package/dist/orchestrator/index.js +62 -38
  11. package/dist/orchestrator/index.js.map +1 -1
  12. package/dist/orchestrator/panel-mcp-http.js +5 -1
  13. package/dist/orchestrator/panel-mcp-http.js.map +1 -1
  14. package/dist/orchestrator/panel-tools.js +105 -5
  15. package/dist/orchestrator/panel-tools.js.map +1 -1
  16. package/dist/orchestrator/qwen-backend.js +1029 -0
  17. package/dist/orchestrator/qwen-backend.js.map +1 -0
  18. package/dist/orchestrator/rgthree-fast-groups-property.js +53 -0
  19. package/dist/orchestrator/rgthree-fast-groups-property.js.map +1 -0
  20. package/dist/services/comfy-view-ref.js +236 -4
  21. package/dist/services/comfy-view-ref.js.map +1 -1
  22. package/dist/services/panel-launcher.js +5 -1
  23. package/dist/services/panel-launcher.js.map +1 -1
  24. package/dist/services/ui-bridge.js +5 -0
  25. package/dist/services/ui-bridge.js.map +1 -1
  26. package/dist/tools/vocabulary.js +1 -1
  27. package/docs/design/panel-surface.txt +1 -0
  28. package/locales/ar/main.json +2 -1
  29. package/locales/en/main.json +2 -1
  30. package/locales/es/main.json +2 -1
  31. package/locales/fa/main.json +2 -1
  32. package/locales/fr/main.json +2 -1
  33. package/locales/ja/main.json +2 -1
  34. package/locales/ko/main.json +2 -1
  35. package/locales/pt-BR/main.json +2 -1
  36. package/locales/ru/main.json +2 -1
  37. package/locales/tr/main.json +2 -1
  38. package/locales/zh/main.json +2 -1
  39. package/locales/zh-TW/main.json +2 -1
  40. package/package.json +1 -1
  41. package/plugin/.mcp.json +2 -3
  42. package/plugin/scripts/launch-server.mjs +130 -0
  43. package/plugin/skills/minimax-h3-video/SKILL.md +270 -0
  44. package/plugin/skills/model-registry/SKILL.md +16 -0
  45. package/plugin/skills/rgthree/SKILL.md +26 -1
@@ -164,7 +164,8 @@
164
164
  "minimax": "O agente em segundo plano não está respondendo — o MiniMax não conseguiu iniciar. Defina MINIMAX_API_KEY a partir de platform.minimax.io; depois Desconectar → Conectar para tentar de novo.",
165
165
  "moonshot": "O agente em segundo plano não está respondendo — o Moonshot (Kimi K3) não conseguiu iniciar. Defina MOONSHOT_API_KEY a partir de platform.kimi.ai; depois Desconectar → Conectar para tentar de novo.",
166
166
  "ollama": "O agente em segundo plano não está respondendo — o Ollama não está acessível. Suba-o com `ollama serve` e baixe o nosso modelo ajustado (`ollama pull artokun/gemma4-comfyui-mcp:e4b` — gemma4 treinado no conjunto de ferramentas do comfyui-mcp, o melhor modelo local da arena; `:12b` para cerca de 8 GB de VRAM); depois Desconectar → Conectar para tentar de novo.",
167
- "pi": "O agente em segundo plano não está respondendo — a CLI do pi não conseguiu executar `pi --list-models`. Instale a partir de https://pi.dev (`curl -fsSL https://pi.dev/install.sh | sh`), configure um provedor (defina uma chave de API de provedor ou rode `pi` uma vez e `/login`); depois Desconectar → Conectar para tentar de novo."
167
+ "pi": "O agente em segundo plano não está respondendo — a CLI do pi não conseguiu executar `pi --list-models`. Instale a partir de https://pi.dev (`curl -fsSL https://pi.dev/install.sh | sh`), configure um provedor (defina uma chave de API de provedor ou rode `pi` uma vez e `/login`); depois Desconectar → Conectar para tentar de novo.",
168
+ "qwen": "O agente em segundo plano não está respondendo — a CLI do Qwen Code não pôde ser iniciada. Verifique se o Qwen Code está instalado (npm i -g @qwen-code/qwen-code) e conectado (rode `qwen` uma vez e conclua o /auth, ou defina DASHSCOPE_API_KEY); depois Desconectar → Conectar para tentar de novo."
168
169
  },
169
170
  "llamacpp_no_jinja": "Seu llama-server está rodando SEM `--jinja`, então a chamada de ferramentas está desativada — qualquer ação do agente falharia. Reinicie-o com suporte a ferramentas: `llama-server -m <model>.gguf --jinja -c 16384` (builds atuais habilitam jinja por padrão; os mais antigos precisam da flag); depois Desconectar → Conectar.",
170
171
  "llamacpp_small_context": "O contexto do seu llama-server é de {tokens} tokens (flag de inicialização -c). A carga de ferramentas do agente precisa de pelo menos 16384 — abaixo disso, turnos longos são truncados silenciosamente. Considere reiniciar com `-c 16384` ou mais.",
@@ -164,7 +164,8 @@
164
164
  "minimax": "Фоновый агент не отвечает — не удалось запустить MiniMax. Задайте MINIMAX_API_KEY с platform.minimax.io, затем «Отключиться» → «Подключиться», чтобы повторить.",
165
165
  "moonshot": "Фоновый агент не отвечает — не удалось запустить Moonshot (Kimi K3). Задайте MOONSHOT_API_KEY с platform.kimi.ai, затем «Отключиться» → «Подключиться», чтобы повторить.",
166
166
  "ollama": "Фоновый агент не отвечает — Ollama недоступна. Запустите её командой `ollama serve` и загрузите нашу дообученную модель (`ollama pull artokun/gemma4-comfyui-mcp:e4b` — gemma4, обученная на наборе инструментов comfyui-mcp, лучшая локальная модель по результатам арены; `:12b` для примерно 8 ГБ видеопамяти), затем «Отключиться» → «Подключиться», чтобы повторить.",
167
- "pi": "Фоновый агент не отвечает — pi CLI не смог выполнить `pi --list-models`. Установите его с https://pi.dev (`curl -fsSL https://pi.dev/install.sh | sh`), настройте провайдера (задайте ключ API провайдера или запустите `pi` один раз и выполните `/login`), затем «Отключиться» → «Подключиться», чтобы повторить."
167
+ "pi": "Фоновый агент не отвечает — pi CLI не смог выполнить `pi --list-models`. Установите его с https://pi.dev (`curl -fsSL https://pi.dev/install.sh | sh`), настройте провайдера (задайте ключ API провайдера или запустите `pi` один раз и выполните `/login`), затем «Отключиться» → «Подключиться», чтобы повторить.",
168
+ "qwen": "Фоновый агент не отвечает — Qwen Code CLI не удалось запустить. Убедитесь, что Qwen Code установлен (npm i -g @qwen-code/qwen-code) и выполнен вход (запустите `qwen` один раз и завершите /auth или задайте DASHSCOPE_API_KEY), затем «Отключиться» → «Подключиться», чтобы повторить."
168
169
  },
169
170
  "llamacpp_no_jinja": "Ваш llama-server запущен БЕЗ `--jinja`, поэтому вызов инструментов отключён — любое действие агента завершится ошибкой. Перезапустите его с поддержкой инструментов: `llama-server -m <model>.gguf --jinja -c 16384` (в свежих сборках jinja включён по умолчанию, старым нужен флаг), затем «Отключиться» → «Подключиться».",
170
171
  "llamacpp_small_context": "Контекст вашего llama-server составляет {tokens} токенов (флаг запуска -c). Инструментам агента нужно не меньше 16384 — ниже этого длинные ходы будут молча обрезаться. Стоит перезапустить с `-c 16384` или больше.",
@@ -164,7 +164,8 @@
164
164
  "minimax": "Arka plandaki aracı yanıt vermiyor — MiniMax başlatılamadı. platform.minimax.io adresinden aldığınız MINIMAX_API_KEY değerini ayarlayın, sonra Bağlantıyı kes → Bağlan ile yeniden deneyin.",
165
165
  "moonshot": "Arka plandaki aracı yanıt vermiyor — Moonshot (Kimi K3) başlatılamadı. platform.kimi.ai adresinden aldığınız MOONSHOT_API_KEY değerini ayarlayın, sonra Bağlantıyı kes → Bağlan ile yeniden deneyin.",
166
166
  "ollama": "Arka plandaki aracı yanıt vermiyor — Ollama'ya erişilemiyor. `ollama serve` ile başlatın ve ince ayarlı modelimizi indirin (`ollama pull artokun/gemma4-comfyui-mcp:e4b` — comfyui-mcp araç takımıyla eğitilmiş gemma4, arenadaki en iyi yerel model; yaklaşık 8 GB VRAM için `:12b`), sonra Bağlantıyı kes → Bağlan ile yeniden deneyin.",
167
- "pi": "Arka plandaki aracı yanıt vermiyor — pi CLI `pi --list-models` komutunu çalıştıramadı. https://pi.dev adresinden kurun (`curl -fsSL https://pi.dev/install.sh | sh`), bir sağlayıcı yapılandırın (bir sağlayıcı API anahtarı ayarlayın ya da `pi` komutunu bir kez çalıştırıp `/login` yapın), sonra Bağlantıyı kes → Bağlan ile yeniden deneyin."
167
+ "pi": "Arka plandaki aracı yanıt vermiyor — pi CLI `pi --list-models` komutunu çalıştıramadı. https://pi.dev adresinden kurun (`curl -fsSL https://pi.dev/install.sh | sh`), bir sağlayıcı yapılandırın (bir sağlayıcı API anahtarı ayarlayın ya da `pi` komutunu bir kez çalıştırıp `/login` yapın), sonra Bağlantıyı kes → Bağlan ile yeniden deneyin.",
168
+ "qwen": "Arka plandaki aracı yanıt vermiyor — Qwen Code CLI başlatılamadı. Qwen Code'un kurulu (npm i -g @qwen-code/qwen-code) ve oturum açılmış olduğundan emin olun (`qwen` komutunu bir kez çalıştırıp /auth adımını tamamlayın ya da DASHSCOPE_API_KEY ayarlayın), sonra Bağlantıyı kes → Bağlan ile yeniden deneyin."
168
169
  },
169
170
  "llamacpp_no_jinja": "llama-server sunucunuz `--jinja` OLMADAN çalışıyor, bu yüzden araç çağırma devre dışı — aracının her eylemi başarısız olur. Araç desteğiyle yeniden başlatın: `llama-server -m <model>.gguf --jinja -c 16384` (güncel sürümlerde jinja varsayılan olarak açıktır, eski sürümlerde bu bayrak gerekir), sonra Bağlantıyı kes → Bağlan.",
170
171
  "llamacpp_small_context": "llama-server bağlamınız {tokens} belirteç (başlatma bayrağı -c). Aracının araç yükü için en az 16384 gerekir; bunun altında uzun turlar sessizce kırpılır. `-c 16384` veya daha yükseğiyle yeniden başlatmayı düşünün.",
@@ -164,7 +164,8 @@
164
164
  "minimax": "后台智能体没有响应 —— MiniMax 无法启动。请设置在 platform.minimax.io 获取的 MINIMAX_API_KEY ,然后依次“断开连接”→“连接”重试。",
165
165
  "moonshot": "后台智能体没有响应 —— Moonshot(Kimi K3)无法启动。请设置在 platform.kimi.ai 获取的 MOONSHOT_API_KEY ,然后依次“断开连接”→“连接”重试。",
166
166
  "ollama": "后台智能体没有响应 —— 连不上 Ollama 。请用 `ollama serve` 启动,并拉取我们的微调模型(`ollama pull artokun/gemma4-comfyui-mcp:e4b` —— 基于 comfyui-mcp 工具套件训练的 gemma4,竞技场里表现最好的本地模型;显存约 8 GB 可用 `:12b`),然后依次“断开连接”→“连接”重试。",
167
- "pi": "后台智能体没有响应 —— pi CLI 无法运行 `pi --list-models` 。请从 https://pi.dev 安装(`curl -fsSL https://pi.dev/install.sh | sh`),配置一个服务商(设置服务商 API 密钥,或运行一次 `pi` 再执行 `/login`),然后依次“断开连接”→“连接”重试。"
167
+ "pi": "后台智能体没有响应 —— pi CLI 无法运行 `pi --list-models` 。请从 https://pi.dev 安装(`curl -fsSL https://pi.dev/install.sh | sh`),配置一个服务商(设置服务商 API 密钥,或运行一次 `pi` 再执行 `/login`),然后依次“断开连接”→“连接”重试。",
168
+ "qwen": "后台智能体没有响应 —— Qwen Code CLI 无法启动。请确认已安装 Qwen Code(npm i -g @qwen-code/qwen-code)并已登录(运行一次 `qwen` 并完成 /auth,或设置 DASHSCOPE_API_KEY),然后依次“断开连接”→“连接”重试。"
168
169
  },
169
170
  "llamacpp_no_jinja": "你的 llama-server 没有带 `--jinja` 运行,因此工具调用被禁用 —— 智能体的每个动作都会失败。请带上工具支持重新启动:`llama-server -m <model>.gguf --jinja -c 16384` (较新的版本默认启用 jinja,旧版本需要显式加这个参数),然后依次“断开连接”→“连接”。",
170
171
  "llamacpp_small_context": "你的 llama-server 上下文为 {tokens} 个 token(启动参数 -c)。智能体的工具负载需要不少于 16384,低于这个值时较长的对话会被悄悄截断。建议用 `-c 16384` 或更高重新启动。",
@@ -164,7 +164,8 @@
164
164
  "minimax": "背景代理程式沒有回應 —— MiniMax 無法啟動。請設定在 platform.minimax.io 取得的 MINIMAX_API_KEY ,然後依序「中斷連線」→「連線」重試。",
165
165
  "moonshot": "背景代理程式沒有回應 —— Moonshot(Kimi K3)無法啟動。請設定在 platform.kimi.ai 取得的 MOONSHOT_API_KEY ,然後依序「中斷連線」→「連線」重試。",
166
166
  "ollama": "背景代理程式沒有回應 —— 連不上 Ollama 。請用 `ollama serve` 啟動,並拉取我們的微調模型(`ollama pull artokun/gemma4-comfyui-mcp:e4b` —— 以 comfyui-mcp 工具套件訓練的 gemma4,競技場中表現最好的本機模型;顯示記憶體約 8 GB 可用 `:12b`),然後依序「中斷連線」→「連線」重試。",
167
- "pi": "背景代理程式沒有回應 —— pi CLI 無法執行 `pi --list-models` 。請從 https://pi.dev 安裝(`curl -fsSL https://pi.dev/install.sh | sh`),設定一個供應商(設定供應商 API 金鑰,或執行一次 `pi` 再執行 `/login`),然後依序「中斷連線」→「連線」重試。"
167
+ "pi": "背景代理程式沒有回應 —— pi CLI 無法執行 `pi --list-models` 。請從 https://pi.dev 安裝(`curl -fsSL https://pi.dev/install.sh | sh`),設定一個供應商(設定供應商 API 金鑰,或執行一次 `pi` 再執行 `/login`),然後依序「中斷連線」→「連線」重試。",
168
+ "qwen": "背景代理程式沒有回應 —— Qwen Code CLI 無法啟動。請確認已安裝 Qwen Code(npm i -g @qwen-code/qwen-code)並已登入(執行一次 `qwen` 並完成 /auth,或設定 DASHSCOPE_API_KEY),然後依序「中斷連線」→「連線」重試。"
168
169
  },
169
170
  "llamacpp_no_jinja": "你的 llama-server 沒有帶 `--jinja` 執行,因此工具呼叫被停用 —— 代理程式的每個動作都會失敗。請帶上工具支援重新啟動:`llama-server -m <model>.gguf --jinja -c 16384` (較新的版本預設啟用 jinja,較舊的版本需要明確加上這個旗標),然後依序「中斷連線」→「連線」。",
170
171
  "llamacpp_small_context": "你的 llama-server 上下文為 {tokens} 個 token(啟動旗標 -c)。代理程式的工具內容需要至少 16384,低於這個值時較長的回合會被無聲截斷。建議用 `-c 16384` 或更高重新啟動。",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "comfyui-mcp",
3
- "version": "0.52.13",
3
+ "version": "0.52.15",
4
4
  "mcpName": "io.github.artokun/comfyui-mcp",
5
5
  "description": "Local-first, agent-native control plane for ComfyUI — MCP server + autonomous sidebar agent that drives your live graph in natural language on ANY LLM: Claude/ChatGPT/Gemini on your subscription (no API key), free local models via Ollama (fully offline), or any hosted model via an OpenAI-compatible endpoint (DeepSeek, GLM, MiMo, OpenRouter). Generate images, video & audio, author and run workflows, manage models and custom nodes; a compact tool-router mode keeps even 4B models effective, and a built-in LLM Arena benchmarks them on real ComfyUI tasks. Also a Claude Code plugin with skills, slash commands, and installer packs. Local, LAN, VPS, or Comfy Cloud.",
6
6
  "homepage": "https://comfyui-mcp.artokun.io/docs",
package/plugin/.mcp.json CHANGED
@@ -1,9 +1,8 @@
1
1
  {
2
2
  "comfyui": {
3
- "command": "npx",
3
+ "command": "node",
4
4
  "args": [
5
- "-y",
6
- "comfyui-mcp",
5
+ "${CLAUDE_PLUGIN_ROOT}/scripts/launch-server.mjs",
7
6
  "--full"
8
7
  ],
9
8
  "env": {
@@ -0,0 +1,130 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * #1447 — plugin MCP server launcher.
4
+ *
5
+ * The plugin's .mcp.json used to be `npx -y comfyui-mcp --full` directly. On a
6
+ * cold npx cache that downloads the whole ~900 MB dependency tree INSIDE the
7
+ * client's MCP handshake timeout; the client kills the attempt, nothing is
8
+ * persisted, and every retry pays the full cost again (the measured `_npx`
9
+ * cache timestamped to a manual run, never to a client attempt).
10
+ *
11
+ * This wrapper is the warm path: when `comfyui-mcp` is installed globally
12
+ * (`npm install -g comfyui-mcp`), it resolves `<npm root -g>/comfyui-mcp/
13
+ * dist/index.js` and runs it with this same node — sub-second start, no
14
+ * registry round-trip, and a killed handshake cannot discard an install that
15
+ * already happened. With no global install it falls back to the original
16
+ * `npx -y comfyui-mcp` behaviour, so the plugin works identically for anyone
17
+ * who never installs globally.
18
+ *
19
+ * STDIO IS THE MCP TRANSPORT. Everything the child says on stdout/stderr is
20
+ * the protocol — the wrapper inherits stdio verbatim and must never write to
21
+ * stdout itself. Diagnostics (there is exactly one, a spawn failure) go to
22
+ * stderr via console.error.
23
+ *
24
+ * Import-safe: nothing here runs on import. The resolution decisions are pure
25
+ * exports so the tests can call the real thing instead of reimplementing it
26
+ * (the #1385 lesson — a test of a copy is a test of nothing).
27
+ */
28
+
29
+ import { execFile, spawn } from "node:child_process";
30
+ import { existsSync } from "node:fs";
31
+ import { join } from "node:path";
32
+ import { pathToFileURL } from "node:url";
33
+
34
+ /** `npm root -g` is normally ~100 ms; 5 s bounds a wedged npm without hanging the handshake. */
35
+ const NPM_ROOT_TIMEOUT_MS = 5000;
36
+
37
+ /**
38
+ * `npm root -g` stdout → the global install's entry point, or null when the
39
+ * output is unusable or the install is absent/incomplete (no dist/index.js).
40
+ * null is not an error here — it means "take the npx path".
41
+ */
42
+ export function globalEntry(npmRootStdout, { exists = existsSync } = {}) {
43
+ const root = typeof npmRootStdout === "string" ? npmRootStdout.trim() : "";
44
+ if (!root) return null;
45
+ const entry = join(root, "comfyui-mcp", "dist", "index.js");
46
+ return exists(entry) ? entry : null;
47
+ }
48
+
49
+ /**
50
+ * Tokens the shell form may contain. The fallback's args come from the
51
+ * plugin's own .mcp.json (today: `--full`), never from user input — but a
52
+ * shell command is string concatenation, so anything outside this alphabet
53
+ * (whitespace, quotes, metacharacters) is REFUSED rather than passed through
54
+ * to be silently misparsed.
55
+ */
56
+ const SHELL_SAFE_TOKEN = /^[A-Za-z0-9@._~:/=-]+$/;
57
+
58
+ /**
59
+ * The spawn spec for the resolved server. `extraArgs` are the plugin's own
60
+ * args from .mcp.json (`--full`), forwarded to whichever server runs so the
61
+ * manifest stays the declarative source of how the server is started.
62
+ *
63
+ * With an entry point we spawn node directly — no shell, no shim resolution.
64
+ *
65
+ * The npx fallback differs by platform:
66
+ * - POSIX: spawn npx with an args array, no shell.
67
+ * - Windows: npx is npx.cmd, which Node refuses to spawn without a shell
68
+ * since the 18.20.2/20.12.2 bat-file fix — but `shell` + an args array
69
+ * triggers DEP0190 (args are concatenated UNESCAPED, and Node warns every
70
+ * launch). So the Windows spec is one pre-validated command string instead:
71
+ * same concatenation, but every token is checked against SHELL_SAFE_TOKEN
72
+ * first, so what the shell parses is exactly what was written.
73
+ */
74
+ export function serverSpec(entry, extraArgs, { platform = process.platform, node = process.execPath } = {}) {
75
+ if (entry) return { command: node, args: [entry, ...extraArgs], shell: false };
76
+ const npxArgs = ["-y", "comfyui-mcp", ...extraArgs];
77
+ if (platform !== "win32") return { command: "npx", args: npxArgs, shell: false };
78
+ for (const token of ["npx", ...npxArgs]) {
79
+ if (!SHELL_SAFE_TOKEN.test(token)) {
80
+ throw new Error(`[comfyui-mcp launcher] refusing to pass unsafe shell token to npx: ${JSON.stringify(token)}`);
81
+ }
82
+ }
83
+ return { command: ["npx", ...npxArgs].join(" "), args: [], shell: true };
84
+ }
85
+
86
+ /** Ask npm where global packages live. String-command form on Windows for the
87
+ * same DEP0190 reason as serverSpec; the command is a fixed literal, so there
88
+ * is nothing to inject. Resolves null on any failure — the npx fallback is
89
+ * the original behaviour, so a failed probe degrades to exactly what the
90
+ * plugin did before this wrapper existed. */
91
+ function probeGlobalRoot() {
92
+ const isWin = process.platform === "win32";
93
+ return new Promise((resolve) => {
94
+ execFile(
95
+ isWin ? "npm root -g" : "npm",
96
+ isWin ? [] : ["root", "-g"],
97
+ { shell: isWin, timeout: NPM_ROOT_TIMEOUT_MS },
98
+ (error, stdout) => resolve(error ? null : stdout),
99
+ );
100
+ });
101
+ }
102
+
103
+ function run(spec) {
104
+ const child = spawn(spec.command, spec.args, { stdio: "inherit", shell: spec.shell });
105
+
106
+ // The client kills the WRAPPER when it shuts the server down; without
107
+ // forwarding, the real server would be orphaned holding stdio. Killing the
108
+ // child on signal makes its `exit` fire, which is what exits the wrapper.
109
+ for (const sig of ["SIGINT", "SIGTERM", "SIGHUP"]) {
110
+ process.on(sig, () => child.kill(sig));
111
+ }
112
+
113
+ child.on("exit", (code, signal) => {
114
+ // A signal death has no code; 1 keeps the client from reading it as clean.
115
+ process.exit(code ?? (signal ? 1 : 0));
116
+ });
117
+ child.on("error", (err) => {
118
+ // Honest failure, loudly: a server that never started must not look like
119
+ // one that exited cleanly. stderr, never stdout — see the header.
120
+ console.error(`[comfyui-mcp launcher] failed to start ${spec.command}: ${err.message}`);
121
+ process.exit(1);
122
+ });
123
+ }
124
+
125
+ const isMain = process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href;
126
+ if (isMain) {
127
+ const extraArgs = process.argv.slice(2);
128
+ const entry = globalEntry(await probeGlobalRoot());
129
+ run(serverSpec(entry, extraArgs));
130
+ }
@@ -0,0 +1,270 @@
1
+ ---
2
+ name: minimax-h3-video
3
+ description: Build MiniMax H3 (Hailuo) local video workflows — native T2V/I2V/R2V nodes, Comfy-Org INT8 weights, turbo LoRAs for 8GB VRAM, 15-second stereo-audio clips, and the official MiniMax prompting guides (cite by link, do not copy).
4
+ globs:
5
+ - "**/*.json"
6
+ ---
7
+
8
+ # MiniMax H3 (Hailuo) — local video
9
+
10
+ This skill teaches the **local-weights** MiniMax H3 path in ComfyUI. It is the
11
+ pilot for `#1155` (Official vs Empirical sources) because MiniMax publishes a
12
+ real prompting guide. **Cite that guide by URL. Do not copy it into this repo.**
13
+
14
+ ## Two products, two cost models — pick one
15
+
16
+ They share a brand and **must not be mixed**.
17
+
18
+ | Path | Nodes | Cost | VRAM | When |
19
+ |---|---|---|---|---|
20
+ | **Local weights** (this skill) | `MiniMaxH3ImageToVideo`, `MiniMaxH3ReferenceToVideo`, `EmptyMiniMaxH3LatentAV`, `MiniMaxH3SigmaShift`, `MiniMaxH3MemoryEfficientSageAttentionPatch` | Free after download | Yes — INT8 + turbo LoRA is the 8 GB story | User wants 4–15 s stereo clips on their GPU |
21
+ | **Partner API** | `MinimaxHailuo03TextToVideoNode`, `MinimaxHailuo03FirstLastFrameNode`, `MinimaxHailuo03ReferenceNode`, `MinimaxTextToVideoNode`, `MinimaxImageToVideoNode`, `MinimaxHailuoVideoNode` | Paid per generation | None | User has a MiniMax / Hailuo API key and does not want local weights |
22
+
23
+ API nodes do not take `MiniMaxH3SigmaShift` or Sage-attention patches. Local
24
+ nodes do not spend API credits. If the user asked for Hailuo *cloud*, stop and
25
+ use the API nodes + their key; do not download 40 GB of weights.
26
+
27
+ `MiniMaxH3Director` is a **third-party** pack (`muse-collective-26/MiniMaxH3-Director`),
28
+ not core. Do not require it for T2V / I2V / R2V.
29
+
30
+ ## License — cite, do not copy
31
+
32
+ Local weights and MiniMax's own documentation sit under the
33
+ [MiniMax H3 Community License](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE).
34
+ `Materials` includes the Documentation. The agreement's Applicable Territory
35
+ **excludes the United States, the EU, the UK, and South Korea**. This skill
36
+ does **not** reproduce MiniMax's `skills/h3-prompt-writing/` SKILL.md or the
37
+ prompting-guide prose. Linking to a public URL is the `#1155` requirement.
38
+
39
+ This is not legal advice. Tell a US/EU/UK/KR user that the *local* path is
40
+ territory-restricted and that the **paid API** is a separate product under
41
+ MiniMax platform terms.
42
+
43
+ ## Prefer the Comfy-Org template over hand-wiring
44
+
45
+ ComfyUI ≥ **0.30.0** (templates in the 0.33 line). These are **core**
46
+ `comfyui-workflow-templates` graphs in the frontend **Template Library →
47
+ Video**, not installer packs and not custom-node `example_workflows`:
48
+
49
+ | Mode | Template Library card | File | Diffusion file |
50
+ |---|---|---|---|
51
+ | T2V / I2V / FL2VA | MiniMax H3: Text to Video / Image to Video | `video_minimax_h3_t2v.json` / `video_minimax_h3_i2v.json` | `minimax_h3_fl2va_pruned_int8_convrot.safetensors` |
52
+ | R2V (omni-reference) | MiniMax H3: Reference to Video | `video_minimax_h3_r2v.json` | `minimax_h3_ref2va_pruned_int8_convrot.safetensors` |
53
+
54
+ `list_packs action:"list_templates"` will **not** list them.
55
+ `enqueue_workflow action:"run_template"` will **not** resolve
56
+ `video_minimax_h3_t2v` / `_i2v` / `_r2v` — that action only loads bundled
57
+ installer packs, and there is no `packs/minimax-h3-*` yet. Do not call it
58
+ until a pack exists. `panel_load_workflow` needs `pack:`, a disk `path:`, or
59
+ an inline UI `graph` — a Template Library basename is none of those.
60
+
61
+ **Load path that actually works:**
62
+
63
+ 1. **Preferred.** Ask the user to open **Template Library → Video → MiniMax H3:
64
+ Text to Video** (or Image to Video / Reference to Video). Pick the local
65
+ `video_minimax_h3_*` cards, **not** the `api_minimax_h3_*` paid partner
66
+ templates.
67
+ 2. **Agent, no UI click.** Fetch the UI JSON from
68
+ https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_t2v.json
69
+ (or `_i2v` / `_r2v`; raw.githubusercontent.com is the same files), save it
70
+ with `save_workflow action:"save"` `filename:"video_minimax_h3_t2v.json"`,
71
+ then `panel_load_workflow path:"video_minimax_h3_t2v.json"`. Same pattern as
72
+ `video-extend` (stage on disk, then `path:`). Do not pass the GitHub URL as
73
+ `path:` or `pack:`.
74
+
75
+ After it lands, retarget the **subgraph's exposed widgets** (prompt, duration,
76
+ `turbo_mode`, megapixels). Official T2V/I2V graphs wrap
77
+ `MiniMaxH3ImageToVideo` inside a subgraph (`type` is a UUID). Do not flatten
78
+ that interior unless you are hand-building.
79
+
80
+ Hand-building the subgraph is slower and easy to get wrong.
81
+
82
+ Comfy tutorial (wiring, not MiniMax's prompt formula):
83
+ https://docs.comfy.org/tutorials/video/minimax/minimax-h3
84
+
85
+ ## Models (Comfy-Org INT8 pack)
86
+
87
+ All from `huggingface.co/Comfy-Org/MiniMax-H3`. Download with
88
+ `download_model` `action:"download"`.
89
+
90
+ | File | Folder | Role |
91
+ |---|---|---|
92
+ | `minimax_h3_fl2va_pruned_int8_convrot.safetensors` | `diffusion_models/` | T2V / I2V / first-last-frame |
93
+ | `minimax_h3_ref2va_pruned_int8_convrot.safetensors` | `diffusion_models/` | R2V only — **different UNet** |
94
+ | `qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | `text_encoders/` | Qwen3-VL-32B encoder, `CLIPLoader` **type=`minimax`** |
95
+ | `minimax_h3_video_vae_fp16.safetensors` | `vae/` | Visual VAE |
96
+ | `minimax_h3_audio_vae_fp32.safetensors` | `vae/` | Stereo audio VAE (32 kHz) |
97
+
98
+ ### Turbo LoRAs (4–8 steps instead of ~20)
99
+
100
+ The Comfy-Org T2V template already switches these on with `turbo_mode`:
101
+
102
+ | Steps | File | Source |
103
+ |---|---|---|
104
+ | 8 | `minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors` | `lightx2v/Minimax-h3-Turbo` |
105
+ | 4 | `minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors` | `Comfy-Org/MiniMax-H3` `loras/` |
106
+
107
+ Kijai conversions live at `Kijai/MiniMax-H3_comfy` (`loras/`) and experimental
108
+ W4A8 at `Kijai/MiniMax-H3-experimental`. Same job (low-step / low-VRAM). Prefer
109
+ the Comfy-Org / lightx2v filenames the template already names; only switch to a
110
+ Kijai file if that is what is on disk.
111
+
112
+ 4-step is faster and softer; **6–8 steps** is the usual sharpness compromise.
113
+
114
+ ## Output spec
115
+
116
+ | Knob | Value |
117
+ |---|---|
118
+ | Duration | 4–15 seconds |
119
+ | Frame rate | **24 fps** (`CreateVideo.fps`) |
120
+ | Audio | Native stereo, decoded by the audio VAE, muxed in `CreateVideo` |
121
+ | Short edge | 768 px native; cap **768×1344**, multiple of **32** |
122
+ | Preview size | `ResolutionSelector` megapixels **0.4** → 864×480 at 16:9 |
123
+ | Full 768p | megapixels **~0.98** → **1344×768** at 16:9 |
124
+
125
+ Duration → frame `length` (Comfy-Org template math, 17-frame blocks):
126
+
127
+ ```
128
+ max(5, round(seconds * 24)) + (5 - (max(5, round(seconds * 24)) % 17)) % 17
129
+ ```
130
+
131
+ That is the `17k+5` grid. Do not invent a WAN-style `4n+1` length.
132
+
133
+ ## Node graph (local T2V / I2V)
134
+
135
+ From the Comfy-Org T2V subgraph (core nodes, not the Markdown notes):
136
+
137
+ ```
138
+ ResolutionSelector (aspect, megapixels, multiple=32) → width, height
139
+
140
+ UNETLoader (fl2va int8)
141
+ ├─ LoraLoaderModelOnly (turbo LoRA) ─┐
142
+ └────────────────────────────────────┤ ComfySwitchNode (turbo_mode)
143
+ ▼
144
+ BasicGuider + BasicScheduler + KSamplerSelect(res_multistep)
145
+ ▼
146
+ CLIPLoader (type=minimax, qwen3vl 32b) → MiniMaxH3ImageToVideo
147
+ VAELoader (video vae) ─────────────────→ prompt, width, height, length
148
+ optional first_frame / last_frame ─────→ → CONDITIONING + LATENT
149
+ ▼
150
+ SamplerCustomAdvanced → LATENT
151
+ ├─ VAEDecode (video vae) → IMAGE
152
+ └─ VAEDecodeAudio (audio vae) → AUDIO
153
+ ▼
154
+ CreateVideo (fps=24) → SaveVideo
155
+ ```
156
+
157
+ `MiniMaxH3ImageToVideo` **is** T2V when both image sockets are empty, I2V with
158
+ `first_frame`, FL2VA with both frames. Do not add a second T2V-only node.
159
+
160
+ R2V replaces the UNet with **ref2va** and the conditioner with
161
+ `MiniMaxH3ReferenceToVideo`. Do not load fl2va into an R2V graph.
162
+
163
+ ### Local-only helpers
164
+
165
+ | Node | Role |
166
+ |---|---|
167
+ | `EmptyMiniMaxH3LatentAV` | Empty audio-video latent when you are not using `MiniMaxH3ImageToVideo`'s built-in latent |
168
+ | `MiniMaxH3SigmaShift` | Flow-matching shift on the local UNet |
169
+ | `MiniMaxH3MemoryEfficientSageAttentionPatch` | Core Sage patch; or KJNodes `Patch Sage Attention KJ` (`sage_attention=auto`) between `UNETLoader` and `BasicGuider` |
170
+
171
+ Sage roughly doubles speed. H3 has mixed dtypes — console lines about falling
172
+ back to pytorch attention on some layers are expected.
173
+
174
+ ## Sampler defaults (Comfy-Org template)
175
+
176
+ | Mode | Sampler | Scheduler | Steps |
177
+ |---|---|---|---|
178
+ | Base (no turbo) | `res_multistep` | `simple` | **20** |
179
+ | Turbo on | `res_multistep` | `simple` | **4–8** (template default turbo steps widget) |
180
+
181
+ Guider is `BasicGuider` (CFG-distilled checkpoint — do not crank CFG). Seed via
182
+ `RandomNoise`.
183
+
184
+ ## Prompting — read the vendor guide, do not paste it here
185
+
186
+ Write the prompt **in the MiniMax H3 node**, not a generic `CLIPTextEncode`.
187
+
188
+ **Official MiniMax guides** (read these; do not copy them into graphs as a
189
+ system prompt dump):
190
+
191
+ - T2VA / I2VA / FL2VA / L2VA:
192
+ https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
193
+ - Full-reference / R2V:
194
+ https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
195
+ - Vendor skill (install separately if the user wants it — we do not bundle it):
196
+ https://github.com/MiniMax-AI/MiniMax-H3 (`npx skills add … --skill h3-prompt-writing`)
197
+
198
+ H3-Context-IR (the hosted prompt rewriter) is **not** in the open weights.
199
+ Local ComfyUI has no IR node. Either write the structured prompt yourself from
200
+ the guide, or call MiniMax's Context-IR API and paste `content.prompt` into the
201
+ local node.
202
+
203
+ Comfy-Org's own template notes (safe to follow, not MiniMax docs):
204
+
205
+ 1. One block covering **look, scene, timed shots, camera, and audio** (dialogue,
206
+ SFX, score).
207
+ 2. Time shots (`[0s-1.5s] Shot 1: …`).
208
+ 3. R2V: name each input in connection order (`<Picture 1>`, `<Video 1>`,
209
+ `<Audio 1>`) and say what job each one does (identity, motion, voice).
210
+ 4. R2V caps (vendor model card, not a guess): ≤9 images, ≤3 videos, ≤3 audio
211
+ clips, ≤12 files mixed; each AV clip 2–15 s.
212
+
213
+ ## 15-second clips and chaining
214
+
215
+ One H3 shot is **at most ~15 s**. Longer pieces are concatenated clips, not a
216
+ bigger `length`.
217
+
218
+ 1. Generate clip N (up to 15 s).
219
+ 2. Confirm the file with `get_image` `action:"list_outputs"` (`kind:"video"`) —
220
+ video nodes often skip `/history`.
221
+ 3. Stage the last frame (or the whole clip) with `upload_image` `action:"stage"`.
222
+ 4. Clip N+1: `MiniMaxH3ImageToVideo.first_frame` = last frame of N, **or** R2V
223
+ with `<Video 1>` as a continuation reference.
224
+ 5. Concat with ffmpeg (`director` skill) or an editor.
225
+
226
+ This is **not** WAN Pusa (`video-extend`). Pusa LoRAs and `flowmatch_pusa` do
227
+ not apply to H3.
228
+
229
+ ## VRAM
230
+
231
+ | Card | Practical setup |
232
+ |---|---|
233
+ | **24 GB+** | INT8 fl2va + Qwen3-VL + both VAEs; 1344×768; 10–15 s; Sage optional |
234
+ | **12–16 GB** | Same INT8 pack; drop megapixels toward 0.4–0.6; turbo LoRA on; Sage |
235
+ | **8 GB** | INT8 + turbo LoRA + Sage + short preview (0.2–0.4 MP, 4–6 s). Minutes per clip. Kijai W4A8 if INT8 still OOMs. |
236
+
237
+ Always `clear_vram` before switching to H3 from WAN / LTX / a checkpoint.
238
+
239
+ ## Gotchas
240
+
241
+ - **`CLIPLoader` type must be `minimax`.** `qwen_image` / `flux` will load the
242
+ wrong encoder layout.
243
+ - **fl2va vs ref2va.** T2V/I2V templates on ref2va (or R2V on fl2va) are garbage
244
+ or a load error.
245
+ - **Turbo off, 4 steps.** The switch defaults off and base steps are 20. Four
246
+ steps without the LoRA is mush.
247
+ - **API node in a local graph.** Costs money and ignores the UNet you downloaded.
248
+ - **WAN frame math.** H3 is 24 fps and `17k+5`, not 16 fps `4n+1`.
249
+ - **Verify video on disk**, then stage — never guess `input/` paths.
250
+ - **ffmpeg** is required for `CreateVideo` / `SaveVideo` / `VHS_VideoCombine`.
251
+ - Desktop/Cloud ComfyUI lags nightly. Missing `MiniMaxH3*` nodes → update to
252
+ ≥0.30.0 (0.33 templates) before hunting custom packs.
253
+
254
+ ## See also
255
+
256
+ - `video-extend` — WAN Pusa temporal continuation (different family)
257
+ - `director` — multi-clip concat after you have 15 s H3 shots
258
+ - `prompt-engineering` — generic CLIP syntax; **H3 does not use it**
259
+ - `triton-sageattention` — installing Sage on Windows
260
+
261
+ No bundled `packs/minimax-h3-*` installer yet — that is why
262
+ `enqueue_workflow action:"run_template"` cannot load these graphs. Use the
263
+ Template Library (or the GitHub fetch → `save_workflow` →
264
+ `panel_load_workflow path:` path above) + `download_model` against
265
+ `Comfy-Org/MiniMax-H3`.
266
+
267
+ ## Sources
268
+
269
+ - **Official:** MiniMax prompting guides https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md (T2VA/I2VA/FL2VA/L2VA) and https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md (full-reference / R2V); vendor repo https://github.com/MiniMax-AI/MiniMax-H3; ComfyUI tutorial + templates https://docs.comfy.org/tutorials/video/minimax/minimax-h3 (wiring, duration grid, INT8 filenames). MiniMax's own `skills/h3-prompt-writing` is linked, not copied — Community License includes Documentation.
270
+ - **Empirical:** local vs partner-API node split and 8 GB turbo-LoRA note from issue #1167 / the reporter's rig; Sage mixed-dtype fallback from the Comfy tutorial; chaining last-frame→next-clip from observed ComfyUI I/O (stage + list_outputs), not a vendor extender.
@@ -70,6 +70,22 @@ Comfy-Org repackages everything: `huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repac
70
70
  | Qwen-Image-Edit 2511 Lightning LoRA | `huggingface.co/lightx2v/Qwen-Image-Edit-2511-Lightning` | `loras/` |
71
71
  | Qwen-Image 2512 Turbo LoRA | `huggingface.co/Wuli-art/Qwen-Image-2512-Turbo-LoRA` | `loras/` |
72
72
 
73
+ ## MiniMax H3 video (see `minimax-h3-video` skill)
74
+
75
+ Comfy-Org INT8 pack: `huggingface.co/Comfy-Org/MiniMax-H3`. Local weights are
76
+ under the MiniMax H3 Community License (territory-restricted). Prefer the
77
+ filenames the Comfy-Org templates already name.
78
+
79
+ | File | Source | Target |
80
+ |---|---|---|
81
+ | `minimax_h3_fl2va_pruned_int8_convrot.safetensors` | `huggingface.co/Comfy-Org/MiniMax-H3/resolve/main/diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors` | `diffusion_models/` |
82
+ | `minimax_h3_ref2va_pruned_int8_convrot.safetensors` | same repo, `diffusion_models/` | `diffusion_models/` |
83
+ | `qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | same repo, `text_encoders/` | `text_encoders/` |
84
+ | `minimax_h3_video_vae_fp16.safetensors` | same repo, `vae/` | `vae/` |
85
+ | `minimax_h3_audio_vae_fp32.safetensors` | same repo, `vae/` | `vae/` |
86
+ | `minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors` | `huggingface.co/lightx2v/Minimax-h3-Turbo/resolve/main/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors` | `loras/` |
87
+ | `minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors` | `huggingface.co/Comfy-Org/MiniMax-H3/resolve/main/loras/minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors` | `loras/` |
88
+
73
89
  ## Z-Image (see `z-image-txt2img` skill)
74
90
 
75
91
  | File | Source | Target |
@@ -71,6 +71,28 @@ All of the following are node properties — set them with `panel_set_property`:
71
71
  | `showNav` | bool | `true` | Per-row jump-to-group arrow. |
72
72
  | `showAllGraphs` | bool | `true` | Include groups that live inside subgraphs. |
73
73
 
74
+ ### matchTitle does not rebuild the toggle list on the first write
75
+
76
+ `panel_set_property` stores `matchTitle` (and `matchColors` / `sort` / …) and
77
+ the reply `from`/`to` is truthful. Fast Groups nodes **do not implement
78
+ `onPropertyChanged`**. The toggle list is rebuilt by rgthree's
79
+ `refreshWidgets()` on a service tick (~8 ms after add, then every ~500 ms),
80
+ and leftover-row removal increments the index while splicing — a 22-group
81
+ list can stall at 13 with non-matching `Enable …` rows still present. A
82
+ never-drawn node can also come back as `widgets:{}` — that means the list
83
+ has **not been built yet**, not that there are no matching groups.
84
+ `panel_query_graph` also keys widgets by name, and every toggle is named
85
+ `RGTHREE_TOGGLE_AND_NAV`, so a built list collapses to one key.
86
+
87
+ **Do this:**
88
+
89
+ 1. Set `matchTitle` **immediately** after `panel_add_node` (before the first
90
+ unfiltered refresh paints every group).
91
+ 2. Re-read with `panel_query_graph {ids:[<id>], fields:'detail'}`.
92
+ 3. If `widgets` is empty **or** the canvas still shows `Enable` rows that do
93
+ not match the regex, **set `matchTitle` again**. Do **not** delete and
94
+ re-add the node — that is slower and still needs a second set.
95
+
74
96
  ### Recipe — make pipeline stages toggleable
75
97
 
76
98
  1. `panel_create_group` per stage, with a **prefixed title** (`STAGE 1 — …`) so one
@@ -94,6 +116,9 @@ All of the following are node properties — set them with `panel_set_property`:
94
116
  unwired. Add nodes one at a time, not as a parallel batch.
95
117
 
96
118
  4. `panel_set_property` → `matchTitle` = `^STAGE`, and `sort` = `alphanumeric`.
119
+ Set them immediately after the add, then re-read the node. If `widgets` is
120
+ empty or leftover `Enable` rows remain, set `matchTitle` again — do not
121
+ delete and re-add.
97
122
 
98
123
  5. Toggle, then verify with `panel_graph_outline` — it tags nodes `[bypass]` / `[mute]`.
99
124
 
@@ -182,4 +207,4 @@ Turning a LoRA **off** (`lora_N.on = false`) is usually safer than clearing it a
182
207
  ## Sources
183
208
 
184
209
  - **Official:** https://github.com/rgthree/rgthree-comfy
185
- - **Empirical:** frontend-only allowlist and properties-not-widgets notes verified against the installed pack and the panel guard.
210
+ - **Empirical:** frontend-only allowlist and properties-not-widgets notes verified against the installed pack and the panel guard. `#1808` matchTitle rebuild: Fast Groups have no `onPropertyChanged`; `refreshWidgets()` leftover removal is `removeWidget(index++)` in `fast_groups_muter.ts`; first unfiltered tick is scheduled from `addFastGroupNode` in `fast_groups_service.ts`.