makaron-cli 0.15.4 → 0.16.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/README.md +2 -2
- package/bin/makaron.mjs +13 -5
- package/package.json +1 -1
- package/skills/makaron/SKILL.md +4 -2
package/README.md
CHANGED
|
@@ -384,7 +384,7 @@ npx makaron-cli edit --image photo.jpg --out result.jpg "make it dramatic"
|
|
|
384
384
|
npx makaron-cli edit --image-model gpt-image-2.5-flare --background transparent --out sticker.png "a magenta star sticker"
|
|
385
385
|
```
|
|
386
386
|
|
|
387
|
-
Options: `--image`, `--image-model gemini|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image`, `--ref <file>` (
|
|
387
|
+
Options: `--image`, `--image-model gemini|gemini-2.1|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image`, `--ref <file>` (model-specific limit), `--aspect <ratio>`, `--background auto|opaque|transparent`, `--out <path>`. Nano Banana 2.1 (`gemini-2.1`) supports up to 14 total input images and `--image-resolution 1K|2K|4K` (default 1K); it uses OpenRouter and does not retry or switch models automatically. Qwen Spicy supports 0–3 input images; legacy `qwen` requests map to it. Pony and WAI are retired. Transparent output routes strictly to GPT Image 2.5 Flare and is returned only when the provider supplies real PNG/WebP alpha.
|
|
388
388
|
|
|
389
389
|
`wan2.7-image` uses Alibaba international for fast, approximately 1K generation and editing (default 6 credits/image). Failed or timed-out Wan requests are not automatically retried or switched to another model. Face identity can change. Example: `makaron edit --image portrait.jpg --image-model wan2.7-image --aspect 16:9 --out stadium.jpg "Place this woman in a baseball stadium, preserving her face."`
|
|
390
390
|
|
|
@@ -429,7 +429,7 @@ Options for `video create`: `--script "..."`, `--script-file <path>`, `--image <
|
|
|
429
429
|
|
|
430
430
|
fal H3 Turbo: use `--video-model minimax-h3-max` for faster-than-real-time generation. It supports exactly 5/10/15 seconds at 480p/768p and defaults to native 768p, with no image for T2V or exactly one `--image` for I2V. It does not accept reference video/audio or multiple images.
|
|
431
431
|
|
|
432
|
-
Seedance 2.5: use `--video-model seedance-2.5` for 4-30 second output at 480p
|
|
432
|
+
Seedance 2.5: use `--video-model seedance-2.5` or `seedance-2.5-eco` for 4-30 second videos generated at 480p then automatically upscaled to 720p, 1080p (default), 2K or 4K. Use `seedance-2.5-native` explicitly for native output at 480p (default) or 720p. `--image` accepts local files or URLs (up to 30), while repeatable `--video` and `--audio` accept up to 10 each. Use `--video-operation generate|edit|extend`, `--extend-direction forward|backward`, `--output-format mp4|mov`, `--web-search`, `--generated-audio` / `--no-generated-audio`, and `--relaxed-content-filter`. Edit/extend require a video reference. Eco delivers MP4; MOV is native-only. The native Evolink route does not expose 4K output.
|
|
433
433
|
|
|
434
434
|
Wan 3.0: choose `--video-model wan-3.0` or the faster `--video-model wan-3.0-prime`. Both support 2-30 second generation, 480p/720p/1080p/2K/4K, and up to 10 images, 5 videos, and 5 audio references. Pass `--video-resolution 2k|4k` to use the matching FlashVSR/Pro endpoint automatically; Pro is not a separate model selector. Use generation mode with feature references; typed edit/extend and `--relaxed-content-filter` are not supported.
|
|
435
435
|
|
package/bin/makaron.mjs
CHANGED
|
@@ -1935,9 +1935,11 @@ Usage:
|
|
|
1935
1935
|
|
|
1936
1936
|
Options:
|
|
1937
1937
|
--image <file|url> Base image to edit. Omit for text-to-image.
|
|
1938
|
-
--ref <file|url> Additional reference image. Repeatable
|
|
1939
|
-
--image-model <id> gemini, gemini-lite, qwen-spicy, openai, gpt-image-2.5-flare, gpt-image-2.5-sunburst, or wan2.7-image. Legacy qwen maps to qwen-spicy.
|
|
1938
|
+
--ref <file|url> Additional reference image. Repeatable; model-specific limit (Nano Banana 2.1: 14 total inputs).
|
|
1939
|
+
--image-model <id> gemini, gemini-2.1, gemini-lite, qwen-spicy, openai, gpt-image-2.5-flare, gpt-image-2.5-sunburst, or wan2.7-image. Legacy qwen maps to qwen-spicy.
|
|
1940
1940
|
--skill <id> enhance, creative, wild, or captions.
|
|
1941
|
+
--image-resolution <id> 1K, 2K, or 4K; Auto chooses Nano Banana 2.1.
|
|
1942
|
+
--nsfw Route NSFW directly to Qwen Spicy.
|
|
1941
1943
|
--aspect <ratio> Output aspect ratio, for example 1:1, 16:9, or 9:16.
|
|
1942
1944
|
--background <mode> auto, opaque, or transparent.
|
|
1943
1945
|
--out <file> Save the generated image to this path.
|
|
@@ -1992,7 +1994,8 @@ Recent model choices:
|
|
|
1992
1994
|
5/10/15s; native 768p default or 480p; no video/audio/multi-image references.
|
|
1993
1995
|
wan-3.0-prime Faster Wan 3.0 tier; 2-30s; 480p through 4k; multimodal refs.
|
|
1994
1996
|
wan-3.0 Wan standard tier with the same public duration/resolution range.
|
|
1995
|
-
seedance-2.5 4-30s;
|
|
1997
|
+
seedance-2.5 Defaults to Eco; 4-30s; 1080p default, 720p/2k/4k; multimodal refs.
|
|
1998
|
+
seedance-2.5-eco 480p generation → ByteDance Fast; 1080p default, 2k/4k.
|
|
1996
1999
|
minimax-h3 4-15s; 768p default or 2k; image/video/audio feature refs.
|
|
1997
2000
|
grok T2V/reference generation plus typed edit/extend.
|
|
1998
2001
|
sync-lipsync-v3 Exactly one video plus one MP3/WAV replacement track.
|
|
@@ -3030,7 +3033,12 @@ if (!command || command === '--help' || command === '-h' || command === 'help')
|
|
|
3030
3033
|
editArgs.referenceImages = editArgs.referenceImages || [];
|
|
3031
3034
|
editArgs.referenceImages.push(imageToArg(args[++i]));
|
|
3032
3035
|
}
|
|
3036
|
+
else if (args[i] === '--image-resolution' && args[i + 1]) {
|
|
3037
|
+
const resolution = args[++i];
|
|
3038
|
+
editArgs.imageResolution = resolution;
|
|
3039
|
+
}
|
|
3033
3040
|
else if (args[i] === '--aspect' && args[i + 1]) editArgs.aspectRatio = args[++i];
|
|
3041
|
+
else if (args[i] === '--nsfw') editArgs.isNsfw = true;
|
|
3034
3042
|
else if (args[i] === '--background' && args[i + 1]) {
|
|
3035
3043
|
const background = args[++i];
|
|
3036
3044
|
if (!['auto', 'opaque', 'transparent'].includes(background)) {
|
|
@@ -3043,7 +3051,7 @@ if (!command || command === '--help' || command === '-h' || command === 'help')
|
|
|
3043
3051
|
else promptParts.push(args[i]);
|
|
3044
3052
|
}
|
|
3045
3053
|
editArgs.editPrompt = promptParts.join(' ');
|
|
3046
|
-
if (!editArgs.editPrompt) { console.error('Usage: makaron edit [--image <file|url>] [--image-model gemini|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image] [--ref <file>] [--aspect <ratio>] [--background auto|opaque|transparent] [--out <file>] "prompt"'); process.exit(1); }
|
|
3054
|
+
if (!editArgs.editPrompt) { console.error('Usage: makaron edit [--image <file|url>] [--image-model gemini|gemini-2.1|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image] [--ref <file>] [--aspect <ratio>] [--background auto|opaque|transparent] [--out <file>] "prompt"'); process.exit(1); }
|
|
3047
3055
|
process.stderr.write('🎨 Generating...\n');
|
|
3048
3056
|
const result = await callMcpTool(baseUrl, headers, 'makaron_edit_image', editArgs);
|
|
3049
3057
|
saveMcpImage(result, outputPath);
|
|
@@ -3123,7 +3131,7 @@ if (!command || command === '--help' || command === '-h' || command === 'help')
|
|
|
3123
3131
|
: ['wan3-prime', 'wan3.0-prime', 'wan30-prime', 'wan-3-prime', 'w3.0-video-prime', 'w3.0-video-prime-pro', 'wan-3.0-prime-pro', 'prime'].includes(videoModel)
|
|
3124
3132
|
? 'wan-3.0-prime'
|
|
3125
3133
|
: (videoModel || 'fal-h3-max');
|
|
3126
|
-
const isSeedance25 =
|
|
3134
|
+
const isSeedance25 = ['seedance-2.5', 'seedance-2.5-eco'].includes(selectedVideoModel);
|
|
3127
3135
|
const isWan30 = selectedVideoModel === 'wan-3.0' || selectedVideoModel === 'wan-3.0-prime';
|
|
3128
3136
|
const isSeedanceModel = selectedVideoModel === 'seedance-fast' || selectedVideoModel === 'seedance-mini' || selectedVideoModel === 'seedance' || isSeedance25;
|
|
3129
3137
|
const isMinimaxH3 = selectedVideoModel === 'minimax-h3';
|
package/package.json
CHANGED
package/skills/makaron/SKILL.md
CHANGED
|
@@ -315,7 +315,7 @@ npx makaron-cli edit --image photo.jpg --out result.jpg "make it dramatic"
|
|
|
315
315
|
npx makaron-cli edit --image-model gpt-image-2.5-flare --background transparent --out sticker.png "a magenta star sticker"
|
|
316
316
|
```
|
|
317
317
|
|
|
318
|
-
Options: `--image`, `--image-model gemini|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image`, `--ref <file>` (
|
|
318
|
+
Options: `--image`, `--image-model gemini|gemini-2.1|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image`, `--ref <file>` (model-specific limit), `--aspect <ratio>`, `--image-resolution 1K|2K|4K`, `--background auto|opaque|transparent`, `--nsfw`, `--out <path>`. Image Model Capability prefers matching requirements, then relaxes output conflicts to deliver an image: GPT Image 2.5 Flare → Nano Banana 2.1 → Qwen Spicy. Omit the model for Auto: an 8:1 banner or an explicit resolution tier selects Nano Banana 2.1. Native ratio presets do not promise exact mathematical ratios or user pixel dimensions. For Agent-assessed NSFW, pass `--nsfw` to route directly to Spicy; never probe another provider first. Unknown IDs use Flare-first Auto. Output conflicts relax before submission: transparent 8:1 normally keeps 8:1 with opaque output; NSFW always uses Spicy regardless of unsupported dimensions. Preserve input content first and explain actual adjustments. See current MCP tool descriptions for the generated capability table. Classic Nano Banana 2 (`gemini`), Lite, Sunburst and Wan are explicit routes; `openai` means Flare and legacy `qwen` means Spicy. Pony/WAI are retired. Transparent output prefers Flare or explicitly selected Sunburst, and may relax alpha when other requirements conflict. Supplier failures never trigger automatic image resubmission or model switching.
|
|
319
319
|
|
|
320
320
|
### `video` — Standalone video tools (no project timeline)
|
|
321
321
|
|
|
@@ -355,7 +355,7 @@ Provider integration contract: every image passed to video generation is a featu
|
|
|
355
355
|
|
|
356
356
|
fal H3 Turbo uses the `minimax-h3-max` selector, supports exactly 5/10/15 seconds at 480p/768p, and defaults to native 768p for faster-than-real-time T2V or one-start-image I2V.
|
|
357
357
|
|
|
358
|
-
Seedance 2.5 uses `--video-model seedance-2.5`
|
|
358
|
+
Seedance 2.5 uses `--video-model seedance-2.5` or `seedance-2.5-eco` for 4-30s videos: generate at 480p, then automatically upscale to 720p, 1080p (default), 2K or 4K in one task. This is upscaled output. Use `seedance-2.5-native` explicitly for unenhanced native output at 480p (default) or 720p; the native Evolink route does not expose 4K. All Seedance 2.5 routes accept up to 30 images + 10 videos + 10 audios, repeatable local/URL references, `--video-operation generate|edit|extend`, `--extend-direction`, and `--web-search`. Eco delivers MP4; native also supports `--output-format mov`. Use `makaron chat` to enhance an existing video without regeneration, or MCP `makaron_upscale_video` with one video URL and `resolution=720p|1080p|2k|4k` (up to 60s).
|
|
359
359
|
|
|
360
360
|
Wan 3.0 exposes two model choices: `--video-model wan-3.0` and the faster `--video-model wan-3.0-prime`. Both support 2-30s generation, native audio, up to 10 images + 5 videos + 5 audios, and 480p/720p/1080p/2K/4K. Pass `--video-resolution 2k|4k` to select the matching FlashVSR/Pro endpoint automatically; Pro is not a separate model selector. Use generation mode with feature references; typed edit/extend and the relaxed content-filter flag are not supported.
|
|
361
361
|
|
|
@@ -514,3 +514,5 @@ send_message "All done!"
|
|
|
514
514
|
FAL video models: **fal H3 Turbo** uses `minimax-h3-max` for single-start-frame I2V/T2V. **FAL H3 Max** uses the new selector `fal-h3-max`: native T2V or image/video/audio reference-to-video, default 768p, optional 480p/1080p, integer 5–15s; at most 9 images / 3 videos / 3 audios / 12 total. Reference video/audio each 2–15s and each modality totals at most 15s. Source-video modifications use generation with feature references, not typed edit/extend. Reference input tokens are billed in addition to output video; query current pricing.
|
|
515
515
|
|
|
516
516
|
The legacy `openai` image-model parameter now resolves to GPT Image 2.5 Flare.
|
|
517
|
+
|
|
518
|
+
Naming Seedance 2.5 in chat defaults to Seedance 2.5 Eco. Use `seedance-2.5-native` only when explicitly requesting native generation; Eco supports 720p/1080p/2K/4K final delivery.
|