makaron-cli 0.15.4 → 0.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "makaron-cli",
3
- "version": "0.15.4",
3
+ "version": "0.16.0",
4
4
  "description": "Give Claude Code a creative agent. Pass complete creative requests and source media to Makaron Chat.",
5
5
  "author": {
6
6
  "name": "Versa AI",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "makaron-cli",
3
- "version": "0.15.4",
3
+ "version": "0.16.0",
4
4
  "description": "Give Codex a creative agent. Pass complete creative requests and source media to Makaron Chat.",
5
5
  "author": {
6
6
  "name": "Versa AI",
package/README.md CHANGED
@@ -384,7 +384,7 @@ npx makaron-cli edit --image photo.jpg --out result.jpg "make it dramatic"
384
384
  npx makaron-cli edit --image-model gpt-image-2.5-flare --background transparent --out sticker.png "a magenta star sticker"
385
385
  ```
386
386
 
387
- Options: `--image`, `--image-model gemini|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image`, `--ref <file>` (up to 3), `--aspect <ratio>`, `--background auto|opaque|transparent`, `--out <path>`. Qwen Spicy supports 0–3 input images; legacy `qwen` requests map to it. Pony and WAI are retired. Transparent output routes strictly to GPT Image 2.5 Flare and is returned only when the provider supplies real PNG/WebP alpha.
387
+ Options: `--image`, `--image-model gemini|gemini-2.1|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image`, `--ref <file>` (model-specific limit), `--aspect <ratio>`, `--background auto|opaque|transparent`, `--out <path>`. Nano Banana 2.1 (`gemini-2.1`) supports up to 14 total input images and `--image-resolution 1K|2K|4K` (default 1K); it uses OpenRouter and does not retry or switch models automatically. Qwen Spicy supports 0–3 input images; legacy `qwen` requests map to it. Pony and WAI are retired. Transparent output routes strictly to GPT Image 2.5 Flare and is returned only when the provider supplies real PNG/WebP alpha.
388
388
 
389
389
  `wan2.7-image` uses Alibaba international for fast, approximately 1K generation and editing (default 6 credits/image). Failed or timed-out Wan requests are not automatically retried or switched to another model. Face identity can change. Example: `makaron edit --image portrait.jpg --image-model wan2.7-image --aspect 16:9 --out stadium.jpg "Place this woman in a baseball stadium, preserving her face."`
390
390
 
@@ -429,7 +429,7 @@ Options for `video create`: `--script "..."`, `--script-file <path>`, `--image <
429
429
 
430
430
  fal H3 Turbo: use `--video-model minimax-h3-max` for faster-than-real-time generation. It supports exactly 5/10/15 seconds at 480p/768p and defaults to native 768p, with no image for T2V or exactly one `--image` for I2V. It does not accept reference video/audio or multiple images.
431
431
 
432
- Seedance 2.5: use `--video-model seedance-2.5` for 4-30 second output at 480p/720p. `--image` accepts local files or URLs (up to 30), while repeatable `--video` and `--audio` accept up to 10 each. Use `--video-operation generate|edit|extend`, `--extend-direction forward|backward`, `--output-format mp4|mov`, `--web-search`, `--generated-audio` / `--no-generated-audio`, and `--relaxed-content-filter`. Edit/extend require a video reference. The Evolink route does not currently expose 4K output.
432
+ Seedance 2.5: use `--video-model seedance-2.5` or `seedance-2.5-eco` for 4-30 second videos generated at 480p then automatically upscaled to 720p, 1080p (default), 2K or 4K. Use `seedance-2.5-native` explicitly for native output at 480p (default) or 720p. `--image` accepts local files or URLs (up to 30), while repeatable `--video` and `--audio` accept up to 10 each. Use `--video-operation generate|edit|extend`, `--extend-direction forward|backward`, `--output-format mp4|mov`, `--web-search`, `--generated-audio` / `--no-generated-audio`, and `--relaxed-content-filter`. Edit/extend require a video reference. Eco delivers MP4; MOV is native-only. The native Evolink route does not expose 4K output.
433
433
 
434
434
  Wan 3.0: choose `--video-model wan-3.0` or the faster `--video-model wan-3.0-prime`. Both support 2-30 second generation, 480p/720p/1080p/2K/4K, and up to 10 images, 5 videos, and 5 audio references. Pass `--video-resolution 2k|4k` to use the matching FlashVSR/Pro endpoint automatically; Pro is not a separate model selector. Use generation mode with feature references; typed edit/extend and `--relaxed-content-filter` are not supported.
435
435
 
package/bin/makaron.mjs CHANGED
@@ -1935,9 +1935,11 @@ Usage:
1935
1935
 
1936
1936
  Options:
1937
1937
  --image <file|url> Base image to edit. Omit for text-to-image.
1938
- --ref <file|url> Additional reference image. Repeatable, up to 3.
1939
- --image-model <id> gemini, gemini-lite, qwen-spicy, openai, gpt-image-2.5-flare, gpt-image-2.5-sunburst, or wan2.7-image. Legacy qwen maps to qwen-spicy.
1938
+ --ref <file|url> Additional reference image. Repeatable; model-specific limit (Nano Banana 2.1: 14 total inputs).
1939
+ --image-model <id> gemini, gemini-2.1, gemini-lite, qwen-spicy, openai, gpt-image-2.5-flare, gpt-image-2.5-sunburst, or wan2.7-image. Legacy qwen maps to qwen-spicy.
1940
1940
  --skill <id> enhance, creative, wild, or captions.
1941
+ --image-resolution <id> 1K, 2K, or 4K; Auto chooses Nano Banana 2.1.
1942
+ --nsfw Route NSFW directly to Qwen Spicy.
1941
1943
  --aspect <ratio> Output aspect ratio, for example 1:1, 16:9, or 9:16.
1942
1944
  --background <mode> auto, opaque, or transparent.
1943
1945
  --out <file> Save the generated image to this path.
@@ -1992,7 +1994,8 @@ Recent model choices:
1992
1994
  5/10/15s; native 768p default or 480p; no video/audio/multi-image references.
1993
1995
  wan-3.0-prime Faster Wan 3.0 tier; 2-30s; 480p through 4k; multimodal refs.
1994
1996
  wan-3.0 Wan standard tier with the same public duration/resolution range.
1995
- seedance-2.5 4-30s; 480p/720p; generate/edit/extend and multimodal refs.
1997
+ seedance-2.5 Defaults to Eco; 4-30s; 1080p default, 720p/2k/4k; multimodal refs.
1998
+ seedance-2.5-eco 480p generation → ByteDance Fast; 1080p default, 2k/4k.
1996
1999
  minimax-h3 4-15s; 768p default or 2k; image/video/audio feature refs.
1997
2000
  grok T2V/reference generation plus typed edit/extend.
1998
2001
  sync-lipsync-v3 Exactly one video plus one MP3/WAV replacement track.
@@ -3030,7 +3033,12 @@ if (!command || command === '--help' || command === '-h' || command === 'help')
3030
3033
  editArgs.referenceImages = editArgs.referenceImages || [];
3031
3034
  editArgs.referenceImages.push(imageToArg(args[++i]));
3032
3035
  }
3036
+ else if (args[i] === '--image-resolution' && args[i + 1]) {
3037
+ const resolution = args[++i];
3038
+ editArgs.imageResolution = resolution;
3039
+ }
3033
3040
  else if (args[i] === '--aspect' && args[i + 1]) editArgs.aspectRatio = args[++i];
3041
+ else if (args[i] === '--nsfw') editArgs.isNsfw = true;
3034
3042
  else if (args[i] === '--background' && args[i + 1]) {
3035
3043
  const background = args[++i];
3036
3044
  if (!['auto', 'opaque', 'transparent'].includes(background)) {
@@ -3043,7 +3051,7 @@ if (!command || command === '--help' || command === '-h' || command === 'help')
3043
3051
  else promptParts.push(args[i]);
3044
3052
  }
3045
3053
  editArgs.editPrompt = promptParts.join(' ');
3046
- if (!editArgs.editPrompt) { console.error('Usage: makaron edit [--image <file|url>] [--image-model gemini|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image] [--ref <file>] [--aspect <ratio>] [--background auto|opaque|transparent] [--out <file>] "prompt"'); process.exit(1); }
3054
+ if (!editArgs.editPrompt) { console.error('Usage: makaron edit [--image <file|url>] [--image-model gemini|gemini-2.1|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image] [--ref <file>] [--aspect <ratio>] [--background auto|opaque|transparent] [--out <file>] "prompt"'); process.exit(1); }
3047
3055
  process.stderr.write('🎨 Generating...\n');
3048
3056
  const result = await callMcpTool(baseUrl, headers, 'makaron_edit_image', editArgs);
3049
3057
  saveMcpImage(result, outputPath);
@@ -3123,7 +3131,7 @@ if (!command || command === '--help' || command === '-h' || command === 'help')
3123
3131
  : ['wan3-prime', 'wan3.0-prime', 'wan30-prime', 'wan-3-prime', 'w3.0-video-prime', 'w3.0-video-prime-pro', 'wan-3.0-prime-pro', 'prime'].includes(videoModel)
3124
3132
  ? 'wan-3.0-prime'
3125
3133
  : (videoModel || 'fal-h3-max');
3126
- const isSeedance25 = selectedVideoModel === 'seedance-2.5';
3134
+ const isSeedance25 = ['seedance-2.5', 'seedance-2.5-eco'].includes(selectedVideoModel);
3127
3135
  const isWan30 = selectedVideoModel === 'wan-3.0' || selectedVideoModel === 'wan-3.0-prime';
3128
3136
  const isSeedanceModel = selectedVideoModel === 'seedance-fast' || selectedVideoModel === 'seedance-mini' || selectedVideoModel === 'seedance' || isSeedance25;
3129
3137
  const isMinimaxH3 = selectedVideoModel === 'minimax-h3';
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "makaron-cli",
3
- "version": "0.15.4",
3
+ "version": "0.16.0",
4
4
  "description": "Talk to Makaron Agent from the terminal — create projects, edit images, generate videos",
5
5
  "type": "module",
6
6
  "scripts": {
@@ -315,7 +315,7 @@ npx makaron-cli edit --image photo.jpg --out result.jpg "make it dramatic"
315
315
  npx makaron-cli edit --image-model gpt-image-2.5-flare --background transparent --out sticker.png "a magenta star sticker"
316
316
  ```
317
317
 
318
- Options: `--image`, `--image-model gemini|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image`, `--ref <file>` (up to 3), `--aspect <ratio>`, `--background auto|opaque|transparent`, `--out <path>`. Qwen Spicy supports 0–3 input images; legacy `qwen` requests map to it. Pony and WAI are retired. Transparent output routes strictly to GPT Image 2.5 Flare and fails instead of returning an opaque fallback. Wan 2.7 Image is an explicit fast ~1K route; do not automatically retry failures/timeouts, and do not promise exact face preservation.
318
+ Options: `--image`, `--image-model gemini|gemini-2.1|gemini-lite|qwen-spicy|openai|gpt-image-2.5-flare|gpt-image-2.5-sunburst|wan2.7-image`, `--ref <file>` (model-specific limit), `--aspect <ratio>`, `--image-resolution 1K|2K|4K`, `--background auto|opaque|transparent`, `--nsfw`, `--out <path>`. Image Model Capability prefers matching requirements, then relaxes output conflicts to deliver an image: GPT Image 2.5 Flare → Nano Banana 2.1 → Qwen Spicy. Omit the model for Auto: an 8:1 banner or an explicit resolution tier selects Nano Banana 2.1. Native ratio presets do not promise exact mathematical ratios or user pixel dimensions. For Agent-assessed NSFW, pass `--nsfw` to route directly to Spicy; never probe another provider first. Unknown IDs use Flare-first Auto. Output conflicts relax before submission: transparent 8:1 normally keeps 8:1 with opaque output; NSFW always uses Spicy regardless of unsupported dimensions. Preserve input content first and explain actual adjustments. See current MCP tool descriptions for the generated capability table. Classic Nano Banana 2 (`gemini`), Lite, Sunburst and Wan are explicit routes; `openai` means Flare and legacy `qwen` means Spicy. Pony/WAI are retired. Transparent output prefers Flare or explicitly selected Sunburst, and may relax alpha when other requirements conflict. Supplier failures never trigger automatic image resubmission or model switching.
319
319
 
320
320
  ### `video` — Standalone video tools (no project timeline)
321
321
 
@@ -355,7 +355,7 @@ Provider integration contract: every image passed to video generation is a featu
355
355
 
356
356
  fal H3 Turbo uses the `minimax-h3-max` selector, supports exactly 5/10/15 seconds at 480p/768p, and defaults to native 768p for faster-than-real-time T2V or one-start-image I2V.
357
357
 
358
- Seedance 2.5 uses `--video-model seedance-2.5` and supports 4-30s at 480p/720p, up to 30 images + 10 videos + 10 audios, repeatable local/URL references, `--video-operation generate|edit|extend`, `--extend-direction`, `--output-format mp4|mov`, and `--web-search`. The Evolink route does not currently expose 4K output.
358
+ Seedance 2.5 uses `--video-model seedance-2.5` or `seedance-2.5-eco` for 4-30s videos: generate at 480p, then automatically upscale to 720p, 1080p (default), 2K or 4K in one task. This is upscaled output. Use `seedance-2.5-native` explicitly for unenhanced native output at 480p (default) or 720p; the native Evolink route does not expose 4K. All Seedance 2.5 routes accept up to 30 images + 10 videos + 10 audios, repeatable local/URL references, `--video-operation generate|edit|extend`, `--extend-direction`, and `--web-search`. Eco delivers MP4; native also supports `--output-format mov`. Use `makaron chat` to enhance an existing video without regeneration, or MCP `makaron_upscale_video` with one video URL and `resolution=720p|1080p|2k|4k` (up to 60s).
359
359
 
360
360
  Wan 3.0 exposes two model choices: `--video-model wan-3.0` and the faster `--video-model wan-3.0-prime`. Both support 2-30s generation, native audio, up to 10 images + 5 videos + 5 audios, and 480p/720p/1080p/2K/4K. Pass `--video-resolution 2k|4k` to select the matching FlashVSR/Pro endpoint automatically; Pro is not a separate model selector. Use generation mode with feature references; typed edit/extend and the relaxed content-filter flag are not supported.
361
361
 
@@ -514,3 +514,5 @@ send_message "All done!"
514
514
  FAL video models: **fal H3 Turbo** uses `minimax-h3-max` for single-start-frame I2V/T2V. **FAL H3 Max** uses the new selector `fal-h3-max`: native T2V or image/video/audio reference-to-video, default 768p, optional 480p/1080p, integer 5–15s; at most 9 images / 3 videos / 3 audios / 12 total. Reference video/audio each 2–15s and each modality totals at most 15s. Source-video modifications use generation with feature references, not typed edit/extend. Reference input tokens are billed in addition to output video; query current pricing.
515
515
 
516
516
  The legacy `openai` image-model parameter now resolves to GPT Image 2.5 Flare.
517
+
518
+ Naming Seedance 2.5 in chat defaults to Seedance 2.5 Eco. Use `seedance-2.5-native` only when explicitly requesting native generation; Eco supports 720p/1080p/2K/4K final delivery.